I’ve been playing around with the latest Gemini video model, and Gemini Omni 1.1 Flash feels like one of those updates that is more about fixing the annoying parts of the first version than completely reinventing it.
The original Gemini Omni Flash only came out in June, and now Google has already moved on to version 1.1. The biggest changes are pretty easy to understand: longer videos, better control over scenes, video references, and a cheaper way to test ideas before spending more on a final render.
So, What Is Gemini Omni 1.1 Flash?
Gemini Omni 1.1 Flash is Google’s multimodal video generation model. You can work with text, images, audio, and video as inputs, rather than being limited to a simple text-to-video prompt.
It can generate short videos with synchronized audio, including dialogue, sound effects, and ambient sounds. It supports 360p, 720p, 1080p, and 4K output, although the 1080p and 4K options are upscaled rather than natively rendered at those resolutions.
The model also supports several different ways of creating video, including text-to-video, image-to-video, reference-to-video, and conversational editing.
That last part is probably what I find most interesting.
The New Features That Actually Matter
The biggest upgrade is video extension.
The previous version was basically limited to 10-second clips. With Omni 1.1 Flash, you can extend clips in 10-second increments up to 40 seconds in total. Even better, the model can look at up to 10 seconds of the previous footage when continuing a scene.
That sounds like a small technical change, but it can make a big difference. Instead of the next shot suddenly changing the character’s position, lighting, or overall environment, the model has more context to work with.
Another feature I really like is first-and-last-frame interpolation. You can give the model a starting frame and an ending frame, and it generates the movement between them. This could be useful for things like camera orbits, zoom shots, and timelapse-style transitions.
There’s also a new 360p draft mode. It’s around 60% faster and costs roughly a third of the regular price, which makes a lot of sense for experimentation. I’d much rather test several ideas at lower quality before paying for a final version.
And Omni 1.1 Flash can now use short video clips as references, instead of relying only on still images. That should make it easier to keep a character, movement style, or visual idea consistent between different shots.
The Conversational Editing Is Probably the Best Part
One thing that makes Omni 1.1 Flash different from a basic text-to-video tool is the way editing works.
Instead of generating a completely new video every time something looks wrong, you can continue the interaction and tell the model what you want to change.
For example, you can ask it to change the background, alter the time of day, or make a character look toward the camera while keeping the rest of the scene relatively consistent.
That feels much closer to actually editing a video than starting over with a new prompt every time.
The same idea applies to extending a scene. Generate a clip, continue it, then continue it again. In theory, you can gradually build a longer sequence instead of trying to describe an entire video in one huge prompt.
Why I Like the 360p Draft Mode
The new draft mode might actually be one of the most useful changes for casual creators.
AI video generation can get expensive when you keep regenerating clips just because the camera movement isn’t quite right or a character doesn’t look the way you expected. With a cheaper 360p draft, you can focus on the idea first instead of worrying about final quality.
My workflow would probably be something like this: generate a few versions at 360p, pick the one that actually works, and only then worry about getting the final video into better quality.
And there’s another option if the final clip still looks a little soft.
Instead of generating everything at the highest resolution from the beginning, you can take a good low-resolution result and enhance it afterward. I’ve been using HitPaw VikPea AI Video Enhancer for this kind of workflow, especially when I want to clean up AI-generated footage and make the details look sharper.
It supports AI video upscaling, so it can be useful when you already have a clip you like but don’t want to keep spending credits regenerating the same scene just to get a higher-resolution version. For me, that makes the whole generate → test → enhance workflow a little more practical.
Is It Actually Better Than the Original?
I think the answer is yes, but mainly because it gives you more control.
The original Omni Flash already had the basic idea of generating and editing video through natural language. Version 1.1 adds the things that were missing when you actually tried to use it seriously: longer scenes, better continuation, frame control, video references, and a cheaper way to experiment.
It still isn’t perfect, though. Complex motion can be inconsistent, and text generated inside videos isn’t always reliable. And if you need genuinely native 4K generation or much longer continuous scenes, there are still reasons to look at other video models.
But for short creative videos, concept development, social content, and workflows where you expect to make lots of small changes, Omni 1.1 Flash looks much more practical than the first version.
For me, that’s probably the biggest takeaway from this update. It doesn’t completely change AI video generation. It just makes the process a lot less frustrating.
I’m curious to see how people actually use the longer scene extension and first/last-frame controls once more creators start experimenting with them.


