I mainly focused on multi-round editing and character consistency.
Ease of use: very low friction.
You can generate directly from a prompt in Gemini or Flow. What matters most isn’t technical parameters, but how clearly you can describe the camera, motion, and mood you want.
You can simply write:
“A person walking down the street.”
But something like:
“Start on a medium shot and slowly push in. The person is wearing a red coat, and neon lights reflect off the wet street after the rain.”
will usually give you a more stable result.
The biggest improvement is that you don’t have to keep rerolling the whole clip.
With earlier AI video tools, if one part was wrong, you often had to regenerate everything.
Gemini Omni feels more like continuing to edit the result you already have.
If the character looks right but you want a different background, a new camera angle, or different lighting, you can simply keep giving instructions instead of starting over.
Character consistency is also much better than before. Across multiple shots, faces, clothing, and positioning are less likely to suddenly change, which matters a lot for short films and advertising.
But complex scenes still break.
Simple gravity, cloth motion, and object movement often look fairly natural. Once you introduce multiple objects colliding or more complicated cause-and-effect relationships, the scene can start to fall apart.
Text is still an obvious weak point.
If a reference image contains complicated text, it may become blurry, distorted, or simply incorrect after generation.
There’s also a roughly 10-second maximum per generation, so a complete short film still needs to be assembled from multiple clips.
What works well
- Conversational editing is genuinely useful.
You can keep refining the same clip instead of regenerating everything from scratch.
- Flexible inputs.
Text, images, audio, and video can all be used together.
- Characters and scenes stay more stable.
Continuity across multiple shots is noticeably easier to maintain.
- Video and audio can be generated together.
That cuts down on separate voiceover and ambient-sound work afterward.
- Strong Google ecosystem integration.
The workflow between Gemini, Flow, and YouTube feels relatively natural.
What doesn’t work as well
- Individual generations are still short.
A roughly 10-second limit is a clear constraint for longer content.
- Complex physics still fail.
Multi-object interaction, collisions, and more complicated causal motion remain unreliable.
- Text generation isn’t dependable.
This is especially noticeable with complicated English or Chinese text.
- Fine-grained control is limited.
Conversational editing is good for broader changes, but it still can’t replace frame-by-frame control in professional editing software.
- You need a paid plan for the full experience.
To properly use the main capabilities, you’ll generally need to move beyond the free tier.
Best for
Content creators and video bloggers
Useful for quickly creating B-roll, creative assets, and short-form clips.
Advertising and marketing teams
A strong fit for concept videos, pitches, and testing visual directions.
Education and explainer creators
Useful for turning abstract ideas into animated demonstrations quickly.
Film and animation pre-production teams
Convenient for testing storyboards, camera movement, and scene ideas.
Not ideal for
Anyone who needs a film-ready final result
It still can’t fully replace live-action production, CG, or professional post-production.
Users who need frame-by-frame control
Conversational editing is better for macro-level changes than precise editing.
Projects that depend heavily on text inside the image
Text rendering still isn’t reliable enough.
Anyone who needs continuous long-form video
For now, you still need to generate multiple clips and stitch them together.
Comments (0)