Gemini Omni Flash on Flixly: Google's Any-Input Video Model
Gemini Omni Flash is Google's first Omni-family model — video from text, an image, three reference images, or an existing clip, with audio generated in the same pass. It's live on Flixly.
TL;DR
Gemini Omni Flash is Google's first Omni-family video model. It generates 4, 6, 8, or 10 second clips at 720p, 1080p, or 4K with audio rendered in the same pass, and accepts a text prompt, a single image, exactly three reference images, or an existing video to transform. All modes are live on Flixly today.
Gemini Omni Flash is Google's first Omni-family video model, and it is live on Flixly — text-to-video, image-to-video, reference fusion, and video-input transformation.
What makes Omni different
- Any input in. A prompt alone, a single image, three reference images fused into one scene, or an existing video you want reshaped. The model was announced at Google I/O 2026 as "create anything from any input, starting with video."
- World knowledge. Omni connects generation to Gemini's understanding of physics, materials, and cause-and-effect — actions have consequences in the scene, objects behave like the things they depict.
- Audio is native. Every generation renders sound with the picture — no toggle, no extra cost, no separate dubbing step.
- Conversational editing. Upload footage and say what should change: swap the environment, alter the action, add objects, move the camera. The rest of the scene stays put.
Specs
| Property | Value |
|---|---|
| Modes on Flixly | Text-to-Video, Image-to-Video, Reference-to-Video (images and video input) |
| Duration | 4, 6, 8, or 10 seconds (with a video input, length follows your clip) |
| Resolution | 720p, 1080p, or 4K — 720p and 1080p cost the same |
| Aspect ratios | 16:9, 9:16 |
| Audio | native in every pass |
| References | up to 7 images, or 1 video |
| Prompt length | up to 20,000 characters |
How to use it
- Open the Video Generator and pick the tab you need — Text, Image, or Reference.
- Pick Gemini Omni Flash from the model selector.
- Set resolution and duration. Billing is per generation: a 4-second draft costs about half a 10-second take, and 1080p costs no more than 720p.
- For a transformation, drop your clip in the Reference tab and describe only the change you want — "make it rain, keep everything else the same" beats re-describing the whole scene.
- For reference fusion, attach exactly three images — a subject, a style, a setting — and describe how they combine.
Pricing notes
Omni Flash is billed per generation, not per second. Duration snaps to the 4/6/8/10 tier you picked, 720p and 1080p share a price, and 4K costs about double. A video-input job is a flat price per generation regardless of length — the estimate on the page always shows the exact cost before you generate.
Try it
Gemini Omni Flash is live in Text to Video, Image to Video, and Reference to Video.
Frequently Asked Questions
What is Gemini Omni Flash?▾
Gemini Omni Flash is the first model in Google's Gemini Omni family, built to create video from any kind of input — text, images, or an existing clip — and to refine results conversationally. It leans on Gemini's world knowledge, so scenes follow real-world physics and logic more closely than prompt-only video models.
How long and what resolution can a clip be?▾
Four, six, eight, or ten seconds, at 720p, 1080p, or 4K. 720p and 1080p cost the same, so 1080p is the sensible default; 4K costs roughly double. Pricing is per generation rather than per second.
Does it generate sound?▾
Yes — audio is native to every generation and there is no toggle. Music, ambience, effects, and speech come out of the same pass as the picture.
What does video input do?▾
Upload a clip in the Reference tab and describe the change — a new environment, different action, added objects, a shifted camera. The model transforms your footage while keeping the scene coherent. With a video input the output length follows your clip, and the job is billed at a flat rate per generation.
How do reference images work?▾
One image animates it (image-to-video). Three images fuse — a subject, a style, a scene — into one video. You can attach up to seven on Flixly; note that some routes require exactly one or exactly three, and Flixly automatically picks a provider that matches what you attached.
Where can I try it on Flixly?▾
Gemini Omni Flash is live in Text to Video, Image to Video, and Reference to Video on the Flixly dashboard. Pick it from the model selector, set resolution and duration, and generate.
