Gemini Omni Workflow to Build Full Campaigns
Follow the exact Gemini Omni steps inside Flixly to produce a full campaign from script to final export using Veo 3.1, Seedance 2.0, and Gemini 3.1 Flash TTS.
TL;DR
Build a 45-second campaign inside Flixly by chaining Text to Video with Veo 3.1, Lip Sync Video, Voice Cloning, and Music Generation. Total cost is quoted before each step runs. The workflow keeps every file in one account and removes export handoffs.
Campaign production costs keep rising
A 60-second social campaign needs script, visuals, voiceover, music, and captions. Hiring separate teams runs $1800-$3200 per piece. One missed sync between video and audio forces a full re-render and another $400 bill.
Where standard toolchains break
Most creators start with a text-to-image tool then export frames to a separate video model. File handoffs create resolution drops from 4K to 1080p. Lip sync must be added later, often requiring a third platform that charges per minute. The result is three invoices and three login pages.
The single-platform method
Flixly keeps every step inside one dashboard. Start with a prompt in the Text to Video tool powered by Veo 3.1. Generate a 15-second base clip at 1080p 30fps. Then route the same clip through Lip Sync Video using Gemini 3.1 Flash TTS for the voice track.
Add background score from the Music Generation tool. The model accepts a 10-second reference stem and extends it to match the video length exactly. Export the final file with Auto Captions baked in at 24-point bold.
Concrete workflow steps
- Write a 45-word script. Paste into Text to Speech to preview timing.
- Feed the script into Text to Video with Seedance 2.0. Set duration to 45 seconds.
- Open the output in Image to Video if you need to extend one shot by 12 seconds.
- Apply Voice Cloning on the generated audio track using a 30-second sample of the brand voice.
- Drop the finished video into Auto Captions for burned-in text.
Model choices and their specs
| Step | Model | Resolution | Max Length | Credit Cost |
|---|---|---|---|---|
| Script to voice | Gemini 3.1 Flash TTS | 48 kHz | 120 s | quoted in-app |
| Base video | Seedance 2.0 | 1080p | 45 s | quoted in-app |
| Lip sync | Kling 3.0 | 1080p | 60 s | quoted in-app |
| Music | Wan 2.7 | 44.1 kHz | 90 s | quoted in-app |
Edge cases and limits
Character consistency across 12 shots drops when you exceed 90 seconds total runtime. Switch to Reference to Video and upload the same face reference image for every generation. Music stems longer than 120 seconds must be split and stitched manually inside the editor.
Next step
Open the Text to Video page and paste your first campaign script to see the full chain run in under four minutes.
Frequently Asked Questions
How many credits does a 45-second campaign use?▾
Every step of a full 45-second piece with voice, video, music, and captions is quoted before it runs, using the models listed in the table above.
Can I reuse the same voice clone across multiple campaigns?▾
Yes. Upload the 30-second sample once. The cloned voice remains available in your account for any future Lip Sync Video or Text to Speech jobs.
What happens if the generated video length does not match the voice track?▾
Trim the longer file inside the Video to Video tool before final export. The platform aligns both tracks automatically after the trim.
Is Gemini 3.1 Flash TTS required or can I use another model?▾
You can substitute ElevenLabs via the Alternatives page, but Gemini 3.1 Flash TTS stays inside the same credit balance and produces output in one fewer click.


