All posts
guides

Veo 3.1 speed optimization for faster output

There is no inference steps setting. The real lever is picking a different Veo, and Fast and Lite give up different things.

By Flixly TeamMay 7, 2026
Veo 3.1 speed optimization for faster output

TL;DR

Veo 3.1 exposes prompt, duration, aspect ratio, resolution and audio generation. There is no inference steps, sampler or guidance setting. The real lever is variant choice: Veo 3.1 at ULTRA does references and audio, Veo 3.1 Fast at PREMIUM drops reference support, and Veo 3.1 Lite at STANDARD keeps references but drops audio. Pick the cheapest variant that still has the capability you need, usually Lite.

There is no inference steps setting. You cannot drop Veo from 50 steps to 22, because that control does not exist.

Veo 3.1 exposes five parameters: prompt, duration, aspect ratio, resolution, and whether to generate audio. That is the whole surface. No sampler, no step count, no CFG.

Which leaves one real lever for speed, and it is a good one: pick a different Veo.

Three Veos, and what each costs you

Model Tier Reference images Audio generation
Veo 3.1 ULTRA Yes Yes
Veo 3.1 Fast PREMIUM No Yes
Veo 3.1 Lite STANDARD Yes No

That table is the whole optimisation problem, and the interesting part is that Fast and Lite give up different things.

Veo 3.1 Fast drops reference-to-video. If your work depends on a character reference, Fast cannot do it at any price, and moving to it will silently ignore the thing holding your character together.

Veo 3.1 Lite keeps references but drops audio generation. It sits at STANDARD, the lowest tier of the three.

So the honest advice is not "use Fast". It is: work out which capability you actually need, then take the cheapest model that still has it.

Choosing correctly

You need a character reference. Veo 3.1 Lite first, since it is two tiers below Veo 3.1 and keeps the capability. Move up only if quality demands it. Do not use Fast; it will drop your reference.

You need generated audio. Fast or full 3.1. Lite has no audio parameter.

You need both reference and audio. Only full Veo 3.1 does both.

You need neither. Lite, and stop there.

Most people reaching for a "speed guide" are on ULTRA-tier Veo 3.1 for work that needs neither references nor audio, which is the expensive way to do a cheap job.

Beyond Veo

If speed genuinely is the priority, the Fast and Turbo variants across the whole catalog exist for exactly this: Seedance 2.0 Fast at STANDARD, which keeps references and first-to-last frame, and Kling 3.0 Turbo at PREMIUM.

Seedance 2.0 Fast is worth singling out. It is STANDARD tier and supports text-to-video, image-to-video, reference-to-video and first-to-last frame — the widest capability set at that tier.

The full catalog is at Models, reachable from Text to Video.

Things that genuinely affect generation time

Not settings, but real:

Duration. More seconds is more work. This is the largest honest factor.

Resolution. Higher costs more time and more credits.

Queue conditions. Load varies, which is why no article can promise "11 seconds per clip". Anyone quoting a fixed generation time measured it once on one day.

Failed attempts. By far the biggest time sink in practice. Three regenerations because the prompt was vague costs more than any model choice. Getting the prompt right first is the real optimisation.

The prompt is where the time actually goes

Since there are no sampler settings, iteration speed is dominated by how quickly you converge on a prompt that works.

Say what the camera does. Unspecified cameras invent movement, and that is the most common reason a clip gets binned and regenerated.

One action per short clip. Multiple actions in five seconds produce a rushed version of all of them.

Describe the end state, so the clip lands rather than cutting mid-motion.

Explore on a cheap model, finish on the expensive one. Get the wording right on Lite or a Fast variant, then run the final once on the model you actually want. This is the single biggest cost and time saving available, and it is a workflow rather than a setting.

What does not exist

No inference steps, no sampler, no CFG or guidance scale. None of the diffusion-level controls appear anywhere.

No frame rate control. You cannot set 24 fps.

No published per-model generation times. Speed varies with load and content, so tables listing seconds-per-clip per model invented them.

One model does expose a seed: Wan 2.7, along with a negative prompt and a prompt-extend option. It is the only entry with those, and worth knowing when repeatability matters.

The short version

Stop looking for a step count. Pick the cheapest Veo that still has the capability you need — usually Lite — and spend the saved effort on the prompt.

Each generation is quoted before it runs, so you can see the difference between the variants directly rather than trusting a table. Pack prices are on the pricing page.

Frequently Asked Questions

How do I lower inference steps on Veo 3.1?▾

You cannot, because that setting does not exist. Veo 3.1 exposes prompt, duration, aspect ratio, resolution and whether to generate audio. There is no step count, sampler, CFG or guidance scale anywhere in the platform.

What is the difference between Veo 3.1, Fast and Lite?▾

Veo 3.1 sits at ULTRA and supports both reference images and audio generation. Veo 3.1 Fast sits at PREMIUM and drops reference-to-video entirely. Veo 3.1 Lite sits at STANDARD, keeps reference support and drops audio generation. Fast and Lite give up different capabilities, so the cheaper option depends on what you need.

Which Veo should I use if I need a character reference?▾

Veo 3.1 Lite, which is two tiers below Veo 3.1 and keeps reference support. Do not use Veo 3.1 Fast for reference work, since it has no reference capability and will simply ignore the image.

What actually affects how long a generation takes?▾

Duration is the largest honest factor, then resolution, then queue load, which varies. Articles quoting a fixed time such as eleven seconds per clip measured it once on one day. In practice the biggest time sink is failed attempts, so getting the prompt right first beats any model choice.

Is there a faster option outside the Veo family?▾

Seedance 2.0 Fast is worth knowing about. It sits at STANDARD tier and supports text-to-video, image-to-video, reference-to-video and first-to-last frame, which is the widest capability set at that tier. Kling 3.0 Turbo at PREMIUM is the other speed-oriented option.

Can I set the frame rate to speed things up?▾

No. There is no frame rate control at generation time, so instructions to keep it at 24 fps for speed are describing a setting that is not there.

What is the best way to save time and credits overall?▾

Explore on a cheap model and finish on the expensive one. Get the prompt working on Lite or a Fast variant, then run the final version once on the model you actually want. That is a workflow rather than a setting, and it saves more than any parameter would.

Tools mentioned in this post

guidesveovideooptimization

Ready to create with guides?

Jump straight into Flixly's AI studio and try guides with 50+ models — free to start.