All posts
tutorials

How to Make Video with AI Tools

Every guide starts with a tool. That's the wrong end. Start with what you already have, and most of the choices disappear.

By Flixly TeamMay 20, 2026
How to Make Video with AI Tools

TL;DR

Choose the tool by what you already have: nothing means text to video, a picture means image to video, a specific person means reference to video across thirteen models, a clip whose movement you want means motion transfer, and long footage means cutting rather than generating. The controls are prompt, duration, aspect ratio and resolution, so the prompt does the work. Finish in order: voice, lip sync, music, upscale, captions.

Every guide on this starts with a tool. That is the wrong end.

Start with what you already have, because that single question decides which tool is correct and eliminates most of the choices immediately.

  • Nothing — text to video.
  • A picture — image to video.
  • A picture of someone specific who must stay themselves — reference to video.
  • A clip whose movement you want — motion transfer.
  • A long video — cut it, do not generate anything.

Pick the wrong branch and you will fight the tool for an hour. Pick the right one and most of the work disappears.

Starting from nothing

Text to Video generates a clip from a description. Forty-four video models are available, and switching between them is a dropdown rather than another subscription.

The controls are prompt, duration, aspect ratio and resolution. That is genuinely the whole surface for most models, plus a negative prompt on some. There is no motion strength, no guidance scale, no seed on all but one model.

So the prompt does the work. Write it as a shot: subject, then action, then camera, then light.

"Barista pours milk into a cup, slow overhead shot, warm window light, camera static."

Two habits worth forming immediately. Say what the camera does, or the model invents movement. And describe one action, because five seconds does not fit three.

Starting from a picture

Image to video animates a still. Usually a faster route to a specific look than describing that look from scratch, since the picture already contains it.

Write about what should move rather than what is in the frame. The model can see the frame.

If you know both the start and end of a shot, first-to-last frame generates the movement between two images you supply. Seedance 2.0 and Seedance 2.0 Fast only, and it is the most predictable option on the platform.

Keeping a person or product consistent

This is the problem that defeats most first attempts, and text will not solve it. "Dark-haired woman in a red jacket" describes millions of people, and the model picks a different one each generation.

Thirteen models accept a character or product reference through reference to video. Showing beats describing, every time.

For a continuous sequence rather than separate shots, the Long Video Generator chains segments so each one continues from the last, carrying both the character references and the previous clip. That is how identity and environment stay stable across a whole piece.

For discrete scenes that cut between locations, the Series Generator is the right shape instead.

Starting from footage you already have

If you have long video, generating new footage is usually the expensive answer to a question you have already solved.

The Shorts Generator takes a YouTube URL, an upload or a transcript, finds clips worth cutting, and renders word-by-word captions. Nothing is generated, so it is dramatically cheaper than creating equivalent footage.

A back catalogue of long video is a reel supply already paid for.

The finishing pass

Same regardless of how you got the footage, and order matters.

Voice firstText to Speech, or Voice Cloning for a consistent voice across a series.

Then lip syncLip Sync matches a mouth to an existing track, so it has to come after the audio is final.

Then musicMusic Generation for a bed that will not get the upload muted for rights.

Then upscale — the Video Upscaler, after the edit rather than before, so you are not paying to sharpen frames you cut.

Captions lastAuto Captions, once picture and audio are both final. Most social video is watched muted, so this is not optional.

Reversing any of these means redoing work. Lip syncing before the voice is settled is the most common and most expensive version of that mistake.

What to expect, honestly

Short clips work, long ones drift. Models hold coherence over a few seconds and lose it over many. Generate short and cut together rather than asking one generation to cover a whole sequence.

The first attempt is rarely the one you use. Explore on Fast and Turbo variants, finish on the full model.

Cost is quoted before each generation, based on model and settings, so you never have to work from someone's out-of-date table. Pack prices are on the pricing page.

Framing is decided at generation. Ask for vertical up front if that is where it is going; cropping afterwards throws away most of the frame.

The shortest useful version

Work out what you already have. Pick the branch that matches. Write the prompt as a shot with the camera specified. Keep it short. Finish in the order above.

The full catalog is at Models, and History keeps everything you generate, which matters more than it sounds the first time someone asks for last week's version.

Frequently Asked Questions

Which AI video tool should I start with?

Decide by what you already have. Nothing means text to video. A picture means image to video. A picture of someone who must stay themselves means reference to video. A clip whose movement you want transferred means motion transfer. Long footage means cutting it with the Shorts Generator rather than generating anything.

What settings can I change when generating video?

Prompt, duration, aspect ratio and resolution on most models, plus a negative prompt or reference images on some. There is no motion strength or guidance scale, and only one of the 37 video generation models exposes a seed. The prompt is doing nearly all the work.

How do I write a good video prompt?

Write it as a shot: subject, then action, then camera, then light. Always say what the camera does, since an unspecified camera invents movement. Describe one action rather than three, because a few seconds does not fit more than one.

How do I keep the same character across shots?

Show rather than describe. Text like "dark-haired woman in a red jacket" describes millions of people and the model picks a different one each time. Thirteen models accept a character reference. For a continuous sequence the Long Video Generator chains segments so each continues from the last, and for discrete scenes the Series Generator is the right shape.

In what order should I add voice, music and captions?

Voice first, then lip sync, then music, then upscale, then captions. Lip sync matches a mouth to an existing audio track, so doing it before the voice is settled means doing it twice. Upscale after the edit so you are not paying to sharpen frames you cut.

Why do longer clips look worse?

Models hold coherence over a few seconds and lose it over many, so drift appears in the middle of long generations. Generate several short clips and cut them together. Short clips also fail more cheaply when an attempt does not work.

How much does generating a video cost?

It depends on model and settings, and the app quotes the exact cost before each generation rather than you working from a published table. Current credit pack prices are on the pricing page.

Tools mentioned in this post

tutorialsvideotext-to-videoworkflow

Ready to create with tutorials?

Jump straight into Flixly's AI studio and try tutorials with 50+ models — free to start.