Images to Video Maker Guide
There is no reference strength setting and no motion strength. Which leaves two levers that matter more: which picture you upload, and what you write about it.

TL;DR
There is no reference strength or motion strength setting. Across all 37 video generation models the parameters are prompt, duration, aspect ratio and resolution, plus the image. So results come from the source picture and the prompt. Choose an image with room to move, clear subject separation and no motion blur, then write about what should happen rather than what is visible, and always say what the camera does. Thirty-five models accept an image input.
There is no reference strength setting. No motion strength either.
Those two numbers appear in nearly every guide on this topic, usually with confident values like 0.82 and 0.55. They do not exist. Across all 37 video generation models the parameters are prompt, duration, aspect ratio and resolution, plus the image itself.
Which changes the advice completely, because if you cannot dial the amount of motion, the only levers left are which picture you upload and what you write about it. Both are more powerful than a slider would have been.
The four things that decide the result
The image. Everything inherits it. A soft, badly lit or awkwardly cropped source produces a soft, badly lit, awkwardly cropped video. This is where most of the quality is decided, before any generation runs.
The prompt. Not a description of the picture, which the model can already see. A description of what should happen: the movement, and what the camera does while it happens.
Duration. Short holds together, long drifts. Models lose coherence over time, so a four to five second clip is where motion quality peaks.
The model. Thirty-five video models accept an image input. They differ in how much they respect the source versus reinventing it.
Write for motion, not for description
The most common mistake is describing the photo.
"A woman in a red coat standing on a bridge at sunset"
The model has the photo. It knows. That prompt spends its whole budget restating what is already visible and says nothing about what should move, so the model invents movement of its own.
"She turns her head slowly to look down the river. Camera holds still. Hair moves in the wind."
That works, because every clause is an instruction. Subject motion, camera behaviour, secondary motion.
Always say what the camera does. "Camera static" is the single most useful phrase in image-to-video, because an unspecified camera drifts, and drift is what makes a still-derived clip look wrong.
Choosing the right starting image
Some pictures animate well and some fight you.
Room to move. A tightly cropped face has nowhere to turn. Leave headroom and space in the direction of intended motion.
Clear subject separation. Where the subject blends into a busy background, models blend them further once things start moving.
Sharp source. Motion blur in the input becomes smeared motion in the output, and no amount of prompting recovers it.
Frontal or three-quarter faces. Profiles have no second eye to work from, so turns tend to invent one.
If you are generating the source image anyway, generate it with the motion in mind. A still composed for animation is a different picture from a still composed as a still.
Which tool for which job
| What you want | Where to go |
|---|---|
| Motion from one picture | Image to Video |
| A subject that stays itself across several shots | Reference to Video, thirteen models |
| A specific start and end point | First-to-last frame, Seedance 2.0 and 2.0 Fast only |
| A specific performance transferred onto a character | Motion Control |
| Moving key art from a still | Motion Poster |
The distinction that matters most: image-to-video invents the motion, motion transfer supplies it. If you already have a clip showing the movement you want, transferring it beats describing it.
Fixing a clip that came back wrong
Since there are no settings to tune, everything here is a change of input or wording.
Subject morphs or melts. Duration is too long for the model. Shorten it, or generate two short clips instead of one long one.
Camera drifts unpleasantly. Add an explicit camera instruction. This fixes more bad clips than anything else.
Face changes partway. The source was too small or too soft for the model to hold. Start from a higher-quality crop, or switch to a reference-to-video model built for identity.
Motion is too subtle. Describe a larger action. "Turns to face the camera" rather than "shifts slightly".
Motion is chaotic. Describe one action instead of several, and say what stays still as well as what moves.
Sound comes last
Generated clips are silent unless the model supports audio and you asked for it.
Add voice with Text to Speech or Voice Cloning, then align mouth movement with Lip Sync. Always in that order — lip sync matches a mouth to an existing track, so generating audio afterwards means redoing it.
Auto Captions last of all, once the picture and audio are final.
The short version
Pick a picture with room to move and enough sharpness to survive motion. Write about what happens, not what is visible. Say what the camera does. Keep it short.
There is no strength value to get right, which sounds like less control and is actually more, because the levers you do have are the ones that matter.
The catalog is at Models, and each generation is quoted before it runs. Pack prices are on the pricing page.
Frequently Asked Questions
What should I set reference strength and motion strength to?▾
Neither exists. Across all 37 video generation models the parameters are prompt, duration, aspect ratio and resolution, plus the image itself. Guides quoting values like 0.82 or 0.55 are describing controls no model exposes. The real levers are which image you upload and what you write about it.
How should I write the prompt for image to video?▾
Describe what should happen, not what is in the picture. The model can already see the image, so restating it wastes the prompt and leaves the motion unspecified. Write the subject's movement, then what the camera does, then any secondary motion such as hair or fabric.
Why does my clip drift or wander?▾
Usually because the camera was left unspecified, so the model invented movement. Adding "camera static" or another explicit camera instruction fixes more bad image-to-video clips than any other change.
What makes a good source image?▾
Room to move in the direction of intended motion, clear separation between subject and background, a sharp source without motion blur, and frontal or three-quarter faces rather than profiles. A tightly cropped face has nowhere to turn, and blur in the input becomes smeared motion in the output.
Why does the face change partway through the clip?▾
The source was too small or too soft for the model to hold identity across frames. Start from a higher-quality crop, or use one of the thirteen reference-to-video models, which are built to keep a subject consistent.
Should I use image to video or motion transfer?▾
Image to video invents the motion from your description. Motion Control transfers a performance from a driving video onto a character image. If you already have a clip showing the movement you want, transferring it beats describing it.
How do I add voice to a generated clip?▾
Generate the silent clip first, add voice with Text to Speech or Voice Cloning, then align the mouth with Lip Sync, and caption last. Lip sync matches a mouth to an existing audio track, so generating the audio afterwards means doing the sync twice.

