Reference to Video Tutorial 2026
A reference image is not a strength value. You either give the model one or you don't, and thirteen of the 37 models can accept it.
TL;DR
Thirteen of the 37 video generation models accept a character reference: Seedance 2.0, 2.0 Fast and 2.5, Veo 3.1 and 3.1 Lite, Wan 2.7, 3.0 and 3.0 Prime, Kling 3.0 Motion Control, MiniMax H3, Gemini Omni Flash, and both Happy Horse models. Kling 3.0 itself does not. There is no reference strength value, so results come from the reference image: frontal or three-quarter, large in frame, evenly lit, neutral, sharp, and the exact same file every time.
A reference image is not a strength value. You either give the model one or you do not.
There is no 0.75 to tune, no landmark error percentage to chase, and no encoder setting to configure. What decides whether a character holds across a clip is the quality of the image you hand over and which model you hand it to.
Both of those you control completely, which is better than a dial would have been.
Which models actually accept a reference
Thirteen of the 37 video generation models. Getting this wrong is the most common reason people conclude reference-to-video "does not work":
Seedance 2.0, 2.0 Fast, and 2.5 · Veo 3.1 and 3.1 Lite · Wan 2.7, 3.0 and 3.0 Prime · Kling 3.0 Motion Control · MiniMax H3 · Gemini Omni Flash · Happy Horse and Happy Horse 1.1
Notably absent: Kling 3.0 itself. It is text-to-video and image-to-video only. If you have been uploading a face to Kling 3.0 and watching it ignore you, that is why — the capability is on Kling 3.0 Motion Control instead.
Start at reference to video.
The reference image is the whole job
Everything inherits this file, so it deserves more care than the prompt.
Frontal or three-quarter. Profiles give the model one eye and half a jaw to work from, so it invents the rest and the invention changes every frame.
Large in frame. A face occupying a small part of a wide shot carries little detail. Crop in.
Evenly lit. Hard shadows get baked in as facial structure. Flat, even light gives the model the actual shape.
Neutral expression. A big smile in the reference tends to persist through shots where it does not belong.
Sharp. Softness in the reference becomes mush in motion, and no prompt recovers it.
If you are generating the reference rather than photographing it, generate it deliberately: a clean portrait built for this purpose beats a frame grabbed from something else.
Reuse the exact same file across every generation. Not a similar image, the same one. Two similar references produce two similar-but-different people, which is precisely the failure you are trying to avoid.
Writing the prompt around the reference
The model can see the face. Do not describe it.
Spending prompt words on "brown eyes, sharp jawline, dark hair" competes with the image rather than supporting it, and can pull the result away from the reference you supplied.
Write about what the character does, where they are, and what the camera does:
"She walks slowly along the harbour wall, looking out to sea. Overcast light. Camera tracks alongside."
Keep the setting and lighting wording consistent across shots, change only the action. That is how a set of clips reads as one piece.
Why identity still drifts
Three real causes, none of which is a missing setting.
Duration. Identity holds over a few seconds and degrades over many. If a face survives four seconds and dissolves by ten, generate shorter clips and cut them together.
Reference quality. Covered above, and it is the most common cause by a wide margin.
Wrong model. If the model does not support references, the reference does nothing at all.
When reference-to-video is not the right tool
A continuous sequence rather than separate shots. The Long Video Generator chains segments so each carries the previous one's finished clip and continues from its final frame, with references riding along throughout. That holds the environment as well as the character, which references alone cannot do.
A specific performance. Motion Control transfers movement from a driving video onto a character image. When you have a clip of the motion you want, transferring beats describing.
Discrete scenes across locations. The Series Generator is shaped for that.
Precise start and end points. First-to-last frame, on Seedance 2.0 and 2.0 Fast only.
What is not real
No reference strength. No value scaling how much the reference is respected, on any model.
No landmark or drift metric. No published figure for facial drift per model, no percentage to keep below. Articles quoting one invented it.
No dedicated reference encoder you can configure, and no per-region masks anywhere.
One model does expose a seed: Wan 2.7, which also takes a negative prompt. It is the only one, which makes it the choice when repeatability matters more than anything else.
A workflow that holds
- Build one strong reference image and save it.
- Pick a model from the thirteen.
- Write about action and camera, never about the face.
- Generate short. Judge identity honestly at the end of the clip, not the start.
- Reuse the identical reference for every subsequent shot.
Step 1 carries the result. Everything after it is repetition.
The catalog is at Models, each generation is quoted before it runs, and pack prices are on the pricing page.
Frequently Asked Questions
Which models accept a reference image?▾
Thirteen of the 37: Seedance 2.0, 2.0 Fast and 2.5, Veo 3.1 and 3.1 Lite, Wan 2.7, 3.0 and 3.0 Prime, Kling 3.0 Motion Control, MiniMax H3, Gemini Omni Flash, and both Happy Horse models. Kling 3.0 itself is text-to-video and image-to-video only, so uploading a face to it does nothing.
What reference strength should I set?▾
There is no reference strength on any model. You either supply a reference or you do not. What decides the result is the quality of the image and which model receives it, both of which you control completely.
What makes a good reference image?▾
Frontal or three-quarter rather than profile, large in frame, evenly lit without hard shadows, a neutral expression, and sharp. Softness in the reference becomes mush in motion and no prompt recovers it. Reuse the exact same file every time, since two similar references produce two similar-but-different people.
Should I describe the character in the prompt too?▾
No. The model can see the face. Describing it competes with the image and can pull the result away from the reference you supplied. Write about what the character does, where they are and what the camera does, keeping setting and lighting wording consistent across shots.
Why does the face still drift?▾
Three causes: the clip is too long, since identity holds over a few seconds and degrades over many; the reference image is weak; or the model does not support references at all. None of these is a missing setting.
When should I use something other than reference-to-video?▾
Use the Long Video Generator for a continuous sequence, since chaining holds the environment as well as the character. Use Motion Control when you have a clip of the movement you want transferred. Use the Series Generator for discrete scenes across locations, and first-to-last frame when you know the exact start and end images.
Is there a drift or landmark accuracy metric?▾
No. There is no published figure for facial drift per model and no percentage to keep below. Articles quoting landmark error rates invented them. One model, Wan 2.7, does expose a seed along with a negative prompt, which makes it the choice when repeatability matters most.

