Why single continuity images break video output
A reference image tells the model who this is. It does not tell the model what happens at second three. Those are different jobs, and references only do the first.

TL;DR
A reference image is an identity anchor, not a per-frame instruction: it tells the model who the subject is, not what happens at a given second. References carry no timestamps, no spacing and no positional weighting. Seedance 2.0 and 2.0 Fast accept up to nine reference images, the largest set available, while most of the thirteen reference-capable models take one. For genuine positional control use first-to-last frame on Seedance 2.0 or 2.0 Fast, and for long sequences chain segments in the Long Video Generator.
The assumption is that one reference image locks a character for the whole clip.
It does not, and understanding why saves a lot of wasted generations.
A reference image tells the model who this is. It does not tell the model what happens at second three. Those are different jobs, and references only do the first one.
What a reference actually does
It is an identity anchor, not a per-frame instruction.
The model reads the reference for the things that make a subject recognisable — face structure, hair, clothing, the shape of a product — and carries that into whatever motion your prompt describes. It is strong at that.
What it does not do is constrain the shot. If your prompt says the character walks across a room, the reference does not decide how they walk, where the camera goes, or what the third second looks like. The prompt does.
So drift across a clip is not usually a reference problem. It is a duration problem, and the fix is a shorter clip.
References carry no timestamps
Worth stating flatly, because a lot of writing on this assumes otherwise.
You cannot place a reference at 0 s, another at 1.8 s and a third at 3.6 s. There is no timestamp field, no spacing interval, no weighting by position. You supply a set of images and the model uses the set.
Which means all the advice about clustering references at peak motion beats, or spacing them at 0.9-second intervals, is describing a control that does not exist.
How many references you can actually supply
This is where models genuinely differ.
Seedance 2.0 and Seedance 2.0 Fast take up to nine reference images. That is the largest reference set available anywhere here, and it is the reason to reach for them when identity really matters.
Most of the others take one. Seedance 2.5, Gemini Omni Flash, Wan 3.0 and Wan 3.0 Prime each accept a single reference image.
Thirteen models accept a reference in total, through reference to video. Note that Kling 3.0 is not one of them — that is Kling 3.0 Motion Control, a separate entry. Uploading a reference to plain Kling 3.0 achieves nothing.
If you have nine references to give, use them on the multiple angles of one subject rather than nine different things. Front, three-quarter and profile of the same character beats nine unrelated frames.
The one positional control that is real
First-to-last frame, on Seedance 2.0 and 2.0 Fast only.
You supply the opening frame and the closing frame, and the model fills between them. This genuinely constrains position, because you have fixed both ends. It is the only mechanism here that does.
It is also the right tool for chaining: shot two starts on shot one's final frame, so the world carries forward rather than being re-imagined.
The mechanism that actually holds a long sequence
If your problem is a character staying stable across a minute rather than across five seconds, references are the wrong layer entirely.
The Long Video Generator chains segments, each continuing from the previous one's final frame. That keeps the environment stable too — lighting, background, props — which a reference image never addresses, because it only ever described the subject.
Making the reference itself better
Since you get one shot at identity, the reference quality matters more than the count.
Generate the reference deliberately. Build it in image to image or text to image, pick the one you actually like, and reuse that exact file everywhere. Two similar references produce two similar-but-different characters.
Neutral beats dramatic. A clearly lit, front-facing frame reads more reliably than a moody three-quarter with half the face in shadow.
Match the aspect ratio of the clip you are generating, so the model is not rescaling before it starts.
Do not describe the face in the prompt once you have supplied a reference. Text and image compete, and text is the lossier of the two. Write about action and camera instead.
What is not real
No motion brush, no per-region masks. You cannot brush the collar or the feet to stabilise them.
No reference timestamps or weighting. Covered above, but it is the most common invented control in this area.
No retention percentages. Figures like "identity retention drops from 94 percent at frame 1 to 61 percent at frame 24" are not measurements of anything. No such metric is published or computed.
No seed on most models. Wan 2.7 is the only video model exposing one, which matters because "generate twice with different seeds to verify" is not available anywhere else.
Kling 3.0 does not take references. Kling 3.0 Motion Control does.
The short version
A reference sets identity, not choreography. There are no timestamps. Seedance 2.0 and 2.0 Fast take up to nine images and most models take one, so pick the model from how much identity control you need.
For position, use first-to-last frame. For length, chain segments rather than stretching one generation.
The catalog is at Models, and each generation is quoted before it runs.
Frequently Asked Questions
Why doesn't one reference image hold a character for a whole clip?▾
Because a reference sets identity, not choreography. It tells the model what makes the subject recognisable, then the prompt decides the motion, camera and everything that happens second by second. Drift across a clip is usually a duration problem rather than a reference problem, and the fix is a shorter clip.
Can I place references at specific timestamps?▾
No. There is no timestamp field, no spacing interval and no weighting by position. You supply a set of images and the model uses the set. Advice about clustering references at peak motion beats or spacing them at 0.9-second intervals describes a control that does not exist.
How many reference images can I supply?▾
Seedance 2.0 and Seedance 2.0 Fast take up to nine, the largest reference set available anywhere here. Most of the others take one, including Seedance 2.5, Gemini Omni Flash, Wan 3.0 and Wan 3.0 Prime. Thirteen models accept a reference in total.
Does Kling 3.0 accept reference images?▾
No. That is Kling 3.0 Motion Control, a separate catalog entry. Uploading a reference to plain Kling 3.0 achieves nothing, which is a common source of the conclusion that reference-to-video does not work.
Is there any way to control position rather than identity?▾
First-to-last frame, on Seedance 2.0 and 2.0 Fast only. You supply the opening and closing frames and the model fills between them, which genuinely constrains position because both ends are fixed. It is also the right mechanism for chaining, since shot two can start on shot one's final frame.
What holds a character stable across a long sequence?▾
Chaining rather than references. The Long Video Generator runs segments where each continues from the previous one's final frame, which keeps lighting, background and props stable too. A reference image never addressed those, because it only ever described the subject.
How do I make the reference image itself better?▾
Generate it deliberately, pick the one you actually like and reuse that exact file everywhere, since two similar references produce two similar-but-different characters. Favour neutral, clearly lit and front-facing over dramatic, match the clip's aspect ratio, and stop describing the face in the prompt once a reference is supplied.
Is there a motion brush for stabilising specific areas?▾
No. There is no motion brush and no per-region masking, so you cannot brush the collar or the feet to hold them steady. Nor are there identity-retention percentages: figures like 94 percent at frame one dropping to 61 percent at frame 24 are not measurements of anything published or computed.

