All posts
guides

First to Last Frame AI for Smooth Video

Two models support first-to-last frame, not eight. Seedance 2.0 and 2.0 Fast. Uploading two frames anywhere else means one gets ignored.

By Flixly TeamApril 25, 2026
First to Last Frame AI for Smooth Video

TL;DR

Only Seedance 2.0 and Seedance 2.0 Fast support first-to-last frame, out of 37 video generation models. Veo 3.1, Kling 3.0 and Wan 2.7 do not, and Nano Banana Pro is an image model with no video output. It is the most deterministic control available because you decide where the shot ends. Pair frames that share a subject, place and lighting, and supply the same image twice for a perfect loop.

Two models support first-to-last frame. Not eight, not "most frontier models" — two.

Seedance 2.0 and Seedance 2.0 Fast. That is the complete list out of 37 video generation models.

Veo 3.1 does not. Kling 3.0 does not. Wan 2.7 does not. Nano Banana Pro is an image model and has no video output at all. If a guide tells you to pick one of those for a start-and-end-frame job, it is describing something that cannot happen.

Knowing this saves the most common wasted hour on this topic: uploading two frames to a model that will quietly ignore one of them.

What it actually does

You supply an opening image and a closing image. The model generates the movement between them.

That makes it the most deterministic control available here. Every other approach asks the model to invent where a shot ends. This one tells it, so the result lands exactly where you decided rather than wherever the generation drifted to.

Use first-to-last frame. Alongside it, both Seedance models also accept reference images, so you can pin identity and endpoints in the same job.

Where it earns its place

Joining two shots seamlessly. Take the final frame of shot one as your start frame. The next clip begins exactly where the last ended, so the cut disappears. This is the cleanest join available without chaining.

Perfect loops. Supply the same image as both start and end. The movement returns home by construction rather than by luck. For anything looping — a motion poster, a background plate, a social clip — this is the reliable method.

A specific transformation. Before and after, closed and open, empty and full. When both states matter and the path between them is the deliverable.

Controlled camera moves. Frame the start and end positions as stills, and the model interpolates a move that hits both marks.

Making both frames work together

The pairing matters more than either image alone.

Same subject, same place. Two unrelated images produce a morph, not a move. The model bridges what it is given, and if the gap is too wide the bridge is a mess.

Consistent lighting. A lighting change between frames gets interpreted as something happening in the scene, and the model will animate that change.

Plausible movement between them. A person standing then sitting works. A person standing then on a different continent does not.

Matching aspect and framing. Mismatched crops make the model resolve two compositions at once.

The most reliable way to get a good pair is to generate the end frame from the start frame with an image editing model, so the two share everything except the thing you want to change.

Duration is the tension

Longer clips give the movement room to look natural. Shorter clips hold coherence better.

If your two frames are close together, a short duration is fine and cleaner. If they are far apart, you need enough time for the transition to be believable, but the further you stretch it the more chance of drift in the middle.

When a long transition looks wrong, the usual fix is to make an intermediate frame and run two shorter jobs instead of one long one.

What is not real

No motion strength. Across all 37 video generation models the parameters are prompt, duration, aspect ratio and resolution, plus references, a negative prompt or first and last frames on some. Values like 0.45 do not exist.

No fixed input resolution. You are not required to supply 1024x576 frames.

No frame rate control, so nothing runs "at 12 fps to save credits".

No latent space, attention layer or noise scheduling controls. Descriptions of flash attention or temporal noise scheduling as things you configure are describing internals you cannot reach, if they exist at all.

Wan 2.7 is the only video model exposing a seed, along with a negative prompt. It does not do first-to-last frame, so those two capabilities cannot be combined.

When something else fits better

A continuous sequence of several shots. The Long Video Generator chains segments, carrying each finished clip into the next. Better than manually pairing frames shot after shot.

Only the start matters. Image to video is simpler and works on 35 models.

A specific performance. Motion Control transfers movement from a driving video.

The short version

Two models, Seedance 2.0 and 2.0 Fast. Two frames that share a subject, a place and a light. Enough duration for the movement to be plausible, and no more.

For loops, give it the same image twice. That single trick is worth the whole feature.

The catalog is at Models, each generation is quoted before it runs, and pack prices are on the pricing page.

Frequently Asked Questions

Which models support first-to-last frame?

Seedance 2.0 and Seedance 2.0 Fast, and no others among the 37 video generation models. Veo 3.1, Kling 3.0 and Wan 2.7 do not support it, and Nano Banana Pro is an image model with no video output at all. Uploading two frames to any of those means one is ignored.

Why use first-to-last frame instead of text-to-video?

Because it is deterministic about the ending. Every other approach lets the model invent where a shot finishes, while this one tells it, so the clip lands exactly where you decided rather than wherever the generation drifted.

How do I make a perfect loop?

Supply the same image as both the start and end frame. The movement returns home by construction rather than by luck, which makes it the reliable method for motion posters, background plates and looping social clips.

What makes a good pair of frames?

Same subject and same place, consistent lighting, a plausible movement between the two states, and matching aspect and framing. Two unrelated images produce a morph rather than a move. The most reliable approach is to generate the end frame from the start frame with an image editing model, so they share everything except what changes.

How long should the clip be?

Long enough for the movement to look plausible and no longer. Close-together frames suit a short duration and stay cleaner; far-apart frames need more time but risk drift in the middle. If a long transition looks wrong, make an intermediate frame and run two shorter jobs instead.

Can I set motion strength or frame rate?

No. Across all 37 video generation models the parameters are prompt, duration, aspect ratio and resolution, plus references, a negative prompt or first and last frames on some. There is no motion strength value, no frame rate control and no fixed input resolution requirement.

Can I combine first-to-last frame with a seed?

No. Wan 2.7 is the only video model exposing a seed, and it does not support first-to-last frame, so the two capabilities cannot be used together.

Tools mentioned in this post

guidesfirst-to-last-frameseedancevideo

Ready to create with guides?

Jump straight into Flixly's AI studio and try guides with 50+ models — free to start.

First to Last Frame: Only Two Models Do It | Flixly