All posts
tutorials

Create Video from Pictures Without Keyframes

You can't upload ten photos and get one video. Image-to-video takes one image per generation, which changes how the whole job is organised.

By Flixly TeamJune 1, 2026
Create Video from Pictures Without Keyframes

TL;DR

Image-to-video takes one image per generation, so there is no multi-photo upload. A slideshow needs no AI at all and is better done in an editor or FFmpeg. Animating stills is one generation per photo: find a prompt on your most typical image, then repeat it unchanged across the set, keeping motion, duration and aspect ratio identical so the clips cut together as a sequence rather than ten unrelated shots.

You cannot upload ten photos and get one video. Image-to-video takes one image per generation.

That single fact reorganises the whole task, because "make a video from my pictures" turns out to be two different jobs and the wrong assumption sends you looking for a batch uploader that does not exist.

A slideshow shows your photos in sequence. The photos stay photos; the video is the container. This needs no AI at all.

Animated stills means each photo becomes a moving clip. That is generation, one image at a time, and then an edit.

If you want a slideshow

Do not use a generation tool. Any editor assembles a sequence of stills with transitions, and FFmpeg does it from a single command at no cost.

Paying per generation to produce something a free tool does better is the most common mistake on this topic.

If you want the photos to move

Now it is one generation per photo, and the work splits into finding a prompt and then repeating it.

Ten photos is ten submissions through image to video. There is no batch upload and no queue manager. What makes it fast is that the prompt is reusable.

Find the prompt on one photo first. Use your most typical image, not your best one. A prompt tuned to the exceptional shot fails on the ordinary ones, and most of a set is ordinary.

Then repeat it unchanged, varying only the image. You have stopped deciding, which is where the speed comes from.

Writing for a set

Consistency across the set matters more than any single clip being perfect, because a viewer sees the sequence.

Same motion, same duration, same aspect ratio. Ten clips that each move differently read as ten unrelated clips.

Say what the camera does. Unspecified cameras drift, and ten independent drifts look like a mistake rather than a style.

Keep the movement small. A slow push or a gentle drift cuts together cleanly. Dramatic moves fight each other at the joins.

"Slow push in. Camera moves, subject stays still. Soft natural light."

Which photos will animate

Some stills fight this and it is worth knowing before you spend on them.

Room in the direction of movement. A tightly cropped subject has nowhere to go.

Clear separation from the background. Subjects that blend get blended further once things move.

Sharp. Blur becomes smearing, and no prompt recovers it.

Frontal or three-quarter faces. Profiles have half a face to work from.

Cutting them together

Generation gives you clips. Assembly is still an edit.

Order them for rhythm rather than chronology, keep each clip short, and let the motion direction alternate so the sequence has some shape. Two clips pushing in back to back feel repetitive; a push followed by a hold does not.

If the whole thing needs to feel continuous rather than like a sequence of separate shots, that is a different tool: the Long Video Generator chains segments so each continues from the previous one's final frame.

Finishing

Order matters, because reversing it means redoing work.

Voice from Text to Speech if there is narration, then Lip Sync only if a person is on screen and only after the audio is final. Music from Music Generation. Captions last through Auto Captions, since most of this is watched muted.

Everything you generate stays in History, which matters when you are working through a set.

What is not real

No multi-image upload for a single generation. One image in, one clip out.

No motion strength. Across all 37 video generation models the parameters are prompt, duration, aspect ratio and resolution, plus references, a negative prompt or first and last frames on some.

No keyframes, easing curves or motion paths. This is not a timeline tool.

No published generation times. Speed varies with load.

If you need a specific subject to stay recognisably itself across every clip, thirteen models accept a reference through reference to video.

The short version

Decide whether you want a slideshow or animated stills. If it is a slideshow, use an editor and pay nothing.

If it is animated stills, find one prompt on one photo, apply it to the rest unchanged, keep the motion small and consistent, then cut them together.

The catalog is at Models, and each generation is quoted before it runs.

Frequently Asked Questions

Can I upload several photos and get one video?

No. Image-to-video takes one image per generation, and there is no batch upload or queue manager. Ten photos means ten submissions, then an edit to assemble them.

Do I need AI to make a slideshow?

No, and you should not pay for it. A slideshow shows your photos in sequence with the video as a container, which any editor assembles and FFmpeg does from a single command at no cost. Paying per generation for something a free tool does better is the common mistake here.

How do I make a set of clips look consistent?

Find the prompt on your most typical photo rather than your best one, then repeat it unchanged across the set, keeping motion, duration and aspect ratio identical. Ten clips that each move differently read as ten unrelated clips, and a viewer sees the sequence before any single clip.

What motion works best when cutting clips together?

Small movements. A slow push or gentle drift cuts cleanly, while dramatic moves fight each other at the joins. Always specify what the camera does, since unspecified cameras drift and ten independent drifts read as a mistake rather than a style.

Which photos animate well?

Ones with room in the direction of intended movement, clear separation between subject and background, a sharp source without blur, and frontal or three-quarter faces rather than profiles. Blur in the input becomes smearing in the output.

Are there keyframes or motion paths?

No. There are no keyframes, easing curves or motion paths, and no motion strength setting. Across all 37 video generation models the parameters are prompt, duration, aspect ratio and resolution, plus references, a negative prompt or first and last frames on some.

How do I make the result feel continuous rather than a sequence of shots?

That is a different tool. The Long Video Generator chains segments so each continues from the previous one's final frame, which holds a world together. Assembling separate image-to-video clips will always read as a sequence of shots, which is fine when that is what you want.

Tools mentioned in this post

tutorialsimage-to-videoworkflowphotos

Ready to create with tutorials?

Jump straight into Flixly's AI studio and try tutorials with 50+ models — free to start.