Now you generate videos with Flixly
Forty photos and an hour. Generating first is how the hour disappears. Decide the shape, find one prompt, then run the set.

TL;DR
Decide aspect ratio and duration before generating anything, since neither can change afterwards without regenerating. Find one motion prompt on a typical photo using a cheap Fast variant, then run the whole set with it unchanged. Review the set as a group rather than clip by clip, because inconsistency reads worse than any single imperfect clip. Finish in order: voice, lip sync, music, upscale, captions.
Forty product photos and an hour. The instinct is to start generating immediately, and that is how the hour disappears.
The order that actually finishes on time is: decide the shape, find one prompt, then run the set. Generating first means discovering on photo thirty that the aspect ratio is wrong.
Decide the shape once
Two things you cannot change afterwards without regenerating everything:
Aspect ratio. Vertical for Reels, Shorts, TikTok. Horizontal for a site hero. Square for a feed grid. Crop later and you throw away most of the frame along with the composition the model chose.
Duration. Four to five seconds is where motion quality peaks. Longer clips drift, and drift on a product shot means a warped label.
Both are set at generation. Decide them before photo one.
Find the prompt on one photo
Not your best photo, your most typical one. A prompt tuned to the exceptional shot fails on the ordinary ones, and most of your forty are ordinary.
What works for products is small and specific:
"Slow push in toward the bottle. Camera moves, product stays still. Reflections shift on the glass."
Decide whether the camera or the product moves. Not both. Two motions resolved at once against a flat source is where products warp.
Say what stays still. As useful as saying what moves.
Protect the label. Text on packaging smears first. Keep movement slow and small.
Use a cheap model while you search for the wording — Seedance 2.0 Fast at STANDARD, or another Fast variant. Once the sentence works, switch to the model you actually want and run it there.
Run the set
Same prompt, same duration, same aspect ratio, changing only the image. Image to video is the tool; thirty-five video models accept an image input.
There is no batch upload and no queue manager, so this is genuinely forty submissions. What makes it fast is that you are not deciding anything any more.
Everything lands in History, which matters at this volume — it is where you go when a client asks for the third version of the blue one.
Review as a set, not one by one
The step people skip, and the one that saves the deliverable.
Forty clips watched individually all seem fine. Watched together, the four where the camera drifted the other way stand out immediately, and inconsistency across a set reads worse to a viewer than any single imperfect clip.
Regenerate the outliers with a slightly firmer camera instruction rather than reshooting the whole run.
Keeping the product identical across clips
If the same product must be recognisably itself in every shot, thirteen models accept a reference through reference to video.
For repeatability specifically, Wan 2.7 is the only video model with a seed, along with a negative prompt. If you want the same camera behaviour across the set rather than forty independent interpretations, that is the single entry offering it.
Finishing
In this order, because reversing it means redoing work:
Voice from Text to Speech, if the piece needs narration.
Lip sync through Lip Sync only if a person is on screen, and only after the audio is final.
Music from Music Generation for a bed that will not get the upload muted for rights.
Upscale with the Video Upscaler after the edit, not before — upscaling first pays to sharpen frames you then cut.
Captions last, through Auto Captions, once picture and audio are both settled.
What is not there
No motion strength. Across all 37 video generation models the parameters are prompt, duration, aspect ratio and resolution, plus references, a negative prompt or first and last frames on some.
No frame rate control. You do not set 24 or 30 fps.
No published generation times. Speed varies with load, so any article promising "42 seconds per job" measured it once.
No reference masks and no per-region controls anywhere.
The hour, realistically
Ten minutes finding the prompt. Thirty running the set. Ten reviewing as a group and regenerating outliers. Ten on captions and delivery.
Most of the risk sits in the first ten minutes, which is why it is worth spending them deliberately rather than starting the run and hoping.
The catalog is at Models, each generation is quoted before it runs, and pack prices are on the pricing page.
Frequently Asked Questions
What should I decide before generating a batch?▾
Aspect ratio and duration. Both are set at generation time and cannot be changed afterwards without regenerating, and cropping later throws away most of the frame along with the composition the model chose. Four to five seconds is where motion quality peaks for product work.
How do I get consistent results across many clips?▾
Find one motion prompt on your most typical photo, not your best one, then apply it unchanged across the set. Use a cheap Fast variant while searching for the wording and switch to the model you want once it works. Wan 2.7 is the only video model with a seed if you need the same camera behaviour rather than independent interpretations.
Why does my product warp when it moves?▾
Usually because both the camera and the product are moving. Decide which one moves and explicitly say the other stays still. Keeping the movement small also protects label text, which is the first thing to smear on packaging.
Is there a batch upload or queue manager?▾
No. Forty images means forty submissions. What makes it fast is that once the prompt is settled you are not making decisions any more, only repeating. Everything generated lands in History, which matters at that volume.
How should I review a large batch?▾
As a group. Clips watched one at a time all seem acceptable, but watched together the few where the camera drifted the other way stand out immediately. Inconsistency across a set reads worse to a viewer than any single imperfect clip.
In what order should I add audio and captions?▾
Voice, then lip sync if a person is on screen, then music, then upscale, then captions. Lip sync matches a mouth to existing audio so it must follow the voice, and upscaling before the edit pays to sharpen frames you then cut.
What settings control the amount of motion?▾
None. Across all 37 video generation models the parameters are prompt, duration, aspect ratio and resolution, plus references, a negative prompt or first and last frames on some. There is also no frame rate control, so you do not set 24 or 30 fps.


