H3 Max Reference-to-Video on Flixly: Consistent Characters from the #1-Ranked Video Model
MiniMax H3 Max now takes references: up to 9 images to lock your subjects and style, plus up to 3 short clips to steer the motion. The #1-ranked model for quality and prompt understanding, now with consistency — live on Flixly.
TL;DR
H3 Max — fal's post-trained MiniMax H3, ranked #1 for overall quality, prompt understanding, and aesthetics — now supports reference-to-video. Give it up to 9 reference images to keep subjects and style consistent, and up to 3 short video clips (15 seconds combined) to steer the motion, then address them in the prompt by order: "Image 1 is the hero, Video 1 is how she runs." Output is 5-15 seconds at 480p or 768p. It's an early fal preview running near real-time, and it's live in the Reference tab of Flixly's Video Generator today.
MiniMax H3 Max — the model ranked #1 for overall quality, prompt understanding, and aesthetics — now takes references. Reference-to-video is live on Flixly in the Video Generator's Reference tab.
What references buy you
The hardest problem in AI video isn't a pretty shot — it's the SAME character in the next shot. Reference-to-video fixes that at the input:
- Cast with images. Attach up to 9 reference images and the model keeps those subjects — a person, a mascot, a product, a style plate — consistent in the new shot.
- Direct with clips. Attach up to 3 short video clips (15 seconds combined) and the model follows their motion: a gait, a camera move, a dance.
- Address them in the prompt. References are cited by order: "Image 1 is the detective. Image 2 is her car. Video 1 is the chase camera move." H3 Max's #1-ranked prompt understanding is exactly what makes this feel like directing.
Specs
| Property | Value |
|---|---|
| Reference images | up to 9 |
| Reference clips | up to 3, ≤15s combined |
| Duration | 5-15 seconds, one-second steps |
| Resolution | 480p or 768p |
| Aspect ratio | six ratios, or adaptive (follows the references) |
| Prompt expansion | disabled / balanced / quality |
| Speed | early fal preview, near real-time at 768p |
How to use it
- Open the Video Generator, switch to the Reference tab, and pick MiniMax H3 Max.
- Add one to three clean images of your subject — front-lit, uncluttered backgrounds work best. Add a motion clip only if the movement matters.
- Write the prompt as direction and cite the references by order: subject, action, camera.
- Draft at 5 seconds, then re-run the take you like at full length.
What it costs
Reference jobs bill the output length plus a small metered amount for the references you attach — the credit estimate on the page reflects the exact settings and attachments before you generate. Images are cheap (your first effectively rides free); reference video is the dear part, so attach a clip when the motion matters and skip it when an image cast is enough.
Try it
H3 Max Reference-to-Video is live in Reference to Video — and the model's Text to Video and Image to Video modes are right beside it.
Frequently Asked Questions
What is reference-to-video on H3 Max?▾
Instead of describing your subject from scratch, you attach references: images that define who or what appears (and the look), and optionally short video clips that define how things move. The prompt then cites them by order — Image 1, Image 2, Video 1 — and the model composes a new shot that keeps those subjects consistent while following your direction.
How many references can I attach?▾
Up to 9 reference images and up to 3 reference clips, with at most 15 seconds of reference video combined. In practice one to three clean images of a subject go a long way — add more only when you need several distinct subjects or a specific style plate.
Do references change the price?▾
Yes, slightly. fal meters reference inputs on top of the output seconds, and the estimate on the page reflects everything you've attached before you generate — images are cheap, reference video is the dear part. Your first reference image effectively rides free.
How is this different from image-to-video?▾
Image-to-video animates your exact frame — the output starts from that pixel-for-pixel composition. Reference-to-video treats your images as casting, not as the first frame: the model stages a new shot with those subjects in it, which is what you want for consistent characters across many clips.
What settings does it support?▾
The same H3 Max controls: 5-15 seconds in one-second steps, 480p or 768p, prompt expansion, and aspect ratio — including an adaptive option where the output follows the references.
Where can I try it?▾
Open the Video Generator on the Flixly dashboard, switch to the Reference tab, and pick MiniMax H3 Max. Add your images and clips, cite them in the prompt, and generate.