All posts
Model Launches

H3 Max Reference-to-Video on Flixly: Consistent Characters from the #1-Ranked Video Model

MiniMax H3 Max now takes references: up to 9 images to lock your subjects and style, plus up to 3 short clips to steer the motion. The #1-ranked model for quality and prompt understanding, now with consistency — live on Flixly.

August 31, 2026

TL;DR

H3 Max — fal's post-trained MiniMax H3, ranked #1 for overall quality, prompt understanding, and aesthetics — now supports reference-to-video. Give it up to 9 reference images to keep subjects and style consistent, and up to 3 short video clips (15 seconds combined) to steer the motion, then address them in the prompt by order: "Image 1 is the hero, Video 1 is how she runs." Output is 5-15 seconds at 480p or 768p. It's an early fal preview running near real-time, and it's live in the Reference tab of Flixly's Video Generator today.

MiniMax H3 Max — the model ranked #1 for overall quality, prompt understanding, and aesthetics — now takes references. Reference-to-video is live on Flixly in the Video Generator's Reference tab.

What references buy you

The hardest problem in AI video isn't a pretty shot — it's the SAME character in the next shot. Reference-to-video fixes that at the input:

  • Cast with images. Attach up to 9 reference images and the model keeps those subjects — a person, a mascot, a product, a style plate — consistent in the new shot.
  • Direct with clips. Attach up to 3 short video clips (15 seconds combined) and the model follows their motion: a gait, a camera move, a dance.
  • Address them in the prompt. References are cited by order: "Image 1 is the detective. Image 2 is her car. Video 1 is the chase camera move." H3 Max's #1-ranked prompt understanding is exactly what makes this feel like directing.

Specs

Property Value
Reference images up to 9
Reference clips up to 3, ≤15s combined
Duration 5-15 seconds, one-second steps
Resolution 480p or 768p
Aspect ratio six ratios, or adaptive (follows the references)
Prompt expansion disabled / balanced / quality
Speed early fal preview, near real-time at 768p

How to use it

  1. Open the Video Generator, switch to the Reference tab, and pick MiniMax H3 Max.
  2. Add one to three clean images of your subject — front-lit, uncluttered backgrounds work best. Add a motion clip only if the movement matters.
  3. Write the prompt as direction and cite the references by order: subject, action, camera.
  4. Draft at 5 seconds, then re-run the take you like at full length.

What it costs

Reference jobs bill the output length plus a small metered amount for the references you attach — the credit estimate on the page reflects the exact settings and attachments before you generate. Images are cheap (your first effectively rides free); reference video is the dear part, so attach a clip when the motion matters and skip it when an image cast is enough.

Try it

H3 Max Reference-to-Video is live in Reference to Video — and the model's Text to Video and Image to Video modes are right beside it.

Frequently Asked Questions

What is reference-to-video on H3 Max?

Instead of describing your subject from scratch, you attach references: images that define who or what appears (and the look), and optionally short video clips that define how things move. The prompt then cites them by order — Image 1, Image 2, Video 1 — and the model composes a new shot that keeps those subjects consistent while following your direction.

How many references can I attach?

Up to 9 reference images and up to 3 reference clips, with at most 15 seconds of reference video combined. In practice one to three clean images of a subject go a long way — add more only when you need several distinct subjects or a specific style plate.

Do references change the price?

Yes, slightly. fal meters reference inputs on top of the output seconds, and the estimate on the page reflects everything you've attached before you generate — images are cheap, reference video is the dear part. Your first reference image effectively rides free.

How is this different from image-to-video?

Image-to-video animates your exact frame — the output starts from that pixel-for-pixel composition. Reference-to-video treats your images as casting, not as the first frame: the model stages a new shot with those subjects in it, which is what you want for consistent characters across many clips.

What settings does it support?

The same H3 Max controls: 5-15 seconds in one-second steps, 480p or 768p, prompt expansion, and aspect ratio — including an adaptive option where the output follows the references.

Where can I try it?

Open the Video Generator on the Flixly dashboard, switch to the Reference tab, and pick MiniMax H3 Max. Add your images and clips, cite them in the prompt, and generate.

Tools mentioned in this post

MiniMaxH3 MaxHailuofalReference to VideoConsistent CharactersAI VideoFlixly

Ready to create with Model Launches?

Jump straight into Flixly's AI studio and try model launches with 50+ models — free to start.