MiniMax H3 Max on Flixly: The #1-Ranked Video Model for Quality and Prompt Understanding
H3 Max is fal's post-trained variant of MiniMax H3, and it ranks #1 for overall quality, prompt understanding, and aesthetics against leading video models. Text-to-video and image-to-video, 5-15 seconds per clip — it's live on Flixly.
TL;DR
MiniMax H3 Max is fal's post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics — and it ranks #1 for overall quality, prompt understanding, and aesthetics in head-to-head comparisons against leading video models. It generates 5-15 second clips at 480p or 768p from a text prompt (six aspect ratios) or a start frame (with an optional end frame), and it runs on a throughput-optimized inference stack, so renders come back fast. Both modes are live on Flixly today.
MiniMax H3 Max is fal's post-trained variant of MiniMax H3, and it is live on Flixly in text-to-video and image-to-video.
Why H3 Max
- Ranked #1 where it counts. In head-to-head comparisons against leading video models, H3 Max ranks first for overall quality, prompt understanding, and aesthetics. That middle one is the practical headline: the model does what the prompt says — subjects, actions, and camera direction land the way you wrote them.
- Tuned for the picture. The post-training targets aesthetics specifically — composition, light, and color grade that look art-directed rather than defaulted.
- Fast renders. The model is co-optimized with a custom inference stack for higher throughput, so iterations come back quickly instead of queueing.
- Keyframe control in image-to-video. Give it a start frame, and optionally an end frame, and the clip animates from one to the other.
Specs
| Property | Value |
|---|---|
| Modes on Flixly | Text-to-Video, Image-to-Video |
| Duration | 5-15 seconds, one-second steps |
| Resolution | 480p or 768p |
| Aspect ratios (t2v) | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Aspect ratio (i2v) | follows the source image |
| End frame | optional in image-to-video |
| Prompt expansion | disabled / balanced / quality |
| Prompt length | up to 5,000 characters |
How to use it
- Open the Video Generator in the dashboard and pick the tab you need — Text or Image.
- Pick MiniMax H3 Max from the model selector.
- Set resolution and duration. Draft at 480p and 5 seconds — it costs ~40% less than 768p — then re-run the prompt you like at 768p and full length.
- Write the prompt as direction, not description: the subject, the action, the camera move. This is the model that rewards being specific — prompt understanding is its #1-ranked axis.
- For image-to-video, upload a start frame; add an end frame if you want the shot to land on a specific composition.
What it costs
H3 Max is billed by output length and resolution, and by nothing else: a 15-second clip costs three times a 5-second one, 768p costs a bit more than half again over 480p, and prompt expansion and end frames add nothing. The credit estimate on the page always reflects the exact settings you've chosen before you generate.
Try it
MiniMax H3 Max is live in Text to Video and Image to Video.
Frequently Asked Questions
What is MiniMax H3 Max?▾
H3 Max is a post-trained variant of MiniMax's H3 video model, tuned by fal for stronger prompt adherence and better aesthetics and co-optimized with a custom inference stack for higher throughput. In head-to-head evaluations against leading video models it ranks #1 for overall quality, prompt understanding, and aesthetics.
How is H3 Max different from MiniMax H3?▾
Same underlying model family, different tuning and different trade-offs. H3 Max is post-trained specifically for following your prompt more faithfully and composing more aesthetic shots, and it renders on a faster stack. The original MiniMax H3 on Flixly still has its own strengths — 2K output and native stereo audio — so pick H3 for maximum resolution with sound, and H3 Max for the best-looking, most prompt-faithful picture.
How long can an H3 Max video be?▾
Anywhere from 5 to 15 seconds, set in one-second steps. Longer clips cost proportionally more because the model is billed by output length, so start short while you iterate on the prompt and extend once the shot is right.
What resolutions and aspect ratios does it support?▾
480p or 768p. In text-to-video you choose from six aspect ratios — 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16. In image-to-video the output simply follows your source image's aspect, which is usually what you want.
What does prompt expansion do?▾
Before rendering, the model can rewrite your prompt into the richer form it composes best from. Balanced is the default; quality spends more effort on the rewrite, and disabled sends your prompt through untouched — useful when you've already engineered it precisely. It doesn't change the cost either way.
Where can I try it on Flixly?▾
H3 Max is live in Text to Video and Image to Video on the Flixly dashboard. Pick it from the model selector, set your resolution and duration, and generate.