Black Forest Labs' FLUX 3 — text-to-video with synchronized audio generated by the model itself. Up to 20 seconds at 1080p, no reference image required.
~2-4 min
curl https://www.flixly.ai/api/v1/generate \
-H "Authorization: Bearer flx_live_..." \
-H "Content-Type: application/json" \
-d '{
"model": "flux-3",
"prompt": "A beautiful sunset over the ocean",
"type": "TEXT_TO_VIDEO"
}'Get your API key from the Developer Portal.
Premium video generation model with native audio
ByteDance's newest video model. Generates up to 30 seconds in a single shot with synchronized audio — dialogue, music, and effects in the same pass — from a prompt, a start frame, or up to 50 image, video, and audio references.
The premium tier of Alibaba's Wan 3.0 — faster turnaround, cinematic motion, and stronger visual continuity. Text, image, and reference modes, native 30-second clips in a single pass, with audio.
Post-trained MiniMax H3, ranked #1 for overall quality, prompt understanding, and aesthetics against leading video models. 5-15 second clips in native 480p, 768p or 1080p from a prompt, a start frame, or up to 9 reference images, 3 motion clips and 3 audio clips. Reference-to-video is generally available: up to 2x faster, with much stronger subject preservation.