Avatar X on Flixly: Talking-Avatar Video That Keeps Your Identity
Avatar X is Mirage's most advanced avatar model — type a script for a lifelike stock presenter, or drive any face from a photo or clip with your own voice track, up to 3 minutes in one take. It's live on Flixly.
TL;DR
Avatar X is Mirage's newest avatar model, built for identity preservation and expressive delivery. In Text to Video you write a script (50-1,500 characters) and a stock presenter speaks it. In Reference to Video you upload a voice track of 3 seconds to 3 minutes and drive any identity — a photo, a short clip, or a stock avatar. Both modes are live on Flixly today, billed by the second of generated video.
Avatar X is Mirage's most advanced avatar model, and it is live on Flixly in both of its modes: script-driven text-to-video and audio-driven reference-to-video.
What makes Avatar X different
- Identity preservation. The face you provide — a photo, a clip, or a stock presenter — stays consistent frame after frame. No drifting features, no uncanny morphing halfway through a take.
- Expressivity. Delivery follows the words: emphasis, pauses, gesture, and facial emotion track the script or the audio rather than looping a canned animation.
- Long takes. A single generation runs up to 3 minutes of speech — enough for a full product pitch, an explainer, or a UGC-style testimonial without stitching clips.
- Two ways in. Type a script and let a stock presenter deliver it, or upload your own voice track and drive any identity with it.
The two modes
Text to Video — write a script. Pick one of 24 stock presenters (portrait for vertical platforms, "(16:9)" variants for landscape, plus four news-desk anchors) and write 50 to 1,500 characters. The presenter speaks it in a matching voice. Every 50 characters is roughly four seconds of speech.
Reference to Video — bring a voice. Upload the audio your avatar should speak — 3 seconds to 3 minutes — and give it an identity: a single photo (9:16 or 16:9, at least 512px on the short edge) or a short clip (1-60 seconds). Or attach audio alone and pick a stock presenter to deliver it. Lips, timing, and expression follow your track.
Specs
| Property | Value |
|---|---|
| Modes on Flixly | Text-to-Video, Reference-to-Video |
| Script length | 50-1,500 characters (~4s per 50 chars) |
| Driving audio | 3 seconds - 3 minutes (mp3, wav, m4a, aac, ogg) |
| Identity reference | one photo OR one clip (1-60s), 9:16 or 16:9 |
| Stock presenters | 24 — portrait, landscape, and news-desk variants |
| Aspect ratios | 9:16 (default), 16:9 |
| Output | MP4 video with voice |
How to use it
- Open the Video Generator and pick the tab you need — Text or Reference.
- Pick Avatar X from the model selector.
- In Text mode: choose a presenter and write your script the way it should be spoken — punctuation shapes the pacing, short sentences land harder on camera.
- In Reference mode: attach your voice track, then add a photo or clip for the identity (or leave it to a stock presenter). Clean, dry audio without background music gives the best lip-sync.
- Generate. The clip length follows your script or audio automatically.
Pricing notes
Avatar X is billed by the second of generated video, and the length follows your input — there is no duration picker. A 30-second voice track costs half of a one-minute one; a 500-character script costs roughly a third of a 1,500-character one. The estimate on the page always shows the exact credit cost for your current script or audio before you generate.
Try it
Avatar X is live in Text to Video and Reference to Video.
Frequently Asked Questions
What is Avatar X?▾
Avatar X is Mirage's most advanced generation model, built for talking-avatar video. Its headline strengths are identity preservation — the face you provide stays that face, frame after frame — and expressivity: delivery, gesture, and emotion that follow the words rather than a loop. It runs in two modes on Flixly: script-driven (Text to Video) and audio-driven (Reference to Video).
How does the script mode work?▾
Write a script of 50 to 1,500 characters and pick one of 24 stock presenters — portrait or landscape variants, plus news-desk anchors. The avatar speaks your script with a matching voice. As a rule of thumb, every 50 characters is about four seconds of speech, so a full 1,500-character script runs a bit over two minutes.
Can I use my own face and voice?▾
Yes — that's Reference to Video. Upload the audio you want spoken (3 seconds to 3 minutes) plus an identity reference: a single photo or a short clip of up to 60 seconds. The avatar keeps that identity and speaks your audio with synced lips and natural movement. You can also drive a stock presenter with your own voice by attaching audio alone.
What are the input requirements?▾
Audio: mp3, wav, m4a, aac or ogg, between 3 seconds and 3 minutes. Identity photo: 9:16 or 16:9, at least 512 pixels on the short edge. Identity clip: 1 to 60 seconds, 9:16 or 16:9. Use a photo or a clip, not both — Flixly picks the clip if you attach two. Scripts run 50 to 1,500 characters.
How is Avatar X priced?▾
By the second of generated video. There's no duration picker — the length follows your script or your audio, and the exact credit cost is shown on the page before you generate: longer scripts and longer voice tracks cost proportionally more.
Where can I try it on Flixly?▾
Avatar X is live in Text to Video and Reference to Video on the Flixly dashboard. Pick it from the model selector, choose a presenter or upload your references, and generate.
