Soundify Guide Using Flixly Tools
The instinct is to make the video first and add sound afterwards. For a lot of work that's the wrong order, and fixing it saves the most time.

TL;DR
Ten video models generate synchronized audio as part of the generation, so for diegetic sound the fastest route is describing it in the video prompt rather than adding it afterwards. Go to a separate pass only for specific music, specific spoken words, or a clip that already exists. Music Generation runs Suno Music with a real editing suite behind it in Music Tools, Text to Speech has six models with cloning on two of them, and lip sync runs last against finished audio. Mixing belongs in an editor, since there is no mixer, EQ or timeline here.
The instinct is to make the video first and add sound afterwards.
For a lot of work here, that is the wrong order, and it is the single change that saves the most time.
Ten models generate the audio with the clip
This is the part people miss. Audio is not always a second step, because a number of video models produce synchronized sound as part of the generation: Veo 3.1 and Veo 3.1 Fast, Kling 3.0, Wan 2.6 Flash, Seedance 2.0, 2.0 Fast and 2.5, FLUX 3, and Wan 3.0 and 3.0 Prime.
The sound arrives already matched to what is on screen, because the model made both at once. No alignment pass, no drift, no scrubbing to check whether the footstep lands on the footfall.
Describe the sound in the video prompt. Not in a separate audio prompt afterwards — in the same sentence as the action:
"Boots on wet tile, echoing corridor, a door clicking shut at the end, low ventilation hum throughout."
Audio generation is a toggle, so check it is on. On Seedance 2.0 it defaults to on; it is worth confirming rather than assuming.
If the sound you need is diegetic — footsteps, doors, impacts, ambience, anything caused by something visible — try this before reaching for anything else. It is the only approach where sync is free rather than earned.
When you do need a separate pass
Native audio is not the answer for everything. Go separate when:
You need a specific piece of music. A model inventing ambience is not going to produce the track you have in mind.
You need specific words spoken. Native audio does not take a script.
The clip already exists and was generated without audio, or came from somewhere else.
Music
Music Generation runs Suno Music, which writes a track from a description.
What is less well known is that there is a real editing suite behind it at Music Tools, and it is more capable than most people expect:
Vocal Separation splits a track into vocal and instrumental stems. Add Vocals and Add Instrumental go the other way. Extend Music lengthens a track that ended too early. Replace Section swaps out a passage without regenerating the whole thing. Boost Music Style pushes a track further toward a genre. Music Cover reinterprets an existing piece. Generate MIDI gives you notes rather than audio. Convert to WAV handles the format.
That set covers the practical problems — a track that is fifteen seconds short, a vocal you want gone, a bridge that does not work — without starting over.
Voice
Text to Speech has six models: OpenAI TTS and OpenAI TTS HD, ElevenLabs Multilingual V2 and ElevenLabs Turbo V2.5, F5-TTS, and Gemini 3.1 Flash TTS.
Turbo and the standard OpenAI model are the quick ones. HD and Multilingual V2 are where you go when the delivery matters.
For a specific voice rather than a stock one, Voice Cloning is supported by ElevenLabs Multilingual V2 and F5-TTS — those two, not the whole list.
Getting mouths to match
Lip Sync runs three models: LatentSync, Sync LipSync and Volcengine Lip Sync.
Order matters. Generate the silent clip first, then run lip sync against the finished audio line. Doing it the other way round means the mouth movement was matched to something you then replaced.
Mixing is an editor's job
There is no mixer here, and this is where the fabricated version of this guide goes furthest wrong.
No dB levels, no faders, no per-stem gain. You cannot set footsteps to -12 dB and hum to -18 dB.
No EQ, no notch filters, no frequency analysis. There is no place to notch 3.2 kHz.
No timeline, no keyframes, no automation. You cannot dip a stem on specific frames.
No sample rate or bit depth settings. You do not choose 48 kHz or 24-bit.
No batch audio queue, no project folders, no version branches, no rollback. None of that exists.
Generate the pieces here, then mix in an editor. Any free editor does levels and EQ properly, and no generative tool is going to beat a fader at being a fader.
A workflow that actually works
- Write the sound into the video prompt and generate on an audio-capable model. For a lot of clips this is the whole job.
- If you need music, generate it in Music Generation and fix its length or structure in Music Tools rather than regenerating.
- If you need dialogue, write it in Text to Speech, using a cloned voice if the speaker matters.
- If a character speaks on screen, run the silent clip through Lip Sync against that finished line.
- Mix in an editor. Balance, trim, and place.
Every generation is quoted before it runs, so the cost of each step is in front of you before you commit to it.
The short version
Ask for the sound in the video prompt first, because ten models will make it in sync for free. Fall back to separate passes only for specific music, specific words, or clips that already exist.
Music Tools is more capable than its name suggests. Lip sync goes last. Mixing goes in an editor.
Start in Text to Video, and the catalog is at Models.
Frequently Asked Questions
Do I have to add sound as a separate step?▾
Often not. Ten video models generate synchronized audio as part of the generation: Veo 3.1 and Veo 3.1 Fast, Kling 3.0, Wan 2.6 Flash, Seedance 2.0, 2.0 Fast and 2.5, FLUX 3, and Wan 3.0 and 3.0 Prime. The sound arrives already matched to what is on screen because the model made both at once, so there is no alignment pass and no drift.
How do I ask for specific sounds in a generation?▾
Describe them in the video prompt itself, in the same sentence as the action rather than in a separate audio prompt afterwards. Something like "boots on wet tile, echoing corridor, a door clicking shut at the end, low ventilation hum throughout". Audio generation is a toggle, so confirm it is on rather than assuming.
When should I generate audio separately instead?▾
Three cases. When you need a specific piece of music, since a model inventing ambience will not produce the track you had in mind. When you need specific words spoken, because native audio does not take a script. And when the clip already exists, whether generated without audio or brought in from elsewhere.
What can I do with music beyond generating a track?▾
More than most people expect. Behind Music Generation, which runs Suno Music, sits an editing suite: Vocal Separation splits vocal and instrumental stems, Add Vocals and Add Instrumental go the other way, Extend Music lengthens a track that ended early, Replace Section swaps a passage without regenerating everything, plus Boost Music Style, Music Cover, Generate MIDI and Convert to WAV.
Which models handle text to speech and voice cloning?▾
Text to Speech has six models: OpenAI TTS and OpenAI TTS HD, ElevenLabs Multilingual V2 and Turbo V2.5, F5-TTS, and Gemini 3.1 Flash TTS. Turbo and standard OpenAI are the quick ones while HD and Multilingual V2 suit delivery that matters. Voice cloning is supported by ElevenLabs Multilingual V2 and F5-TTS only, not the whole list.
What order should lip sync come in?▾
Last. Generate the silent clip first, then run lip sync against the finished audio line using one of the three models, LatentSync, Sync LipSync or Volcengine Lip Sync. Doing it the other way round means the mouth movement was matched to audio you then replaced.
Can I set mix levels, EQ or sample rates?▾
No. There is no mixer, no dB levels or faders, no EQ or notch filtering, no timeline or keyframe automation, and no sample rate or bit depth settings. There is also no batch audio queue, project folder, version branch or rollback. Generate the pieces here and mix in an editor, where a fader is a fader.



