Auto Captions for Silent Films
Auto-captions listens, it does not watch. A silent film has nothing to transcribe, so here is what actually works instead.
TL;DR
Auto-captions transcribes speech, so silent footage produces no captions at all rather than poor ones. No tool writes captions by watching the action. The working routes are to narrate the film with Text to Speech or Voice Cloning and caption that narration, to write intertitles as title cards, or to caption your own commentary track. Real caption options are size, five styles including Word pop, and ten languages, priced by clip length.
Auto-captions listens. It does not watch.
That one sentence resolves the whole question, because a silent film has nothing to listen to. There is no dialogue track to transcribe, so pointing captioning software at one produces nothing. Not bad captions: no captions.
No tool captions a silent film by watching the action. Anything claiming to detect events frame by frame and write intertitles for them is describing software that does not exist.
What you can do is give the film an audio track first, then caption that. Here is how.
What auto-captions actually does
It takes video that contains speech, transcribes it, and burns the text into the picture.
The real options are:
Size. Small, Medium or Large.
Style — Word pop, which reveals words one at a time and holds attention best on muted autoplay; karaoke; sentence; boxed; or minimal.
Language. English, Spanish, French, German, Italian, Portuguese, Dutch, Japanese, Chinese or Korean.
Cost scales with the length of the clip, and the app quotes it before the job runs.
There is no caption timeline editor, no per-word timing panel, and no accuracy percentage published anywhere. Guides quoting a figure like 98% invented it.
The actual workflow for silent footage
Three routes, depending on what you want the finished piece to be.
Narrate it, then caption the narration
The most common answer, and the one that produces something people will actually watch.
Write the narration yourself. Generate it with Text to Speech, or use Voice Cloning if the piece belongs to a series and the voice should stay consistent across episodes.
Lay that audio over the footage, then run Auto Captions on the result. Now there is speech to transcribe, and the captions match it exactly, because they were made from it.
Write intertitles instead
Silent films had a visual language for this, and it still works. Cards between shots, in period type, carrying dialogue or narration.
This is a design job rather than a transcription one. Generate title cards as images and cut them in. It suits archival restorations, where burned-in modern captions would look wrong against the grain and framing of the original.
Caption a commentary track
If the piece is a video essay over silent footage, your voice discussing what is on screen, that is just a normal captioning job. The audio is your commentary, and auto-captions handles it like any other speech.
If the footage has speech but you cannot hear it
Worth separating from genuinely silent film, because the fix is different.
Damaged or very quiet audio is a transcription problem, not a captioning one. Clean the audio first. Transcription quality tracks audio quality closely, and no captioning tool recovers words that are not audible in the source.
Footage where people are visibly speaking but the recording was never made has no solution beyond lip-reading, which is a human specialism and not something any current model does reliably.
Things people conflate with this
Translating an existing video. Video Translator handles a video that already has speech and needs another language.
Re-syncing a mouth to new audio. Lip Sync aligns mouth movement in an existing video to a target audio track, which is the tool for dubbing. Volcengine Lip Sync supports multilingual dubbing specifically.
Cutting clips with captions already applied. The Shorts Generator finds clips in a longer video and renders word-by-word captions, positioned top, centre or bottom.
None of these captions a silent film either. They all need audio.
Why captions are worth the step
Most social video is watched muted. An uncaptioned clip is a silent clip to most of the audience, which for archival footage is an odd irony. The film was silent, the caption is what gives it a voice.
Word pop is the style to pick for social. Revealing one word at a time is what holds a scrolling viewer, and it is the reason that style exists rather than being a stylistic preference.
For anything longer than a social cut, sentence or minimal reads better and gets in the way of the picture less.
The short version
Auto-captions needs speech. A silent film has none, so you add some.
Narrate and caption the narration for most purposes. Write intertitles when the period look matters. Caption your commentary if the piece is an essay.
What you cannot do is point captioning software at silent footage and get text describing what happens. That capability does not exist here or anywhere, whatever a comparison table tells you.
Costs depend on clip length and are quoted before each job. Pack prices are on the pricing page.
Frequently Asked Questions
Can AI caption a silent film by watching the action?▾
No. Captioning software transcribes speech from an audio track. Silent footage has nothing to transcribe, so it produces no captions rather than inaccurate ones. Any tool claiming to detect events frame by frame and write intertitles for them is describing a capability that does not exist.
So how do I caption silent footage?▾
Give it audio first. Write narration and generate it with Text to Speech, or use Voice Cloning if the piece belongs to a series, lay that over the footage, then run auto-captions on the result. The captions then match the narration exactly, because they were made from it.
What caption options can I actually set?▾
Size (Small, Medium or Large), style (Word pop, karaoke, sentence, boxed or minimal) and language, with English, Spanish, French, German, Italian, Portuguese, Dutch, Japanese, Chinese and Korean supported. There is no caption timeline editor or per-word timing panel.
Which caption style should I use for social video?▾
Word pop. Revealing one word at a time is what holds a scrolling viewer on muted autoplay, which is why the style exists rather than being a matter of taste. For longer pieces, sentence or minimal reads better and obscures less of the picture.
How accurate are the captions?▾
No accuracy figure is published, and articles quoting one to a decimal place invented it. Transcription quality tracks audio quality closely, so clean audio produces good captions and damaged or very quiet audio produces poor ones regardless of the tool.
What if the footage has speech I cannot hear clearly?▾
That is a transcription problem rather than a captioning one. Clean the audio first, since no captioning tool recovers words that are not audible in the source. Footage where people are visibly speaking but no recording was made has no solution beyond human lip-reading.
How much do auto captions cost?▾
Cost scales with the length of the clip, and the app quotes it before the job runs. Current credit pack prices are on the pricing page.



