Captions and subtitles are not the same
Captions serve viewers watching muted or with hearing loss. Subtitles translate speech for viewers who do not share the language. EzScribe produces either from the same transcript.
Paste a video link and turn spoken audio into translated captions. Export WebVTT for web players and social uploads, or copy caption text straight into your post.
Powered by Whisper
#1 in speech-to-text accuracy
It writes out what is said in a video and translates it into the caption text your viewers read — on a player, on a social post, or as an accessibility track.
Browser video players read WebVTT natively, so a caption file in that format drops into a player without conversion.
WebVTT is the caption format an HTML video element loads without any extra tooling.
Editing software usually prefers SRT. Both come out of the same transcript.
Export TXT when you want the wording with no timing attached at all.
WebVTT comes first because that is what web players and most social uploads accept. The same captions are available timed or untimed.
The format HTML5 players expect, and the one most social platforms accept when you upload a caption file.
For when the captions have to go back through an editor before the video ships.
Caption text also fills descriptions, pinned comments and post copy, without rewatching the video.
Generate caption text from a public video's audio, translate it, and export the format your player or post needs.
Give viewers something to read whether they are watching muted, hard of hearing, or reading in another language.
Most social video starts without sound, so on-screen text decides whether the first few seconds land.
Provide caption text for viewers with hearing loss, and proofread it before publishing rather than shipping a raw draft.
Post the same video to each market with captions in the language that market actually reads.
Caption text also fills descriptions, pinned comments and post copy without watching the video again.
The caption language is detected automatically, so you choose only the language your audience reads. Export one version per market from the same video.
Captions come from the spoken audio of a video you are allowed to process. Here is what that covers, and what it does not.
Pick the tool that matches your source: paste a link from a social platform, or translate a transcript, subtitle file, or caption.
Paste a video link you own or are authorized to process, and give every viewer something to read — muted, hard of hearing, or in another language.
Translate captionsVideo links · 100+ languages · WebVTT, SRT and text exports
Answers about captions versus subtitles, burned-in text, exports for web players, languages and accuracy.
Captions carry what is heard for viewers watching without sound or with hearing loss. Subtitles translate speech for viewers who do not share the language. EzScribe can produce either from the same transcript.
Paste a public video URL you own or are authorized to process, wait for the caption text, then pick the language you want it translated into.
No. EzScribe listens to the audio track rather than reading text off the picture, so burned-in captions are not detected. Anything spoken aloud is still transcribed and translated.
No. Caption text is generated from the audio, so it works even when the platform offers none and does not depend on auto-caption availability.
WebVTT for HTML5 players and web captions, SRT for editing software, and TXT for plain reading. The export follows the view you have open, original or translated.
More than 100, with the source language detected automatically from the caption text.
Accuracy depends on audio clarity and on machine translation, so proofread the text before using it for accessibility or publishing it to an audience.
Captions here are what viewers read on the post or player. The subtitle translator focuses on timed files for a video editing workflow, and the transcript translator on readable text.