88 free credits on sign-up

Auto subtitle generator, in SRT

Paste a link, upload a file or record from your mic. You get a subtitle file — SRT or WebVTT, timed to the speech — to import alongside your video, not captions burned into the picture.

Generate subtitles from a video or audio file

Paste a link and we will fetch the audio

Try:

YouTube, TikTok, Instagram, Facebook, X and Spotify links

Only submit media you own, can access publicly, or are authorized to process. EzScribe does not redistribute source media.

Accurate AI Transcription
100+ Languages
Saved to Your Account
Private & Secure

Powered by Whisper

#1 in speech-to-text accuracy

How to generate subtitles for a video

Three steps, and the only one that takes real time is the one you do not do yourself.

  1. 01

    Add your video or audio

    Paste a link from a supported platform, upload a file from your device, or record straight from the microphone.

  2. 02

    Whisper transcribes and times it

    Speech is transcribed and split into segments, each with a start and end taken from the audio.

  3. 03

    Export SRT or WebVTT

    Download the format your editor or player wants. Plain TXT is there too if you only need the words.

SRT, WebVTT or plain text

All three come out of one job, as sidecar files you import alongside the video rather than captions rendered into it. If all you need is an SRT generator, take the first one and ignore the rest.

SRT

The universal one. CapCut, Premiere Pro, DaVinci Resolve, Final Cut and YouTube Studio all import SRT generator output directly, no conversion in between.

WebVTT

The same idea for the web, read natively by HTML5 players. Take this one when the video is going on a page you control.

Plain TXT

The words with the timings stripped out, for when the subtitle file was never the point and you only wanted a readable script.

What the AI subtitle generator gives you

Free credits on sign-up

Create an account and the first transcripts are on us. After that it is pay-as-you-go — no subscription, and a job that fails is never charged.

Results in seconds

Whisper large-v3-turbo speech recognition turns a minute of audio into text in a few seconds. Long files are transcribed in parts and stitched automatically.

100+ languages

English, Spanish, Portuguese, Hindi, Arabic, Japanese and dozens more — detected automatically, and translatable into any of the 100+ in the picker.

TXT, SRT & WebVTT export

Copy plain text for notes and articles, or download timestamped subtitle files for video work.

Saved to your account

Every transcript stays in My Transcripts — reopen it, export it again in another format, or delete it. Nothing depends on keeping the tab open.

Works without captions

We transcribe the actual audio track, so videos with no captions, or with captions burned into the picture, work fine.

What people use the subtitle generator for

Four reasons to generate subtitles rather than key them in by hand.

Captions for social video

Most short-form is watched on mute. An auto subtitle generator pass takes minutes instead of the hour it takes to type a track by hand.

Accessibility on your own site

A WebVTT track attached to an HTML5 player is the baseline requirement for a video that anyone can follow without sound.

Course and training video

Generate subtitles once and the lesson becomes searchable as well as watchable, which is usually the harder half of the job.

A base for translated tracks

Make the original track here, then translate it with the timings left alone. One source track becomes a dozen language versions.

How the subtitle generator works

Subtitles are a transcript plus timing. The subtitle generator runs speech recognition over the audio track and keeps the start and end of every sentence, then packages the result as numbered SRT blocks or a web-native WebVTT file. Because the timing comes from the speech model rather than being spread evenly across the runtime, most tracks import without nudging.

The auto subtitle generator takes the video however you have it: paste a public YouTube, TikTok, Instagram, X or Facebook link, upload the file itself, or record straight from your microphone. The output is identical either way.

A video timeline with timed subtitle segments lined up under the speech

Sidecar subtitles, not burned-in captions

A subtitle generator can hand you one of two things: a file that sits beside the video, or captions painted permanently into the frames. This one makes the file, which is the less obvious choice and usually the better one. A sidecar track stays editable after the video is exported, a viewer can switch it off, a player can restyle it, and a search engine can read it. Burned-in captions are pixels, and pixels cannot be corrected, translated or turned off.

If you do want them burned in, the subtitle generator is still the first step: generate the SRT here, drop it on the timeline in CapCut, Premiere Pro or DaVinci Resolve, and render. Leaving the burn to your editor is deliberate rather than a gap — font, size, position and safe margins are design decisions that depend on the video, and no automatic pass gets them right for a full-bleed vertical and a 16:9 lecture at the same time.

Translated tracks, and words with no timings

To put a subtitle file you already have into another language with its timings left alone, the video subtitle translator does that instead — it rewrites the text inside the file and leaves every timecode where it was. If you want the words with no timings at all, the MP4 and audio pages run the same engine and return plain prose.

Generate your subtitle track

One source in, a timed SRT or WebVTT out. No manual keying, no nudging timecodes line by line.

Generate subtitles

Sign in to start. New accounts get free credits.

Subtitle generator — FAQ

The questions that come up before the first export.

Add the video — paste a link, upload the file, or record audio in the browser. The speech is transcribed and split into timed segments, and you download the result as SRT or WebVTT.

SRT. Both take it directly, as do Final Cut, DaVinci Resolve and YouTube. Use WebVTT when the file is going into an HTML5 player on a web page.

Each segment carries the start and end taken from the audio during transcription, so the timings follow the speech rather than being spread evenly across the runtime. Overlapping music or crosstalk can still shift a boundary by a fraction of a second.

No. You get a subtitle file to import alongside your video. Burning them in is a render, and that belongs in your editor where you control the font, position and safe margins.

Yes. Transcribe first, then translate the result into any of 100+ languages from the same screen. The timings carry across, so the translated track drops in where the original did.

Long recordings are fine — the ceiling is 12 hours per job. Credits are charged per minute of measured audio, so a long file costs more than a short one but does not need splitting.

This page makes a subtitle track from scratch, out of the audio. The Video Subtitle Translator takes subtitles that already exist and puts them into another language. If you have no subtitle file yet, you want this page.

Yes. A transcription job has to be billed to someone and filed somewhere, so it needs an account. New accounts get free credits, and your transcripts stay in your account afterwards.