Skip to content

Video to SRT subtitles

Drop a video in and download a ready-to-use .srt subtitle file. The speech model runs in this browser tab, so the video stays on your machine — no upload, no queue, no watermark, no per-minute charge.

Cues are split at sentence boundaries and wrapped to two readable lines, so the result looks like subtitles rather than a transcript with timecodes bolted on.

Starting the engine…

Drop an audio or video file

or , paste with Ctrl + V, or

  • MP3
  • MP4
  • WAV
  • M4A
  • MOV
  • WEBM
  • OGG
  • FLAC
  • AAC

The file is read by this page and never uploaded — no account, no size limit but your own RAM.

Preset: SRT output, balanced English model.

How it works

  1. Open this page — the tool is already set up for “Video to SRT”.
  2. Drop the file in, browse for it, or paste it with Ctrl+V.
  3. The model runs on your own device; the transcript appears as it goes.
  4. Correct anything misheard, then download text, subtitles or notes.

What makes a subtitle file usable

Most free auto-subtitle tools do the recognition and stop there, which produces cues that are technically valid and practically unreadable: a single cue holding thirty seconds of speech, or lines so long they wrap three deep over the picture. Broadcast practice — the convention behind the BBC and Netflix style guides everyone borrows from — is roughly two lines of about 42 characters, held long enough to read.

This tool applies that automatically. Over-long segments are split, preferring a sentence boundary near the middle and distributing the timing across the pieces; each cue is then wrapped to two lines at word boundaries. You can still correct any line by hand before downloading, and your edits go into the file.

Where the .srt file goes

Save it beside the video with the same base name (talk.mp4 and talk.srt) and VLC, IINA and most smart TVs will pick it up automatically. Editors — Premiere, DaVinci Resolve, Final Cut, CapCut, Shotcut — all import SRT directly onto a caption track. YouTube, LinkedIn and Vimeo accept it as a subtitle upload, which is worth doing even when they offer automatic captions: your version is corrected, theirs is not.

If you need WebVTT for an HTML5 player or a streaming manifest, switch the format — it is the same transcript written with a different timestamp separator and a header line.

Accessibility, honestly

Automatic subtitles are a strong start and a weak finish. They mishear names, they do not mark who is speaking, and they will not describe the door slamming off-screen. For subtitles that genuinely serve deaf and hard-of-hearing viewers, the machine pass is step one and a human read-through is step two — this tool exists to make step one free and private, not to remove step two.

The editor is built for that read-through: play the video in one window, click any line here to jump the audio, fix it inline, re-export. It is much faster than typing captions from scratch, which is the realistic alternative.

Things that save a re-run

Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.

Frequently asked questions

Is the SRT ready to use straight away?

It is a valid, correctly numbered SubRip file that any player will load. Whether it is ready to publish depends on the audio — read it through and fix names and technical terms first. That is true of every automatic captioning tool, including the paid ones.

Can I get VTT instead?

Yes — switch the download format to WebVTT. It is the same content with dot-separated timestamps and a WEBVTT header, which is what HTML5 video and most web players expect.

Do the subtitles include who is speaking?

Only if you label them. Whisper does not identify speakers, and inventing labels would be worse than omitting them. You can assign a speaker to a line and every line after it in one click, and those labels are written into the cues.

How long can the video be?

There is no imposed limit. The audio has to fit in memory, so an hour-long video is straightforward and a three-hour one will work but use a lot of RAM. Progress and partial results appear as it goes.