Skip to content

Transcribe a podcast episode

Drop an episode in and get a full transcript, timestamped show notes and subtitle files — free and unlimited, because it runs on your machine rather than someone’s metered API.

Podcast transcripts are the cheapest search traffic a show can get, and the accessibility case for them is straightforward.

Starting the engine…

Drop an audio or video file

or , paste with Ctrl + V, or

  • MP3
  • MP4
  • WAV
  • M4A
  • MOV
  • WEBM
  • OGG
  • FLAC
  • AAC

The file is read by this page and never uploaded — no account, no size limit but your own RAM.

Preset: balanced English model, timestamped output.

How it works

  1. Open this page — the tool is already set up for “Podcasts”.
  2. Drop the file in, browse for it, or paste it with Ctrl+V.
  3. The model runs on your own device; the transcript appears as it goes.
  4. Correct anything misheard, then download text, subtitles or notes.

Transcripts are how podcasts get found

Audio is invisible to search engines. An episode page with a full transcript is indexable, quotable and linkable, and it turns every topic mentioned in ninety minutes of conversation into something a search can land on. For an independent show it is usually the single highest-leverage thing available, and the cost of the machine draft here is zero.

It is also the accessible version. A published transcript serves deaf and hard-of-hearing listeners, people in noisy environments, non-native speakers and anyone who would rather skim than listen — and unlike a summary, it does not decide for them which parts mattered.

Working the output into show notes

The timestamped export doubles as a chapter list: scan the paragraph starts, keep the six or seven that mark a topic change, and you have chapter markers with real timings. The Markdown export drops straight into a static site or a CMS with the structure intact.

For the episode page itself, the plain-text export is usually right — paragraphs grouped at natural pauses, no timecode clutter. If you want quotable pull-outs, the click-to-play editor makes checking a quote against the audio a two-second job.

Music, intros and multiple hosts

Musical intros and stingers are where Whisper invents text — it is trained to output speech, so over music it produces something plausible. Repeated identical lines are filtered automatically; anything left is obvious at the top of the transcript and takes a second to delete.

Multi-host shows record on separate tracks more often than not. If you have those, transcribing each track separately gives you genuine speaker separation with no guessing. From a single mixed file, use the manual labels — podcast turns are long, so it is quick.

Things that save a re-run

Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.

Frequently asked questions

Is there a limit on episode length?

No imposed limit. The audio has to fit in your device’s memory, so a ninety-minute episode is comfortable on a laptop. Nothing is metered and nothing is charged.

Can I get chapter timestamps out of it?

The timestamped export gives every paragraph a start time; picking the ones that mark topic changes gives you chapters in a couple of minutes. There is no automatic topic detection — that would be guessing.

How do I handle a two-host show from one mixed file?

Label speakers manually — click a line, name the host, and it applies until the next label. Podcast turns are long, so it is usually a handful of clicks per episode.

Will the transcript be good enough to publish?

Not without a read-through. Studio-quality podcast audio transcribes well, but guest names, company names and technical terms will need fixing, and a published transcript with mangled names reflects on the show. Budget twenty minutes of editing per hour of audio.