Transcribe a podcast episode
Drop an episode in and get a full transcript, timestamped show notes and subtitle files — free and unlimited, because it runs on your machine rather than someone’s metered API.
Podcast transcripts are the cheapest search traffic a show can get, and the accessibility case for them is straightforward.
Language and translation need the Multilingual model.
Drop an audio or video file
or , paste with Ctrl + V, or
- MP3
- MP4
- WAV
- M4A
- MOV
- WEBM
- OGG
- FLAC
- AAC
The file is read by this page and never uploaded — no account, no size limit but your own RAM.
Preset: balanced English model, timestamped output.
Download
Click a line to jump the audio there. Click the text to fix a word — your edits go into every download.
How it works
- Open this page — the tool is already set up for “Podcasts”.
- Drop the file in, browse for it, or paste it with Ctrl+V.
- The model runs on your own device; the transcript appears as it goes.
- Correct anything misheard, then download text, subtitles or notes.
Transcripts are how podcasts get found
Audio is invisible to search engines. An episode page with a full transcript is indexable, quotable and linkable, and it turns every topic mentioned in ninety minutes of conversation into something a search can land on. For an independent show it is usually the single highest-leverage thing available, and the cost of the machine draft here is zero.
It is also the accessible version. A published transcript serves deaf and hard-of-hearing listeners, people in noisy environments, non-native speakers and anyone who would rather skim than listen — and unlike a summary, it does not decide for them which parts mattered.
Working the output into show notes
The timestamped export doubles as a chapter list: scan the paragraph starts, keep the six or seven that mark a topic change, and you have chapter markers with real timings. The Markdown export drops straight into a static site or a CMS with the structure intact.
For the episode page itself, the plain-text export is usually right — paragraphs grouped at natural pauses, no timecode clutter. If you want quotable pull-outs, the click-to-play editor makes checking a quote against the audio a two-second job.
Music, intros and multiple hosts
Musical intros and stingers are where Whisper invents text — it is trained to output speech, so over music it produces something plausible. Repeated identical lines are filtered automatically; anything left is obvious at the top of the transcript and takes a second to delete.
Multi-host shows record on separate tracks more often than not. If you have those, transcribing each track separately gives you genuine speaker separation with no guessing. From a single mixed file, use the manual labels — podcast turns are long, so it is quick.
Things that save a re-run
- Delete the invented text over the musical intro before publishing — it is always at the very top.
- If you record hosts on separate tracks, transcribe each one for real speaker separation.
- Publish the transcript on the episode page, not as a download — the point is that it is indexable.
Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.
Frequently asked questions
Is there a limit on episode length?
No imposed limit. The audio has to fit in your device’s memory, so a ninety-minute episode is comfortable on a laptop. Nothing is metered and nothing is charged.
Can I get chapter timestamps out of it?
The timestamped export gives every paragraph a start time; picking the ones that mark topic changes gives you chapters in a couple of minutes. There is no automatic topic detection — that would be guessing.
How do I handle a two-host show from one mixed file?
Label speakers manually — click a line, name the host, and it applies until the next label. Podcast turns are long, so it is usually a handful of clicks per episode.
Will the transcript be good enough to publish?
Not without a read-through. Studio-quality podcast audio transcribes well, but guest names, company names and technical terms will need fixing, and a published transcript with mangled names reflects on the show. Budget twenty minutes of editing per hour of audio.