Turn a voice memo into text
Dictated a note and want it in writing? Drop the memo in — or record straight into the page — and it is transcribed on your device in seconds.
This is the easy case for speech recognition: one voice, close to the microphone. The Fast model handles it almost instantly.
Language and translation need the Multilingual model.
Drop an audio or video file
or , paste with Ctrl + V, or
- MP3
- MP4
- WAV
- M4A
- MOV
- WEBM
- OGG
- FLAC
- AAC
The file is read by this page and never uploaded — no account, no size limit but your own RAM.
Preset: Fast model — plenty for close-microphone dictation, and it finishes almost instantly.
Download
Click a line to jump the audio there. Click the text to fix a word — your edits go into every download.
How it works
- Open this page — the tool is already set up for “Voice memos”.
- Drop the file in, browse for it, or paste it with Ctrl+V.
- The model runs on your own device; the transcript appears as it goes.
- Correct anything misheard, then download text, subtitles or notes.
Record here, or bring a file
The microphone button records inside the page and transcribes when you stop, so a spoken thought becomes text without a file ever existing on disk. Nothing is streamed anywhere while you record — the audio accumulates in the tab and is discarded when you close it.
If the memo already exists, drop it in: .m4a from iPhone Voice Memos, .mp3 or .m4a from Android recorders, .ogg from a messaging app. All of them decode natively.
Why the smallest model is enough here
The Fast tier is a 45 MB download and transcribes several times faster than real time even on modest hardware. On close-microphone single-speaker speech — which is what a voice memo is — the gap between it and the larger models is small, because the audio simply is not ambiguous.
Where it falls behind is exactly what a memo is not: distant microphones, several voices, background noise. If a particular memo comes out rough, switching to Balanced and re-running takes one click, and the model is cached from then on.
Getting it somewhere useful
Copy to clipboard is usually the fastest route — the plain-text export is already grouped into paragraphs rather than one line per phrase, so it pastes cleanly into a task manager, an email or a document.
For notes that belong in a knowledge base, the Markdown export arrives with a heading and timestamped paragraphs, ready for Obsidian, Notion or Logseq without reformatting.
Things that save a re-run
- Speak in natural sentences with real pauses — punctuation is predicted from the speech, not from spoken commands.
- Recording in a car gives the model mostly road noise; park first if the note matters.
- For a long rambling memo, use the timestamped export so you can jump to the part worth keeping.
Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.
Frequently asked questions
Can I record directly rather than uploading a file?
Yes. The microphone button records in the page and transcribes when you press stop. Nothing is streamed while recording, and the audio is discarded when the tab closes.
How fast is it?
On the Fast model, several times quicker than real time — a two-minute memo takes seconds once the model is cached. The first run has to download the model, which is 45 MB for that tier.
Does saying “new paragraph” work?
No. The model predicts punctuation from natural speech rather than accepting spoken commands, so those words appear as text. Speak normally and pause between thoughts.
Is anything kept after I close the tab?
Only the speech model, which is cached so you do not have to download it again. Your audio and your transcript are not stored anywhere and are gone when the tab closes — export anything you want to keep.