Skip to content

M4A and voice memos to text

Drop an .m4a file in — the format the iPhone Voice Memos app and most phone recorders produce — and get a transcript without sending it anywhere.

This is the format people most often want transcribed and most often should not upload: interviews recorded on a phone, notes dictated in a car, a conversation someone agreed to record but not to publish.

Starting the engine…

Drop an audio or video file

or , paste with Ctrl + V, or

  • MP3
  • MP4
  • WAV
  • M4A
  • MOV
  • WEBM
  • OGG
  • FLAC
  • AAC

The file is read by this page and never uploaded — no account, no size limit but your own RAM.

Preset: balanced English model, plain-text output.

How it works

  1. Open this page — the tool is already set up for “M4A to text”.
  2. Drop the file in, browse for it, or paste it with Ctrl+V.
  3. The model runs on your own device; the transcript appears as it goes.
  4. Correct anything misheard, then download text, subtitles or notes.

Getting the file off an iPhone

In Voice Memos, tap the recording, tap the three dots and choose Share, then save to Files or AirDrop it to your computer — you get an .m4a. From the Files app you can also drag it straight into this page in Safari on a Mac or iPad. On Android, most recorder apps export .m4a or .mp3 and both work.

If the file arrives with no extension or an unfamiliar one, drop it in anyway. The browser sniffs the container rather than trusting the name, so a mislabelled file usually decodes fine.

Phone recordings are harder audio than they look

A phone on a table between two people is one of the harder cases for any speech model: the microphone is far from both speakers, the room adds reverberation, and phones apply aggressive noise processing tuned for calls rather than transcription. Expect more errors than from a headset recording, and expect them clustered on names and on the person further from the phone.

Two things help more than changing model: putting the phone closer to whoever is quieter, and asking people not to talk over each other. Neither is available after the fact — but if you are about to record something you will need transcribed, they are worth thirty seconds of setup.

Dictated notes

For a memo you dictated yourself, the Fast model is usually enough and finishes almost instantly — it is a close-microphone, single-speaker recording, which is exactly the easy case. Use the Markdown export if the notes are heading into Obsidian or Notion; it arrives with a heading, the source file name and timestamped paragraphs.

You can also skip the file entirely and record straight into the page with the microphone button, which produces the same result without a file ever existing on disk.

Things that save a re-run

Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.

Frequently asked questions

Does this work with iPhone Voice Memos files?

Yes — Voice Memos exports .m4a (AAC), which every browser decodes. Share the memo to Files or AirDrop it to a computer, then drop it in.

Can I record directly instead of uploading a file?

Yes. The microphone button records in the page and transcribes when you stop. The recording is held in the tab and never written to a server; closing the tab discards it.

Why is my phone recording less accurate than expected?

Distance and room noise, mostly. A phone lying on a table is far from both speakers and picks up reverberation, and phone microphones apply processing tuned for calls. The Balanced or Multilingual model helps a little; microphone placement helps a lot more.

Is the audio uploaded to check the format?

No. Format detection and decoding both happen in the browser using its own decoders. No part of the file is sent anywhere at any point.