MOV and iPhone video to text
Drop a .mov file — the format iPhones and QuickTime produce — and the audio is transcribed in this tab. The video is never uploaded and the picture is never read.
Useful for filmed interviews, site walkthroughs, recorded statements and anything shot on a phone that you need in writing but cannot hand to a cloud service.
Language and translation need the Multilingual model.
Drop an audio or video file
or , paste with Ctrl + V, or
- MP3
- MP4
- WAV
- M4A
- MOV
- WEBM
- OGG
- FLAC
- AAC
The file is read by this page and never uploaded — no account, no size limit but your own RAM.
Preset: balanced English model. Switch to SRT for subtitles.
Download
Click a line to jump the audio there. Click the text to fix a word — your edits go into every download.
How it works
- Open this page — the tool is already set up for “MOV to text”.
- Drop the file in, browse for it, or paste it with Ctrl+V.
- The model runs on your own device; the transcript appears as it goes.
- Correct anything misheard, then download text, subtitles or notes.
iPhone video specifics
iPhones record .mov with H.264 or HEVC video and AAC audio. Browsers decode the AAC audio track without needing the video codec at all, so an HEVC clip that will not preview in some desktop players still transcribes here — the tool only asks for sound.
Files straight off a phone are large. Transferring the original rather than a "share" copy is worth it: iOS sometimes re-encodes on share, and while that rarely hurts the audio, it can subtly change timings you might later be matching subtitles against.
Filmed statements and walkthroughs
A phone video is often the record of something that happened once: a damaged property, a site inspection, a witness describing what they saw. A written transcript with timecodes turns that into something searchable and quotable, and doing it locally means the footage — which shows faces, addresses and surroundings — never goes to a third party.
Use the timestamped text export for this. Each paragraph carries the time it starts, so a note can reference "at 4:12" and anyone can jump straight there in the original video.
Wind, distance and outdoor audio
Outdoor phone footage is the hardest audio this tool sees. Wind across the microphone destroys speech information outright, and there is nothing a model can recover from it. Distance is nearly as bad: at three metres in the open, an iPhone microphone picks up more environment than voice.
If you can influence the recording, get within a metre and shield the microphone with a hand or a body. If you cannot, expect a partial transcript and use it as an index into the footage rather than a substitute for it.
Things that save a re-run
- Transfer the original file rather than a shared copy to keep timings exact.
- For a long walkthrough, the timestamped export doubles as a chapter list.
- If the clip has no speech in parts, expect gaps — that is correct behaviour, not a failure.
Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.
Frequently asked questions
Do I need the video codec to be supported?
No. Only the audio track is decoded, so an HEVC clip works as long as your browser can read the AAC audio inside the container — which it can.
Can I transcribe straight from my iPhone?
Yes, in Safari on iOS you can pick a video from Files or the photo library. It works, but phones have far less memory than laptops, so use the Fast model and keep clips short. A laptop is a much better experience for anything long.
Is the footage uploaded to extract the audio?
No. Extraction is done by your browser’s own media decoder inside the page. Nothing is transmitted at any stage.
Why is outdoor audio so much worse?
Wind noise and distance remove the speech information before it ever reaches the file. No speech model can recover what was not captured — this is a recording problem rather than a transcription one.