Translate audio into English text
Drop a recording in any of 99 languages and get English text straight out — translated in a single pass by a model running in your browser.
No upload, no account, no per-minute charge. The language is detected automatically unless you pick one.
Language and translation need the Multilingual model.
Drop an audio or video file
or , paste with Ctrl + V, or
- MP3
- MP4
- WAV
- M4A
- MOV
- WEBM
- OGG
- FLAC
- AAC
The file is read by this page and never uploaded — no account, no size limit but your own RAM.
Preset: multilingual model, output set to translate into English.
Download
Click a line to jump the audio there. Click the text to fix a word — your edits go into every download.
How it works
- Open this page — the tool is already set up for “Translate to English”.
- Drop the file in, browse for it, or paste it with Ctrl+V.
- The model runs on your own device; the transcript appears as it goes.
- Correct anything misheard, then download text, subtitles or notes.
One pass, not two
This is not transcription followed by machine translation. Whisper’s multilingual checkpoints were trained on a translation task directly: given foreign-language speech, produce English text. The model goes from audio to English in a single decode, which avoids compounding a transcription error into a translation error.
The direction is fixed. Whisper translates into English only — there is no setting to produce Spanish from English speech, or any other target language. That is a property of the model, not a limitation of this site.
What the quality is actually like
It varies enormously with the source language, and it is best to be blunt about it. Speech in the major European languages — Spanish, French, German, Italian, Portuguese, Dutch — produces English that reads reasonably well and conveys the substance. Languages that are linguistically distant from English or thinly represented in the training data produce something closer to a gist: you will understand what the recording is about, but the phrasing will be crude and nuance will be lost.
Across all of them, this is a working translation. It is enough to triage a recording, decide whether it matters, and identify the two minutes worth paying a human translator for. It is not enough to publish, to sign, or to rely on in any setting where the precise words matter.
Detection and mixed-language audio
Language detection reads the first few seconds, so a clip that opens in one language and continues in another can be identified wrongly. For a short clip, or one that begins with music or an English introduction, set the language explicitly instead of relying on detection.
Genuinely bilingual recordings — where speakers alternate mid-sentence — are the hardest case and no setting fixes them. Translation output sometimes handles these better than transcription, simply because everything ends up in one language either way.
Things that save a re-run
- Set the language explicitly for short clips or anything starting with music — detection reads only the opening seconds.
- Use it to triage: find the part that matters, then get that part translated properly.
- Translation only goes into English; there is no other target language.
Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.
Frequently asked questions
Which languages can it translate from?
The 99 languages Whisper’s multilingual checkpoints support. Quality varies a great deal between them — major European languages are strong, thinly-represented languages produce a rough gist.
Can it translate English into another language?
No. Whisper translates only into English. Producing speech or text in another target language is a different kind of model and is not part of this tool.
Is it transcribing first and then translating?
No — it goes from audio directly to English text in one pass, which is what the model was trained to do. That avoids a transcription error being compounded by a translation error.
Is the recording uploaded to be translated?
No. The multilingual model is downloaded to your browser and everything happens on your device. Nothing is transmitted.