German audio to text
Drop a German recording in and get a German transcript, produced on your own device — no upload, no account, no data leaving the machine.
German is one of Whisper’s stronger languages, and long compound nouns are generally written out correctly rather than split.
Language and translation need the Multilingual model.
Drop an audio or video file
or , paste with Ctrl + V, or
- MP3
- MP4
- WAV
- M4A
- MOV
- WEBM
- OGG
- FLAC
- AAC
The file is read by this page and never uploaded — no account, no size limit but your own RAM.
Preset: multilingual model, language set to German.
Download
Click a line to jump the audio there. Click the text to fix a word — your edits go into every download.
How it works
- Open this page — the tool is already set up for “German to text”.
- Drop the file in, browse for it, or paste it with Ctrl+V.
- The model runs on your own device; the transcript appears as it goes.
- Correct anything misheard, then download text, subtitles or notes.
Why local processing gets asked about here
German-speaking organisations tend to ask harder questions about where data is processed than most, and works councils, data protection officers and university ethics boards routinely rule out sending recordings to external services. The technical position of this tool is simple to state in that conversation: the audio is never transmitted, because the site has no upload endpoint and the model runs in the browser.
That is a statement about the software, not a compliance opinion — whether it satisfies a particular internal policy is for whoever owns that policy. But it is a materially different starting point from a cloud transcription service, and it is verifiable: transcribe with the network disconnected and it still works.
Compounds, umlauts and ß
Umlauts and ß are produced correctly, and compound nouns are generally written as single words rather than split into pieces — the model has seen enough German to know that Datenschutzbeauftragter is one word. Very long or invented compounds specific to an organisation are less reliable and are worth checking.
Capitalisation of nouns follows standard German orthography. Where a transcript looks wrong, it is far more often a misheard proper noun than a spelling convention error.
Dialect is the limit
Standard German transcribes well. Strong regional dialect — Bavarian, Swiss German, broad Austrian — is much weaker, because the training data is dominated by standard forms. Swiss German in particular often comes out as an approximate standard-German rendering rather than what was actually said.
If a recording is heavily dialectal, expect to correct a great deal, and consider whether a transcript is the right artefact at all.
Things that save a re-run
- Check organisation-specific compound nouns and acronyms — they are the predictable errors.
- Strong dialect, especially Swiss German, is where accuracy drops sharply.
- Transcribe with the network off to demonstrate to a colleague that nothing is sent.
Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.
Frequently asked questions
Are umlauts and ß handled correctly?
Yes. The model produces standard German orthography including umlauts, ß and noun capitalisation.
Does it handle Swiss German or Bavarian?
Poorly compared with standard German. Training data is dominated by standard forms, so strong dialect is often rendered as an approximate standard-German version rather than what was said.
Is the recording processed in Germany or the EU?
It is processed on your own computer, wherever that is. No audio is transmitted to any server in any country, which is usually the point of the question.
Can it translate German into English?
Yes — set the output to translate and it produces English text directly from the German speech. It does not translate in the other direction.