Japanese audio to text
Drop a Japanese recording in and get a transcript in normal mixed kanji and kana, produced in your browser with nothing uploaded.
Japanese is reasonably well supported, with the usual caveats about names and homophones.
Language and translation need the Multilingual model.
Drop an audio or video file
or , paste with Ctrl + V, or
- MP3
- MP4
- WAV
- M4A
- MOV
- WEBM
- OGG
- FLAC
- AAC
The file is read by this page and never uploaded — no account, no size limit but your own RAM.
Preset: multilingual model, language set to Japanese.
Download
Click a line to jump the audio there. Click the text to fix a word — your edits go into every download.
How it works
- Open this page — the tool is already set up for “Japanese to text”.
- Drop the file in, browse for it, or paste it with Ctrl+V.
- The model runs on your own device; the transcript appears as it goes.
- Correct anything misheard, then download text, subtitles or notes.
Kanji, kana and homophones
Output is ordinary written Japanese — kanji where kanji is conventional, kana elsewhere — rather than kana-only or romaji. The model chooses characters from context, which is exactly where Japanese speech recognition is hardest: the language is full of homophones, and picking the wrong kanji produces text that reads as a different word entirely while looking perfectly fluent.
Personal and company names are the worst case, because name readings are irregular even for humans. Check them; do not assume that a plausible-looking name is the right one.
Spacing and punctuation
Japanese is written without spaces, and the transcript follows that convention. Sentence-final punctuation (。and 、) is predicted from the rhythm of the speech, so a speaker who runs clauses together produces longer sentences in the text.
Because there are no word boundaries to anchor on, corrections are a little more fiddly than in a space-separated language — the click-to-play editor helps, since you can hear exactly the segment you are fixing.
Translation to English
Setting the output to translate produces English directly from Japanese speech. Japanese-to-English is a large linguistic distance and the result is the roughest of the language pairs offered here — expect the gist rather than the nuance, and expect politeness levels and subject omission to be handled crudely.
For understanding what a recording is about, it works. For anything else, use it to find the passage and get a human translation of that part.
Things that save a re-run
- Verify every personal and company name — irregular readings are the model’s weakest point.
- Homophone errors read fluently, so proofread by listening rather than by scanning the text.
- Translation to English is rough; use it to locate the important passage, not to render it.
Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.
Frequently asked questions
Does it output kanji or just kana?
Normal mixed kanji and kana, as Japanese is ordinarily written. It selects characters from context, which is also where its mistakes come from.
Why do the errors look like real words?
Japanese has many homophones, so a wrong character choice produces a fluent-looking but incorrect word. Proofread by listening rather than by reading — the text will not look obviously wrong.
How good is Japanese to English translation?
Rough. It conveys the gist and handles politeness levels and omitted subjects crudely. Use it to find the important passage, then get that part translated properly if it matters.
Is any audio sent to a server?
No. The model runs entirely in your browser.