WebM to text
WebM is what browsers, OBS, Loom-style screen recorders and most web apps produce. Drop one in and it is transcribed locally, with no upload step in between.
It is the native format of the web, so it decodes instantly and reliably here — including the Opus audio that browser recorders default to.
Language and translation need the Multilingual model.
Drop an audio or video file
or , paste with Ctrl + V, or
- MP3
- MP4
- WAV
- M4A
- MOV
- WEBM
- OGG
- FLAC
- AAC
The file is read by this page and never uploaded — no account, no size limit but your own RAM.
Preset: balanced English model, plain-text output.
Download
Click a line to jump the audio there. Click the text to fix a word — your edits go into every download.
How it works
- Open this page — the tool is already set up for “WebM to text”.
- Drop the file in, browse for it, or paste it with Ctrl+V.
- The model runs on your own device; the transcript appears as it goes.
- Correct anything misheard, then download text, subtitles or notes.
Screen recordings and demos
A screen recording is a document that happens to be a video: someone explaining a process while showing it. Turning it into text gives you something searchable, quotable and pasteable into documentation — and doing it locally matters, because screen recordings routinely capture internal dashboards, customer names, ticket queues and half-open email clients.
The timestamped export works well here: each paragraph is anchored to a moment in the recording, which makes it straightforward to turn a twenty-minute walkthrough into a written procedure with references back to the video for the fiddly parts.
Opus audio and why it is a good input
WebM files from browsers almost always carry Opus audio, which is designed for speech and stays intelligible at low bitrates. That makes it a genuinely good input for recognition — better, at the same file size, than the equivalent MP3.
The microphone recorder built into this page produces WebM/Opus for the same reason. Whether you record here or bring a file, the path through the model is identical.
System audio versus microphone
Screen recorders can capture your microphone, the system audio, or both mixed together. Mixed tracks are harder: music stings, notification sounds and video playback inside the recording all become things the model tries to interpret as speech. If your recorder can produce a microphone-only track, that transcribes noticeably more cleanly.
Where a mixed track is all you have, expect some invented text over the musical sections. Repeated identical lines are filtered out automatically, but read the boundaries around any music.
Things that save a re-run
- Record a microphone-only track when the tool allows it — mixed system audio costs real accuracy.
- OBS "capture audio" defaults often include desktop sound; check before a long recording.
- WebM from browsers is Opus, which is efficient — there is no reason to convert it to WAV first.
Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.
Frequently asked questions
Does WebM audio-only work as well as WebM video?
Yes. Only the audio track is used in either case, so an audio-only .webm and a screen recording behave identically.
Are screen recordings safe to transcribe here?
The file is never uploaded, so nothing that is visible on screen or audible in the recording is disclosed to anyone. That is the reason to use a local tool for work recordings in the first place.
Why does music in my recording produce strange text?
Whisper is trained to produce speech, so over music or long silence it sometimes invents a plausible phrase. Repeated identical lines are stripped automatically; anything left over is easy to spot and delete in the editor.
Can I record here instead of using a screen recorder?
You can record audio directly with the microphone button. Screen capture itself is out of scope — this tool works with the file your recorder produces.