Skip to content

Transcribe medical dictation on your own device

Dictated notes turned into text inside your browser. Nothing is uploaded, so patient-identifying audio is not handed to a transcription vendor to get a first draft.

For clinicians who dictate and then edit, and for anyone who needs a written version of recorded clinical audio without adding a processor to the chain.

Starting the engine…

Drop an audio or video file

or , paste with Ctrl + V, or

  • MP3
  • MP4
  • WAV
  • M4A
  • MOV
  • WEBM
  • OGG
  • FLAC
  • AAC

The file is read by this page and never uploaded — no account, no size limit but your own RAM.

Preset: balanced English model, plain-text output.

How it works

  1. Open this page — the tool is already set up for “Medical dictation”.
  2. Drop the file in, browse for it, or paste it with Ctrl+V.
  3. The model runs on your own device; the transcript appears as it goes.
  4. Correct anything misheard, then download text, subtitles or notes.

One fewer party in the chain

Clinical dictation names patients, describes conditions and frequently mentions family members. Every cloud transcription service that touches it becomes another organisation holding that material, another contract to maintain, another retention period to track and another surface in a breach report. That is manageable at institutional scale and disproportionate for a clinician who just wants a draft of what they dictated.

This removes the transmission entirely: the model downloads to your browser and runs on your machine. There is no upload endpoint on the site, so there is no processor relationship to establish and no vendor copy to account for. Your existing obligations around consent, storage and records are unchanged.

Clinical vocabulary is the weak point

Whisper is a general speech model, not a medical one. Drug names, dosages, anatomical terms and abbreviations are exactly where it errs, and it errs confidently — a mis-transcribed dose looks identical to a correct one. Purpose-built medical ASR products are trained on clinical corpora and do better on precisely this vocabulary; that is a real advantage and it is worth being clear about.

The practical consequence: treat the output as a draft to be read, and check every number, drug name and abbreviation against the audio. The click-to-play editor is built for that — click a line and hear it, correct it inline, move on.

Dictation records well

The good news is that dictation is the easy case acoustically: one speaker, close to the microphone, speaking deliberately. That is where even the small models perform strongly, and the Fast tier often suffices for a first pass and finishes almost instantly.

Dictating punctuation aloud ("comma", "full stop") does not work — Whisper predicts punctuation from the speech itself rather than taking commands, so spoken markers appear as words. Speaking in natural sentences with real pauses produces better punctuation than trying to control it.

Things that save a re-run

Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.

Frequently asked questions

Is this HIPAA compliant?

Compliance depends on your whole workflow and is assessed by your privacy officer or counsel, not by a tool’s marketing page. The relevant technical fact is that the audio is never transmitted, so no business-associate relationship arises for the transcription step. Consent, storage, access control and retention remain entirely your responsibility.

Is it as accurate as a medical transcription service?

Not on clinical vocabulary. Purpose-built medical ASR is trained on clinical language and handles drug names and abbreviations better. This is a general model, so expect to correct terminology — and always to check numbers.

Can I dictate directly into the page?

Yes — the microphone button records in the tab and transcribes when you stop. The recording is never written to a server and is discarded when you close the tab.

Does saying “comma” or “new paragraph” work?

No. The model predicts punctuation from natural speech rather than accepting spoken commands, so those words appear as text. Speak in ordinary sentences and let it punctuate.