630 MB · Windows 10 or 11, 64-bit · faster with an NVIDIA GPU
What it looks like
Click where you'd normally type, hold the key, and talk. That's the whole thing.
New Message
To: Sam
Subject: Friday delivery
Hi Sam — can you push Friday's delivery tothe afternoon? The morning crew is short andI'd rather not have the truck waiting. Thanks.
Hi Sam, can you push Friday's delivery to the afternoon?
Hold right ⌥Hold right CtrlClick into any text boxListening — just talkYour words, typed in for you
1
Click where you want the words to go — an email, a text, a form. Anywhere you can type.
2
Hold right ⌥right Ctrl and say it out loud, in your normal voice.
3
Let go. The words appear right where your cursor was. Nothing to copy or paste.
Getting started
1
Drag it to Applications
Open the download and drag Dictation across. Then open it.
2
Let it set up once
The first launch downloads the speech model — about 2.3 GB. It shows a progress window and only does this once.
3
Say yes to three permissions
Microphone to hear you, Accessibility to place the text, Input Monitoring to notice the hotkey. macOS asks for each.
4
Hold right-option and speak
Let go and your words are typed at the cursor. Double-tap right ⌥ instead to go hands-free, then tap again to stop.
1
Unzip it wherever you keep programs
There's no installer. Unzip the folder, keep its contents together, and run Dictation.exe from inside it.
2
Tell Windows to run it anyway
It isn't code-signed, so SmartScreen will say the publisher is unknown. Choose More info, then Run anyway. It asks once.
3
Let it set up once
The first launch downloads the speech engine and shows a progress window. It only does this once. On most PCs that's about 500 MB; if your machine has an NVIDIA graphics card it takes the larger, more accurate engine instead — about 2.9 GB.
4
Hold right-Ctrl and speak
Let go and your words are pasted at the cursor. Double-tap right Ctrl instead to go hands-free, then tap again to stop. No permission prompts — Windows doesn't need them for this.
The orb tells you what it's doing
There's no window. A small circle sits in the bottom-right corner of your screen, level with the Dock.
FadedIdle, waiting
RedListening — it grows with your voice
GreenTranscribing what you said
The tray icon tells you what it's doing
There's no window. A small glyph sits in the notification area, down by the clock.
GreyIdle, waiting
RedListening
AmberTranscribing what you said
BlueLoading the speech model, just after launch
Right-click the glyph for the menu: what it's doing, the last thing it
typed, your input device, which key is the hotkey, and the vocabulary file.
Windows likes to hide new tray icons — drag it out of the overflow arrow
once and it stays put.
It follows your microphone
Pick your input under Input device in the
menutray menu.
AirPods, a desk mic, the built-in oneA headset, a desk mic, the one in your webcam
— switch whenever, including mid-session.
Audio hardware moves around constantly —
AirPods connect, a headset changes profile, Continuity hands a call overa headset connects, Windows changes the default device, a meeting app grabs the mic.
The app re-checks what is actually plugged in rather than trusting what was
there when it started, and captures at whatever rate your microphone
prefers instead of insisting on its own.
It formats lists for you
Say a list out loud and you get a list. Say a sentence and you get a sentence.
“One, call the bank. Two, send the invoice. Three, book the flight.”
arrives as a numbered list. Announce a count — “I have four
things to tell you” — and the four things that follow arrive
as bullets.
It would rather do nothing than guess.
Numbered lists only form when the numbers you speak start at one and climb
without gaps. Bulleted lists only form when the count you announced matches
the number of things you actually said. Anything else is left exactly as
you dictated it — because mangled text costs more to fix than
unformatted text does. This is plain pattern matching in the app, not a
language model, so nothing is sent anywhere to make it work.
Before you download
Needs an Apple Silicon Mac — M1 or newer. An Intel Mac won't run it.
Needs macOS 13 (Ventura) or later.
Set aside 2.6 GB — the app plus the speech model it fetches on first launch.
Transcription runs on your Mac's own neural engine, through Apple's MLX
framework. There is no server, no account, and no API key — because there
is nothing to connect to.
Needs Windows 10 or 11, 64-bit.
Wants an NVIDIA GPU. With one it transcribes in float16, about a
second for a spoken sentence. Without one it falls back to the CPU in
int8 — it still works, it's just slower. Nothing to configure either way;
the app picks at launch and says which in its log.
Set aside about 3.9 GB — 1.0 GB unzipped, plus the 2.9 GB
speech model it fetches on first run.
Transcription runs locally through faster-whisper on CTranslate2, using
the large-v3 model — the big one,
because the smaller ones mangled real surnames past the point the
vocabulary could recover them. There is no server, no account, and no
API key — because there is nothing to connect to.
Why this is private
Your speech is turned into text entirely on your own machine, by a
model that lives in the app. No audio is uploaded. No transcript is
uploaded. There is no account to create and no company on the other end of
it — including us.
Your audioNever leaves the MacPC. Held in memory, discarded after transcribing.
Your textGoes to your clipboard and into whatever you're typing. Nowhere else.
Works offlineTurn the Wi-Fi off and it behaves identically.
The one exception, stated plainly: the
first time you open the app it downloads the speech model —
about 2.3 GBabout 2.9 GB,
once. That is the only time it uses the network. After that it never
connects again, and it does not check for updates, send analytics, or
report usage.