Hold a key.
Talk.

Your words appear wherever you're typing — email, Slack, a document, anywhere.

On-device transcription. Your voice never leaves your MacPC.

Download for Mac Download for Windows

264 MB · Apple Silicon · macOS 13 or later

630 MB · Windows 10 or 11, 64-bit · faster with an NVIDIA GPU

What it looks like

Click where you'd normally type, hold the key, and talk. That's the whole thing.

1

Click where you want the words to go — an email, a text, a form. Anywhere you can type.

2

Hold right ⌥right Ctrl and say it out loud, in your normal voice.

3

Let go. The words appear right where your cursor was. Nothing to copy or paste.

Getting started

1

Drag it to Applications

Open the download and drag Dictation across. Then open it.

2

Let it set up once

The first launch downloads the speech model — about 2.3 GB. It shows a progress window and only does this once.

3

Say yes to three permissions

Microphone to hear you, Accessibility to place the text, Input Monitoring to notice the hotkey. macOS asks for each.

4

Hold right-option and speak

Let go and your words are typed at the cursor. Double-tap right ⌥ instead to go hands-free, then tap again to stop.

1

Unzip it wherever you keep programs

There's no installer. Unzip the folder, keep its contents together, and run Dictation.exe from inside it.

2

Tell Windows to run it anyway

It isn't code-signed, so SmartScreen will say the publisher is unknown. Choose More info, then Run anyway. It asks once.

3

Let it set up once

The first launch downloads the speech engine and shows a progress window. It only does this once. On most PCs that's about 500 MB; if your machine has an NVIDIA graphics card it takes the larger, more accurate engine instead — about 2.9 GB.

4

Hold right-Ctrl and speak

Let go and your words are pasted at the cursor. Double-tap right Ctrl instead to go hands-free, then tap again to stop. No permission prompts — Windows doesn't need them for this.

The orb tells you what it's doing

There's no window. A small circle sits in the bottom-right corner of your screen, level with the Dock.

FadedIdle, waiting
RedListening — it grows with your voice
GreenTranscribing what you said

The tray icon tells you what it's doing

There's no window. A small glyph sits in the notification area, down by the clock.

GreyIdle, waiting
RedListening
AmberTranscribing what you said
BlueLoading the speech model, just after launch
Right-click the glyph for the menu: what it's doing, the last thing it typed, your input device, which key is the hotkey, and the vocabulary file. Windows likes to hide new tray icons — drag it out of the overflow arrow once and it stays put.

It follows your microphone

Pick your input under Input device in the menutray menu. AirPods, a desk mic, the built-in oneA headset, a desk mic, the one in your webcam — switch whenever, including mid-session.

Audio hardware moves around constantly — AirPods connect, a headset changes profile, Continuity hands a call overa headset connects, Windows changes the default device, a meeting app grabs the mic. The app re-checks what is actually plugged in rather than trusting what was there when it started, and captures at whatever rate your microphone prefers instead of insisting on its own.

It formats lists for you

Say a list out loud and you get a list. Say a sentence and you get a sentence.

One, call the bank. Two, send the invoice. Three, book the flight.” arrives as a numbered list. Announce a count — “I have four things to tell you” — and the four things that follow arrive as bullets.

It would rather do nothing than guess. Numbered lists only form when the numbers you speak start at one and climb without gaps. Bulleted lists only form when the count you announced matches the number of things you actually said. Anything else is left exactly as you dictated it — because mangled text costs more to fix than unformatted text does. This is plain pattern matching in the app, not a language model, so nothing is sent anywhere to make it work.

Before you download

Needs an Apple Silicon Mac — M1 or newer. An Intel Mac won't run it.
Needs macOS 13 (Ventura) or later.
Set aside 2.6 GB — the app plus the speech model it fetches on first launch.

Transcription runs on your Mac's own neural engine, through Apple's MLX framework. There is no server, no account, and no API key — because there is nothing to connect to.

Needs Windows 10 or 11, 64-bit.
Wants an NVIDIA GPU. With one it transcribes in float16, about a second for a spoken sentence. Without one it falls back to the CPU in int8 — it still works, it's just slower. Nothing to configure either way; the app picks at launch and says which in its log.
Set aside about 3.9 GB — 1.0 GB unzipped, plus the 2.9 GB speech model it fetches on first run.

Transcription runs locally through faster-whisper on CTranslate2, using the large-v3 model — the big one, because the smaller ones mangled real surnames past the point the vocabulary could recover them. There is no server, no account, and no API key — because there is nothing to connect to.

Why this is private

Your speech is turned into text entirely on your own machine, by a model that lives in the app. No audio is uploaded. No transcript is uploaded. There is no account to create and no company on the other end of it — including us.

Your audioNever leaves the MacPC. Held in memory, discarded after transcribing.
Your textGoes to your clipboard and into whatever you're typing. Nowhere else.
Works offlineTurn the Wi-Fi off and it behaves identically.
The one exception, stated plainly: the first time you open the app it downloads the speech model — about 2.3 GBabout 2.9 GB, once. That is the only time it uses the network. After that it never connects again, and it does not check for updates, send analytics, or report usage.