Otobong Okoko
Back to lab
macOS AppStatus: Beta

Tauker

Push-to-talk dictation for the Mac. Hold a key, speak, release, and finished text lands in whatever app you are using. Your voice is transcribed on the Mac and the recording never leaves it.

Tech Stack

  • Swift 6
  • SwiftUI
  • AppKit
  • Apple SpeechAnalyzer
  • Accessibility API
  • XcodeGen
  • Ollama (optional)

Problem

Dictation tools like Wispr Flow are fast and accurate, and part of how they get there is by sending your audio to the cloud along with context read from your screen. My research notes on Wispr record five kinds of context collected from the Mac and uploaded with the recording. For notes, client work and half-formed thoughts, I did not want that trade.

macOS 26 shipped SpeechAnalyzer, Apple's new on-device transcription API. The question was whether a native app built on it could feel as quick as the cloud tools while keeping the audio on the machine.

Approach

Tauker is a menu-bar app. Hold Fn, or another key you choose, and a small floating indicator with a waveform shows that it is listening. Release, and the text is inserted into the frontmost app. Escape cancels. Audio streams to the recogniser while the key is held, so letting go only costs a final flush instead of a whole transcription pass.

Apple's recogniser returns unpunctuated words, so cleanup is its own stage. A local refiner removes filler, capitalises sentences, handles spoken layout commands, builds lists and resolves self-corrections. It always runs and needs no network. An optional AI cleanup pass is off by default. If you switch it on you pick the provider, including a local Ollama model, and only the text is ever sent, never the recording. The prompt tells the model that its input came from a speech recogniser and not from someone typing, which is what lets it fix sound-alike errors such as "entropy" for "Anthropic".

Insertion was the hardest part. The app writes through the Accessibility API and then reads the value back, because Electron apps and Chrome accept the write and silently ignore it. When the read-back fails it falls back to paste. Per-app overrides cover the stubborn cases, and a newline policy keeps a dictated line break from running a command in Terminal. The indicator is a non-activating panel that can never become the key window. If it took focus, the text field you were dictating into would lose it.

Privacy is enforced, not promised. A CI gate fails the build if transcript text ever appears in the logs, which record timings, outcomes and character counts and never the words. API keys live in the login keychain. The test suite has 203 tests, including recorded WAV fixtures that run through the real speech engine and a small fixture app that verifies each insertion strategy against a native text view, a secure field and a web view.

The build itself was an experiment in working with AI. I ran several Claude Code agents at once under a written file-claim protocol so they would not edit the same files, and commissioned clean-room reviews of the codebase that scored it against a rubric. The score went from 2.8 to 4.6 out of 5 over successive reviews, across 86 commits in a little over two weeks.

Key Learnings

  • Native APIs report success more often than they succeed. Reading back an Accessibility write is the only way to know the text arrived.
  • A privacy claim is only as strong as the check that enforces it. Failing the build on a logged transcript turned a promise into a property of the codebase.
  • Latency is a design material. Streaming while the key is held, and choosing lite models that finish in roughly 600 to 750 ms over ones that take 2.5 to 3 s, is what makes dictation feel instant.
  • Off by default is a product decision. The on-device path had to be good enough alone, so the cloud option stays a choice and never a requirement.
  • Several AI agents can work one codebase if ownership is written down. The file-claim protocol did the job a team's working agreements normally do.