Skip to content
KekosoDownload

Speech in, text out. Nothing leaves your Mac.

Voice to text on your Mac: dictate into any app, drop in a recording or a YouTube link, record a call and get both sides as text, and hand your AI agents a transcription server that runs locally. One purchase — no cloud, no subscription, no account.

macOS 14.4 or later · Apple silicon · $29 once, not per month · 7-day trial

Mail
Kekoso
New Message
Subject:
Filing timeline

Quick note before the deposition: the client confirmed the timeline, so we can file on Thursday.

Listening⌥ Space

What it does

Four kinds of speech. One app, one machine.

Speech to text on a Mac usually means paying three or four services to cover it — one for dictation, one for calls, one for converting files, an API key for everything else. Kekoso does all four, and the voice recognition runs on hardware you already own.

Dictation

Talk into any app

Press the hotkey, say it, and the text appears where your cursor already was — Mail, Slack, Cursor, a browser form. No per-app integration, no window to switch to.

Files and links

Anything with a voice in it

Drop in an audio or video file — MP3, M4A, WAV, MP4 — a voice recording off your phone, or paste a link to a YouTube video you have the rights to. You get a clean transcript with timestamps, exportable as TXT, SRT or VTT. Hours of recording converted to text on your own machine, with nothing uploaded to convert it.

Calls

Calls, without a bot in the room

Record from your side and get a transcript already split in two: your microphone on one track, everyone else on the other. Nobody has to admit a notetaker to the meeting, and “who said what” is a property of the recording rather than a guess.

Agents

A transcription server for your tools

Kekoso runs a local MCP server with twelve tools. Claude Code, Claude Desktop and Codex transcribe through it, and manage your vocabulary and models — no API key, no per-minute billing, no round trip to anyone else.

The distinction

Bot-free is not the same as local.

Nearly every competitor captures audio on your machine and then sends it somewhere else to be transcribed. That is a different promise than the one it sounds like. Here are the two routes your voice can take.

Most "private" dictation apps

  1. 1

    Your microphone

    Audio captured on device

  2. 2

    Their servers

    Uploaded for recognition

  3. 3

    Their storage

    Transcript retained, sometimes for training

Recording locally is not the same as processing locally.

Kekoso

  1. 1

    Your microphone

    Audio captured on device

  2. 2

    Your Neural Engine

    Model runs on your Mac

  3. 3

    Your text field

    Text inserted, audio discarded

No account, no upload, no server to subpoena.

How it works

Three seconds from thought to text.

Talk to text is the whole loop, and it is three steps long: press, speak, and the words are already in the field you were typing in. No window to switch to, no file to upload, no queue to wait in.

  1. 01

    Press the hotkey

    Hold to talk, or tap to toggle. Works in every app, including the ones that block other tools. Your cursor stays exactly where it was.

  2. 02

    Say it out loud

    The model is already warm in memory, so the first word is never clipped. A small panel shows the level so you know it heard you.

  3. 03

    The text is already there

    Inserted into the field you were in, with your own vocabulary applied, and written to history at the same moment — so nothing is lost if the paste misses.

The rest of it

What else Kekoso does.

Dictation is where most people start. These are the parts that make it worth keeping on the machine.

Files and links

A three-hour recording, transcribed while you make coffee

Drag in an MP3, an MOV, a voice memo, a screen recording. Or paste a link to a YouTube video you have the rights to. Everything after that happens on your Mac — the audio-to-text conversion included — so there is no upload, no queue and no per-minute meter running.

  • Audio and video: mp3, m4a, wav, aiff, ogg, opus, mp4, mov and the rest macOS already opens
  • A YouTube link for material you own or have permission to use, or a watched folder that queues new files by itself
  • Export as TXT, SRT or VTT with timestamps — or send a file in from Finder’s “Open With”
Kekoso’s Transcribe screen: a drop zone for audio and video files, a field for a YouTube link and a watched folder

Call recording

Who said what, because it was never one track

Turn recording on before a call and Kekoso captures two streams: your microphone and everything the other side sends to your speakers. Both are transcribed separately and merged back into one timeline, so “You” and “Others” is a fact about the recording, not a guess made afterwards.

  • No bot joins the call, and nobody has to approve one
  • Echo from your speakers is subtracted, so your side is not doubled
  • The transcript exports with speaker labels — hand it to your agent for the summary
Kekoso’s Recordings screen with a finished call recording ready to transcribe

Agents

Give Claude or Codex a pair of ears

Kekoso runs an MCP server on a local socket. One click installs it into Claude Code, Claude Desktop or Codex / ChatGPT desktop, and transcription becomes something your agent can just do — no API key to leak, no per-minute bill, no audio leaving the machine.

  • Twelve tools: transcribe a file or a YouTube link, read the history, read and edit the vocabulary, list, install and remove models
  • Deleting anything needs your click in Kekoso — the agent can ask, it cannot remove
  • Every call is written to a journal you can read, and the whole server has one off switch
Kekoso’s MCP screen: the tools toggle and Claude Code, Claude Desktop and Codex listed as connected clients

Vocabulary

Teach it the words only you use

Client names, drug names, internal acronyms, the library nobody spells right. Add a word once and every transcript after that writes it your way — dictation, files, calls and the agent’s calls alike.

  • Two kinds of rule: fix the spelling of a word it already hears, or swap one word for another
  • Import a list you already keep — CSV, encoding and separator detected for you
  • Applies to every language, not only English
Kekoso’s Vocabulary screen showing spelling fixes and word replacements, openai to OpenAI and git hub to GitHub

Models

Pick the trade-off yourself

Five models from three families, downloaded on demand and kept outside the app bundle. Each card shows measured accuracy and speed on the same benchmark, so “fast” and “accurate” are numbers you compare, not adjectives we chose.

  • Parakeet for speed, Whisper for reach — 99 languages, SenseVoice for CJK
  • Word error rate and real-time factor on every card, with the caveats spelled out
  • Set the recognition language per model, or leave it on Auto
Kekoso’s Models library with Parakeet and Whisper models, each showing word error rate and speed

Make it yours

Set it up once, the way you work.

A dictation tool is on your keyboard all day. Every part of that — the key you press, the microphone it listens to, whether it makes a sound, how long it remembers — is yours to decide.

  • Five assignable shortcuts: start/stop, cancel, push-to-talk, a mouse button, and call recording
  • Pick the microphone yourself, and get told when macOS quietly switches it mid-call
  • Recording panel in full, mini, or hidden entirely — plus light, dark or system theme
  • Sound cues you can turn off, launch at login, and a retention limit that clears old transcripts for you
Kekoso’s Configuration screen: microphone input, light and dark theme, recording window styles and keyboard shortcuts
Kekoso’s Home screen: words dictated, number of dictations, words per minute and a bar chart of the week

Your week, counted

How much you actually talk

Words dictated, times you spoke and your speaking rate, for today, this week, this month or all of it — with a bar chart of the days you kept up and the days you did not.

Kekoso’s History screen with a saved transcript, the app it was pasted into and copy and delete buttons

Nothing lost

Every transcript, still on your Mac

Each dictation is saved with the app it went into — including anything that failed to paste, so your words survive a wrong window. Copy an entry again, delete one, or let retention clear them on a schedule.

Details

The small things you notice on day two.

None of these sell an app. All of them decide whether you still have it installed a month later.

Works in every app

Mail, Slack, Notion, Cursor, the browser. The text lands in whatever field your cursor was already in, so there is no per-app integration to wait for.

Push-to-talk or toggle

Hold the key for a quick sentence, tap it for a long one. Five shortcuts in total, including a mouse button and one for call recording.

Warm model, no cold start

The model stays loaded, so recognition begins the instant you press the key rather than a second and a half later.

Honest microphone handling

Pick the input you actually want, and get warned when macOS has quietly switched you to the wrong one mid-call.

It knows when macOS blocks it

In a password field the system blocks every input tool, correctly. Kekoso tells you it is paused instead of silently dropping your words.

No words invented in the silence

Speech models famously write subtitles credits over a quiet passage. Lines with no sound behind them are struck out before you ever see them.

Ninety-nine languages

Whichever the model supports — up to 99 with Whisper. Set one per model, or leave it on Auto and let it decide.

Updates that check themselves

Signed updates over Sparkle, checked automatically or on demand from the menu bar, and switched off entirely if you prefer.

Privacy

Take the unplug test.

Don't trust a privacy page — including this one. Turn off Wi-Fi, pull the ethernet cable, and dictate a paragraph. Kekoso keeps working, at full quality, because the model was never anywhere else. Try the same test on anything else you're considering.

  • No account. Nothing to sign up for, nothing to log in to.
  • No audio or transcript leaves the machine — not for recognition, not for analytics.
  • The recognition pipeline makes no network calls at all. Not even optional ones.
  • Files, calls, dictation and the MCP server all use the same local models.
  • History is a plain text file in your Application Support folder. Open it, back it up, delete it.
  • The MCP server listens on a local socket, is off until you turn it on, and never removes anything without your click.

The one exception, stated plainly

Pasting a YouTube link means Kekoso fetches that video from YouTube, so that one request goes to Google. Nothing of yours is sent: no audio of yours, no transcript, no identifier — and once the file is on your Mac, transcription is as local as everything else. Whether you may take a copy of a particular video is between you and YouTube’s terms; the feature is meant for material you own or have permission to use.

verify.sh
# Every open network connection Kekoso holds
$ lsof -i -a -p $(pgrep -x Kekoso)

(no output)

# Airplane mode on. Dictate anyway.
$ networksetup -setairportpower en0 off
Wi-Fi is off. Recognition still runs on-device.

Point Little Snitch or Lulu at it if you prefer a GUI. You will find the same nothing.

Languages

Ninety-nine languages, and the words no model knows.

Recognition runs in up to ninety-nine of them, depending on the model you pick. But the words that decide whether a transcript is usable are rarely ordinary ones: they are product names, acronyms and the names of the people you work with — and every model gets those wrong in the same predictable ways.

Set a fixed language if you always dictate in one, or let it detect — with detection switched off when you need it, because on short phrases every detector guesses wrong.

What you said

Push the OAuth fix to GitHub and tell Renée on Slack.

Two product names, an acronym and a colleague whose name carries an accent.

What you usually get

Push the oauth fix to git hub and tell Renee on slack.

Lower-cased, split in half, stripped of its accent. Every time, in every transcript.

What Kekoso writes

Push the OAuth fix to GitHub and tell Renée on Slack.

Add each one to your vocabulary once — every transcript after that spells it your way.

Interface

The app itself speaks ten of them.

Kekoso follows macOS by default, or you pick a language in settings and only Kekoso changes. Recognition languages are chosen separately, per model — the two never fight over each other.

  • English
  • Deutsch
  • Español
  • Français
  • Italiano
  • Português
  • Русский
  • 简体中文
  • 日本語
  • 한국어

The maths

What this normally costs, every month, forever.

Speech-to-text got split into four products, each with its own subscription, each uploading your audio to a different company.

The usual stack

Dictation app
$8–15/mo
Meeting notetaker
$10–30/mo
Transcription service for files
$10–25/mo
Speech-to-text API for your agents
$0.006/min

$336–840a year, indefinitely

Rough figures from typical plans, not a specific competitor. Add up your own — the point stands either way.

Kekoso

$29once

  • All four jobs, one app.
  • No per-minute billing, however much you transcribe.
  • No bill arrives if you stop using it for three months.
  • Nothing to cancel, because there is no account.

Pays for itself somewhere in the first month.

Pricing

Buy it once. It's a tool, not a tenancy.

Four jobs that usually arrive as four subscriptions, running on hardware you already paid for.

$29once

7-day free trial. No card to start.

  • Dictation, files and links, call recording and the MCP server — all of it
  • No per-minute billing, no monthly cap, no usage dashboard
  • Use it on up to 3 of your own Macs
  • Free updates through the whole 1.x line
  • Works offline indefinitely; no license check to phone home weekly
  • 30-day refund, no argument, no questionnaire

macOS 14.4 or later · Apple silicon · payment handled by Lemon Squeezy

Questions

The things people ask before buying.

Does it really work with no internet at all?
Yes. After the models are downloaded on first launch, you can stay offline permanently — dictation, files, calls and the MCP server all run locally. The one thing that needs the network is fetching a YouTube link, because the video has to come from YouTube.
What can it transcribe besides my own voice?
Any audio or video file macOS can open — MP3, WAV, M4A, MOV, MP4, voice memos, screen recordings — plus YouTube links to material you own or have permission to use. Length is limited by your patience and disk space, not by a plan tier, because nothing is metered.Transcribing video to text, format by format
Can it transcribe a voice recording from my phone?
Yes. AirDrop the voice memo across, or drop the file in from anywhere in Finder — m4a, mp3, wav, whatever your recorder produced. Voice recording transcription runs the same way everything else does here: locally, with timestamps, and without the file leaving the machine. A lecture you recorded, an interview, a note you left yourself in the car — same path, no upload.Transcribe a voice recording on your Mac
How does the MCP server work?
Kekoso runs a local MCP server on your machine. Add it to the config of Claude Code or any other MCP client, and transcription becomes a tool your agent can call. There is no API key, no per-minute billing and no upload — the agent gets the same models the app uses.
Are my recordings stored anywhere?
Dictation audio is held in memory while it is recognised and discarded afterwards. Files you transcribe stay wherever you put them; we never copy them anywhere. Transcripts go into a plain local file you can read, copy from and delete — entry by entry or all at once, and automatically once they pass the retention limit you set.
Is on-device recognition actually as good as the cloud?
For dictation and clean recordings, yes — modern on-device models are at parity, and Kekoso keeps the model warm so it usually beats a round trip to a server. In genuinely hard audio, a large cloud model can still edge ahead. We would rather tell you that than pretend otherwise.Voice to text on a Mac: four ways, compared
Which models do I get, and how big are they?
Five, downloaded on demand and removable at any time: two Parakeet models (227 MB and 507 MB, fast, English or 25 languages), Whisper small and large-v3-turbo (216 MB and 646 MB, up to 99 languages) and SenseVoice for Chinese, Cantonese, Japanese and Korean. Every card shows its measured word error rate and speed, so you choose the trade-off instead of trusting an adjective.
In which languages does the app itself run?
English, German, Spanish, French, Italian, Portuguese, Russian, Simplified Chinese, Japanese and Korean. It follows macOS unless you pick one in settings. Recognition languages are set separately, per model.
Why does it need Accessibility permission?
That is the macOS API that lets an app place text into the field you are typing in. Without it, Kekoso could recognise your speech but not deliver it anywhere. It is requested only at the point where insertion first matters, not at the start of onboarding.
What about password fields?
macOS deliberately blocks all input tools while a secure field is focused, and that is correct behaviour. Kekoso detects secure input and tells you why it is paused instead of silently dropping your words.
Do I need consent to record a meeting?
Often yes, and that is on you — the rules differ by country and by state. Kekoso records only from your side, shows a visible indicator while it does, and never joins the call as a participant. What it removes is the bot in the room, not your obligation to say you are recording.
Which Macs are supported?
macOS 14.4 or later · Apple silicon. That covers every MacBook Air and MacBook Pro with Apple silicon, along with the Mac mini, Mac Studio, iMac and Mac Pro. On-device recognition leans on the Neural Engine, which is why Intel Macs are not supported.
Is this a subscription?
No. $29 once, 3 of your own machines, free updates through the 1.x line. Major versions every year or two may be a paid upgrade, at a discount if you already own the previous one.

Support

Write to us. A person answers.

Every email goes to the person who builds Kekoso. Something broken, something missing, a word your model keeps mangling — all of it is worth sending.

[email protected]

A bug we do not hear about is a bug we cannot fix

Kekoso sends no crash reports and no usage data — that is the point of it. The trade-off is that we only learn something is broken when you tell us. The Support screen in the app writes the version, your macOS and the active model into the email for you.

Ideas get read, and some get built

The roadmap is not a closed document. If the thing you need is one setting away, say so — a good part of what is in the app now started as somebody’s email.

Try it on your own words.

7 days, every feature unlocked, no card and no account. If it does not earn its place in your week, delete it and nothing follows you.

macOS 14.4 or later · Apple silicon