Seven Things I Say to My Mac Instead of Typing

I talk to my Mac more than I type to it now. Not in a science fiction way. I hold a key, say a sentence, let go, and the sentence is sitting in Slack before I have picked up my coffee.

The app is called Yubi. I built it in Tokyo this month. It lives in the menu bar, it is 6 MB, and it does two jobs with two keys.

Hold the right Option key and talk: your words are typed wherever the cursor is. Hold the right Command key and talk: she looks at the window in front of you and helps with it. She points at the button you cannot find, or presses it for you.

I used to think one key would be enough and the app could work out what you meant. It can, most of the time. Most of the time is not good enough for something you press fifty times a day, so there are two keys and the app never has to guess.

Why talk at all

A Stanford study compared speech recognition with the phone keyboard on an iPhone. In English, speech came out at 153 words per minute against 52 for the keyboard, 2.93 times faster. In Mandarin, with Pinyin input, it was 123 against 43.

That was a lab, on a phone, with short messages. A good typist on a real keyboard closes part of that gap. Not all of it, and most of what I write in a day is short messages. That is exactly the case the study measured.

The other reason is less measurable. When I type, I edit while I write. When I talk, I say what I mean and fix it afterward, which is faster and usually sounds more like me.

Seven things I actually say to it

These are the seven scenes on the site. You can play every one of them in the browser before you download anything.

1. Running late, in Slack. "Running five minutes late, blame the trains." That lands as a casual one liner without a full stop. Yubi keeps a tone per app, so a chat message stays a chat message and an email gets full sentences. A terminal gets the words exactly as spoken, with no capital letter added.

2. The polite follow up, in Mail. I say something like "uh hi Maya thanks for today um can you send the new estimate by Friday." What goes in is a clean email with the fillers gone. This is the optional cleanup pass. It is off by default, it sends text only and never audio, and every send is listed in the app so you can see exactly what left.

3. Japanese, in Notes. "明日の買い物、卵とねぎと味噌." I live between English and Japanese all day. The dictation handles both, and it runs on the Mac itself.

4. An extremely specific coffee order, in Messages. "Iced oat latte, extra shot, no syrup. The pastry is your call." This one is not clever. It is the kind of message you would never bother to type carefully, and now you do not have to.

5. Where is Dark Mode? Right Command, then "Where do I turn on Dark Mode?" in System Settings. She reads the window, finds Appearance, and points at Dark. She does not change it. You asked where, so she shows you where.

6. Back to today, in Calendar. "Click Today for me." This time you asked her to do it, so she presses it. Small, but it is the difference between an assistant that talks about your screen and one that can touch it.

7. Send, with a yes, in Mail. "Send this for me." She does not just send it. She stops and asks: text leaves this Mac, send it? Yes sends. No keeps it as a draft.

The yes is the whole design

Number seven is the one I care about most.

Anything near a word that means money, publishing or something gone for good needs a yes first. Send, post, buy, pay, delete, book, and the Japanese equivalents like 送信 and 購入. That check lives in the app's own code, not in the model's judgment. It does not matter how confident the model is. If the button says send, you get asked.

And one yes buys exactly one action. Say yes to sending one email and she cannot send three more on the strength of it.

That sounds slow, but most things you ask for are not risky at all: open a menu, find a setting, go back to today. The yes only shows up where you would want a second look anyway.

What she can see

  • Your voice never leaves the Mac. Transcription happens on the machine. Zero seconds of audio are uploaded.
  • She never takes a screenshot and never asks for Screen Recording. To help with a window she reads its accessibility tree, the same structured list of buttons and labels that VoiceOver uses.
  • She refuses to read password managers at all, whatever you ask.
  • No analytics and no trackers in the app.

Reading the accessibility tree instead of pixels has a limit, and I would rather be honest about it. Some apps draw their own canvas and expose almost nothing to that tree. When she cannot read enough of a window, she tells you, instead of guessing at a button that might not be there.

She does not talk back in a robot voice

Yubi has a face and a voice. The face is a manga style character in the corner of the screen. The voice is a set of short recorded Japanese lines, the little "okay" and "done" moments.

What she actually answers shows up as a caption under her portrait. I tried the synthetic system voices. They made a friendly character sound like a train announcement, so they are not in the app, and they will not be until something sounds right.

What it costs

Dictation is free and unlimited on every plan, because it runs on your Mac and costs me nothing.

Screen help is counted in tasks, one task per spoken request. Free gives you 10 a month with no card, just an email link to sign in. Personal is ¥1,480 a month for 60 tasks. Pro is ¥2,980 a month for 150.

It needs macOS 26 on Apple Silicon. It is not on the Mac App Store, because the App Store sandbox does not let an app press buttons in other apps, which is half of what Yubi is for. It is a direct download, signed with my Developer ID and notarized by Apple.

Try it

The site is in English and Japanese, and you can play all seven scenes there first. When you are ready, download Yubi for Mac, hold the right Option key, and say the first thing you were about to type.

Sources

← All posts