
You decide before you see the answer
AI models keep getting better. But how often people actually reach for them hasn't kept up. No benchmark shows that gap, because the gap isn't in the answers.
It's in the moment before them. Say you hit a small task - something you could finish in ten minutes. Should you ask the AI? Without noticing, you run the math: explaining the situation will take a few minutes. Switching over there will break your focus. And the answer might miss the point anyway, which means you'd check it, fix it, or redo it. Added up, that's more than ten minutes. So you don't ask. You just do it.
That math has three costs in it: explaining it, stopping your work, and getting an answer you can't use. A smarter model only fixes the third one. The first two don't care how good the model is - they're product problems. And your day is mostly made of those ten-minute things.
The question we ask instead
So when we ask ourselves whether Miller is getting better, we don't ask how much it can do. We ask the question you'd ask about a person: is this good to work beside?
A good desk-mate comes down to two things. They know you. And they're there.
It knows you - and knows when it doesn't
Knowing you isn't the same as remembering what you said. We've written about that before. When you ask loosely - "pull this together with that thing from before" - you've left out what, how much, and in what shape. A generic AI fills those gaps with what a generic professional would want, and hands back something well-written that isn't what you wanted. Miller fills them with your defaults: the format you use every week, the project you're actually in.
But that kind of knowing has to be kept honest, in three ways.
It has to be you - not a plausible average of people like you. Average-based answers are right just often enough to be dangerous.
It has to be the you of this week. Understanding a person goes stale: ask for a weekly update and last month's dead project is still in it. That's what stale knowing looks like.
And it has to know where it stops. This one matters most. A wrong answer that sounds confident doesn't just fail to save your ten minutes - it costs you extra, because now you're checking and fixing too. So when Miller isn't sure which document or decision you mean, it asks. An AI that admits uncertainty is saving you time. An AI that guesses fluently is spending it while looking helpful.
It's there - in two ways
The first way is physical: within reach the moment you need it. That's why Miller lives in the notch, which we've also written about. One thing worth adding: what matters isn't speed, it's continuity. A window that opens instantly still breaks what you were doing. The notch doesn't remove milliseconds. It removes the break.
The second way is harder to build, because it isn't about the tool at all. It's about you. The feeling that someone is beside you doesn't come from them helping right now - it comes from being sure they'd answer if you asked. A coworker you ask once a day can still feel beside you all day. That certainty is the real product.
And certainty is built out of something unglamorous: predictability. Three rules.
It answers every time you call. Ten okay answers out of ten build more trust than nine brilliant ones out of ten. People don't remember a tool's best day. They remember the one time it failed them when it mattered.
It's the same every time. If what worked Monday fails on Wednesday, you quietly stop trying by Friday. Consistency beats occasional brilliance.
It never shows up uninvited. We actually built versions of Miller that spoke first. It didn't feel like presence - it felt like one more thing to read. The line isn't whether it moves first, it's whether you invited the move. A recap that arrives Friday at nine because you asked for Fridays at nine is presence. A surprise notification is a chore.
Trust is set by the worst day
Here's the uncomfortable part. Trust doesn't add up evenly. One confidently wrong answer outweighs a long run of good ones - because a good answer saves you ten minutes, but a bad one costs the work plus the checking, and you often find out late. After one of those, the tool becomes "a thing that might work." And a thing that might work is not what you reach for when you're busy.
Which means this kind of trust can't be shipped in a release or claimed in a blog post. It's made of one thing only: days that didn't go wrong. We're early. We haven't banked enough of those days yet, and we know it. Miller is thinnest in your first week - which is exactly the week you decide whether to keep it. That's the hardest problem in front of us, and it's where most of our work goes.
The knock
There is one part of trust you can pull forward, though. Before a tool has earned your certainty, it can at least borrow a gesture your body already trusts.
Think about how you get the attention of someone who isn't looking at you. There are only two ways: say their name, or knock - tap the desk, tap their shoulder, let the sound carry it. Every piece of software ever built went with the first way. We spent our time on the second.
That's where tap-tap came from: double-tap, and Miller looks up. And the three rules above turned into engineering requirements on day one - it has to answer every time you tap, the same way every time, and never when you didn't. How we built that is the next post.
To be clear about what we're not doing: we're not building software that pretends to be a person. We want the human gesture to work, not the human impersonation. Miller isn't meant to become a relationship you open the app for. It's meant to be a presence beside your work.
How we'll know it's working
So our metric isn't how much smarter Miller got this month. It definitely isn't how much time you spend in it - we'd rather that number went down. It's this: did the math you run before asking come out differently, a little more often? Did more of the ten-minute things get handed over instead of quietly done alone?
No benchmark reports that number. That's why we wrote down what we're aiming at - so we can tell whether we're getting closer.



