October 2024. My feed was wall-to-wall ChatGPT screenshots, everyone posting their cleverest prompt and the “great” answer it got back. I’d already jumped to Claude Pro before any of that started. It wasn’t a dramatic realization. It just felt better. Not because it was the cooler choice. Because when I handed off a task, it wandered off less, made up fewer details just to keep talking, and answered the actual question instead of writing something that read well but missed the point.
Then Cursor showed up. Then Augment, then Windsurf. All of them are just a shell around the same decision: which model is underneath, because that’s what decides whether you wake up to a finished product or a mess.
I picked up a few part-time clients around then too. Not because I’d suddenly gotten good at “AI.” Because I was willing to pay for the model to run long enough that a client saw a real result instead of a five-minute demo.
While most people were still deciding whether Claude Pro was worth the money, I was already on Max, $200 a month. A client of mine ran Max at $100. Same scene every time: you’re mid-task, it hits the ceiling, and stops. There’s no clever third option. You either walk away and wait for the reset, or you pay more. Sounds like a punchline. It’s actually just the real cost of treating a model like staff instead of a toy.
The Singapore client
This guy was based in Singapore, the type who always shows up with a new tool. For a while it was Gemini. “Look, it says something smart here.” I never argued with smart. Sure, it’s smart. I argued with trustworthy. Gemini talks fluently, answers fast, sounds reasonable, and hallucinates a lot. I don’t build on something that talks a good game but won’t own the outcome of a single line of code. Claude is the only one I trust with that.
So he bought Claude at $100 and sat down for one seven-hour session, straight through, that same night. Woke up to a working site. Not a placeholder. Something he could actually point people to. He was thrilled. Told me right away that getting this before would’ve meant hiring an outside team for a year, and even then they’d probably still get it wrong.
That line stuck with me. Not because AI is replacing teams. It isn’t, not really. It’s because the gap between “someone is working on it” and “it’s actually what you meant” has always cost something. Good teams still drift from the brief. Long processes drift with whoever’s hands are on the keyboard that week. One seven-hour night with a model you trust closed that gap. Nothing magic about it. Just fewer rounds of explaining yourself.
A while later he came back with Codex. “This one’s good too.” I said no. I only trust Claude. I’d tried Codex myself, not because it’s bad, but because I wasn’t ready to hand it something big and walk away from the screen.
A week after that, he was back. “Yeah. You were right.” No leaderboard involved. Just a guy who used both for a week and figured out on his own which one he could actually rely on versus which one just sounded good in the room.
People who live in code still trust Claude. They still use other things
The engineers I know, the ones actually shipping under deadline for difficult clients, almost all of them trust Claude for the load-bearing work. That doesn’t stop them from opening five other tabs. Gemini for something that needs to be written fast. Another model for brainstorming. Codex for one specific chunk of code it happens to be good at. None of that is a contradiction. Trust was never about using only one thing. Trust is about who you call when the work has to land and you don’t have time to watch every step.
Think of it like commanding a hundred soldiers. You don’t ask who gives the best answer in the briefing room. You ask which one, when you’re not there, doesn’t go off-script, doesn’t quietly rewrite the plan, doesn’t half-finish the job and then talk you through the “process.” That one gets your trust. Doesn’t matter if it’s weaker somewhere else. It gets the job done when nobody’s watching.
There are a few types you learn to spot fast, in models and in people.
There’s the hype-chaser. A new favorite every month, always a fresh screenshot, rarely a product that’s still standing a week later.
There’s the lopsided one. Beautiful prose, falls apart the moment logic gets involved. A gorgeous code snippet next to an architecture that collapses under its own weight. Wins the demo, loses the maintenance phase.
There’s the one that can genuinely do the work but only with a leash. A workflow, explicit rules, a checklist for every step. Skip one line and it drifts. It’s usable. It’s also exhausting, and you didn’t sign up to spend your day writing procedure documents for one task.
And then there’s the other kind. You say one sentence, “handle this for me,” and it comes back done, no loose threads. It doesn’t need you standing over it. Doesn’t need the same context repeated a fourth time. Doesn’t leave you bracing for a morning where it quietly changed direction because of one stray line in an old document somewhere.
Every model on the market right now falls into one of these buckets. No benchmark draws that map. Only using the thing does.
The doctrine isn’t a leaderboard
People love asking which model is “stronger.” It’s a convenient question. It’s also the wrong one. Stronger on which benchmark? Stronger for whose skill level? Stronger when you write a careful, engineer-grade prompt, or when you just toss out a rough idea the way an actual client would?
Better question: which model, in your hands, with your habits, produces something you can ship without babysitting it the whole way.
Talking well isn’t enough. One brilliant answer isn’t enough. Being the model of the month definitely isn’t enough. Trust means you hand off the work and go to sleep, and the morning doesn’t involve cleanup. It means your client never hears you explain why this draft is “basically there.” It means you’re comfortable letting it run for seven hours unsupervised.
Claude isn’t perfect. Nothing is. It’s slow on some things. Stubborn on others. There are tasks another model just handles better. Anyone doing this for real knows that, which is exactly why they don’t treat it like a religion. They trust it, and trusting it doesn’t stop them from keeping specialists around for specific jobs, the same way a commander keeps units for specific terrain.
The one place I don’t budge is wherever the work actually has to ship. I’m not swapping models there because someone’s timeline is flattering a different one this week. A client of mine wanted to switch once. It took exactly one week of using the alternative for him to come back on his own. Not stubbornness on my part. The time bill just doesn’t lie.
Max at $200 sounds steep if you’re comparing it to a chat subscription. It’s cheap if you’re comparing it to a year of outside contractors, misread briefs, and endless rework. Pro feels expensive to someone who still treats AI as a hobby. Max is a bargain to someone who treats it as a shift worked.
Run out of quota, you stop or you pay for more. That’s part of the deal too. No model hands you infinite output for free. What you’re buying is a shift’s worth of work. When the shift ends, you either extend it or call it a night. That’s how staff works. It’s not how a free app is supposed to work, and pretending otherwise is how people burn out on the wrong tool.
There are a hundred models out there. Only a handful earn actual trust, and the ones that do don’t necessarily win every side-by-side comparison people post online. They win the one comparison that matters if you’re doing real work for real clients: the job finishes, the intent doesn’t drift, and you’re not the one staying up to watch it.
Pick the one that gets the job done. Leave the rankings to people with time to argue about them.