I Pay $200 a Month for Claude. Anthropic Just Swapped My Default Model.
On Thursday, while I was out driving, Anthropic released Claude Opus 5 and made it the default model on Claude Max. Which means the $200 a month I pay now runs on a different brain than it did on Wednesday. Nobody asked me. I'm not exactly complaining. But when your daily driver gets swapped overnight, you want to know what changed.
So I went and read everything. Here's what actually changed.
What Anthropic shipped
Opus 5 is Anthropic's fourth model in under two months, after Fable 5, Sonnet 5, and Opus 4.8. That cadence is its own story, but I'll get to that.
The pricing stayed put: $5 per million input tokens, $25 per million output, identical to Opus 4.8. That's half the token price of Fable 5, Anthropic's top-tier model. It's live everywhere, including Claude Code, which is where I actually spend my life.
The headline number: Artificial Analysis, the independent benchmarking shop that tested the model before release, puts Opus 5 at 61 on its Intelligence Index. That's effectively tied with Fable 5 at 60, and ahead of GPT-5.6 Sol at 59 and Kimi K3 at 57. On Frontier-Bench, an agentic coding benchmark, it scored 43.3% against Opus 4.8's 21.1, per The Decoder's reporting. More than double in one generation. On ARC-AGI 3, the novel problem-solving test, it hit 30.2% where the next-best model managed 7.8.
I want to be careful here, because I just wrote a post about Kimi K3 impressing me, and I meant it. One index, one week, is not a verdict on anything. Opus 5 has not been out long enough for anyone to honestly say it's better at everything. Anthropic itself says Fable 5 is still the right pick for the longest, most complex autonomous tasks. But the gap story cuts both ways: K3 sits at 57 on that index and Opus 5 just posted 61. Whatever gap closing I was feeling two weeks ago, Anthropic didn't get the memo.
The number nobody is leading with
Here's the part that matters to me, and it's not the benchmark crown.
Artificial Analysis found that Opus 5 answers more often when it's uncertain. Its hallucination rate jumped 14 points to 50%. Read that again. On their factual reliability test, half the time it doesn't know something, it guesses anyway.
In a chat window, that's annoying. You ask a question, you get a confident wrong answer, you catch it, you move on. But I don't use Claude in a chat window most of the time. I run agents. My TrustHub publishing pipeline, my research jobs, my morning radar, all of it runs unsupervised while I'm doing other things. The entire value of an agent is that it does the right thing when nobody is watching. And "the right thing" very often is saying "I don't know, flagging this for the human" instead of inventing an answer and acting on it.
A model that's smarter on average but more willing to bluff is a real tradeoff for that kind of work. This isn't a dealbreaker. Accuracy improved too, up 7 points on the same test. But I'll be watching how Opus 5 behaves inside my agent runs a lot more closely than I'll be watching its leaderboard position.
The cost story is the quiet win
The thing that might actually matter most for regular people building with this stuff: Artificial Analysis measured Opus 5 at about 26% lower cost per task than Fable 5, at basically the same intelligence. Anthropic's own benchmark has it within half a percent of Fable 5's peak coding score at half the cost per task.
If you pay for AI by the token, and I do for some of my setups, that's the difference between an agent you run on everything and an agent you only run when it matters. Cost decides what gets automated. Not benchmarks.
A strange week to ship a model
I can't separate this launch from the week it landed in. Two days ago I wrote about the White House claiming the model I run every day was stolen from Claude. Last week it was OpenAI's agent breaking out of its sandbox and ending up running code on Hugging Face's servers. And right in the middle of all that, Anthropic ships what it calls its most aligned model ever, with safety safeguards that it says intervene about 85% less often than Fable 5's did.
Make of that what you will. I'm just noting the timing is not subtle.
What I'm actually watching
I haven't put Opus 5 through my real work yet. It became my default two days ago, and most of that time I was driving passengers around LA, not stress-testing a frontier model. What I know from living on Opus 4.7 and 4.8 is that the distance between "tops a benchmark" and "my pipeline still works on Monday" is real, and you only learn it by using the thing.
Three labs are now sitting within two points of each other at the top, shipping new models every few weeks. Two years ago a frontier model release was an event. Now my default model changes while I'm asleep and I find out from a benchmarking site.
That might be the most 2026 sentence I've ever written. But here we are, and honestly, the new brain seems pretty good. I'll let you know if it hallucinates its way through my cron jobs.
0 Comments
No comments yet. Be the first to reply!