Hubs AI and Technology Post
Join TrustHub to participate — every member is ID-verified
Sign Up Free
0

Kimi Let Me Down in March. Two Hours With K3 Changed My Mind.

Back in March I wrote a post called "I Tried Every Chinese AI Model. Here's My Honest Take as a Non-Engineer." If you missed it, the short version is: I was not impressed. I ran DeepSeek and Kimi as the reasoning engine inside my agent setup — the thing that's supposed to take a task, figure out the steps, and execute them. And they fell apart. Give them a clear, specific instruction and they were fine. The moment anything was ambiguous, they'd loop on the same step, make weird decisions, or just give up. I said all of that publicly, and I meant it.

So when Moonshot dropped Kimi K3 this week, I rolled my eyes a little. Another Chinese model launch, another round of hype. I've been watching this cycle for a year.

Then I spent two hours actually using it. I owe Kimi an apology.

What actually launched

Kimi K3 is a 2.8 trillion-parameter model, which makes it the largest open-weight model ever released. For comparison, DeepSeek's V4 Pro — the previous Chinese leader — is 1.6 trillion. K3 has a 1 million-token context window, it's built for long-horizon agentic coding (exactly the use case where K2 broke down on me), and the API is live now. If you want to download the weights and run it yourself, those drop on July 27th.

Moonshot's own blog admits it still trails the most powerful proprietary models overall. Companies always say that part quietly and the benchmark part loudly, so I mostly ignore company blogs. What got my attention is what the independent testers found. Arena.ai, which ranks models based on real people judging real outputs, has K3 first in web interface building — above Claude Fable 5, and 17 spots above Moonshot's own previous model. Vals AI put it second overall, behind Fable 5 but ahead of OpenAI's GPT-5.6 Sol. Artificial Analysis found it comparable to GPT-5.5 and Opus 4.8 on complex multi-step tasks.

I've been burned by cherry-picked Chinese benchmark claims before, so I don't say this lightly: when three independent evaluators all put a model in the same conversation as the best American models, that's not hype anymore. That's a result.

Two hours in

Here's what I can tell you that the benchmarks can't. I've been running K3 inside Hermes, the agent framework I use on my desktop. When I give it a task, I can watch its chain of reasoning in real time — what it's thinking about, how it's breaking the problem down, the actual thought process before it acts.

It thinks for a lot longer than K2 ever did. And the reasoning is night and day from what I was fighting with back in March. The looping, the weird decisions, the giving up — I haven't seen any of it. I've been throwing real work at it, the same kind of vague, figure-it-out-yourself tasks that broke its predecessor, and it's been handling them.

I'm not going to say it's better than Fable. But on the things I actually do every day, it's similar enough that I keep forgetting which model I'm talking to. And that brings me to the part I can't stop thinking about.

A third of the price

K3's API pricing is $3 per million input tokens and $15 per million output tokens — 30 cents for input if your context is cached. Fable 5, per Anthropic's own pricing page, is $10 per million input and $50 per million output. So K3 costs roughly a third of the best proprietary model, while trading punches with it.

I pay $200 a month for Claude and I build with AI agents every single day, so this is not abstract for me. A model this capable at a third of the price changes the math for everyone building on top of this stuff. And it's only going to get cheaper from here. It always does.

The part nobody talks about

Here's the thing that really gets me. Most people are never going to need a model this capable. Most people right now are using AI about as advanced as Google — maybe they have it write a paper or an email. I'd bet 99% of people still don't know what an MCP server is. I'd bet 99% of people don't know these models can reach into an actual physical computer and do real work on it — read files, run commands, build things while you're not watching.

Most people are completely in the dark about what these models can already do, and K3 can do things most people probably won't try even ten years from now. Not because they can't. Because they just don't care about this stuff.

But for those of us who are nerding out about it — the people actually building agents, wiring up tools, watching these models reason through real work — it's insane. A few months ago I wrote about Anthropic accusing Moonshot of distilling Claude. That fight is about to get a lot louder, because the open model giving their flagship a real fight costs a third of the price and is running inside my own setup right now.

To be clear about what I'm claiming and what I'm not: K3 beat Fable on one public leaderboard, in one category. It has not been out long enough for anyone to honestly say it's better at everything, and I'm not saying that. What I am saying is that at a third of the price, it doesn't need to be better at everything. Even close is enough. The value proposition is already insane.

The comfortable story was that Chinese labs trail American ones by about six months. I've repeated that line myself. If the gap isn't already closed, it is definitely not six months anymore. It is much, much, much smaller than that. And almost nobody outside our little corner of the internet has noticed yet.

Two hours. That's all it took to change my mind about a model I publicly trashed. If you're on the fence about trying it, the API costs less than the coffee you'd drink while testing it.

0 Comments

Log in to join this hub and comment.

No comments yet. Be the first to reply!