The Kimi K3 language model overtook Claude and ChatGPT: what makes it special and what it costs

5 min read

On 16 July 2026 Moonshot AI released Kimi K3, and within less than a day it climbed to first place in Arena's frontend coding board, above Claude Fable 5, Anthropic's strongest model.

This isn't another routine release. It's the largest open model ever released, and it arrives at a significantly lower price than the strongest models from Claude and ChatGPT. Anyone following the field immediately understands what that means: the race hasn't paused for a moment.

In this article we'll explain what Kimi K3 is, show the three posts on X that lit up the conversation, compare prices in plain terms, and talk about the thing that really matters: why, in a world with this much competition, you shouldn't tie yourself to a single language model.

A glowing glass podium: the Kimi logo in first place, the Claude logo second and the OpenAI logo third
The podium order according to the Artificial Analysis index: Kimi K3 at 57, Claude Opus 4.8 at 56, and GPT-5.6 Terra at 55.

What is Kimi K3?

The Chinese company Moonshot AI released the model on 16 July 2026. It's an MoE (mixture of experts) model with 2.8 trillion parameters, but it activates only 16 experts out of 896 for each token. In plain terms: the model is enormous, but at any given moment only a small part of it is working, which keeps it efficient and fast.

Its context window is 1,048,576 tokens, roughly a million, meaning you can feed it an entire code repository or hundreds of pages of documents in a single request. Most importantly: the company promised to release the model's weights by 27 July 2026, which would make it the largest open model in existence. Open means anyone can download it and run it on their own servers.

Why is everyone talking about this now?

Arena, which ranks models by comparing real outputs, announced that Kimi K3 had taken first place in the frontend coding board with 1679 points, overtaking Claude Fable 5. That's a 17-place jump from the previous version, Kimi k2.6, which sat at number 18.

In the post above, Arena breaks it down: the model took first place in 6 of 7 frontend categories, among them branding and marketing, design from a reference, data and analytics, consumer products, simulations and content-creation tools. The only category where it stayed second is gaming, where Claude Fable 5 still leads.

Does it really compete with the strongest models?

The numbers are impressive. On GPQA Diamond, which tests doctorate-level science questions, the model scored 93.5%, the highest published result for an open model at the time. On BrowseComp, which tests search and web navigation ability, it scored 91.2%, the best score published on that tracker at release. On Terminal-Bench 2.1 it scored 88.3%, and on MCP Atlas it reached 84.2%.

The independent Artificial Analysis Intelligence Index gives Kimi K3 a score of 57, above Claude Opus 4.8 (around 56), above GPT-5.6 Terra (55), and at the same level as Google's Gemini 3.1 Pro.

It's important to say this honestly. According to the benchmarks the company itself published, the model beats Claude Opus 4.8 and GPT-5.5 in most cases, but loses to Claude Fable 5 and GPT-5.6 Sol, the two strongest models on the market right now. In other words it isn't "the smartest model in the world", it's a model that joins the front row at a far lower price.

In this post, viewed more than 287,000 times, the user puts Kimi K3 against GPT-5.6 Sol in a comparison video. His claim: both models reach the same result, but the route differs, and in his view Kimi is more creative. He adds that the difference in design taste is so pronounced that if you swapped the name and said it was Fable 5, people would believe it.

This post cites a test in which Kimi K3 and Claude Opus 4.8 were given the same task: build a 3D scene of an armoury. The video shows the gaps in level of detail, in textures and in lighting between the two models. That's exactly the kind of test you can't learn from a score table.

What does it cost against the competition?

Here's the interesting part. The prices below are per million tokens (units of text), input (what you send the model) and output (what the model returns):

ModelInputOutput
Kimi K3$3$15
Claude Sonnet 5$3$15
GPT-5.6 Terra$2.50$15
Claude Opus 4.8$5$25
GPT-5.6 Sol$5$30
Claude Fable 5$10$50

In plain terms: Kimi K3 costs exactly what Claude Sonnet 5 does, the cheapest of the strong models. Against Claude Fable 5, which it overtook in the frontend board, you pay under a third on output. Against GPT-5.6 Sol, OpenAI's flagship, it's half the price on output.

There's another important detail in the official pricing: if you repeatedly send the same context (the same long document, say), the input price drops to $0.30 per million tokens, ten times less. For a business running thousands of requests against the same knowledge base, that difference adds up to real money.

So why shouldn't you tie yourself to one model?

Look at the timeline. Anthropic released Claude Sonnet 5 on 30 June. OpenAI released GPT-5.6 on 9 July. Moonshot released Kimi K3 on 16 July. Three frontier releases within less than three weeks. And the model that sat at number 18 jumped to first place.

That is exactly why you shouldn't build an entire business around one model. Anyone who fell in love with a particular model and spent months tuning to it alone discovers every few weeks that their competitor is working with a better or cheaper tool. Tomorrow another model launches, and it happens again.

So what do you actually do? First, choose tools that let you swap models without rewriting everything. Second, test each model on your real tasks rather than on score tables, because as you saw in the videos above, the real difference only shows up in the work. Third, match the model tier to the task tier: a simple task doesn't need the most expensive model, and a critical task doesn't need the cheapest. And yes, it's entirely fine to run two models in parallel and compare results.

Frequently asked questions

Is Kimi K3 really stronger than Claude and ChatGPT?
Answer: it depends on the task. In Arena's frontend coding board it sits in first place, having overtaken Claude Fable 5, and on the Artificial Analysis index it scores 57 against Claude Opus 4.8's 56. On the other hand, according to Moonshot's own benchmarks, it still loses to Claude Fable 5 and GPT-5.6 Sol on other tests.

Can you download Kimi K3 and run it yourself?
Answer: Moonshot promised to release the full weights by 27 July 2026. Until then it's available only through their site and their API.

What does Kimi K3 cost against the other models?
Answer: it costs $3 per million input tokens and $15 per million output tokens, exactly like Claude Sonnet 5. That's under a third of Claude Fable 5's price, and about half the price of GPT-5.6 Sol on output.

What actually is an open model?
Answer: an open model is one whose weights are available to download, so you can run it on private servers without sending information to the company that built it. That's critical for organizations with sensitive data.

How do you avoid getting stuck with one model?
Answer: choose tools that support swapping models, test every new model on your real tasks, and match the model tier to the task tier instead of running everything on the same model.

Good luck,
TodoAI

Avihai and Eitan Elnekave, hands-on help adopting AI and automation in your business

Want hands-on help adopting AI in your business?

At TodoAI we help organizations and business owners pick the right tool, embed it in the team, and win back hours every week.

Message us on WhatsApp
TodoAI's WhatsApp group, join the community

Our busy group, tips and updates

Join the group and be among the first to hear about the hottest AI tools, tips and news, straight to WhatsApp and free.

Join the community
MakePodcast, creating professional podcast videos with AI

Create professional podcast videos with AI

The MakePodcast platform produces short podcast videos for your brand, with your own host and your own script. No camera, no studio, no editor.

Get early access

Enjoyed it? Share it

Help more people find this article

Topics in this article

9 tags
  • Kimi K3
  • Moonshot AI
  • open language model
  • model comparison
  • Claude
  • ChatGPT
  • model pricing
  • artificial intelligence
  • TodoAI