Moonshot AI released Kimi K3 on July 16 and put the weights on Hugging Face eleven days later. It is a 2.8-trillion-parameter model with 104 billion active per token, a context window of one million, and native vision. You can call it through Moonshot’s API, or download the weights and find hardware for a model this size.
The README that ships with them carries a comparison table: K3 against Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5 and GLM-5.2, every model at its maximum setting, every score Moonshot’s own. Forty-five benchmarks sit in four sections — reasoning and knowledge, coding, agentic work, vision — and forty of them report a score for all three of K3, Fable 5 and Sol.
Benchmark scores
| Benchmark | Section | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| HLE-Full | Reasoning | 43.5 | 53.3 | 44.5 |
| AA-LCR | Reasoning | 74.7 | 70.0 | 73.7 |
| FrontierSWE | Coding | 81.2 | 86.6 | 71.3 |
| SWE-Marathon | Coding | 42.0 | 35.0 | 39.0 |
| MCPMark-Verified | Agentic | 94.5 | 87.4 | 92.9 |
| OSWorld 2.0 | Agentic | 58.3 | 66.1 | 62.6 |
| OmniDocBench | Vision | 91.1 | 89.8 | 85.8 |
| BabyVision w/ python | Vision | 85.7 | 90.5 | 88.9 |
Each of the three finishes first somewhere, and by very different amounts. Fable 5 takes HLE-Full by nearly ten points and BabyVision by close to five. Sol takes CritPt at 32.3, where K3 scores 23.4. K3 takes MCPMark-Verified at 94.5 and SWE-Marathon at 42.0. FrontierSWE spreads the three over fifteen points, from Fable 5 at 86.6 down to Sol at 71.3.
Other rows barely move at all. SpreadsheetBench 2 has K3 at 34.8 and Fable 5 at 34.7. OfficeQA Pro has K3 at 63.3 and Sol at 63.2. Terminal-Bench 2.1 puts all three within 0.8 of each other, at 88.3, 88.0 and 88.8.
Counted as wins and losses, the answer moves with the comparator. Against Sol, K3 is ahead on 27 of the forty comparable rows. Against Fable 5 it is ahead on 15, behind on 24, and level on one — ZeroBench (pass@5), where both score 23.0. Section by section against Fable 5, that runs two-two in reasoning, three-six in coding, seven-twelve in agentic and three-four in vision, with the tie falling in vision.
The size of a lead moves with the comparator too. MCPMark-Verified gives K3 7.1 points on Fable 5 and 1.6 on Sol. BrowseComp gives it 3.2 and 0.8. AA-LCR gives it 4.7 and 1.0, and AA-LCR is the only row in the reasoning section where K3 finishes above Sol. SWE-Marathon puts K3 seven points above Fable 5’s 35.0, with Claude Opus 4.8, also in the table, at 40.0.
Twenty-two of the forty-five rows are agentic work, so most of any win-and-loss count is decided there: browser search, MCP and tool calls, OS control, spreadsheets and office documents, and a run of domain agents — Harvey Lab-AA, CorpFin v2, Finance Agent v2, Legal Research Bench. Two of those rows report Elo ratings instead of percentages. The 61 points between K3’s 1686 and Fable 5’s 1747 on GDPval-AA v2 are not the same unit as the 61 points that would separate two percentage scores, and nothing in the table converts between them.
Price and speed
K3 lists at $3 per million input tokens and $15 per million output, with cached input at $0.30. Fable 5 is $10 and $50 — the same rate Anthropic kept for Pro and Team Standard in July while folding Fable 5 back into the top plans. Sol is $5 and $30. Against those two, K3 is the cheap one.
A comment on the Hacker News launch thread reads:
pricing is $3/$15 for 1M tokens (cache $0.3), which is extremely high for a Chinese open-weight model
The thread under it drifted into whether per-token prices can be compared at all. Tokenizers differ, so the same page of text becomes a different number of tokens depending on which model is counting it, and more than one person wanted a price per page or per byte instead.
Artificial Analysis measures K3’s output at 34.9 tokens per second, 57th among the 101 models on its list, running on hosted infrastructure. One Hacker News post is titled Run Kimi K3 using 29 GB of RAM at 0.50 tok/s — thirty tokens a minute, which is what the downloadable weights come to on a machine that size.
The model is under three weeks old and a version bump can move any of these numbers. If you have run K3 next to Fable 5 or Sol on the same job, which rows held up? And if something here is wrong, tell me and I will fix it.

