Kimi K3: second only to Fable 5 on AA-Briefcase
By Dillip Chowdary • Jul 22, 2026 • Source: HN Claude/Codex/Fable
I'll pull the Artificial Analysis article and the HN thread so the paragraphs stick to reported names, rankings, and mechanics—no invented figures.**Kimi K3**, a **2.8T-parameter** model from **Moonshot AI**, ranks second only to **Claude Fable 5** on **AA-Briefcase**, Artificial Analysis’s agentic knowledge-work benchmark. It posts an **AA-Briefcase Elo of 1543**, a **+727** jump from **Kimi K2.6** (816) and just behind Fable 5 at **1574**. On the broader **Artificial Analysis Intelligence Index**, K3 scores **57**, in line with **Opus 4.8** and **GPT-5.5**. Artificial Analysis published the write-up on **July 21, 2026**, after K3’s release the prior week.
**AA-Briefcase** scores models on a private set of realistic tasks over thousands of complex input files. Deliverables include spreadsheets, presentations, and UI mock-ups. The single Elo blends correctness, analytical quality, and presentation quality. K3’s **rubric pass rate is 51%**, second to Fable 5 at **56%** and ahead of **Claude Sonnet 5 (max, 42.3%)** and **GPT-5.6 Sol (max, 41.8%)**. Its **analytical quality Elo is 1754**, slightly above Fable 5’s **1744**. **Presentation Elo is 1471**, under GPT-5.6 Sol (max, **1660**) and Opus 4.8 (max, **1492**). Cost and latency are steep: about **$10.57 per task**, **83 turns**, **120k output tokens**, and **56.4 minutes** average time per task. Pricing is **$3/$15 per 1M input/output tokens**, with a **90%** cache discount.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For builders shipping agent workflows that produce real artifacts, the split matters. K3 is strong where analytical correctness drives the outcome and weaker where polished slides or mock-ups dominate. High turn count and output volume also raise latency and bill size. That trades off against models that finish fewer turns and spend less per task, even if their Elo is lower.
On AA-Briefcase Elo, K3 leads **GPT-5.6 Sol (max, 1501)**, **Claude Sonnet 5 (max, 1388)**, and **Claude Opus 4.8 (max, 1347)**, with only Fable 5 ahead. Versus **K2.6**, output tokens rose from **42k** to **120k** and turns from **54** to **83**, driving roughly a **10x** rise in cost per task. Average time is about **2.5x** Fable 5 and about **3.8x** **Grok 4.5 (high)**. K3 also uses more turns than Fable 5 (**67**) and GPT-5.6 Sol max (**50**).
Watch whether Moonshot cuts turn count, output tokens, or first-party API speed enough to shrink the **$10.57** and **56.4-minute** averages without giving up the **1543** Elo. For production, compare K3 on analysis-heavy Briefcase-style jobs against Fable 5 for quality-at-cost and against faster, cheaper stacks when wall-clock and spend matter more than a few Elo points.
Advertisement
🔎 More interesting news
- OpenAI's models broke containment and cyberattacked Hugging Face — what enterprises need…
- Governments, companies, nonprofits should invest in free, open source AI [pdf]
- Grok Build is open source
- Windows 0-day drops the same day Microsoft releases record number of patches
- Today's full Tech Pulse briefing →