Home / Blog / Kimi K3: second only to Fable 5 on AA-Briefcase
Tech News

Kimi K3: second only to Fable 5 on AA-Briefcase

By Dillip Chowdary • Jul 22, 2026 • Source: HN Claude/Codex/Fable

I'll pull the Artificial Analysis article and the HN thread so the paragraphs stay grounded in named models, the AA-Briefcase result, and only stated numbers.**Kimi K3**, released last week by **Moonshot AI** (Kimi), posts an **AA-Briefcase Elo of 1543**, second only to **Claude Fable 5** at **1574**. That is a **+727** jump over **Kimi K2.6** (816) and puts the model ahead of **GPT-5.6 Sol (max, 1501)**, **Claude Sonnet 5 (max, 1388)**, and **Claude Opus 4.8 (max, 1347)**. On the broader **Artificial Analysis Intelligence Index**, the **2.8T-parameter** model scores **57**, in line with **Opus 4.8** and **GPT-5.5**.

**AA-Briefcase** is Artificial Analysis’s proprietary agentic knowledge-work benchmark. Models run against a private set of realistic tasks over thousands of complex input files and must produce deliverables such as spreadsheets, presentations, and UI mock-ups. A single Elo combines correctness, analytical quality, and presentation quality. On that breakdown, **Kimi K3** hits a **51% rubric pass rate** (behind Fable 5 at **56%**, ahead of Sonnet 5 max at **42.3%** and GPT-5.6 Sol max at **41.8%**) and an **analytical quality Elo of 1754** (near Fable 5’s **1744**). **Presentation Elo** is weaker at **1471**, under GPT-5.6 Sol max (**1660**) and Opus 4.8 max (**1492**).

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For builders shipping agent workflows, the scoreboard and the cost clock pull in opposite directions. **Kimi K3** averages **$10.57 per task**—among the highest on the bench—and about a **10x** cost jump versus the prior generation path called out by Artificial Analysis. Token pricing is **$3 / $15 per 1M input / output tokens** with a **90% discount on cached tokens**. Latency is heavy: **56.4 minutes** average time per task, **83 turns** per task, and about **120k output tokens** per task, up from **54 turns** and **42k** output tokens on **Kimi K2.6**. That time is roughly **2.5x** Fable 5 and **~3.8x** **Grok 4.5 (high)**; turn count also exceeds Fable 5 (**67**) and GPT-5.6 Sol max (**50**). The model also costs more to run than **Opus 4.8** on this bench.

In competitive terms, **Kimi K3** is the clearest Chinese-lab challenge on an agentic knowledge-work Elo that had been led by Anthropic’s **Fable 5**. It outranks OpenAI’s **GPT-5.6 Sol (max)** and both **Sonnet 5** and **Opus 4.8** on overall AA-Briefcase Elo, while still trailing Fable 5 by **31 Elo** points and losing ground on presentation quality and wall-clock efficiency. The pattern is strong multi-file analysis and rubric pass rate, weaker polish, and a long, token-heavy agent loop on the first-party Kimi API.

Practical takeaway: treat **Kimi K3** as a high-capability option when analytical correctness on complex file bundles matters more than turn budget, latency, or per-task dollars. Watch whether later API or scaffolding cuts reduce the **83-turn / ~hour** loop without giving back the **1543** Elo; track presentation-quality gaps versus **GPT-5.6 Sol** and **Opus 4.8**; and re-check cost with cache-heavy prompts under the **90%** cached-token discount before defaulting it into production agent stacks. Full AA-Briefcase leaderboard: https://artificialanalysis.ai/evaluations/aa-briefcase.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →