Home / Blog / Kimi K3: second only to Fable 5 on AA-Briefcase
Tech News

Kimi K3: second only to Fable 5 on AA-Briefcase

By Dillip Chowdary • Jul 22, 2026 • Source: HN Claude/Codex/Fable

I'll pull the Artificial Analysis article and the HN thread so the paragraphs stick to reported names, rankings, and mechanics—no invented figures.**Kimi K3**, a **2.8T-parameter** model from **Moonshot AI**, ranks second only to **Claude Fable 5** on **AA-Briefcase**, Artificial Analysis’s agentic knowledge-work benchmark. It posts an **AA-Briefcase Elo of 1543**, a **+727** jump from **Kimi K2.6** (816) and just behind Fable 5 at **1574**. On the broader **Artificial Analysis Intelligence Index**, K3 scores **57**, in line with **Opus 4.8** and **GPT-5.5**. Artificial Analysis published the write-up on **July 21, 2026**, after K3’s release the prior week.

**AA-Briefcase** scores models on a private set of realistic tasks over thousands of complex input files. Deliverables include spreadsheets, presentations, and UI mock-ups. The single Elo blends correctness, analytical quality, and presentation quality. K3’s **rubric pass rate is 51%**, second to Fable 5 at **56%** and ahead of **Claude Sonnet 5 (max, 42.3%)** and **GPT-5.6 Sol (max, 41.8%)**. Its **analytical quality Elo is 1754**, slightly above Fable 5’s **1744**. **Presentation Elo is 1471**, under GPT-5.6 Sol (max, **1660**) and Opus 4.8 (max, **1492**). Cost and latency are steep: about **$10.57 per task**, **83 turns**, **120k output tokens**, and **56.4 minutes** average time per task. Pricing is **$3/$15 per 1M input/output tokens**, with a **90%** cache discount.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For builders shipping agent workflows that produce real artifacts, the split matters. K3 is strong where analytical correctness drives the outcome and weaker where polished slides or mock-ups dominate. High turn count and output volume also raise latency and bill size. That trades off against models that finish fewer turns and spend less per task, even if their Elo is lower.

On AA-Briefcase Elo, K3 leads **GPT-5.6 Sol (max, 1501)**, **Claude Sonnet 5 (max, 1388)**, and **Claude Opus 4.8 (max, 1347)**, with only Fable 5 ahead. Versus **K2.6**, output tokens rose from **42k** to **120k** and turns from **54** to **83**, driving roughly a **10x** rise in cost per task. Average time is about **2.5x** Fable 5 and about **3.8x** **Grok 4.5 (high)**. K3 also uses more turns than Fable 5 (**67**) and GPT-5.6 Sol max (**50**).

Watch whether Moonshot cuts turn count, output tokens, or first-party API speed enough to shrink the **$10.57** and **56.4-minute** averages without giving up the **1543** Elo. For production, compare K3 on analysis-heavy Briefcase-style jobs against Fable 5 for quality-at-cost and against faster, cheaper stacks when wall-clock and spend matter more than a few Elo points.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →