Home / Blog / Kimi K3: second only to Fable 5 on AA-Briefcase
Tech News

Kimi K3: second only to Fable 5 on AA-Briefcase

By Dillip Chowdary • Jul 22, 2026 • Source: HN Claude/Codex/Fable

I'll pull the Artificial Analysis article so the paragraphs stay grounded in reported names and rankings only.Moonshot AI’s Kimi K3 posts an AA-Briefcase Elo of 1543 on Artificial Analysis’s agentic knowledge-work benchmark, second only to Claude Fable 5 at 1574. That is a +727 jump from Kimi K2.6’s 816 and puts K3 ahead of GPT-5.6 Sol (max, 1501), Claude Sonnet 5 (max, 1388), and Claude Opus 4.8 (max, 1347). On the broader Artificial Analysis Intelligence Index, the 2.8T-parameter model scores 57, in the same band as Opus 4.8 and GPT-5.5. Artificial Analysis published the AA-Briefcase write-up on July 21, 2026; the piece was also linked from Hacker News.

AA-Briefcase is a proprietary agentic benchmark: models work a private set of realistic tasks over thousands of complex input files and must produce deliverables such as spreadsheets, presentations, and UI mock-ups. A single Elo blends correctness, analytical quality, and presentation quality. Kimi K3’s rubric pass rate is 51%, second to Fable 5’s 56% and ahead of Sonnet 5 (max, 42.3%) and GPT-5.6 Sol (max, 41.8%). Its analytical quality Elo is 1754, comparable to Fable 5’s 1744, while presentation quality lags at 1471 versus Sol’s 1660 and Opus 4.8’s 1492.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For builders, the score split is the operational signal. K3 is competitive on correctness and analysis for long-horizon knowledge work, but weaker on polished presentation assets. It also averages 56.4 minutes and about $10.57 per task, with roughly 83 turns and 120k output tokens per task—up from 54 turns and 42k output tokens on K2.6. First-party Kimi API pricing is $3/$15 per 1M input/output tokens, with a 90% discount on cached tokens. Latency and spend dominate if you run this class of work in a tight loop.

Competitively, AA-Briefcase ranks K3 just behind Anthropic’s Fable 5 and above OpenAI’s GPT-5.6 Sol (max) and the other listed Claude max runs. Cost and time undercut the leaderboard story: K3 costs more than Opus 4.8 to run on this suite, averages nearly an hour per task—about 2.5x Fable 5 and about 3.8x Grok 4.5 (high)—and sits among the most expensive models on the board, driven by token pricing, high output volume, and more turns (83 versus 67 for Fable 5 and 50 for Sol max).

Practical takeaway: treat Kimi K3 as a strong analysis/agent engine for briefcase-style work if you can budget ~$10 and ~an hour per hard task, then plan a separate pass or model for presentation polish. Watch turn count, cache hit rate, and output-token growth versus K2.6 when you re-bench in your own harness; those three numbers will decide whether the Elo gain survives production cost and wall-clock limits.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →