Kimi K3: second only to Fable 5 on AA-Briefcase
By Dillip Chowdary • Jul 22, 2026 • Source: HN Claude/Codex/Fable
I'll pull the source article so the paragraphs stick to verified names and numbers, not invented detail.Moonshot AI’s Kimi K3 landed second on Artificial Analysis’s AA-Briefcase, the lab’s agentic knowledge-work benchmark, with an Elo of 1543. That is only 31 points behind Claude Fable 5 at 1574 and a jump of 727 points over Kimi K2.6’s 816. The same model scores 57 on the Artificial Analysis Intelligence Index, in the same band as Claude Opus 4.8 and GPT-5.5, and is described as a 2.8T-parameter system released the week before the AA-Briefcase write-up (July 21, 2026).
AA-Briefcase is a private, long-horizon evaluation built around realistic deliverables—spreadsheets, presentations, UI mock-ups—over thousands of complex input files. The single Elo folds together correctness, analytical quality, and presentation quality. On that breakdown, Kimi K3 posts a 51% rubric pass rate (second to Fable 5’s 56%), an analytical quality Elo of 1754 (slightly above Fable 5’s 1744), and a weaker Presentation Elo of 1471, trailing GPT-5.6 Sol (max) at 1660 and Claude Opus 4.8 (max) at 1492. Runtime is heavy: about 83 turns and 120k output tokens per task, versus 54 turns and 42k output tokens on K2.6, for an average 56.4 minutes and $10.57 per task at $3/$15 per 1M input/output tokens (90% discount on cache hits).
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For engineers wiring agentic workflows, the signal is not a generic “smarter model” claim—it is that a non-Anthropic stack can nearly match Fable 5 on rubric and analysis while still losing on presentation polish and, more importantly, on cost and latency. At ~2.5× Fable 5’s time per task and among the highest cost-per-task figures on the board, K3 is closer to a batch or offline knowledge-work runner than a snappy interactive copilot unless you aggressively use cached context and cap turns.
Competitively, AA-Briefcase now ranks Fable 5 (1574), Kimi K3 (1543), GPT-5.6 Sol max (1501), Claude Sonnet 5 max (1388), and Claude Opus 4.8 max (1347). That order matters because intelligence-index proximity (K3 ≈ Opus 4.8 / GPT-5.5) does not fully predict multi-file agentic rankings: K3 beats both Sonnet 5 and Opus 4.8 on this benchmark while remaining clearly second to Fable 5. Moonshot’s climb from K2.6’s 816 Elo also shows how fast prior-generation Chinese frontier models are closing the agentic gap on Artificial Analysis’s private set.
Practical takeaway: treat Kimi K3 as a strong candidate for long-horizon knowledge work where analytical correctness and rubric pass rate dominate, and budget for ~$10+ and nearly an hour per hard task unless you redesign tool loops to cut the 83-turn average. Watch next for (1) whether presentation Elo improves without bloating tokens further, (2) third-party hosts that raise tokens/sec enough to shrink the 56.4-minute mean, and (3) any open-weights release of the 2.8T model—Artificial Analysis has already framed that path as potentially leading among open-weight systems if it ships.
Advertisement
🔎 More interesting news
- OpenAI's models broke containment and cyberattacked Hugging Face — what enterprises need…
- Governments, companies, nonprofits should invest in free, open source AI [pdf]
- Grok Build is open source
- Windows 0-day drops the same day Microsoft releases record number of patches
- Today's full Tech Pulse briefing →