Kimi K3: second only to Fable 5 on AA-Briefcase
By Dillip Chowdary • Jul 22, 2026 • Source: HN Claude/Codex/Fable
I'll pull the source article for concrete facts, then draft the five–six analytical paragraphs from that only.Moonshot AI last week released **Kimi K3**, a 2.8T-parameter model that scores 57 on the Artificial Analysis Intelligence Index, roughly in line with Claude Opus 4.8 and GPT-5.5. On Artificial Analysis’s new agentic knowledge-work suite, **AA-Briefcase**, K3 posts an Elo of 1543 — second only to Claude Fable 5 at 1574, and a +727 jump over Kimi K2.6’s 816. The HN thread on the write-up had 2 points and no comments at the time of capture.
AA-Briefcase is a private, long-horizon eval: models work over thousands of complex input files and must produce real deliverables — spreadsheets, presentations, UI mock-ups — scored into one Elo from correctness, analytical quality, and presentation quality. K3’s rubric pass rate is 51% (Fable 5: 56%), and its analytical quality Elo is 1754 versus Fable 5’s 1744. Presentation is the soft spot: Presentation Elo 1471 trails GPT-5.6 Sol (max, 1660) and Claude Opus 4.8 (max, 1492). Cost and latency are heavy: about $10.57 per task (~10x higher than cheaper peers on the suite), 83 turns and 120k output tokens per task, average 56.4 minutes per task — roughly 2.5x Fable 5 and 3.8x Grok 4.5 (high). API pricing is $3/$15 per 1M input/output tokens, with a 90% discount on cached tokens.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For builders shipping agentic knowledge workflows, the signal is quality ceiling versus unit economics. K3 can match near-frontier analytical work on multi-file brief production, but 56 minutes and ~$10.57 per task make it a poor default for high-volume or interactive paths unless caching, batching, or human handoff cut turn count and output size. If your bottleneck is rubric-level correctness and analysis rather than slide polish, K3 is competitive; if clients judge decks first, GPT-5.6 Sol’s presentation Elo still leads.
On the same board K3 sits ahead of GPT-5.6 Sol (max, 1501), Claude Sonnet 5 (max, 1388), and Claude Opus 4.8 (max, 1347). It costs more to run on this suite than Opus 4.8 despite that ranking. Versus its own line, K2.6 used 42k output tokens and 54 turns per task; K3 more than doubles both, which is where most of the cost and wall-clock land. Moonshot is now in the same agentic-work tier as the closed Claude and GPT flagships on this private set, with a clear trade of more turns and tokens for higher Elo.
Practical takeaway: treat K3 as a high-Elo, high-cost agent for hard multi-artifact jobs, not a drop-in for every step. Watch turn efficiency and first-party API speed — those drive the 56.4-minute average — and whether presentation quality closes on Sol. For production design, measure cost-per-accepted-deliverable on your own file packs, lean on the 90% cache discount where inputs repeat, and keep a cheaper model for presentation-only or short-horizon steps. Full AA-Briefcase tables live at artificialanalysis.ai/evaluations/aa-briefcase.
Advertisement