Kimi K3: second only to Fable 5 on AA-Briefcase
By Dillip Chowdary • Jul 22, 2026 • Source: HN Claude/Codex/Fable
I'll pull the source article so the paragraphs stick to reported facts—names, rankings, and mechanics—without inventing figures.Kimi K3 from Moonshot AI ranks second only to Claude Fable 5 on AA-Briefcase, Artificial Analysis’s agentic knowledge work benchmark. The model posts an AA-Briefcase Elo of **1543**, behind Fable 5 at **1574** and ahead of GPT-5.6 Sol (max, **1501**), Claude Sonnet 5 (max, **1388**), and Claude Opus 4.8 (max, **1347**). That score is a **+727** jump over Kimi K2.6 (**816**). Separately, the **2.8T**-parameter model scores **57** on the Artificial Analysis Intelligence Index, in the same band as Opus 4.8 and GPT-5.5. Artificial Analysis published the AA-Briefcase write-up on July 21, 2026; the related Hacker News thread had **2** points and **0** comments at the time of the summary.
AA-Briefcase runs models against a fully private set of realistic knowledge-work tasks over thousands of complex input files. Required outputs include spreadsheets, presentations, and UI mock-ups. The single AA-Briefcase Elo combines correctness, analytical quality, and presentation quality. On the objective slice, Kimi K3’s rubric pass rate is **51%**, second to Fable 5 at **56%** and ahead of Sonnet 5 max (**42.3%**) and GPT-5.6 Sol max (**41.8%**). Analytical quality Elo is **1754**, roughly level with Fable 5 at **1744**. Presentation Elo is weaker at **1471**, below Sol max (**1660**) and Opus 4.8 max (**1492**). Cost and latency are the other half of the result: about **$10.57** per task, **83** turns, **120k** output tokens, and **56.4** minutes average time per task—roughly **2.5x** Fable 5 and **3.8x** Grok 4.5 (high). API pricing is **$3/$15** per 1M input/output tokens with a **90%** discount on cached tokens.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For engineers building agentic workflows that produce real office deliverables, the split profile matters more than a single headline rank. Kimi K3 is competitive where the job is to parse dense inputs and produce correct, analytically sound work product. It is less competitive where polished presentation is the gate. The high turn count and output volume show a model that spends more steps and tokens closing tasks, which is useful when autonomy and depth beat speed, and expensive when you are metering thousands of multi-file jobs.
In market terms, open-weight Kimi K3 now sits in the same AA-Briefcase tier as closed frontier systems from Anthropic and OpenAI, with only Fable 5 clearly ahead on the combined Elo. It outranks Opus 4.8 and GPT-5.6 Sol on this agentic knowledge suite even while Artificial Analysis notes it costs more than Opus 4.8 to run and averages nearly an hour per task. Relative to its own prior generation, the K2.6 → K3 leap on both Elo and token/turn intensity is large; relative to Fable 5, the gap is small on analysis and larger on presentation, cost, and wall-clock time.
Practical takeaway: treat Kimi K3 as a strong candidate for agent stacks that score on correctness and analysis of multi-file knowledge work, and budget for higher per-task spend and long runtimes—especially if presentation quality is part of the acceptance bar. Watch whether later Kimi API speedups, caching, or scaffolding cut the **83**-turn / **56.4**-minute profile without giving back the **1543** Elo; that is the lever that decides if second place on AA-Briefcase is deployable at production volume or stays a benchmark win.
Advertisement
🔎 More interesting news
- OpenAI's models broke containment and cyberattacked Hugging Face — what enterprises need…
- Chick-fil-A discloses data breach after credential stuffing attacks
- Claude Code on desktop now works with the iOS simulator
- Show HN: Public-safe skin packs for the Codex desktop app
- Today's full Tech Pulse briefing →