Home / Blog / Kimi K3: second only to Fable 5 on AA-Briefcase
Tech News

Kimi K3: second only to Fable 5 on AA-Briefcase

By Dillip Chowdary • Jul 22, 2026 • Source: HN Claude/Codex/Fable

I'll pull the source article for concrete facts on Kimi K3 and AA-Briefcase so the paragraphs stay specific and don't invent numbers.Moonshot AI last week released **Kimi K3**, a 2.8T parameter model that scores 57 on the Artificial Analysis Intelligence Index, in the same band as Opus 4.8 and GPT-5.5. On **AA-Briefcase**, Artificial Analysis’s agentic knowledge-work benchmark, Kimi K3 posts an Elo of 1543 — second only to Claude Fable 5 at 1574, and a +727 jump over Kimi K2.6 at 816. The same run puts it ahead of GPT-5.6 Sol (max, 1501), Claude Sonnet 5 (max, 1388), and Claude Opus 4.8 (max, 1347).

AA-Briefcase scores models on a fully private dataset of realistic tasks over thousands of complex input files. Deliverables include spreadsheets, presentations, and UI mock-ups; a single Elo combines correctness, analytical quality, and presentation quality. On those axes, Kimi K3 hits a 51% rubric pass rate (behind only Fable 5 at 56%) and an analytical quality Elo of 1754 (near Fable 5’s 1744). Presentation is the weak leg: Presentation Elo 1471 trails GPT-5.6 Sol (max, 1660) and Opus 4.8 (max, 1492). Cost and latency are high: about $10.57 per task at $3/$15 per 1M input/output tokens (90% cache discount), ~83 turns and ~120k output tokens per task, and 56.4 minutes average time per task — roughly 2.5× Fable 5 and 3.8× Grok 4.5 (high).

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For engineers building agentic document and knowledge workflows, the split is the useful signal. Kimi K3 can match or beat top frontier models on objective rubric pass and analytical quality when the job is multi-file synthesis into structured deliverables. It is less competitive when the output must look polished. At ~$10.57 and nearly an hour per task, with more turns than Fable 5 (67) or GPT-5.6 Sol max (50), it is a quality-first option, not a default for high-volume pipelines unless caching and turn budgets are tightly controlled.

Market context is crowded at the top of AA-Briefcase. Fable 5 still leads overall Elo; GPT-5.6 Sol is close on aggregate and stronger on presentation; Sonnet 5 and Opus 4.8 sit lower on this benchmark despite brand weight. Kimi K3’s Intelligence Index score of 57 keeps Moonshot in the same general intelligence tier as Opus 4.8 and GPT-5.5, while AA-Briefcase shows a much larger step-up from K2.6 than a small index delta would suggest. The competitive pressure is on Chinese and Western labs alike to post private-set agentic scores, not only public chat benchmarks.

Practical takeaway: trial Kimi K3 on multi-file analytical work where spreadsheet or analysis correctness matters more than slide polish, and measure cost and wall-clock against Fable 5 and GPT-5.6 Sol on your own task mix. Watch presentation quality, turn count, and first-party API speed if latency budgets are tight; full AA-Briefcase leaderboards at artificialanalysis.ai/evaluations/aa-briefcase are the place to track whether later Kimi or competitor releases close the Fable 5 gap without the hour-scale runtime.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →