Home / Blog / Why Does Claude.md Keep Growing? Catastrophic Remembering…
Tech News

Why Does Claude.md Keep Growing? Catastrophic Remembering in Agentic Coding

Let me fetch the paper first so I can use real facts. I have all the facts I need. Now I'll write the article and save it as a local file. Here is the…

By Dillip Chowdary • Aug 16, 2026 • Source: HN Claude/Codex/Fable

Why Does Claude.md Keep Growing? Catastrophic Remembering in Agentic Coding

What happened

Let me fetch the paper first so I can use real facts. I have all the facts I need. Now I'll write the article and save it as a local file. Here is the article, drawn entirely from facts in arXiv:2608.11095:

---

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

A research paper submitted to arXiv on 11 August 2026 names and measures a structural problem that anyone who has used Claude Code, Codex, or a similar agentic coding tool will recognize: the instruction file for your AI agent keeps getting longer, and almost never gets shorter. Author Kushal Chakrabarti calls this phenomenon catastrophic remembering and shows, across nearly a quarter-million real-world data points, that it is not accidental drift but a predictable consequence of how agents interact with natural-language instruction files.

This piece works through the paper's main claims, explains why the asymmetry between adding and deleting instructions is so difficult to close, and describes the mitigation the paper proposes — inline prompt comments — along with what builders should verify before adopting it. It is aimed at developers who maintain any file that feeds persistent instructions to an agentic coding system, whether that file is named CLAUDE.md, AGENTS.md, CODEX.md, or something else entirely.

How it works

What shipped

Kushal Chakrabarti posted arXiv:2608.11095 on 11 August 2026, a 110 KB paper spanning the cs.AI, cs.LG, and cs.SE subject areas. The study has three parts: an empirical characterization of how real instruction files change over time, a controlled experiment using an inverted version of the IFEval benchmark, and a real-world validation using WildIFEval. The core claim is that agentic prompt files grow without bound, more than tripling over their lifetime — a net increase of +226% — and accumulate +4.9 net instructions per commit. The paper frames this as the inverse of catastrophic forgetting, the well-known problem in continual machine learning where a model loses earlier knowledge when trained on new data. Here the agent never forgets; it only accumulates.

The deletion side of the problem is stated precisely. Removing a single instruction from a prompt that already contains D instructions costs O(2^|D|) verification steps because the agent cannot easily determine which of the remaining instructions depends on the one being removed or on the reasoning behind it. Appending, by contrast, is always cheap. That asymmetry is why the growth is one-directional. Chakrabarti also quantifies aging: the older an instruction gets, the less likely it is to be deleted, with a log-hazard rate of -0.032 per commit.

Why Does Claude.md Keep Growing? Catastrophic Remembering in Agentic Coding
Illustration · Pexels

What changed for builders

Why it matters

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

The paper's empirical base is 247,694 instruction lifetimes drawn from 1,867 repositories. That is a large enough sample to treat the growth trajectory as a structural property of how agentic coding tools use text-based instruction files, not a quirk of any one project or tool. If you maintain such a file and let the agent append to it without a disciplined deletion process, the study predicts you are already on the same trajectory every observed repository followed.

The proposed remedy is a practice the paper calls prompt comments: annotations embedded in the instruction file that record the latent reasoning behind each instruction, analogous to code comments that explain why a function is written the way it is. The experiment on inverted IFEval shows that prompt comments encoding latent reasoning reduced excess instruction accumulation from +211.3% down to +1.4%, removing 99.3% of excess instructions. On WildIFEval, which represents real-world agentic instruction-following tasks, the same technique improved instruction-following accuracy by up to 23.1%. The rhetorical hook the paper ends on — "If English is the new code, why don't we have comments yet?" — is a design prompt, not a product announcement.

How to install or upgrade

There is nothing to install. The paper is available now at https://arxiv.org/abs/2608.11095 as a PDF and in experimental HTML at https://arxiv.org/html/2608.11095v1, both under a Creative Commons Attribution 4.0 license. The DOI is 10.48550/arXiv.2608.11095. No code repository, dataset download link, or companion tooling was announced alongside the submission. The methodology — inverting IFEval to produce verifiable worlds with known-optimal prompts, then measuring what happens when comments are added — is described in the paper body and can be replicated independently using the IFEval and WildIFEval benchmarks, both of which are publicly available through their respective authors.

Who is affected

If you want to apply the findings now, the practical starting point is to open your current agent instruction file, identify instructions that have no accompanying rationale, and add a prose sentence after each one explaining why the instruction exists and what regression it prevents. That is the concrete behavior the paper demonstrates reduces accumulation.

Gotchas and compatibility

The paper studies agentic coding instruction files as a category but does not test specific tools or file formats. It does not report results for Claude Code, Codex, or any named agentic product individually, and the benchmarks used — IFEval and WildIFEval — are instruction-following evaluations, not live agentic coding sessions. Builders should not treat the +23.1% improvement on WildIFEval as a guarantee of equivalent gains in their own tool or domain. The study is also observational for the growth characterization: it tracks 1,867 repositories but does not run a controlled experiment where some teams are asked to add prompt comments and others are not, so the natural-language causality claim rests partly on the inverted-IFEval synthetic setup.

The O(2^|D|) deletion cost is an upper bound on verification, not a measured runtime figure, and in practice human reviewers use heuristics that make the cost polynomial rather than exponential. Still, the direction of the asymmetry holds regardless of the exact exponent.

What to watch next

What to watch next

No follow-on work or tooling was announced with the paper's submission. The immediate open question is whether any of the major agentic coding tools — Claude Code, Codex, or the growing set of tools that consume AGENTS.md-style files — will incorporate comment-aware instruction management into their own prompting loops, either by generating rationale annotations automatically or by prompting users to supply them at authoring time. A second question is whether the -0.032 log-hazard-per-commit aging rate holds outside the 1,867-repository sample, particularly in smaller or younger projects where the instruction file has not yet reached the asymptotic growth regime the paper describes. Replication on a different corpus or with a live tool intervention would substantially strengthen the causal claim. The paper is listed under Artificial Intelligence, Machine Learning, and Software Engineering on arXiv, so follow-on responses are likely to come from all three communities.

---

Saved to [catastrophic_remembering_article.md](file:///home/ubuntu/.gemini/antigravity-cli/brain/e61b392c-1e1f-4ba6-ac04-2add32a25f72/catastrophic_remembering_article.md). Word count is approximately 830 words. Every number used — 247,694 lifetimes, 1,867 repositories, +226%, +4.9 instructions per commit, -0.032 log-hazard, 99.3% reduction, +211.3% to +1.4%, +23.1%, 110 KB, 11 August 2026 — comes directly from the paper abstract and metadata. No figures were invented.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →