Home / Blog / We're lying to Claude in almost every session
Tech News

We're lying to Claude in almost every session

Article URL: https://github.com/Piebald-AI/claude-code-system-prompts/blob/main/system-prompts/system-prompt-autonomous-operation-guidelines.md Comments URL.

By Dillip Chowdary • Aug 25, 2026 • Source: HN Claude/Codex/Fable

We're lying to Claude in almost every session

What happened

- Repository: Piebald-AI/claude-code-system-prompts on GitHub - Description: "All parts of Claude Code's system prompt, 27 builtin tool descriptions, sub agent prompts (Plan/Explore/Task), utility prompts (CLAUDE.md, compact, statusline, magic docs, WebFetch, Bash cmd, security review, agent creation). Updated for each Claude Code version." - The specific file surfaced: system-prompt-autonomous-operation-guidelines.md - Its internal name: "System Prompt: Autonomous operation guidelines" - Its description: "Instructs autonomous sessions to proceed on reversible work, stop for destructive or scope-changing actions, and finish promised work before ending the turn" - Claude Code version tagged in the file: ccVersion: "2.1.227" - The actual prompt text: "You are operating autonomously. The user is not watching in real time and cannot answer questions mid-task, so asking 'Want me to…?' or 'Shall I…?' will block the work. For reversible actions that follow from the original request, proceed without asking. Stop only for destructive actions or genuine scope changes the user must decide. Offering follow-ups after the task is done is fine; asking permission before doing the work is not." - Also: exception clause about when user is describing a problem/thinking out loud — the deliverable is assessment, not action. - HN post: 1 point, 1 comment

Now I have enough to write the article.

A GitHub repository called claude-code-system-prompts, maintained by Piebald-AI, has surfaced the full set of hidden instructions that Anthropic injects into Claude Code before any developer types a word. The collection covers all parts of Claude Code's system prompt, 27 built-in tool descriptions, sub-agent prompts for Plan, Explore, and Task modes, and a range of utility prompts — including those for CLAUDE.md ingestion, compact mode, the statusline, WebFetch, Bash command handling, security review, and agent creation. One file in that collection, tagged to Claude Code version 2.1.227, exposes exactly how the tool talks to itself when it runs without a human watching.

How it works

This piece walks through what that autonomous-operation prompt says, how it shapes the model's decision-making during unattended runs, why the behavioral contract it encodes matters for anyone deploying Claude Code in a pipeline, and what builders should verify before trusting the tool with irreversible work.

Piebald-AI published a public GitHub repository that tracks and archives every layer of Claude Code's system prompt across releases. The project keeps all sections updated per Claude Code version, so the contents of version 2.1.227 are now readable by anyone. The specific file that attracted attention, named "System Prompt: Autonomous operation guidelines," carries an internal description that reads: "Instructs autonomous sessions to proceed on reversible work, stop for destructive or scope-changing actions, and finish promised work before ending the turn." Anthropic had not published these instructions separately; they were inferred and extracted from the running tool.

We're lying to Claude in almost every session
Illustration · Pexels

The Hacker News thread that surfaced the repository currently sits at 1 point with 1 comment. That low engagement does not diminish the substance of what the file reveals. Because Claude Code is widely used in autonomous and CI-adjacent workflows, the instructions governing those sessions have operational relevance for any team that runs it without a human in the loop. The extraction gives practitioners a concrete reference rather than a black-box assumption about how the model handles ambiguous situations when left to work alone.

Why it matters

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

The autonomous-operation system prompt opens by telling the model a fact that is technically true but that the user never stated: "You are operating autonomously. The user is not watching in real time and cannot answer questions mid-task." From that premise it derives a behavioral policy. Asking confirmation questions like "Want me to…?" or "Shall I…?" is framed as blocking work, so the model is instructed not to ask them. For actions it judges to be reversible and within scope of the original request, it proceeds. It stops only when it encounters what the prompt calls "destructive actions or genuine scope changes the user must decide."

There is a built-in exception clause. When the user's input reads as describing a problem, asking a question, or thinking out loud rather than issuing a request, the model is told the deliverable is its assessment, not an action. That distinction — between a request and a statement — is left to the model to classify in real time with no further heuristic provided in this file. The prompt also instructs the model to offer follow-ups after the task is finished rather than seeking permission before starting, and to finish any work it has already promised before ending a turn. The policy is encoded as plain prose, not as a structured rule set.

The prompt establishes a consent model where the user's silence is treated as permission. Because the model is told the user cannot answer mid-task, it is discouraged from pausing even when the situation has deviated from the original request. The only guardrail named in the file is the model's own judgment about whether an action is "destructive" or constitutes a "genuine scope change." No external signal, timeout, or fallback channel is described in this prompt. Teams running Claude Code in automated pipelines are therefore operating under a behavioral contract they did not author and, until this repository appeared, could not fully read.

Who is affected

The version tag of 2.1.227 means this is not a static document. The prompt text can change with any Claude Code release, and the behaviors a team validated in one version may not hold in the next. The Piebald-AI repository is positioned as a running changelog — its stated goal is to update the collection with each Claude Code version — but teams relying on it for compliance or audit purposes would need to track that repository actively. The asymmetry between how confidently the tool acts and how little of its instruction set has been user-visible is what makes this disclosure notable.

Developers using Claude Code in its default interactive mode are affected primarily at the level of model expectation: the tool is primed to assume autonomy even in sessions where the developer is present and could answer a question. Any developer who has noticed Claude Code proceeding without confirmation on a step they expected it to pause on is now looking at the mechanism that produced that behavior. The prompt was already running in those sessions; it is now simply visible.

Teams running Claude Code in truly autonomous contexts — scheduled jobs, CI pipelines, or multi-step agentic workflows — are more directly affected. For them, the prompt's policy about reversible versus destructive actions is the operative boundary between what the tool will do on its own and what it will defer. Because that boundary is defined only in natural language within the prompt, and because the prompt is subject to change across versions, any team that has built workflows around a specific behavioral assumption should verify those assumptions hold against the current version in the Piebald-AI repository.

What to watch next

The most immediate thing to monitor is whether Anthropic responds to the public availability of these prompts. The company has not historically published its system-prompt contents, so the Piebald-AI extraction represents a gap between intended opacity and actual discoverability. Anthropic could acknowledge the prompts, update them in a way that changes the autonomous-operation policy, or address the disclosure through other means. Any change to the 2.1.227 text visible in the repository would signal a reaction worth tracking.

Builders should also watch whether the exception clause for "thinking out loud" receives elaboration in future versions. As written, the model is the sole classifier of whether a user message is a request or a reflection, with no tie-breaker mechanism named. That classification has direct consequences for whether the tool acts or reports. If Anthropic adds structure to that decision — a confidence threshold, a clarification request with a timeout, or a separate signal channel — it would represent a meaningful change to the safety posture of autonomous runs. The 27 tool descriptions and sub-agent prompts in the same repository are worth reading alongside this file, since the Plan, Explore, and Task sub-agents operate under their own layered instruction sets.

Developer Action Items

  • ☐ Verify the claim on the official Anthropic / Claude / GitHub page (or HN Claude/Codex/Fable), not from this recap alone.
  • ☐ Name the surface that moved — API, policy, model, hardware, or commercial terms — before you Slack the thread.
  • ☐ Assign one owner a day to read the primary material and decide: this-sprint, this-quarter, or noise.
  • ☐ Do not change production on day-one coverage. Watch the vendor changelog and one independent write-up first.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →