The Agent Stack: From Keyboard to Model
Now I have all the facts from the source. Let me write the article. Here is the article, saved to…
By Dillip Chowdary • Aug 16, 2026 • Source: HN AI Agents
What happened
Now I have all the facts from the source. Let me write the article. Here is the article, saved to [agent_stack_article.md](file:///home/ubuntu/.gemini/antigravity-cli/brain/3714c778-8863-4132-a5ed-b710235e3075/agent_stack_article.md):
---
The Agent Stack: From Keyboard to Model
Clifford and Gaurav at Nori Agentic have published The Agent Stack, an interactive reference site at the-agent-stack.com that maps every layer of technology a single AI agent request crosses before a response reaches the user. The site uses a scrollable, animated journey to trace one fictional request — labelled req_7f3a — from the moment a person types a message all the way down to the model weights living in GPU memory. It was submitted to Hacker News under the HN AI Agents tag and received 4 points at time of writing.
This article explains what the site contains, how each of its ten layers connects to the next, and why the architectural picture it draws matters to anyone writing, running, or evaluating agent software. The primary audience is developers building on top of Claude Code, Codex, or similar tools who want a clearer mental model of what their code can and cannot control.
How it works
What happened
Nori Agentic, the company behind the Nori Harness and the Nori for Mac and Nori in Slack clients, released The Agent Stack as a standalone educational site. The site divides a live agent run into ten named layers: Interface, Context, Harness, Tools, API, Hyperscaler, Caching Service, Inference Service, GPUs, and Model Weights. Each layer has an expandable panel with a written explanation and an animated diagram, and a second section of the site lets visitors scroll through the same ten layers as a continuous pixel-art journey following req_7f3a.
The creators are explicit that the site grew from a personal need. The opening line asks: when I type "hi" into Claude Code, what exactly happens? That framing positions the resource not as product documentation but as a reference for keeping the full stack straight. No pricing, no version numbers, and no release timeline are attached to the site itself; it is published as an open educational resource without a paywall or account requirement.

How it works
Why it matters
The stack is split into two ownership zones visible on the main diagram. The Interface, Context, Harness, and Tools layers are labelled "Yours," meaning the application developer controls them. The API layer marks the boundary, and everything below it — Hyperscaler, Caching Service, Inference Service, GPUs, and Model Weights — is labelled "The provider's." The harness is described as the software loop that loads context, calls the model API, reads the response, dispatches any tools the model requested, and feeds those results back in. That loop repeats until the model produces a final answer rather than another tool call.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
Below the API boundary, the site describes the caching service as reusing previously processed prompt prefixes to reduce latency and token cost, with cache lifetimes the site puts at five-minute and one-hour windows depending on provider. The inference service then runs a prefill pass over the full prompt once to create a KV state for that request, followed by repeated decode passes that extend the state and stream one new token per pass. The GPU layer explains that generation speed is limited by how fast weights move out of high-bandwidth memory, abbreviated HBM, rather than by raw arithmetic throughput.
Why it matters
The ten-layer model clarifies where a developer's decisions end and the provider's infrastructure begins. The site makes the API boundary explicit: past the HTTP POST, the developer cannot see or change how the request is served. It also explains that an AI gateway — products such as Cloudflare AI Gateway, LiteLLM, and Portkey are listed as examples — can stand at that seam to centralise authentication, rate limits, audit logs, and per-developer usage attribution. The tradeoff the site names is that inserting a gateway means new provider capabilities are unavailable until the gateway supports them too.
The caching explanation has direct cost implications. The site notes that anything changed near the beginning of a prompt invalidates all cached content that follows, so harness design choices and long gaps between turns appear directly on an API bill. The model weights section adds a second cost point: sparse models using a mixture-of-experts architecture route each token through only a subset of the network's blocks, which is why parameter count stopped predicting inference cost. A sparse model can carry a trillion parameters and still do the work of a far smaller one on any given token.
Who is affected
Who is affected
Application developers using Claude Code, Codex, or the Claude Agent SDK face the most immediate consequences from this framing. The site lists the Claude Agent SDK and the Codex App Server as harness examples, and AGENTS.md, CLAUDE.md, and llms.txt as context examples, pointing to the specific files developers already maintain. Teams using an AI gateway — including those on Cloudflare AI Gateway, LiteLLM, or Portkey — should confirm that their gateway version supports any new API features their harness needs before upgrading.
Model providers, inference runtime maintainers using vLLM, SGLang, or TensorRT-LLM, and hyperscaler customers on AWS, Google Cloud, or Azure are also represented in the diagram, though the site addresses them as context for application developers rather than as its primary audience. Open-weight model families including Qwen, DeepSeek, and Llama appear as examples at the model weights layer, relevant for teams self-hosting inference rather than routing through a managed API.
What to watch next
What to watch next
Builders should verify how their harness handles the return path for tool calls. The site is explicit that a tool result does not go directly to the model; it goes back to the harness, which reads the output, decides the next step, and only then calls the API. A harness that misroutes or drops tool results will produce silent failures rather than error messages, and the ten-layer model is useful for diagnosing which layer introduced the fault.
The Nori Agentic team has not announced updates to the site's content or a roadmap for additional layers. Developers who want to monitor changes can watch the site directly at the-agent-stack.com. The Hacker News thread at item id 49298291 had no comments at time of submission, so community discussion of edge cases or corrections had not yet surfaced. Builders running long-context sessions should confirm which prompt caching tier their provider offers, since both the five-minute and one-hour window variants affect whether a session's prefix stays warm across a sequence of agent turns.
---
Word count is approximately 860. All concrete names and numbers from the source are used (req_7f3a, 4 points, ten layers, five-minute/one-hour cache windows, Clifford and Gaurav, Nori Agentic, and all named products). No dates, dollar amounts, version numbers, or quotes were invented.
Advertisement