Context Engineering: The Discipline Behind Reliable AI Agents

Key Takeaways
- Context engineering is the practice of deciding what information an agent has in its context window at every step of a task: instructions, retrieved code, tool definitions, memory, and the results of earlier steps. The window is only the visible part. It depends on a knowledge layer, the code graph, tickets, docs, and team conventions, that the agent can query through tools while it runs.
- Prompt engineering vs context engineering comes down to scope. A prompt is one input written once. Context is assembled continuously while the agent runs, and most production failures trace back to that assembly rather than to the prompt.
- Bigger context windows do not solve the problem. Research across 18 models shows accuracy drops as inputs grow, even on simple tasks, so context window management is about choosing and compressing what goes in.
- AI agent memory has to be scoped and skeptical. Memory that persists across runs makes agents better over time, and a confident wrong memory makes them worse faster than anything else.
- Overcut treats context as part of the orchestration layer. It automatically injects the related repositories and tickets from the organization's structure and systems, splits workflows into deterministic and agentic steps that each start with their own context and the previous step's output, and runs workflows that keep organizational knowledge current.
Most teams find the limit of prompting the same way. An agent does well in a demo, gets wired into a real repository, and then starts failing in ways that are hard to reproduce. It edits the wrong module. It ignores a convention that was spelled out in the instructions. It retries a failing test with the same fix it already tried. Someone opens the prompt, adds a sentence in capital letters, and the failure moves somewhere else.
When we trace failures like these back, the prompt is rarely the cause. The agent made a reasonable decision based on what was in front of it, and what was in front of it was wrong, stale, missing, or buried under forty thousand tokens of test output. The fix lives in the system that decides what the model sees. That system has a name now, context engineering, and for any team running agents on real engineering work it matters more than the wording of any single instruction.
The Gap Between Prompting and Production-Grade Agent Reliability
Prompting was built for a single exchange. You write instructions, the model answers, you read the answer. An agent working a ticket runs dozens or hundreds of turns. Every tool call, file read, and test run gets appended to the window, so by the middle of a real task the instructions a team wrote are a small fraction of what the model is reading.
That is where reliability goes. Chroma's context rot study tested 18 models, including GPT-4.1, Claude 4, and Gemini 2.5, and found performance changes significantly with input length, even on simple retrieval tasks. A single distractor, a passage that looks relevant but is not, measurably lowered accuracy. Anthropic's engineering team describes the same thing as an attention budget: every token added competes for the model's attention with every other token, so more context is not free even when it fits.
Drew Breunig gave the failure modes useful names. Context poisoning is when a hallucination lands in the window and gets referenced again and again. Context distraction is when the accumulated history crowds out what the model knows from training. Context confusion comes from superfluous material shaping the answer, and context clash from new information contradicting old. None of these can be fixed by rewording a prompt, because the prompt is not where they happen. They happen in the running history, which no one wrote and no one reviewed.
This is the gap between an agent that works and an agent a team can depend on. Prompting controls the first few thousand tokens. Reliability depends on the rest.
The Discipline Behind Reliable AI Agents
The term context engineering took hold in mid-2025. Shopify's Tobi Lütke described it as providing all the context needed for a task to be plausibly solvable by the model, and Andrej Karpathy called it the art and science of filling the context window with the right information for each step. Anthropic's working definition is the most operational: strategies for curating and maintaining the optimal set of tokens during inference.
The word context gets stretched in two directions, so here is how we use it. In this post, context means everything the model reads when it makes a decision at a given step: the instructions, the files and tickets it pulled in, the tool definitions, its notes, and the history of the task so far. That is narrower than what people usually mean by organizational context, which is the codebase and its dependency graph, the ticket and incident history, the runbooks, and the conventions a team has built up over years. We treat the two as layers of one system. The knowledge layer is everything an agent could know. The window is what it knows right now. Context engineering is the work of moving the right slice from one to the other at the right moment.
That is also why we do not think of context as a window problem alone. An agent with a perfectly curated window and no way to look anything up is stuck with whatever someone chose for it before the task started, and nobody can predict every file, caller, or past incident a real ticket will need. Agents that hold up on real work have runtime access to knowledge through tools: code search, a graph of how modules and symbols depend on each other, the ticket and incident history, the team's written procedures. Those tools let the agent decide what to fetch based on what it learned a minute ago, and the quality of that access sets a ceiling on everything else. A code graph that answers "who calls this function" in one query keeps hundreds of files out of the window that the agent would otherwise have to open to find out.
The phrase "for each step" carries most of the weight. Prompt engineering asks how to phrase an instruction. Context engineering asks what the model should be looking at right now, what should have been removed, what should be fetched only if needed, and what should survive into the next step or the next run. The first is a writing problem. The second is a systems problem, with state, budgets, retrieval, and failure handling, and it gets solved in code and infrastructure rather than in a text box.
We think the distinction matters because it changes who owns reliability. When a team treats agent failures as prompt problems, the fix is a person editing instructions until the symptom goes away. When a team treats them as context problems, the fix is a change to how context gets assembled, and that change applies to every future run.
The Components of the Context Engineering System
A context engineering system has a handful of parts. Each one answers a different question about what the model sees.

Instructions and specifications. The stable layer: the role, the task, the constraints, and the definition of done. This is where prompt engineering still lives, and it should be short. Anything that only applies to some tasks belongs in a later layer, loaded when relevant.
Retrieval. The code, tickets, docs, and logs the task needs, and the tools the agent uses to reach them. The pattern that works for agents is just-in-time retrieval: the agent holds lightweight references, file paths and ticket IDs, and uses search and graph queries to load content when it needs it rather than having everything stuffed in up front. The window only ever holds a slice of what the team knows, so how much of that knowledge retrieval can reach, and how precisely, matters as much as what the agent is handed at the start.
Tools. Tool definitions are context too, and they add up quickly. Factory's deferred context engine keeps only compact tool metadata in the prompt and loads full schemas when a task calls for them. They measured an average 50.8% input token reduction in sessions with more than 100 hidden tools, and found that only 5.4% of sessions actually executed an MCP tool at all.
Memory. AI agent memory comes in two kinds. Working memory is the agent's notes within a task: the plan, what it tried, what it learned, written outside the window so it survives compaction. Persistent memory carries lessons across runs, like the fact that a particular test suite is flaky and should be rerun once before anyone investigates. Persistent memory is the most powerful component on this list and the most dangerous one, because a wrong lesson gets applied with confidence on every run after it.
Window management. Context window management is what happens as a task grows. Compaction summarizes the history and restarts with the summary. Tool outputs get trimmed to the parts that matter, so a failing test run contributes the failure rather than ten thousand lines of passing tests. Subtasks get handed to sub-agents that work in a clean window and return a condensed result.
Observability. A record of exactly what the model saw at each step. Without it, debugging an agent means guessing, and every failure turns back into a prompt edit.
How Context Engineering Works Inside Software Development Agent Workflows
Take a ticket moving from intake to a merged pull request, the kind of chained flow we wrote about in the agentic software development lifecycle.
At intake, the context is the ticket itself plus what surrounds it: linked issues, the reporter's earlier reports, the owning team, recent incidents on the same service. A triage agent that sees only the ticket text guesses. One that sees the surrounding record classifies.
At planning, the agent needs a map of the repository, not the repository. Module boundaries, ownership, the conventions the team has written down, and the skills that describe how this team does a migration or adds an endpoint. The plan it produces becomes a compact artifact that the next stage can load in place of the whole reasoning trail that produced it.
At implementation, window management takes over. The agent loads files as it needs them, runs tests, and reads failures. Windows fill fastest here, and trimming tool output and keeping working notes outside the window decide whether the agent finishes or loses the thread halfway through.
Review is where we make a deliberate choice. The review agent does not inherit the implementer's history. It gets the diff, the ticket, the plan, and the team's review standards, in a clean window. An implementer's context is full of its own reasoning, including the parts that were wrong, and a reviewer that reads that reasoning tends to agree with it. Separating the windows gives the review an independent view of the change, which is the whole point of having one.
At handoff, each stage passes forward a distilled result and writes back what it learned. In Overcut, that memory is scoped to the workflow step rather than shared globally, and it is biased toward skepticism: a memory that led an agent astray loses more weight per bad run than a helpful one gains per good run. A lesson has to keep proving itself to stay in the window. We think that tradeoff is right, because a missing memory costs a little time and a wrong one costs a lot.

How Overcut Assembles and Optimizes Context
Everything above applies to any team building agents. Here is how we apply it in Overcut, and why.
Agents start with the organization already mapped. When a workflow runs on a ticket, Overcut automatically injects what the organization's structure and systems say is related to it: the repositories the work touches, the linked and related tickets, and the knowledge attached to them. The agent does not spend its first twenty turns working out which repository owns a service or whether the bug was reported before. We made this automatic because discovery is where agents quietly go wrong. An agent that picks the wrong repository produces work that looks fine and is aimed at the wrong place. Injecting the relationships up front, while leaving the contents to be fetched on demand, gives the agent a correct starting map without filling its window.
Workflows are separate steps, and each one starts clean. A workflow in Overcut is a chain of steps, and not every step needs a model. Fetching a ticket, checking out a branch, running the test suite, or applying a label are deterministic steps that run as code, cost no tokens, and behave the same way every time. Steps that need judgment, like planning, implementing, and reviewing, are agentic. Every step has access to the output of the steps before it, and every step starts with its own context instead of inheriting the running history of the one before. The clean review window described earlier is one instance of this rule. We think it is one of the most effective optimizations in the system, because it caps window growth at the size of a single step, and it keeps one step's mistakes from becoming the next step's assumptions. Moving mechanical work into deterministic steps also removes a whole category of context: no tool definitions, no retries, and no verbose output for work that never needed a model in the first place.
The knowledge layer is maintained by workflows too. Runtime access is only as good as what it reaches, and organizational knowledge goes stale constantly. Repositories get split, ownership moves, conventions evolve, and a runbook written last year describes a system that no longer exists. In Overcut, keeping that knowledge current is itself a workflow. The same mix of deterministic and agentic steps that works tickets can run on a schedule or on events to refresh what the organization knows about its own code and systems, so the context agents draw on improves as the organization changes instead of drifting away from it. We will go deeper into how that works in an upcoming post.
What Teams Need to Do Differently to Engineer Context
Moving from prompting to context engineering changes what a team treats as the product of its agent work.
Treat context sources as maintained assets. The docs, ownership maps, and written procedures an agent reads are inputs to a production system now, and stale ones produce confident wrong work. The teams that get consistent results are the ones that version their skills and conventions alongside their code and fix the source when an agent goes wrong, instead of patching the prompt.
Debug by reading the context first. When an agent fails, the first question should be what it saw at the step where it went wrong. Most failures that look like model mistakes turn out to be retrieval that missed something or history that should have been trimmed.
Design the handoffs. Every boundary between steps or agents is a decision about what carries forward. Pass artifacts, like plans, diffs, and summaries, rather than raw transcripts. The ticket is a good carrier for this, which is one reason we argued that the real interface for AI development is the ticketing system rather than a new IDE.
Budget the window per step. Decide how much of the window instructions, tools, retrieval, and history each get, and enforce it. A token budget that nobody sets gets set by whichever tool returns the longest output.
Some of this is still hard. There is no reliable automatic way to tell which past lesson is still true, and compaction still occasionally drops the one detail that turns out to matter three steps later. We expect both to improve. What we do not expect to change is where reliability comes from. Models will keep getting better, and the teams that get the most out of them will be the ones that control what those models see.
FAQs
Is context engineering the same as retrieval augmented generation?
No. Retrieval augmented generation is one technique inside context engineering: fetching documents and inserting them into the prompt. Context engineering also covers what to leave out, how tool definitions are loaded, how memory persists across runs, how history is compacted, and what gets passed between agents. A team can run a solid RAG pipeline and still have unreliable agents if the running history grows unchecked or stale memories keep resurfacing.
How does context engineering scale as agent workflows grow?
It scales by keeping each window small rather than making one window bigger. As workflows add steps and agents, each step should receive a purpose-built context and pass forward a compact artifact. That keeps per-step token cost roughly flat as the workflow grows. The harder scaling problem is governance: knowing which sources, memories, and tools each step can reach, and keeping those scopes current as the codebase and team change.
What happens when an agent exhausts its context window?
Without management, the request fails or the oldest content gets truncated, often including the original instructions. Well-built agents compact before that point, summarizing the history and continuing from the summary, with the plan and key decisions saved to working notes outside the window. The risk is that compaction loses a detail that matters later, so what gets preserved should be decided deliberately rather than left to a generic summarizer.
Can context engineering compensate for a weaker underlying model?
Partly. Good context narrows the problem so a smaller model can handle tasks it would otherwise fail, and many teams route simpler steps to cheaper models once each step's context is well scoped. It cannot add reasoning ability the model lacks. For long, ambiguous tasks the stronger model still wins. Poor context hurts every model, though, so better context usually improves results more than a model upgrade alone.
How do you measure whether context engineering is working?
Track outcomes per workflow step rather than per session: task success rate, human rejection rate, retries, and rework after merge. Add context-level signals, including tokens per step, how often compaction triggers, and how often retrieved content is never referenced. When a step fails, check whether the needed information was in the window. A falling share of failures caused by missing or wrong context is the clearest sign the system is improving.



