AI-Ready Data Yohan Kim

Context Engineering for AI Agents: What to Keep at Each Step

Analyst selecting relevant evidence for context engineering for AI agents.

Imagine a support agent that retrieves the current refund policy, checks the customer’s order, and then gives an answer consistent with an older rule in its conversation history. The correct policy was available, but it did not guide the answer. To understand why, the team needs to inspect what the model received at that step.

Context engineering for AI agents is the practice of selecting, organizing, and updating the information a model receives at each step of a task. That includes instructions, tool definitions, retrieved material, and working state, not just the user’s prompt.

Investigating a failed run therefore requires more than finding the right document. The team needs to examine the input to the relevant model call and trace it back to the source material used at the time.

Build context around the next decision

An agent’s context window is a working input, not an archive of everything the system knows. A tool may have access to a document without that document being included in the next model call. Conversely, an obsolete passage may remain in the conversation after the source has changed.

Microsoft’s context-engineering guide distinguishes instructions, knowledge, tools, conversation history, and user preferences. The practical implication is to assemble these inputs for the current task rather than assume that a longer transcript contains a better answer.

For a refund decision, the agent needs the applicable policy, relevant order facts, the customer’s request, and the limits of its authority. It may need a lookup tool to resolve a missing fact. It does not necessarily need every policy search result or every earlier troubleshooting exchange.

The application’s design should make those distinctions explicit. A source excerpt is evidence to assess. A working note is a summary that may need checking. A tool definition describes an available operation; it does not, by itself, authorize the agent to perform it.

A larger window does not settle what belongs in it

In Lost in the Middle, Liu and colleagues tested multi-document question answering and key-value retrieval. Performance often depended on where relevant information appeared, with weaker results when it was placed in the middle of long inputs. Those experiments concern the models and tasks tested in 2023, not a universal failure rate for current agents.

The engineering lesson is to test how the agent uses context, not merely whether the input fits. A larger context window still leaves the team deciding which facts to include, which instructions take precedence, and which earlier conclusions to discard.

Anthropic’s engineering guidance describes retrieving information as needed, compacting long histories, keeping structured notes, and isolating focused work in subagents. Each approach needs testing against the task: compaction can remove a critical exception, while a subagent’s summary can omit the uncertainty behind its conclusion.

Choose the smallest change that addresses the observed failure. If the needed passage was never retrieved, rewriting the summary will not fix retrieval. If it was retrieved but an obsolete rule remained authoritative, adding more documents may leave the conflict unresolved.

Trace one wrong answer before redesigning the agent

Consider a hypothetical support workflow. A customer asks whether an opened accessory can be returned. An earlier conversation summary says that all opened products are ineligible. The current policy contains an exception for this accessory category. The agent retrieves the exception but still drafts a refusal.

Start at the model call that produced the refusal. Compare the policy excerpt with the saved summary and the instructions governing policy precedence. Do not infer what the model saw from the documents currently in the search index.

If the exception never entered the call, inspect retrieval and filtering. If it entered without the category definition needed to apply it, inspect the selected passage and its surrounding context. If both the exception and the outdated summary were present, inspect how the application marks superseded information. These are different interventions, even though they produce the same wrong answer.

A compact working record for this example could read:

Current task: determine return eligibility for the accessory, then draft a reply. Applicable evidence: policy version 12, accessory exception, and the order’s product category. Superseded assumption: the earlier blanket exclusion for opened products. Unresolved fact: whether the return falls within the permitted period. Action boundary: no refund execution is authorized.

This illustrative record lets another engineer see what remains unresolved without rereading the entire conversation. It is not a production prompt, and stating an action boundary does not enforce it. The application must control which actions the agent can perform.

Preserve a route back to the evidence

Working context and an investigation record serve different purposes. The first helps the agent decide what to do next. The second helps a person establish what happened.

Next decision
Selected inputsPolicy excerptOrder factsWorking summary
Retained evidence
Source versionPolicy version usedOrder lookupSummary used in the call

Microsoft recommends compact inspection records with identifiers, hashes, selection information, and policy labels instead of indiscriminately logging sensitive prompts or tool results. For this workflow, a useful record would connect the decision step to the selected policy version, the order lookup, and the summary used in that call.

Decide what must be recoverable under the organization’s access and retention rules. An identifier or hash can help establish identity, but it cannot recreate content that was never retained. Where reconstruction is required and permitted, preserve an authorized way to retrieve the relevant source version or input artifact. Keep restricted data in controlled storage rather than duplicating it into general debugging logs.

Freshness and historical reconstruction also require different checks. Refreshing retrieval can supply a newer policy for today’s decision. It does not establish which policy yesterday’s agent actually received. A current source and a preserved historical input are both useful, but they answer different questions.

Test the context change against the task

For the hypothetical refund workflow, compare the existing context assembly with a revised version on the same approved test cases. Hold the model, tool behavior, and evaluation criteria constant where possible. Include cases with a superseded policy, an applicable exception, missing order information, and a request that requires escalation.

Judge whether the agent reaches the correct eligibility decision, identifies missing facts, and respects the action boundary. Track unnecessary lookups and input length as secondary measures. A shorter context is not an improvement if it removes the exception the agent needs.

Inspect failures individually before assigning the cause to data or context. Tool errors, application logic, and model behavior can also explain an incorrect result. If several conditions changed together, record the uncertainty instead of crediting the new context design alone.

Connect context decisions to the data used in the run

When the investigation points to the underlying data, the next step is to test whether changing that data improves the specified task under comparable conditions.

Syntitan, CUBIG’s AI-Ready Data Platform, addresses this data-readiness question. Its published workflow starts with a defined task, compares results before and after data refinement, and links subsequent runs to a versioned data state. The agent application still handles context selection and execution tracing. Recording the data state supports investigation; it does not guarantee the same answer on a later run.

For the refund workflow, the immediate job is to establish why the applicable exception did not guide the answer. If the problem lies in the source data, compare the original and revised data under the same test conditions before deciding what to change in production.

Before changing the model, test whether the data behind the agent’s context is fit for the task. Explore Syntitan’s approach to task-specific data validation and run tracking.

Explore Syntitan, CUBIG’s AI-Ready Data Platform

FAQ

Should the entire tool response stay in the context window?

Not automatically. Keep the fields needed for the next decision and a reference to the retained evidence where permitted. Test any filtering against cases in which an omitted field could change the answer. The full response may belong in restricted evidence storage rather than every subsequent model call.

How much context should an agent retain?

There is no single token budget that suits every workflow. Evaluate the model and task using realistic histories, relevant exceptions, and tool responses. Compare task outcomes alongside input length and latency rather than treating the smallest context as the goal.

Does a saved conversation prove which data the agent used?

Only to the extent that it captures the relevant inputs and their provenance. A transcript may omit retrieval filters, transformed tool output, or source versions. Verify what the application actually records before relying on it for an investigation.

Should an agent resolve conflicting policies by choosing the newest one?

Not by date alone. Applicability and authority matter as well. A newer policy for another region may not govern the case. Define precedence in the application and route unresolved conflicts through the approved review process.