Context Engineering
A long task fills the model’s context: a 9,000-token log from one tool, forty steps of tool calls. By the end of this page you will have seen the two things the runtime does about it, both on by default: a large tool result is saved to the workspace and the model gets a preview it can search, and a run’s messages are trimmed — or summarized — before a model call would go over your budget.Every output on this page is what the code printed when it was run with
gpt-5.4-mini. The model’s wording, and which artifact tool it picks, will
differ on your run.A large tool result, offloaded
This tool returns a 400-line log — about 9,200 tokens. Nothing is configured:workspace/artifacts/ folder.)
The model was sent 163 tokens instead of 9,201, found the two errors with
search_artifact, and the full log stayed on disk. Without the instruction
to search, the model on another run called read_artifact and loaded the
whole log back into its context — say in the system instruction how you want
large results handled.
A long run, trimmed before each model call
Here the budget is made tiny on purpose, so it is crossed on a six-item task:sliding_window counts messages, and the run acts once it holds more than 6.
- 4). With
"strategy": "truncate"they would simply be dropped. Each reduction is in the trajectory, with digests of the messages before and after.
How it works
Offloading. After a tool returns, its result is measured. Overthreshold_bytes or threshold_tokens, the full result is written to
workspace/artifacts/ (or your S3/R2 workspace) and the model’s observation
becomes the preview above: the first max_preview_lines lines, cut to
max_preview_tokens. Results of the workspace file tools and the artifact
tools themselves are never offloaded. While offloading is on, the model has
four tools to reach the rest:
Context management. Before every model call, the run’s messages are
measured: in
token_budget mode, tokens against threshold_percent of
value; in sliding_window mode, the number of messages (not counting the
system prompt) against value. Over it, the messages between the system prompt
and the preserve_recent latest are dropped (truncate), or summarized by
one extra model call into a single message (summarize_and_truncate). A
tool call and its results are kept or dropped together. If the summary call
fails, the run truncates instead and carries on.
Set value below your model’s real context window: the check happens before
the call, so the provider never sees an oversized prompt.
Options
enable_subagents, context management and
offloading are switched on for the lead and every worker, whatever you set,
keeping your other values. Every setting is in the
agent settings reference.
When things go wrong
ValueError: context_management.preserve_recent must be at least 4, got 2
ValueError: context_management.preserve_recent must be at least 4, got 2
The latest four messages are always kept, so the model never loses the
call it just made and its result. Use 4 or more.
ValueError: context_management.strategy must be one of {'truncate', 'summarize_and_truncate'}, got 'summarize'
ValueError: context_management.strategy must be one of {'truncate', 'summarize_and_truncate'}, got 'summarize'
The two strategies are
truncate and summarize_and_truncate. The same
kind of message comes for mode (token_budget or sliding_window),
threshold_percent (1–100), and a value or tool_offload threshold
that is not positive: tool_offload.threshold_tokens must be positive, got 0.
The order of the names inside {...} can change from run to run.ValueError: Invalid memory mode: last_n. Must be one of {'sliding_window', 'token_budget'}.
ValueError: Invalid memory mode: last_n. Must be one of {'sliding_window', 'token_budget'}.
memory_config.mode is checked when the agent starts — on its first
run(), not when it is built. Use sliding_window or token_budget.The model reads the whole artifact back
The model reads the whole artifact back
An offloaded result is only saved if the model does not load all of it
again. Tell it how to use large results (“search the artifact instead of
reading all of it”), or raise
threshold_bytes and threshold_tokens for
results the model does need whole.The agent forgets something from early in a long run
The agent forgets something from early in a long run
With
truncate, what falls out of the window is gone for that run. Use
summarize_and_truncate, raise value, or have the agent write what it
must keep to a workspace file (Workspace files),
which no trimming touches.Next
Memory
Session history across runs, in the database you choose, with summaries.
Workspace files
Where offloaded results and the agent’s own files live.
Sub-agents
Split a big task so no single context has to hold it all.
Observability
Every compression and offload is in the run’s trajectory.