Skip to main content

Context Engineering

A long task fills the model’s context: a 9,000-token log from one tool, forty steps of tool calls. By the end of this page you will have seen the two things the runtime does about it, both on by default: a large tool result is saved to the workspace and the model gets a preview it can search, and a run’s messages are trimmed — or summarized — before a model call would go over your budget.
Every output on this page is what the code printed when it was run with gpt-5.4-mini. The model’s wording, and which artifact tool it picks, will differ on your run.
There are three layers. The first is covered on the Memory page.

A large tool result, offloaded

This tool returns a 400-line log — about 9,200 tokens. Nothing is configured:
(The saved-to path is shortened here; it is the absolute path of your workspace/artifacts/ folder.) The model was sent 163 tokens instead of 9,201, found the two errors with search_artifact, and the full log stayed on disk. Without the instruction to search, the model on another run called read_artifact and loaded the whole log back into its context — say in the system instruction how you want large results handled.

A long run, trimmed before each model call

Here the budget is made tiny on purpose, so it is crossed on a six-item task: sliding_window counts messages, and the run acts once it holds more than 6.
The total is right although the early prices had left the context: before each of steps 4 to 7, the older messages were replaced by one summary, keeping the system prompt and the 4 most recent messages (8 → system prompt + summary
  • 4). With "strategy": "truncate" they would simply be dropped. Each reduction is in the trajectory, with digests of the messages before and after.

How it works

Offloading. After a tool returns, its result is measured. Over threshold_bytes or threshold_tokens, the full result is written to workspace/artifacts/ (or your S3/R2 workspace) and the model’s observation becomes the preview above: the first max_preview_lines lines, cut to max_preview_tokens. Results of the workspace file tools and the artifact tools themselves are never offloaded. While offloading is on, the model has four tools to reach the rest: Context management. Before every model call, the run’s messages are measured: in token_budget mode, tokens against threshold_percent of value; in sliding_window mode, the number of messages (not counting the system prompt) against value. Over it, the messages between the system prompt and the preserve_recent latest are dropped (truncate), or summarized by one extra model call into a single message (summarize_and_truncate). A tool call and its results are kept or dropped together. If the summary call fails, the run truncates instead and carries on. Set value below your model’s real context window: the check happens before the call, so the provider never sees an oversized prompt.

Options

A partial dictionary is merged with the defaults. With enable_subagents, context management and offloading are switched on for the lead and every worker, whatever you set, keeping your other values. Every setting is in the agent settings reference.

When things go wrong

The latest four messages are always kept, so the model never loses the call it just made and its result. Use 4 or more.
The two strategies are truncate and summarize_and_truncate. The same kind of message comes for mode (token_budget or sliding_window), threshold_percent (1–100), and a value or tool_offload threshold that is not positive: tool_offload.threshold_tokens must be positive, got 0. The order of the names inside {...} can change from run to run.
memory_config.mode is checked when the agent starts — on its first run(), not when it is built. Use sliding_window or token_budget.
An offloaded result is only saved if the model does not load all of it again. Tell it how to use large results (“search the artifact instead of reading all of it”), or raise threshold_bytes and threshold_tokens for results the model does need whole.
With truncate, what falls out of the window is gone for that run. Use summarize_and_truncate, raise value, or have the agent write what it must keep to a workspace file (Workspace files), which no trimming touches.

Next

Memory

Session history across runs, in the database you choose, with summaries.

Workspace files

Where offloaded results and the agent’s own files live.

Sub-agents

Split a big task so no single context has to hold it all.

Observability

Every compression and offload is in the run’s trajectory.