Stores and Scale
An agent keeps its state in a few stores, each with its own job. By the end of this page you will have a run that pauses in one process and is resumed by another, over one SQLite file, and you will know which store to change when one process becomes several, and several become several servers.Every output on this page is what the code printed when it was run. The
model’s wording and the ids will differ on your run.
Pause in one process, resume in another
Three files: the agent’s definition, a process that runs until a policy asks for approval, and a process that finds the waiting run, approves it and resumes it.1
One definition, shared
pip install "omnicoreagent[postgres]".2
Process one: run until it pauses, then exit
3
Process two: find waiting runs, approve, resume
get_run_trajectory joined the first
process’s trace (suspended) and its own (completed), because both ran in
the same directory and so shared ./workspace/telemetry/. What was left on
disk:
agent.db holds three tables: messages (session history), run_states
(run records) and budget_states (budget counters). The traces are the
workspace/telemetry/ files.
list_runs and get_run both give each approval’s arguments: the call as
the model made it, for the person deciding.Which store holds what
The memory store
Session memory, run records and budget counters share one store:
Records and counters are saved with compare-and-swap on every backend, so two
processes cannot both take over one run or both spend the last of a budget.
Build one router per process and share it between agents: an agent built
without one gets its own in-memory store, which no other agent sees.
History windows and summaries: Memory.
Traces
Each finished run’s trace moves from a write-ahead log to an archive: one body per trace, and an index to find them.
The trace log stays on local disk even when the workspace is S3 or R2. A
process that reads another’s run needs its traces: run both in one directory
(as above), or share the archive.
Processes that run at the same time each keep their own log and share the
archive: every process then answers for every finished trace, and stream
positions stay unique across them.
archive_failures with last_archive_error.
Offloaded payloads
A trace value larger thanmax_payload_bytes (64,000) is truncated, unless
offload_large_payloads is True: then it is kept whole and the trace points
to it (telemetry://payload/...). It is kept in the workspace’s
telemetry/payloads (or beside storage_path), or with
offload_target: "object_storage" in the workspace’s bucket, where every
server can read it. await agent.read_telemetry_payload(reference) reads one.
Background tasks
BackgroundAgentManager(task_store=...) keeps tasks, schedules, runs, attempts
and leases:
Or a dict, such as
{"backend": "sql", "url": "postgresql://..."}. Several
managers on one durable store divide the runs between them, and a run a dead
worker held is taken over when its lease lapses. More in Background
agents.
One process, several processes, several servers
Why more processes: one process runs the runtime’s own work on one core, so
it tops out near 15 runs a second of that work. Concurrent runs in one process
mostly wait on the model, so it serves many at once; past that ceiling, add
processes. A resumed or recovered run can land on any of them, provided they
share the stores above.
Retention
None keeps everything. Each window is applied once, when an agent starts its
first run; await agent.prune_runs() and await agent.prune_telemetry() apply
it now. Session history is not pruned by age (windows and summaries keep it
short), and neither is background run history. Details:
how long a run’s evidence is kept.
Close what you open
Whoever creates a router owns it:
agent.cleanup() leaves it open, because
other agents may share it. Build one per process and close it at shutdown, or
close it in a finally, as ask.py does. Routers that name the same SQL
database share one connection pool: closing one releases its share, and the
pool closes with the last.
Options
Every telemetry key: Telemetry settings.
Every agent key: Agent settings. The manager:
BackgroundAgentManager.
When things go wrong
Process two cannot find the run: LookupError: No run …
Process two cannot find the run: LookupError: No run …
The two processes do not share a memory store. Most often the store is
in memory, sometimes without you asking for it: With
MemoryRouter("sql")
with no DATABASE_URL set falls back to an in-memory store, with a
RuntimeWarning (SQL memory selected but DATABASE_URL is not set; using in_memory (nothing survives a restart)). Print the router to check:DATABASE_URL set it says store=DatabaseMessageStore. The same
holds for redis without REDIS_URL and mongodb without MONGODB_URI.ImportError: SQL database memory requires optional dependency 'sqlalchemy'
ImportError: SQL database memory requires optional dependency 'sqlalchemy'
sql store needs SQLAlchemy, SQLite included.ValueError: Invalid memory store type
ValueError: Invalid memory store type
"sql"; DATABASE_URL picks the
database. A background task store says the same in its own words:
InvalidTaskStoreError: task_store must be 'in_memory', 'sql', 'redis', or 'mongodb'.ValueError: telemetry archive_target must be local or object_storage
ValueError: telemetry archive_target must be local or object_storage
"s3" or "r2" is not a target: set "object_storage" and give the
agent an S3 or R2 workspace.The bodies never reach the bucket
The bodies never reach the bucket
archive_target: "object_storage" with a local workspace keeps the bodies
on local disk, without an error. With telemetry_config["strict"] = True
it is refused:offload_target: "object_storage" behaves the same way.The run is found, but its trajectory is empty
The run is found, but its trajectory is empty
get_run_trajectory shows trace_kept: False for a segment whose trace
this process cannot read: another process wrote it to a different
directory and nothing shares the archive, or it was pruned
(retention_days). Share the archive, or run in the same directory.Next
Durable runs
Records, recovery after a crash, steering, and statuses.
Background agents
Tasks and schedules on a durable task store.
OmniServe
Serve the agent over HTTP, with the same stores.
Agent in a container
Package the agent with its keys and stores.
Memory
Session history, windows and summaries.
Telemetry settings
Every trace, archive and retention key.