Skip to main content

Stores and Scale

An agent keeps its state in a few stores, each with its own job. By the end of this page you will have a run that pauses in one process and is resumed by another, over one SQLite file, and you will know which store to change when one process becomes several, and several become several servers.
Every output on this page is what the code printed when it was run. The model’s wording and the ids will differ on your run.

Pause in one process, resume in another

Three files: the agent’s definition, a process that runs until a policy asks for approval, and a process that finds the waiting run, approves it and resumes it.
1

One definition, shared

SQLite goes through SQLAlchemy: pip install "omnicoreagent[postgres]".
2

Process one: run until it pauses, then exit

3

Process two: find waiting runs, approve, resume

Process two never saw the request; it found the run in the store, with the call’s arguments, and finished it. get_run_trajectory joined the first process’s trace (suspended) and its own (completed), because both ran in the same directory and so shared ./workspace/telemetry/. What was left on disk:
agent.db holds three tables: messages (session history), run_states (run records) and budget_states (budget counters). The traces are the workspace/telemetry/ files.
list_runs and get_run both give each approval’s arguments: the call as the model made it, for the person deciding.

Which store holds what

The memory store

Session memory, run records and budget counters share one store: Records and counters are saved with compare-and-swap on every backend, so two processes cannot both take over one run or both spend the last of a budget. Build one router per process and share it between agents: an agent built without one gets its own in-memory store, which no other agent sees. History windows and summaries: Memory.

Traces

Each finished run’s trace moves from a write-ahead log to an archive: one body per trace, and an index to find them. The trace log stays on local disk even when the workspace is S3 or R2. A process that reads another’s run needs its traces: run both in one directory (as above), or share the archive. Processes that run at the same time each keep their own log and share the archive: every process then answers for every finished trace, and stream positions stay unique across them.
To try it on one machine, stand in a SQLite file for the database and a directory for the bucket, and start two servers at once:
Each server lists both runs, its own from its log and the other’s from the shared archive. What was left on disk:
An archive that cannot be written to never fails a run: the trace stays in the log, and the store counts it in archive_failures with last_archive_error.

Offloaded payloads

A trace value larger than max_payload_bytes (64,000) is truncated, unless offload_large_payloads is True: then it is kept whole and the trace points to it (telemetry://payload/...). It is kept in the workspace’s telemetry/payloads (or beside storage_path), or with offload_target: "object_storage" in the workspace’s bucket, where every server can read it. await agent.read_telemetry_payload(reference) reads one.

Background tasks

BackgroundAgentManager(task_store=...) keeps tasks, schedules, runs, attempts and leases: Or a dict, such as {"backend": "sql", "url": "postgresql://..."}. Several managers on one durable store divide the runs between them, and a run a dead worker held is taken over when its lease lapses. More in Background agents.

One process, several processes, several servers

Why more processes: one process runs the runtime’s own work on one core, so it tops out near 15 runs a second of that work. Concurrent runs in one process mostly wait on the model, so it serves many at once; past that ceiling, add processes. A resumed or recovered run can land on any of them, provided they share the stores above.

Retention

None keeps everything. Each window is applied once, when an agent starts its first run; await agent.prune_runs() and await agent.prune_telemetry() apply it now. Session history is not pruned by age (windows and summaries keep it short), and neither is background run history. Details: how long a run’s evidence is kept.

Close what you open

Whoever creates a router owns it: agent.cleanup() leaves it open, because other agents may share it. Build one per process and close it at shutdown, or close it in a finally, as ask.py does. Routers that name the same SQL database share one connection pool: closing one releases its share, and the pool closes with the last.

Options

Every telemetry key: Telemetry settings. Every agent key: Agent settings. The manager: BackgroundAgentManager.

When things go wrong

The two processes do not share a memory store. Most often the store is in memory, sometimes without you asking for it: MemoryRouter("sql") with no DATABASE_URL set falls back to an in-memory store, with a RuntimeWarning (SQL memory selected but DATABASE_URL is not set; using in_memory (nothing survives a restart)). Print the router to check:
With DATABASE_URL set it says store=DatabaseMessageStore. The same holds for redis without REDIS_URL and mongodb without MONGODB_URI.
Every sql store needs SQLAlchemy, SQLite included.
PostgreSQL, MySQL and SQLite are all "sql"; DATABASE_URL picks the database. A background task store says the same in its own words: InvalidTaskStoreError: task_store must be 'in_memory', 'sql', 'redis', or 'mongodb'.
"s3" or "r2" is not a target: set "object_storage" and give the agent an S3 or R2 workspace.
archive_target: "object_storage" with a local workspace keeps the bodies on local disk, without an error. With telemetry_config["strict"] = True it is refused:
offload_target: "object_storage" behaves the same way.
get_run_trajectory shows trace_kept: False for a segment whose trace this process cannot read: another process wrote it to a different directory and nothing shares the archive, or it was pruned (retention_days). Share the archive, or run in the same directory.

Next

Durable runs

Records, recovery after a crash, steering, and statuses.

Background agents

Tasks and schedules on a durable task store.

OmniServe

Serve the agent over HTTP, with the same stores.

Agent in a container

Package the agent with its keys and stores.

Memory

Session history, windows and summaries.

Telemetry settings

Every trace, archive and retention key.