Background Agents
BackgroundAgentManager runs your agents away from any request: on demand,
once at a set time, every N seconds, or on a cron schedule. Every run is
recorded (its status, attempts, events and answer) in a task store, and
with a durable store a run queued by one process is finished by another.
Every output on this page is what the code printed when it was run. The
model’s wording, the ids and the times will differ on your run.
A first background run
register_agentnames an agent;register_tasksays which agent runs which query, and when.run_now(..., wait=True)queued the run and, as no worker was started, ran it here and returned it finished.- Each run gets its own folder in the workspace:
the query reaches the agent with the folder’s path and asks for the final
result in
output.md, notes inscratchpad/, logs inlogs/. The manager writesrun.json(the run) andevents.jsonl(its events) beside it. The model chose to callwrite_filehere; ask for “no files” in the instruction if you only want the answer. result_previewis the agent’s answer, cut at 1,000 characters.- The agent’s run id is the background run id, so
agent.get_run_trajectory(run_id)shows what the model did, step by step (Durable runs).
Schedules
start() runs the worker loop until
shutdown().
Every schedule but
manual also takes start_at, end_at,
jitter_seconds (a fixed offset, the same for every worker) and
misfire_policy: what happens to times missed while no manager was running.
run_once (the default) queues one run for the missed times, then continues
from now; skip_missed runs nothing for a time missed by more than a minute
and waits for the next; queue_all queues a run for each one missed.
Want an interval task to run at once too? Call run_now after registering.
pause_task stops a schedule (runs already queued go on) and resume_task
starts it again; run_now still works on a paused task.
Overlap: a task that is due while it is running
A run waiting for a person (below) still holds its task: under
skip_if_running, new runs are skipped until it is decided.
One manager runs one run at a time.
allow_parallel lets a task’s runs
go side by side only when several workers share the store
(below).Retries and timeouts
timeout_secondson the task is each attempt’s deadline (none by default). An attempt past it endstimeout.timeout_secondsonrun_now(andwait_for_run,run_until_terminal) is only how long you wait; the run is returned as it is then.- An attempt that raises, or whose agent run ends with an error, or that
times out, is retried while retries are left: the run goes
retrying, thenqueuedagain after the delay. When none are left it endsfailedortimeout, with theerror. A retry is a new attempt from the first step, not a resume: the earlier attempt’s tool calls can run again, so retry only tasks whose tools are safe to repeat (see what is guaranteed).
run_now(wait=True) returns when the run ends, at timeout_seconds, or when
an attempt failed and its retry is scheduled for later: the run comes back
queued, with its queued_at, rather than holding you through the whole
backoff. A started manager runs the retry when it is due; on a manager that
is not started, nothing does, so start() it when a task retries with a
delay.A run that waits for a person
A background agent under governance can pause for an approval or a budget. The run then waits, holding no worker, until someone decides on the agent andresume_run queues it again. Its
next attempt continues the same run; it does not start over.
"budgets": {"request": [{"meter": "model_cost_usd", "limit": 0.0005}]} and
max_tokens 300, a one-line tagline task printed:
- Decide with the agent:
resolve_approval,grant_budgetordeny_budget. After a denial,resume_runlets the run end cleanly. cancel_run(run_id)ends a waiting run instead: it iscancelled, and so is the agent’s own record of it.- With the manager started,
resume_runis enough: a worker picks the run up.run_until_terminalruns it here, for a manager that is not started. - A scheduled task’s runs share one session by default (
session_policy{"mode": "task"}), so asessionbudget with the defaulttotalwindow adds up across runs and, once spent, pauses (by default) every run after it. For a limit per run, use arequestbudget, or give each run its own session withsession_policy={"mode": "run"}. - In another process, build the agent with the same memory store to decide (Approvals: from another process).
Reading a run
Events come from the manager process that emitted them, or from the run’s
events.jsonl in the workspace, so a finished run’s events can still be read
after a restart. A run’s statuses are queued, claimed, running,
retrying; it ends completed, failed, cancelled, timeout or
skipped, or waits in awaiting_approval or awaiting_budget. The
events page lists every event name.
Task stores
The task store keeps agents’ specs, tasks, schedule state, runs, attempts, leases and cancel flags. It is not the agent’s memory: conversations and run records stay in the agent’s memory store.
Use SQL for local durability with SQLite, or
PostgreSQL or MySQL when several processes share one store. Use Redis when your deployment already operates
Redis, with persistence on and no eviction of task-store keys. Use MongoDB
when MongoDB is your durable operational store. A dict also takes
prefix
(tables or keys, so deployments can share one database) and
connect_timeout. Background run history is not pruned.
A queued run survives the process
With a durable store, a queued run left by one process is finished by another over the same task store. Two scripts share one definition:1
Process one queues a run and exits
2
Process two finds it and runs it
replace=True is needed to register it again; a worker that
never registers it rebuilds the agent from that spec, without local tools.
Register in every process an agent that has local tools.
Several workers, one store
Each started manager is a worker. Workers on one store claim runs with a lease (30 seconds, renewed as the run goes), so no run is taken twice. Here four runs of anallow_parallel task, queued into Redis, then two worker
processes started side by side:
session_policy {"mode": "task"}), so parallel runs read each other’s
turns. For runs that must not see each other, give the task
session_policy={"mode": "run"}. If a worker dies, its
lease runs out and another worker takes the run back, continuing from the
agent’s checkpoint; list_attempts shows the lost attempt with the reason
lease_expired and the attempt that took over with recovery. More in Stores and scale.
Starting and shutting down
start()prepares the store, loads each registered agent’s model client (so the first run does not pay for that import), and starts the worker loop. Each turn it takes back runs whose lease ran out, queues due schedules, and runs one queued run.- Without
start(), nothing is scheduled;run_now(wait=True)andrun_until_terminalrun a queued run in your process. shutdown()stops the loop, stops the runs it is executing, and closes the store. Each stopped attempt is recorded asfailedwith the errorworker shutdown. A task with retries left puts the run back in the queue (queued, for the next worker); otherwise the run endsfailedwith that error. It is notcancelled: onlycancel_runcancels. Callshutdown()before the process exits.
governance_engine on the manager, a task is bound to the policy it
was created under. When the policy changes, the task can still be paused or
deleted but not run: at its next due time its schedule is paused, and
get_task_status shows why in schedule_state.paused_reason. Any change to
the policy counts, a budget number included. Register the task again with
replace=True (over HTTP, POST /background/tasks with "replace": true):
the new policy decides whether it may (it needs background.task.update),
binds it to that policy, and its schedule runs again. A task a person paused
stays paused.
Over HTTP
OmniServe registers the served agent with a background manager and exposes it under/background/tasks/:
Options
register_task takes:
BackgroundAgentManager(...) takes task_store, memory_router,
workspace, telemetry_store, governance_engine, worker_id and
lease_seconds (30). Change a task later with update_task(task_id, patch).
Every method is in the BackgroundAgentManager reference.
When things go wrong
AgentNotFoundError: Agent not found
AgentNotFoundError: Agent not found
The task names an agent not registered in this store:Call
register_agent first, with the same id.AgentAlreadyRegisteredError / TaskAlreadyRegisteredError
AgentAlreadyRegisteredError / TaskAlreadyRegisteredError
The id is already in the store; with a durable store, from an earlier
process:Pass
replace=True to overwrite it.TaskNotFoundError: Task not found
TaskNotFoundError: Task not found
run_now on a task that does not exist, or has enabled=False:ValidationError: a schedule that cannot run
ValidationError: a schedule that cannot run
Each schedule type checks its fields:Time zones are IANA names:
Africa/Lagos, not Lagos. Ids are checked
too: task_id must contain only letters, numbers, '_', '-', and '.'.ValueError: only a run in awaiting_approval or awaiting_budget can be resumed
ValueError: only a run in awaiting_approval or awaiting_budget can be resumed
resume_run is for a waiting run:run_now.InvalidTaskStoreError
InvalidTaskStoreError
An unknown backend, or one missing its address:PostgreSQL is
"sql" with a postgresql:// URL.My second run came back skipped
My second run came back skipped
The default
overlap_policy is skip_if_running, and a run of the task
was still active, or waiting for a person. Use queue_next, or decide
the waiting run.An interval task has not run yet
An interval task has not run yet
Its first run is one interval after it was registered, and only while a
manager is started.
get_task_status(task_id)["schedule_state"]["next_due_at"]
says when.Next
BackgroundAgentManager reference
Every method, run status and overlap policy.
Durable runs
Run records, trajectories, resuming after a crash.
Stores and scale
Choosing stores and running several workers.
OmniServe
Background tasks over HTTP.
Agent in a container
Package the agent and its worker.
Approvals
Deciding what a paused run asked for.