Skip to main content

Background Agents

BackgroundAgentManager runs your agents away from any request: on demand, once at a set time, every N seconds, or on a cron schedule. Every run is recorded (its status, attempts, events and answer) in a task store, and with a durable store a run queued by one process is finished by another.
Every output on this page is what the code printed when it was run. The model’s wording, the ids and the times will differ on your run.

A first background run

  • register_agent names an agent; register_task says which agent runs which query, and when.
  • run_now(..., wait=True) queued the run and, as no worker was started, ran it here and returned it finished.
  • Each run gets its own folder in the workspace: the query reaches the agent with the folder’s path and asks for the final result in output.md, notes in scratchpad/, logs in logs/. The manager writes run.json (the run) and events.jsonl (its events) beside it. The model chose to call write_file here; ask for “no files” in the instruction if you only want the answer.
  • result_preview is the agent’s answer, cut at 1,000 characters.
  • The agent’s run id is the background run id, so agent.get_run_trajectory(run_id) shows what the model did, step by step (Durable runs).

Schedules

It was 23:03:25 UTC; 09:00 on Monday in Lagos is 08:00 UTC. Schedules run only while a manager is started: start() runs the worker loop until shutdown(). Every schedule but manual also takes start_at, end_at, jitter_seconds (a fixed offset, the same for every worker) and misfire_policy: what happens to times missed while no manager was running. run_once (the default) queues one run for the missed times, then continues from now; skip_missed runs nothing for a time missed by more than a minute and waits for the next; queue_all queues a run for each one missed. Want an interval task to run at once too? Call run_now after registering. pause_task stops a schedule (runs already queued go on) and resume_task starts it again; run_now still works on a paused task.

Overlap: a task that is due while it is running

A run waiting for a person (below) still holds its task: under skip_if_running, new runs are skipped until it is decided.
One manager runs one run at a time. allow_parallel lets a task’s runs go side by side only when several workers share the store (below).

Retries and timeouts

  • timeout_seconds on the task is each attempt’s deadline (none by default). An attempt past it ends timeout.
  • timeout_seconds on run_now (and wait_for_run, run_until_terminal) is only how long you wait; the run is returned as it is then.
  • An attempt that raises, or whose agent run ends with an error, or that times out, is retried while retries are left: the run goes retrying, then queued again after the delay. When none are left it ends failed or timeout, with the error. A retry is a new attempt from the first step, not a resume: the earlier attempt’s tool calls can run again, so retry only tasks whose tools are safe to repeat (see what is guaranteed).
run_now(wait=True) returns when the run ends, at timeout_seconds, or when an attempt failed and its retry is scheduled for later: the run comes back queued, with its queued_at, rather than holding you through the whole backoff. A started manager runs the retry when it is due; on a manager that is not started, nothing does, so start() it when a task retries with a delay.

A run that waits for a person

A background agent under governance can pause for an approval or a budget. The run then waits, holding no worker, until someone decides on the agent and resume_run queues it again. Its next attempt continues the same run; it does not start over.
A budget works the same way. With "budgets": {"request": [{"meter": "model_cost_usd", "limit": 0.0005}]} and max_tokens 300, a one-line tagline task printed:
  • Decide with the agent: resolve_approval, grant_budget or deny_budget. After a denial, resume_run lets the run end cleanly.
  • cancel_run(run_id) ends a waiting run instead: it is cancelled, and so is the agent’s own record of it.
  • With the manager started, resume_run is enough: a worker picks the run up. run_until_terminal runs it here, for a manager that is not started.
  • A scheduled task’s runs share one session by default (session_policy {"mode": "task"}), so a session budget with the default total window adds up across runs and, once spent, pauses (by default) every run after it. For a limit per run, use a request budget, or give each run its own session with session_policy={"mode": "run"}.
  • In another process, build the agent with the same memory store to decide (Approvals: from another process).

Reading a run

Events come from the manager process that emitted them, or from the run’s events.jsonl in the workspace, so a finished run’s events can still be read after a restart. A run’s statuses are queued, claimed, running, retrying; it ends completed, failed, cancelled, timeout or skipped, or waits in awaiting_approval or awaiting_budget. The events page lists every event name.

Task stores

The task store keeps agents’ specs, tasks, schedule state, runs, attempts, leases and cancel flags. It is not the agent’s memory: conversations and run records stay in the agent’s memory store. Use SQL for local durability with SQLite, or PostgreSQL or MySQL when several processes share one store. Use Redis when your deployment already operates Redis, with persistence on and no eviction of task-store keys. Use MongoDB when MongoDB is your durable operational store. A dict also takes prefix (tables or keys, so deployments can share one database) and connect_timeout. Background run history is not pruned.

A queued run survives the process

With a durable store, a queued run left by one process is finished by another over the same task store. Two scripts share one definition:
1

Process one queues a run and exits

2

Process two finds it and runs it

The store keeps the agent’s spec (model, instruction, settings, MCP servers), so replace=True is needed to register it again; a worker that never registers it rebuilds the agent from that spec, without local tools. Register in every process an agent that has local tools.

Several workers, one store

Each started manager is a worker. Workers on one store claim runs with a lease (30 seconds, renewed as the run goes), so no run is taken twice. Here four runs of an allow_parallel task, queued into Redis, then two worker processes started side by side:
Each worker took two runs, side by side. One run asked for the text instead of translating it: runs of one task share a memory session by default (session_policy {"mode": "task"}), so parallel runs read each other’s turns. For runs that must not see each other, give the task session_policy={"mode": "run"}. If a worker dies, its lease runs out and another worker takes the run back, continuing from the agent’s checkpoint; list_attempts shows the lost attempt with the reason lease_expired and the attempt that took over with recovery. More in Stores and scale.

Starting and shutting down

  • start() prepares the store, loads each registered agent’s model client (so the first run does not pay for that import), and starts the worker loop. Each turn it takes back runs whose lease ran out, queues due schedules, and runs one queued run.
  • Without start(), nothing is scheduled; run_now(wait=True) and run_until_terminal run a queued run in your process.
  • shutdown() stops the loop, stops the runs it is executing, and closes the store. Each stopped attempt is recorded as failed with the error worker shutdown. A task with retries left puts the run back in the queue (queued, for the next worker); otherwise the run ends failed with that error. It is not cancelled: only cancel_run cancels. Call shutdown() before the process exits.
With a governance_engine on the manager, a task is bound to the policy it was created under. When the policy changes, the task can still be paused or deleted but not run: at its next due time its schedule is paused, and get_task_status shows why in schedule_state.paused_reason. Any change to the policy counts, a budget number included. Register the task again with replace=True (over HTTP, POST /background/tasks with "replace": true): the new policy decides whether it may (it needs background.task.update), binds it to that policy, and its schedule runs again. A task a person paused stays paused.

Over HTTP

OmniServe registers the served agent with a background manager and exposes it under /background/tasks/:
Every route, and the task-store environment variables, are on the OmniServe page.

Options

register_task takes: BackgroundAgentManager(...) takes task_store, memory_router, workspace, telemetry_store, governance_engine, worker_id and lease_seconds (30). Change a task later with update_task(task_id, patch). Every method is in the BackgroundAgentManager reference.

When things go wrong

The task names an agent not registered in this store:
Call register_agent first, with the same id.
The id is already in the store; with a durable store, from an earlier process:
Pass replace=True to overwrite it.
run_now on a task that does not exist, or has enabled=False:
Each schedule type checks its fields:
Time zones are IANA names: Africa/Lagos, not Lagos. Ids are checked too: task_id must contain only letters, numbers, '_', '-', and '.'.
resume_run is for a waiting run:
To run a finished task again, call run_now.
An unknown backend, or one missing its address:
PostgreSQL is "sql" with a postgresql:// URL.
The default overlap_policy is skip_if_running, and a run of the task was still active, or waiting for a person. Use queue_next, or decide the waiting run.
Its first run is one interval after it was registered, and only while a manager is started. get_task_status(task_id)["schedule_state"]["next_due_at"] says when.

Next

BackgroundAgentManager reference

Every method, run status and overlap policy.

Durable runs

Run records, trajectories, resuming after a crash.

Stores and scale

Choosing stores and running several workers.

OmniServe

Background tasks over HTTP.

Agent in a container

Package the agent and its worker.

Approvals

Deciding what a paused run asked for.