Skip to main content

OmniServe

OmniServe puts one agent behind an HTTP API: a run as JSON or as a live stream, the run’s record and trajectory afterwards, a person’s approval mid-run, and background tasks, with auth, rate limits and a request timeout. One command, omniserve run --agent my_agent.py, or OmniServe(agent).start() in your own script.
Every output on this page is what the server returned when it was run. The model’s wording, the ids and the timings will differ on your run. Each route’s fields are in the generated HTTP API reference, under HTTP API (OmniServe) in the navigation.

Serve an agent

Save this as my_agent.py. OmniServe takes the module’s agent, or calls its create_agent() when there is no agent:
Serve it (on port 8771 here; the default is 8000):
Wait for /ready, then ask. Loading the model library can take a minute on a busy machine; the agent_name in the answer confirms it is your server:
Two model calls: one to call get_order, one to answer. session_id is the conversation: the next request with "session_id": "alice" remembers this one. Leave it out and each request starts a new session (its id is in the response). Keep run_id: everything about this run is under /runs/{run_id}. Interactive docs for your server are at /docs (Swagger) and /redoc.
No agent file yet? omniserve quickstart serves a plain agent, by default --provider openai --model gpt-5.4-mini. Asked to “Say hello in five words, in plain text”, it answered success | Hello there, hope you’re well.

Stream a run

POST /run takes the same body and answers with Server-Sent Events: the run’s telemetry as it happens, the answer as it is written (text_delta), and a final complete with the same fields as /run/sync. A small client:
Every tool call, whatever its kind, arrives as tool_requested and ends with tool_observation. What ran in between is named by its kind: tool_call and tool_result for your own tools, workspace_read and workspace_write for the workspace file tools. A client that listens only for tool_result misses every file the agent reads or writes. A stream always opens with event: session ("status": "started") and closes with event: session ("status": "ended"). In between, for a run with no tool call, curl -N saw these events, in order:
  • text_delta is live only: it is never stored, so a replay does not have it.
  • Check status on complete before treating the run as a success (it can be error, awaiting_approval or awaiting_budget).
  • A failure arrives as event: error with error, session_id and run_id.
  • Closing the connection cancels the run, whatever the request timeout (below).

Look at a run afterwards

The record never includes the saved conversation. For the whole story, GET /runs/{run_id}/trajectory returns segments (one per stretch of a run that paused and resumed), tool_calls, approvals and totals; here, among the totals:
The stored events of one run, inside a shared session, as JSON, as a compact trace, or replayed and followed over SSE:
Use trace_id for one exact trace, run_id for everything one execution did. /telemetry/events returns at most 200 events by default and /telemetry/traces 100; ask for fewer with limit. More in Observability.

A person in the loop

With governance on, a rule that says ask pauses the run and /run/sync returns awaiting_approval. This agent (refund_agent.py) asks before any refund, and has a slow report tool:
Nothing was refunded. A person decides, then the run resumes:
Send "decision": "deny" with a note the model will read, or approve with "arguments": {...} to run an edited call. Everything else about approvals is on Approvals. A run that ran out of budget pauses the same way with awaiting_budget: read it with GET /runs/{run_id}/budget, decide with POST /runs/{run_id}/budget ({"decision": "grant", "approver": "alice", "amount": 5}), then resume. A running run can be steered (POST /runs/{run_id}/steer, {"message": "..."}) or stopped at its next step (POST /runs/{run_id}/interrupt, then resume); see Durable runs.

Background tasks

The served agent is registered for background runs at startup, as agent default. On the same server:
The second call waited for the run and returned it (abridged):
With {"wait": false} (the default) it returns the queued run at once. A schedule is manual, once, interval or cron. Then:
The run’s events were background_run_queued, background_run_claimed, background_run_started, background_run_completed; while a run is running it also records a background_run_heartbeat at intervals (a longer run has several). Its workspace held run.json and events.jsonl. Under governance the agent’s policy also decides what an operator asks here: creating, changing, pausing, resuming or deleting a task is background.task.<action>, starting or cancelling a run is background.run.<action>. The policy above allows create and start only, so a DELETE is refused with a 403 (below). Allow them for the operators you trust (policy reference).

A task store that survives restarts

Tasks, runs, attempts, leases and retries live in the background task store, which is not conversation memory (that stays in the agent’s MemoryRouter). The default is in-memory: fine for development, gone on restart. Pick one backend for a deployment:
  • Use SQL for local durability with SQLite, or PostgreSQL or MySQL when several processes share one store.
  • Use Redis when your deployment already operates Redis with persistence and no eviction for task-store keys.
  • Use MongoDB when MongoDB is your durable operational store (it writes with majority concern).
With a durable store, a queued run survives: OmniServe can queue a run, stop, start again with the same task store, and the new worker claims and completes it. The server in the next section kept a cron task across a restart.

Serve from your own code

OmniServe(agent, config=...) gives you the same server from a script. server.py, with a path prefix, a token and a SQLite task store:
(The last line is the status’s task_id, schedule.expression and schedule_state.next_due_at, after the restart.) The prefix moves the agent’s routes; /docs, /redoc, /openapi.json and /prometheus stay at the root. server.start() blocks; await server.start_async() runs inside your event loop, and server.get_app() returns the FastAPI app to mount elsewhere or test.
OMNICOREAGENT_SERVE_* and OMNICOREAGENT_BACKGROUND_* environment variables always override what the code or the CLI flags set. A stray OMNICOREAGENT_SERVE_PORT in the environment wins over port=8776.

Your own routes beside the agent’s

The agent file (or OmniServe(..., routers=[...], public_paths=[...])) can add an application’s own pages and endpoints, served on the same origin behind the same middleware:
Everything else these routers serve needs the bearer token when auth is on. apps/steward is a complete application served this way.

How it works

1

Startup

The server connects the agent’s MCP servers, warms the model client, starts the background manager (registering the agent as default) and its worker. The server answers no request, /health included, until startup has finished, which can take up to a minute on a busy machine. After that, /health says the process is up, and /ready says true while the agent is initialized and its MCP servers are connected.
2

Each request

Middleware, outermost first: CORS, request logging, the rate limit, the bearer token, the request timeout, and a last catch that turns an uncaught error into a JSON 500. A policy refusal becomes a 403.
3

A run

/run/sync and /run call agent.run(...) with the session and a new run_id, and record a request trace next to the agent’s own. The run is durable like any other: its record lives in the agent’s memory store and its traces under the workspace (./workspace/telemetry/).
4

Shutdown

The background manager stops, then agent.cleanup() closes its connections.
One server process serves one agent instance, so workers must be 1. To scale, run several processes (or containers) behind a load balancer, sharing a durable memory store and task store; the rate limit is counted per process. Stores and scale has the details.

Routes

Every route, its body and its response fields are in the HTTP API reference. Paths below are relative to api_prefix (empty by default).

Runs

Sessions and events

Telemetry

The trace detail routes return 404 when no trace matches.

Health and operations

/ready’s mcp_servers gives each configured MCP server’s status (connected, disconnected, failed, not_connected) and last error, so one failed server among several is visible. An agent without MCP servers does not need one to be ready.

Background

Mounted unless background is turned off (OMNICOREAGENT_BACKGROUND_ENABLED=false).

Authentication

Off by default. Turn it on with a token:
Then every route but /health, /ready, /prometheus, /docs, /redoc, /openapi.json and your public_paths needs the header:
There is one token per server; for per-user identity, put OmniServe behind a gateway that authenticates users. CORS is open (*) by default, without credentials; set OMNICOREAGENT_SERVE_CORS_ORIGINS to your front end’s origin (credentials are sent only to origins listed there). Behind a proxy, list its address in OMNICOREAGENT_SERVE_TRUSTED_PROXIES so the rate limit sees each client’s address, not the proxy’s.

Configuration

Flags on omniserve run: --agent, --host, --port, --workers (must be 1), --auth-token, --rate-limit N (requests per minute per client IP), --cors-origins, --no-docs. Everything else is an environment variable or an OmniServeConfig field of the same name, lower-cased without the prefix (OMNICOREAGENT_SERVE_REQUEST_TIMEOUT is request_timeout). With none set, the defaults serve on 0.0.0.0:8000 with no auth, and the background API and its worker run with an in-memory task store. A durable task store; pick one backend:
SQL needs omnicoreagent[postgres] (SQLAlchemy, for SQLite too), Redis omnicoreagent[redis], MongoDB omnicoreagent[mongodb]. omniserve config --show prints the configuration the environment gives; omniserve config --env-example prints a template. Every option is in the CLI reference and configuration.

The request timeout

OMNICOREAGENT_SERVE_REQUEST_TIMEOUT (300 seconds) bounds each request:
  • /run/sync and /runs/{run_id}/resume answer 504 when the run takes longer. The run is stopped; its record and trace remain.
  • /run (SSE) sends event: error with "error": "Request timed out" and closes the stream.
  • /background/tasks/{task_id}/run with "wait": true waits a little less than the timeout, then answers 504 with the run_id. The run is not stopped: it carries on in the background. A run that pauses for a person or a budget is answered at once, 200 with awaiting_approval or awaiting_budget.
For work longer than a request should last, use a background task.

Deploy with Docker

It also writes a .dockerignore in the directory you run it from (the build context), if there is none, that keeps credential files (.env, .env.*, *.pem, *.key, id_rsa*, .netrc, credentials*.json and the like, in any subfolder) and the configured workspace folder (at the root only) out of the image (the image copies the build context); an existing one that does not exclude **/.env is named. The generated Dockerfile installs omnicoreagent[serve], pinned to the release that generated it (==0.5.0), on python:3.12-slim, copies the build context to /app, runs as a non-root user, checks /health, and keeps the workspace in /tmp/workspace (gone with the container unless you mount a volume). Its HEALTHCHECK allows a start period of 10 seconds (--start-period=10s); a slow host may need longer before /health first answers, so raise it in the Dockerfile or with docker run --health-start-period=60s. Add the extras your agent needs ([serve,postgres] for a SQL memory store): a comment above the install line in the Dockerfile says so. A persistent volume, a store for runs and readiness checks are on Running the agent in a container.

When things go wrong

Auth is on and the request had no Authorization: Bearer ... header, or the wrong token:
The agent’s policy refused what the request asked, here a task DELETE the policy does not allow:
The body names the refused capability, the reason_code and the matched_rule_ids (none here: no rule matched). Add an allow rule for that capability (background.task.delete), for the operators you trust.An ask nobody can answer inside the request is refused the same way, and the body names the capability and the rule. Under interactive-dev, every background request asks:
An ask outranks an allow, and interactive-dev asks about anything no rule matches: for an agent served to operators you trust, take the ask rule out and allow the capability.
With both, POST /background/tasks returns 200.
A run’s record lives in the agent’s memory store: with the default in-memory store, a restart forgets every run. Use a durable memory store.
The request does not fit the run’s state:
An approval reads already approved (or denied) once decided, and already used once a resume has applied the decision. Decide every approval before resuming; send "replace": true to overwrite a task.
/runs/{run_id}/steer also answers 422 when the injection guardrail blocks the message.
With --rate-limit 3, the fourth request in a minute:
Every request but the public paths counts, rejected ones (a 401) included.
A /run/sync under a 1-second timeout, with a model slower than that:
The run is stopped and its record says timeout (GET /runs/{run_id}); the run_id finds its trajectory. POST /runs/{run_id}/resume answers a timeout the same way.A background run waited on with "wait": true:
That run finished later (GET /background/runs/{run_id} said completed). Raise OMNICOREAGENT_SERVE_REQUEST_TIMEOUT, or queue long work without waiting.
The configuration is checked before the agent is loaded:
A Redis or MongoDB task store without its address is refused the same way: Error: Invalid OmniServe config: Redis background task store requires OMNICOREAGENT_BACKGROUND_TASK_STORE_URL (..._URI for MongoDB).
An exception while importing the file is reported as Error loading agent file: ....
The model key is read when the agent is built; without it the server exits:
Set LLM_API_KEY in the server’s environment. A key the provider rejects starts fine, and each run ends with status error, termination_reason provider_error.
Another process holds the port. The port is checked before the agent is loaded, so this comes at once:
Pick another --port, or stop what holds it.

Next

Running the agent in a container

An image, a volume, a store for runs, health checks.

Durable runs

Run records, pause and resume, recovery after a crash.

Stores and scale

Memory and task stores, and many processes sharing them.

Background agents

Schedules, retries, leases and the task store.

Approvals

Everything about a person in the loop.

HTTP API reference

Every route, body and response field.