OmniServe
OmniServe puts one agent behind an HTTP API: a run as JSON or as a live stream, the run’s record and trajectory afterwards, a person’s approval mid-run, and background tasks, with auth, rate limits and a request timeout. One command,omniserve run --agent my_agent.py, or OmniServe(agent).start()
in your own script.
Every output on this page is what the server returned when it was run. The
model’s wording, the ids and the timings will differ on your run. Each
route’s fields are in the generated HTTP API reference,
under HTTP API (OmniServe) in the navigation.
Serve an agent
Save this asmy_agent.py. OmniServe takes the module’s agent, or calls
its create_agent() when there is no agent:
/ready, then ask. Loading the model library can take a minute on
a busy machine; the agent_name in the answer confirms it is your server:
get_order, one to answer. session_id is the
conversation: the next request with "session_id": "alice" remembers this
one. Leave it out and each request starts a new session (its id is in the
response). Keep run_id: everything about this run is under
/runs/{run_id}.
Interactive docs for your server are at /docs (Swagger) and /redoc.
Stream a run
POST /run takes the same body and answers with Server-Sent Events: the
run’s telemetry as it happens, the answer as it is written (text_delta),
and a final complete with the same fields as /run/sync. A small client:
tool_requested and ends with
tool_observation. What ran in between is named by its kind: tool_call and
tool_result for your own tools, workspace_read and workspace_write for the
workspace file tools. A client that listens only for tool_result misses every
file the agent reads or writes.
A stream always opens with event: session ("status": "started") and
closes with event: session ("status": "ended"). In between, for a run
with no tool call, curl -N saw these events, in order:
text_deltais live only: it is never stored, so a replay does not have it.- Check
statusoncompletebefore treating the run as a success (it can beerror,awaiting_approvalorawaiting_budget). - A failure arrives as
event: errorwitherror,session_idandrun_id. - Closing the connection cancels the run, whatever the request timeout (below).
Look at a run afterwards
GET /runs/{run_id}/trajectory returns segments (one per stretch of a run
that paused and resumed), tool_calls, approvals and totals; here,
among the totals:
trace_id for one exact trace, run_id for everything one execution
did. /telemetry/events returns at most 200 events by default and
/telemetry/traces 100; ask for fewer with limit. More in
Observability.
A person in the loop
With governance on, a rule that saysask
pauses the run and /run/sync returns awaiting_approval. This agent
(refund_agent.py) asks before any refund, and has a slow report tool:
"decision": "deny" with a note the model will read, or approve with
"arguments": {...} to run an edited call. Everything else about approvals
is on Approvals.
A run that ran out of budget pauses the same
way with awaiting_budget: read it with GET /runs/{run_id}/budget, decide
with POST /runs/{run_id}/budget ({"decision": "grant", "approver": "alice", "amount": 5}),
then resume. A running run can be steered (POST /runs/{run_id}/steer,
{"message": "..."}) or stopped at its next step (POST /runs/{run_id}/interrupt,
then resume); see Durable runs.
Background tasks
The served agent is registered for background runs at startup, as agentdefault. On the same server:
{"wait": false} (the default) it returns the queued run at once. A
schedule is manual, once, interval or cron. Then:
background_run_queued, background_run_claimed,
background_run_started, background_run_completed; while a run is running
it also records a background_run_heartbeat at intervals (a longer run has
several). Its workspace held
run.json and events.jsonl.
Under governance the agent’s policy also decides what an operator asks here:
creating, changing, pausing, resuming or deleting a task is
background.task.<action>, starting or cancelling a run is
background.run.<action>. The policy above allows create and start only,
so a DELETE is refused with a 403 (below). Allow them for the operators
you trust (policy reference).
A task store that survives restarts
Tasks, runs, attempts, leases and retries live in the background task store, which is not conversation memory (that stays in the agent’sMemoryRouter). The default is in-memory: fine for development, gone on
restart. Pick one backend for a deployment:
- Use SQL for local durability with SQLite, or PostgreSQL or MySQL when several processes share one store.
- Use Redis when your deployment already operates Redis with persistence and no eviction for task-store keys.
- Use MongoDB when MongoDB is your durable operational store (it writes with majority concern).
Serve from your own code
OmniServe(agent, config=...) gives you the same server from a script.
server.py, with a path prefix, a token and a SQLite task store:
task_id, schedule.expression and
schedule_state.next_due_at, after the restart.) The prefix moves the
agent’s routes; /docs, /redoc, /openapi.json and /prometheus stay at
the root. server.start() blocks; await server.start_async() runs inside
your event loop, and server.get_app() returns the FastAPI app to mount
elsewhere or test.
Your own routes beside the agent’s
The agent file (orOmniServe(..., routers=[...], public_paths=[...])) can
add an application’s own pages and endpoints, served on the same origin
behind the same middleware:
How it works
1
Startup
The server connects the agent’s MCP servers, warms the model client,
starts the background manager (registering the agent as
default) and
its worker. The server answers no request, /health included, until
startup has finished, which can take up to a minute on a busy machine.
After that, /health says the process is up, and /ready says true
while the agent is initialized and its MCP servers are connected.2
Each request
Middleware, outermost first: CORS, request logging, the rate limit, the
bearer token, the request timeout, and a last catch that turns an
uncaught error into a JSON
500. A policy refusal becomes a 403.3
A run
/run/sync and /run call agent.run(...) with the session and a new
run_id, and record a request trace next to the agent’s own. The run
is durable like any other: its record lives in the agent’s memory store
and its traces under the workspace (./workspace/telemetry/).4
Shutdown
The background manager stops, then
agent.cleanup() closes its
connections.workers must be 1.
To scale, run several processes (or containers) behind a load balancer,
sharing a durable memory store and task store; the rate limit is counted per
process. Stores and scale has the
details.
Routes
Every route, its body and its response fields are in the HTTP API reference. Paths below are relative toapi_prefix (empty by default).
Runs
Sessions and events
Telemetry
The trace detail routes return
404 when no trace matches.
Health and operations
/ready’s mcp_servers gives each configured MCP server’s status
(connected, disconnected, failed, not_connected) and last error, so
one failed server among several is visible. An agent without MCP servers
does not need one to be ready.
Background
Mounted unless background is turned off (OMNICOREAGENT_BACKGROUND_ENABLED=false).
Authentication
Off by default. Turn it on with a token:/health, /ready, /prometheus, /docs, /redoc,
/openapi.json and your public_paths needs the header:
*) by default, without
credentials; set OMNICOREAGENT_SERVE_CORS_ORIGINS to your front end’s origin
(credentials are sent only to origins listed there). Behind a proxy, list its
address in OMNICOREAGENT_SERVE_TRUSTED_PROXIES so the rate limit sees each
client’s address, not the proxy’s.
Configuration
Flags onomniserve run: --agent, --host, --port, --workers (must
be 1), --auth-token, --rate-limit N (requests per minute per client IP),
--cors-origins, --no-docs. Everything else is an environment variable
or an OmniServeConfig field of the same name, lower-cased without the
prefix (OMNICOREAGENT_SERVE_REQUEST_TIMEOUT is request_timeout). With
none set, the defaults serve on 0.0.0.0:8000 with no auth, and the
background API and its worker run with an in-memory task store.
A durable task store; pick one backend:
omnicoreagent[postgres] (SQLAlchemy, for SQLite too), Redis
omnicoreagent[redis], MongoDB omnicoreagent[mongodb]. omniserve config --show prints the configuration the environment gives;
omniserve config --env-example prints a template. Every option is in the
CLI reference and configuration.
The request timeout
OMNICOREAGENT_SERVE_REQUEST_TIMEOUT (300 seconds) bounds each request:
/run/syncand/runs/{run_id}/resumeanswer504when the run takes longer. The run is stopped; its record and trace remain./run(SSE) sendsevent: errorwith"error": "Request timed out"and closes the stream./background/tasks/{task_id}/runwith"wait": truewaits a little less than the timeout, then answers504with therun_id. The run is not stopped: it carries on in the background. A run that pauses for a person or a budget is answered at once,200withawaiting_approvalorawaiting_budget.
Deploy with Docker
.dockerignore in the directory you run it from (the build
context), if there is none, that
keeps credential files (.env, .env.*, *.pem, *.key, id_rsa*, .netrc, credentials*.json
and the like, in any subfolder) and the configured workspace folder (at the root only) out of the
image (the image copies the build context); an existing one that does not exclude
**/.env is named. The
generated Dockerfile installs omnicoreagent[serve], pinned to the
release that generated it (==0.5.0), on
python:3.12-slim, copies the build context to /app, runs as a non-root
user, checks /health, and keeps the workspace in /tmp/workspace (gone
with the container unless you mount a volume). Its HEALTHCHECK allows a
start period of 10 seconds (--start-period=10s); a slow host may need
longer before /health first answers, so raise it in the Dockerfile or with
docker run --health-start-period=60s. Add the extras your agent
needs ([serve,postgres] for a SQL memory store): a comment above the
install line in the Dockerfile says so. A persistent volume, a
store for runs and readiness checks are on
Running the agent in a container.
When things go wrong
403 Denied by strict policy
403 Denied by strict policy
The agent’s policy refused what the request asked, here a task The body names the refused An ask outranks an allow, and With both,
DELETE
the policy does not allow:capability, the reason_code and the
matched_rule_ids (none here: no rule matched). Add an allow rule for
that capability (background.task.delete), for the operators you trust.An ask nobody can answer inside the request is refused the same way,
and the body names the capability and the rule. Under interactive-dev,
every background request asks:interactive-dev asks about anything no
rule matches: for an agent served to operators you trust, take the ask
rule out and allow the capability.POST /background/tasks returns 200.404 No run / Task not found
404 No run / Task not found
409 still waiting, already decided, already finished
409 still waiting, already decided, already finished
The request does not fit the run’s state:An approval reads
already approved (or denied) once decided, and
already used once a resume has applied the decision. Decide every
approval before resuming; send "replace": true to
overwrite a task.422 a body that does not validate
422 a body that does not validate
/runs/{run_id}/steer also answers 422 when the injection guardrail
blocks the message.429 Rate limit exceeded
429 Rate limit exceeded
With Every request but the public paths counts, rejected ones (a
--rate-limit 3, the fourth request in a minute:401)
included.504 Request timed out
504 Request timed out
A The run is stopped and its record says That run finished later (
/run/sync under a 1-second timeout, with a model slower than that:timeout
(GET /runs/{run_id}); the run_id finds its trajectory. POST /runs/{run_id}/resume answers a timeout the same way.A background run waited on with "wait": true:GET /background/runs/{run_id} said
completed). Raise OMNICOREAGENT_SERVE_REQUEST_TIMEOUT, or queue long
work without waiting.The server will not start: Invalid OmniServe config
The server will not start: Invalid OmniServe config
The configuration is checked before the agent is loaded:A Redis or MongoDB task store without its address is refused the same
way:
Error: Invalid OmniServe config: Redis background task store requires OMNICOREAGENT_BACKGROUND_TASK_STORE_URL (..._URI for
MongoDB).Agent file must define an 'agent' variable or 'create_agent()' function
Agent file must define an 'agent' variable or 'create_agent()' function
Error loading agent file: ....Application startup failed: LLM_API_KEY not found
Application startup failed: LLM_API_KEY not found
The model key is read when the agent is built; without it the server
exits:Set
LLM_API_KEY in the server’s environment. A key the provider
rejects starts fine, and each run ends with status error,
termination_reason provider_error.OmniServe cannot listen on host:port: it is in use
OmniServe cannot listen on host:port: it is in use
Another process holds the port. The port is checked before the agent is
loaded, so this comes at once:Pick another
--port, or stop what holds it.Next
Running the agent in a container
An image, a volume, a store for runs, health checks.
Durable runs
Run records, pause and resume, recovery after a crash.
Stores and scale
Memory and task stores, and many processes sharing them.
Background agents
Schedules, retries, leases and the task store.
Approvals
Everything about a person in the loop.
HTTP API reference
Every route, body and response field.