Changelog
Every release, with its notes, is also on GitHub Releases.0.5.1
October 2026 · Python 3.12+ Fixes only, no new features: the known issues from 0.5.0, and three more found running the steward on 0.5.0. Most code runs unchanged. What now behaves differently: a custom strict policy no longer blocks the sandbox’s own file sync;local refuses network-off and image when the agent is built; two agents sharing one telemetry store must agree on retention_days; under capture="full", credentials in command text are redacted.
Fixed
-
Copying many files into a fresh Docker sandbox is one archive. Each file was its own
mkdirandput_archive(about 0.3 s a file on the test machine), so 200 workspace files took 61 s; they now go in as one tar stream, in 0.8 s. The same files are copied, with the same paths and modes, and the sandbox user owns them. One path outside the working directory now refuses the whole copy before anything is written. -
The generated
.dockerignoreis no longer broad and narrow at once. It left out any folder namedworkspace(including a Python package of that name). It now leaves out only the workspace folder you configured (OMNICOREAGENT_WORKSPACE_DIR, default./workspace), anchored at the build context’s root. It also leaves out, in any folder,.env,.env.*,*.pem,*.key,*.p12,id_rsa*,id_ed25519*,.netrc,.npmrc,.pypirc,credentials*.json,.awsand.ssh(before,.env*only). An existing.dockerignoreis still kept as it is. -
The
localprovider refusesnetwork_policy: denyandimagewhen the agent is built. It refused them at the first command, after a person had been asked to approve the call. Asandbox_manifestthat asks for either, or says nothing about the network (the default isdeny), or no manifest at all, now fails at construction with the same message the other providers give. Its other behaviour is unchanged. -
A sandbox command is in the record under
capture="full". Thesandbox_exec_startedevent carried only the program name andargc, andsandbox_exec_completedkept the exit code in its metadata but not besidestdoutandstderr. Underfull, the start event now holds the command text, with credentials in it (--api-key=...,TOKEN=...,Authorization: Bearer ...) redacted; under the default capture it still holds the name andargconly. The outcome’soutputnow carriesexit_codeunder both. -
Two agents sharing one telemetry store must agree on
retention_days. The store is one object per file in a process, so the second agent silently got the first one’s window (a warning was logged and its value ignored). A second, different value is now refused when that agent opens the store (initialize()), with a message naming both values; the same value is fine, andNone(keep every trace) counts as a value. Give such agents the sameretention_days, or astorage_patheach. -
A folder operation under an ask rule asks once. Deleting, moving or clearing a folder that
holds a file covered by an ask rule was refused, and the message said the policy “does not
allow” it. The operation now asks a person once and lists the covered files (the first ten, then
“and N more”; the pending approval carries them as
covered_files). Approving lets the operation run, and the covered files count as approved for that operation only; denying changes nothing. A deny rule on any file inside still refuses the whole operation, without asking. Where nobody can be asked (approval_mode="fail") the message says the call needs a person’s approval. -
An approval’s command text follows the same secret rule as a sandbox command. Under
capture="full"the approval events in the trace kept a shell command’s summary verbatim, soprintf token=...carried its credential. They now keep the text withkey=valuecredentials redacted, as the sandbox command’s own record does; under the default capture the summary is still replaced by its count. The approval shown to the person and the run’s own record are unchanged. -
read_skill_fileno longer reads a sibling skill’s files. The requested path is resolved, symlinks followed, and must sit inside the skill’s own folder; otherwise it is refused with a message naming the skill. -
A strict policy no longer silently denies the sandbox’s own file listing. After a sandbox
command, the runtime lists the sandbox’s files to copy what changed back to the workspace
(
sandbox.workspace.sync). A custommode: strictpolicy with no rule for it denied the listing on every command, and nothing came back. With no rule of its own the listing now follows the decision for its command (allowed, or approved), and the trace says so (follows_command). A rule you write for it still decides, and each file copied back is still decided by the workspace file rules. When a rule refuses the listing, the model is told why instead of “the sandbox was lost”. -
The
tool_callsbudget counts a call once. A call that asked for approval was counted when it asked and again when it ran after approval and resume, so the meter ran ahead of the calls made by the number of approved writes (44 against 41 in one run, 40 against 37 in another). A call is now counted when it is allowed to run: one that is waiting, denied or refused by a rule is not counted, and a call the budget cannot afford still does not run. - An approved call the budget then refuses is not asked about again. When an approved call resumed and the budget had nothing left, the person’s approval was already marked used, so after a grant the same call asked a second time. The approval is now kept until the call runs.
-
After a crash and resume, the dead segment no longer shows
running. When a run resumes, its earlier trace segments that never ended are closed with the new trace statusinterruptedand an end time equal to their last event. Listings filtered bystatus=runningno longer return them, and the run’s own record is unchanged.interruptedis a new value of the tracestatusfilter. -
A worker whose lead crashed no longer stays
running. When a lead run ends (finished after a resume, cancelled or abandoned withabandon_run), the workers it started withspawn_subagentsthat are stillrunningwith a lapsed lease are marked with the new run statusabandoned, and the error names the lead’s run. A worker whose lease is alive is never touched.abandonedis also a value of thestatusfilter oflist_runsandGET /runs. -
A very small budget no longer pauses on every call, once a grant names its amount. Nothing in
the ledger is per call: a grant raises the run’s limit and later calls draw on it (a limit of
one tool call plus a grant of 3 runs four calls). The pause came from the default: a grant with
no
amountis the shortfall of the one call that stopped, so a run that stops one call at a time asked again at each. That default is unchanged (a grant never gives more than was asked), and the budget docs now say to name anamountfor the work still to do.
0.5.0
September 2026 · Python 3.12+ Give your agent real work. Keep control. Every action checked before it runs. Every run survives a crash without silently redoing anything. Every step on the record. From this release an agent is governed unless it says otherwise.Changed: governance is on by default
An agent with no governance settings is governed by thepermissive-dev
profile. Its own tools, workspace, memory, code mode, skills, sub-agents and
background runs are allowed, and nothing pauses for a person unless the agent is
given a sandbox with network (then a person is asked first). Refused unasked:
reading raw secrets, unrestricted host files and network, installing packages,
and shell commands on the host. The MCP servers you configure are
trusted: starting them and calling their tools is allowed by a named rule. There
is no budget until you set one. governance_config={"enabled": False} turns it
off. Enabling governance
yourself without naming a profile still means interactive-dev, as before, so
no one’s policy loosens. See Upgrading, “From 0.4 to 0.5.0”.
New: worker profiles
List the kinds of worker a lead may spawn inworker_profiles, each with its
own model, reasoning effort, tools, MCP servers, steps and deny or ask rules,
and the lead picks one for each worker
(Sub-agents).
A profile only narrows what the lead has. Every worker spends the lead’s
budgets, and each spawn records the profile, model and effort it ran with.
Changed: a simple policy, the sandbox is the boundary
- Command rules are prefix rules. A rule names a command by its first
words (
{"prefix": ["git", "push"]}, or{"program": "rm"}). Only a plain chain of commands is split; anything else (a redirect,$(...), a variable,eval,sh -c…) is one unreadable command, never allowed by a rule and, where the policy has command rules, asked about in every mode (refused instrict).args_any,redirect,envand globs in rules are refused with the new form. This replaces 0.4.4’s parser, which was never published. - A sandbox setting the provider does not enforce is refused when the agent is built, naming the provider and the setting. Examples are a workspace mount on a hosted provider, a sandbox lifetime, or a host allowlist on Daytona. Each provider is listed in What stops harm.
filesystem_policyis removed. Its path lists were approved and then applied by no provider; a manifest naming it is refused with what to use.- Modal and Vercel check a sandbox’s network from inside, as E2B and Daytona do, and refuse one that should be offline but is not.
Changed: OmniServe defaults
OMNICOREAGENT_SERVE_CORS_CREDENTIALS is false by default and never applies
with * origins. X-Forwarded-For is read only from the proxies in
OMNICOREAGENT_SERVE_TRUSTED_PROXIES. /prometheus counts per route, not per
path.
Changed: what a trace keeps by default
Undertelemetry_config={"capture": "default"}, a governed agent’s tool and
delegation arguments are recorded by name, not value. Every agent is now
governed, so under that preset this applies to all of them. The capture you
get without setting one, full, keeps the values (personal data redacted).
Fixed
- A folder delete, move or clear is checked for each file in it: a deny on
deleting
prod/*did not stopdelete_file("prod"). A workspace path through a link is refused. - Spawned workers spend the lead’s budgets; they each had their own.
- OmniServe: error bodies never carry a registered credential, and the token is compared in constant time.
- A governed run’s evidence under full capture was reported incomplete.
Every delegation’s policy request masked its task as
[REDACTED], which the trajectory counted as a capture gap. The task text is left out of that record instead (the delegation call keeps it);task_lengthstays.
Docs
- A new landing page that leads with what is different, and a
comparison with an Agno column and a row for what happens
to an interrupted tool call after a crash, each cell from the project’s own
docs, read 2026-09-29. Pydantic AI’s
cost_limitcell is corrected: it is checked after each response.
Known issues
These come from the 0.5.0 release testing. None breaks a core promise on a normal path. Each is planned for 0.5.1.- The first run in a fresh process starts slowly. Loading the model client can take 15–60 s on
a cold or loaded machine. Measured in 0.5.1, it is
import litellm(about 3,000 modules; about 9 s of CPU on a warm bytecode cache and about 20 s on a cold one, on the test machine), which OmniCoreAgent does not control. It is already off the event loop, on a worker thread, and shows in the trace as amodel.client.loadspan; the headless--timeoutstarts after it. What is ours is light:import omnicoreagenttakes about 10 ms, building an agent and initializing it well under a second, withlitellmandopenainot yet loaded. OmniServe and background workers load it while they start, so only a short script or a headless run pays for it on its first run. - A very small budget pauses on every call. With a limit as low as one tool call, each call pauses on its own, so a person grants each one in turn.
read_skill_filecan read a sibling skill’s files. Its path check is by prefix, so a sibling whose name starts with the skill’s own name (../my-skill-extra/filefrommy-skill) is not refused. Paths outside the skills folder are. Fixed in 0.5.1.
0.4.4
Not published: tagged and tested, never uploaded. Its fixes ship in 0.5.0, and its command rules were replaced there by prefix rules. Filming OmniCoreAgent as it really runs (an agent killed mid-payment, a command denied, a budget topped up from another process) found two gaps. This release closes them.New
- Rules on shell commands.
Every command reached the policy as
sh, so no rule could namerm -rf. Aprocess.execrule can now carry acommandmatcher (program,prefix,args_any). The command is parsed first, so quoting, paths,$(...),sudo,sh -c '...',xargsandfind -execdo not hide what runs. Any matching deny or ask applies; an allow needs every command proven; a command whose effect cannot be proven (eval,... | sh) is never allowed by one. Rules carryexamples, checked when the policy loads. Tested against the bypasses that defeated text denylists elsewhere, and against a real shell.
Fixed
- An approval for a command now shows the command, and is bound to it. A
person asked to approve a command on the host saw
shand an argument count, and the approval did not name the command: approvinglslooked the same as approvingrm -rf ~. It now lists the commands it would run and applies to that exact text. budget_statusandabandon_runwork from a process that never ran a query.budget_statusanswered[], andabandon_runleft the run’s budget holds on the day’s counter, when called from a second process.- An unknown key in a policy rule is a clear
PolicyLoadErrornaming the rule and the key, not a bareTypeError.
Changed
- New dependencies:
tree-sitterandtree-sitter-bash(MIT, binary wheels), to parse commands.
0.4.3
September 2026 · Python 3.12+ An outside review read these docs and asked the question every durable runtime must answer exactly: what happens when the process dies halfway through a tool that has side effects? Answering it, we found one case where the answer was wrong, and fixed it.Fixed
- A call that finished while another call of the same turn was still running is no longer run again after a crash. Results of a turn’s calls are saved once every call of the turn has finished. A call that completed while a sibling was still running, when the process died, had its state saved but not its result, and recovery ran it again whatever its idempotency: a card could be charged twice. Now it runs again only if its tool is idempotent; otherwise the model is told “This call finished (it reported success), but its result was lost … do not call it again without checking.”
New
- On Python 3.10 or 3.11, installing fails loudly. A guard release, 0.3.10, exists for exactly those Pythons: pip and uv choose it over 0.3.9 (which cannot build an agent) and stop with a message saying how to get Python 3.12. It never shows as the latest release.
Docs
- What is guaranteed:
what a recovered run repeats and what it never does, per call and model-call
state (exactly once, at most once, at least once), one owner, how budgets
count across a crash, sub-agents and background retries. Four places that
claimed more than the code are corrected, and the recovery sample now also
catches
RunStateConflict. - The navigation leads with what is different: Make It Safe, Run It and See and Improve now come before Build.
- Short paths work:
/docs/quickstart,/docs/installand others redirect instead of returning 404. - One tagline: The governed runtime for Python agents you can let act.
0.4.2
September 2026 · Python 3.12+ Every section of these docs was rewritten by running its examples, and then tested by agents that had only the docs and a build of this release: they built a refunds desk, an ops service, and an evaluation loop, and every feature they used worked. What they and the writers found in the code is fixed here. Most of it is fixes; five settings now do what their names say, and those are listed first.Before you upgrade
- A durable memory store needs its URL.
MemoryRouter("sql")withoutDATABASE_URL(rediswithoutREDIS_URL,mongodbwithoutMONGODB_URI) raises aValueErrornaming the variable. It fell back to memory, so runs silently did not survive a restart or resume elsewhere.MemoryRouter("in_memory"), the default, is unchanged. - Sandbox settings no provider keeps are refused when the agent is
built:
denied_hosts(except withhttp),environment.secret_refs, and a host allowlist on Daytona. They were accepted and silently not enforced. cancel_previousinterrupts the run it replaces, within a heartbeat, instead of letting it finish its attempt and then discarding it.misfire_policy: the default is nowrun_once, which behaves as the old default did.skip_missednow skips: nothing runs for a time missed by more than a minute.- Trace retention removes abandoned traces: a trace still
runningwith no activity for longer thanretention_days, whose process died. A trace waiting for a person is never removed. - OmniServe: a timed-out
/run/syncor resume answers504with adetailobject that names the run (message,run_id, …), wheredetailwas a string. A403for a refused request addscapability,reason_codeandmatched_rule_ids.
New
workspace_config["files_dir"]: the file tools work in a directory you already have; the runtime’s own files stay inworkspace_dir. Workspace files- Harbor trials use the task’s own directory for the agent’s files,
whatever the task set (they pointed at the agent’s workspace, and a model
spent steps on “Path not found”).
harbor doctortakes--akasrundoes. Harbor - A denial is a decision in the evidence: a
denywithdenied_by(the person, orsystemfor an approval nobody decided in time), as an approval is anallowwithapproved_by. - The run header says which privacy boundaries redact (
harness.privacy). await memory_router.close()closes MongoDB, Redis and SQL connections.
Fixed
Privacy and governanceprivacy_config["redact_model_io"]now keeps personal data from the model provider. It was accepted but never applied.- An approval past its expiry no longer strands its run: one nobody decided
in time is a refusal on resume, and a decision made in time stands however
late the resume. A person’s denial reaches the model as theirs, note
included; with
approval_mode="fail", the model is told the call was refused. ls("."), andgloborgrepwith no path, list the whole workspace; they were refused as invalid arguments.- Clearer messages: a strict-policy denial says no rule matched; a YAML policy says YAML is not supported.
- A call that crosses a
model_tokensbudget is recorded in full, and the next call is stopped before it is made. The crossing call’s tokens were lost and a resume made the call again.
- A finished trace read back by a new process was marked incomplete; it reads back complete.
- Pruned traces stay pruned (they came back on the next start), and
prune()counts archived traces. - A resumed run’s totals count a step once; a run with nothing priced says
Nonefor its cost everywhere. training_records()carriesfinal_answer, and every step hasresumed.- An exporter’s failure keeps its reason in the trace.
- A timed-out run’s record says
timeout, as its trace does. list_runs()approvals carry the call’sarguments.- The MongoDB store closes its connections; a store built per request leaked a pool each time. SQL stores release their shared pool.
omniserve runwith a Redis or MongoDB task store and no URL is a clean config error.- The prompt-injection guardrail logs through the
omnicoreagent.guardrailslogger instead of printing to stdout.
- A headless run’s
result.jsoncounts the whole run’s usage. harbor doctorfails a development runtime nothing can build into a wheel.
Docs
Build, Make It Safe, Run It and See and Improve rewritten page by page, with new pages for Policies, Approvals, Budgets, Privacy and credentials, Stores and scale, Read a run, and Outcomes and training records; a comparison with other agent frameworks; and the generated reference now shows each method’s whole description.0.4.1
September 2026 · Python 3.12+Changed: finished run records are removed after 30 days
Every run keeps a record in the memory store; until now they were kept forever. From 0.4.1 a finished run’s record (completed, failed, blocked, cancelled, timed out) is removed 30 days after the run started. A run waiting for a person or a resume is never removed. To keep every record, setagent_config["run_retention_days"] = None. Any other number of days
works too; await agent.prune_runs() prunes on demand and reports what it
removed, and the retention status has a runs section.
This is the run’s record (status, approvals, usage), not its traces: the
step-by-step trajectory lives in the traces, still kept 7 days by
telemetry_config["retention_days"]. When a run’s traces are gone and its
record is not, get_run_trajectory now says so (trace_kept: False,
traces_missing) instead of reporting zero tokens and no cost.
How long a run’s evidence is kept
New
base_urlinmodel_config. Point any provider at another endpoint that speaks its API: vLLM, LM Studio, a gateway, a proxy. Sent with each of the agent’s calls; two agents can use two endpoints. Models
Fixed
- Agents sharing one
ToolRegistryrebound each other’s file tools, so a lead could read a child’s workspace (“File not found” on a file it had listed). Each agent adds its tools to its own copy; your registry is not changed. - A file tool that failed (“File not found”, “File already exists”) was recorded as a success; it is the tool’s error now, with the same text.
get_latest_trace(session)returned a delegated child’s trace; it returns the agent’s own.- A
delegate_<name>tool offered the model the child’stags,provenanceand private parameters; it offers the task only. - An approval for a call made inside a
run_codeprogram carried no arguments; the run’s approval record has them. - A partial
memory_config({"mode": ...}) failed withKeyError; it keeps the other defaults.switch_memory_storekeeps the window and summary settings. - A tool parameter named like a secret (
api_key) had its schema redacted, which marked the run’s evidence partial.
Docs
- The Build section rewritten page by page, every example run; navigation by what a reader is trying to do.
- How OmniCoreAgent compares, each cell sourced.
- Skills:
allowed-toolsis read and grants nothing; the policy decides. - Settings explained one by one: the telemetry
record_*switches (which are master switches and which narrow them), the guardrail, archive, offload and timeout settings that shared one description,memory_config’s default window, and which limit to use amongmax_steps,request_limit,total_tokens_limitand budgets. - How long a run’s evidence is kept: traces (7 days) and run records (30 days), side by side.
0.4.0
September 2026 · Python 3.12+ 0.4 rebuilds OmniCoreAgent as a runtime for agents that act: an agent that can run commands, spend money and change things needs a policy over what it may do, a person to ask, a sandbox, a budget, a record of everything it did, and a run that survives a crash. 0.4 has each of them, on by default where that is safe and one setting away where it is not. Coming from 0.3? Upgrading to 0.4 lists, as measured, what runs unchanged, what was removed and what took its place, and the defaults that changed.New
- Governance. A policy — allow, ask, deny — over every capability: each
tool, each MCP server, the sandbox, the network, delegation, background
runs. Built-in profiles;
askpauses the run for a person. Security model - Durable runs. Every run has a record. It pauses for an approval or a
budget top-up and resumes where it stopped (
resolve_approval,resume);interrupt,steerandabandon_run; a run whose process died continues from its checkpoint, and a call interrupted mid-way is reported, never repeated. Durable runs - Budgets in dollars, tokens, calls and sandbox seconds, per request, session, agent or application.
- Sandboxed execution. An
executetool whose commands run in Docker, E2B, Modal, Daytona or Vercel — orlocalwhere a container is already the boundary — with no network unless the policy allows it, and without your process’s keys and tokens. Sandbox providers - Code mode. A
run_codetool: the model writes a short Python program that calls your tools, run in Monty, every call inside it governed and traced. Code mode - Telemetry. One trace per run, readable end to end, kept on local disk
by default; trajectories,
get_run_trajectoryacross a run’s segments, export to OTLP, LangSmith, Opik or JSONL. Outcomes that arrive later (record_outcome) andtraining_records()for evaluators and trainers. Observability - Privacy. Emails, phone numbers, SSNs and card numbers are redacted from the record; the run itself works with the real data.
- Native tool calls for every provider, streamed (
agent.stream(),on_event=), with provider continuation data kept across turns. - Sub-agents spawned under the same policy and budgets, each with a
linked trace;
sub_agents=[...]exposes agents asdelegate_*tools. - Skills (
.agents/skills), AGENTS.md files, workspace files on local disk, S3 or R2. - Background agents rebuilt on a durable task store (memory, SQL, Redis, MongoDB): cron, intervals, leases, retries, recovery after a restart.
- OmniServe gains approvals, budgets, runs, background tasks and traces over REST and SSE. OmniServe
- The
omnicoreagentcommand:runfor one headless run (CI, evaluations);harborfor Terminal-Bench through Harbor, withharbor doctorandharbor results. CLI reference - The reference, generated from the code: every method, setting, command and HTTP route. Reference
Changed
- Python 3.12 or later.
- Defaults:
max_steps50 (was 15),tool_call_timeout180 s (was 30), context management and tool offload on, the prompt-injection guardrail on (guardrail_mode="full"), workspace files on, telemetry on. get_tracealso takestrace_id=andrun_id=.- An
agent_configkey that does not exist is refused with the settings that do, instead of a bareTypeError.
Removed
SequentialAgent,ParallelAgent,RouterAgent,DeepAgent,OmniAgent.EventRouter,event_router=and the event-store methods, replaced by telemetry.BackgroundOmniCoreAgent,BackgroundTaskScheduler,APSchedulerBackend,TaskRegistry, and thebackgroundextra.agent_config["memory_tool_backend"]and thememory_*tools, replaced by workspace files.- The top-level retry and circuit-breaker helpers;
omniserve run --reload.
Fixed
Found by four agents that built applications from these docs alone, before the release:-
A resumed run’s evidence is whole: the call a person approved appears in
its own step, as approved by them; the run’s totals and its training
record include every segment;
get_run()shows what each approval is for. -
An ask on delegation (
delegate_<name>) pauses the lead, as documented; it was a tool error. -
Full capture keeps governed tool arguments, and the calls a
run_codeprogram makes have theirs. -
budget_statuscounts a person’s grant, as enforcement does. - A run’s deadline holds when the work does not stop when told.
- A host allowlist a sandbox provider cannot enforce is refused when the agent is built.
-
0.3.9 was published without
omnicoreagent.core.workspace, soOmniCoreAgent(...)failed on import. A test now fails if any source file of the package is left out of the wheel.