Skip to main content

Changelog

Every release, with its notes, is also on GitHub Releases.

0.5.1

October 2026 · Python 3.12+ Fixes only, no new features: the known issues from 0.5.0, and three more found running the steward on 0.5.0. Most code runs unchanged. What now behaves differently: a custom strict policy no longer blocks the sandbox’s own file sync; local refuses network-off and image when the agent is built; two agents sharing one telemetry store must agree on retention_days; under capture="full", credentials in command text are redacted.

Fixed

  • Copying many files into a fresh Docker sandbox is one archive. Each file was its own mkdir and put_archive (about 0.3 s a file on the test machine), so 200 workspace files took 61 s; they now go in as one tar stream, in 0.8 s. The same files are copied, with the same paths and modes, and the sandbox user owns them. One path outside the working directory now refuses the whole copy before anything is written.
  • The generated .dockerignore is no longer broad and narrow at once. It left out any folder named workspace (including a Python package of that name). It now leaves out only the workspace folder you configured (OMNICOREAGENT_WORKSPACE_DIR, default ./workspace), anchored at the build context’s root. It also leaves out, in any folder, .env, .env.*, *.pem, *.key, *.p12, id_rsa*, id_ed25519*, .netrc, .npmrc, .pypirc, credentials*.json, .aws and .ssh (before, .env* only). An existing .dockerignore is still kept as it is.
  • The local provider refuses network_policy: deny and image when the agent is built. It refused them at the first command, after a person had been asked to approve the call. A sandbox_manifest that asks for either, or says nothing about the network (the default is deny), or no manifest at all, now fails at construction with the same message the other providers give. Its other behaviour is unchanged.
  • A sandbox command is in the record under capture="full". The sandbox_exec_started event carried only the program name and argc, and sandbox_exec_completed kept the exit code in its metadata but not beside stdout and stderr. Under full, the start event now holds the command text, with credentials in it (--api-key=..., TOKEN=..., Authorization: Bearer ...) redacted; under the default capture it still holds the name and argc only. The outcome’s output now carries exit_code under both.
  • Two agents sharing one telemetry store must agree on retention_days. The store is one object per file in a process, so the second agent silently got the first one’s window (a warning was logged and its value ignored). A second, different value is now refused when that agent opens the store (initialize()), with a message naming both values; the same value is fine, and None (keep every trace) counts as a value. Give such agents the same retention_days, or a storage_path each.
  • A folder operation under an ask rule asks once. Deleting, moving or clearing a folder that holds a file covered by an ask rule was refused, and the message said the policy “does not allow” it. The operation now asks a person once and lists the covered files (the first ten, then “and N more”; the pending approval carries them as covered_files). Approving lets the operation run, and the covered files count as approved for that operation only; denying changes nothing. A deny rule on any file inside still refuses the whole operation, without asking. Where nobody can be asked (approval_mode="fail") the message says the call needs a person’s approval.
  • An approval’s command text follows the same secret rule as a sandbox command. Under capture="full" the approval events in the trace kept a shell command’s summary verbatim, so printf token=... carried its credential. They now keep the text with key=value credentials redacted, as the sandbox command’s own record does; under the default capture the summary is still replaced by its count. The approval shown to the person and the run’s own record are unchanged.
  • read_skill_file no longer reads a sibling skill’s files. The requested path is resolved, symlinks followed, and must sit inside the skill’s own folder; otherwise it is refused with a message naming the skill.
  • A strict policy no longer silently denies the sandbox’s own file listing. After a sandbox command, the runtime lists the sandbox’s files to copy what changed back to the workspace (sandbox.workspace.sync). A custom mode: strict policy with no rule for it denied the listing on every command, and nothing came back. With no rule of its own the listing now follows the decision for its command (allowed, or approved), and the trace says so (follows_command). A rule you write for it still decides, and each file copied back is still decided by the workspace file rules. When a rule refuses the listing, the model is told why instead of “the sandbox was lost”.
  • The tool_calls budget counts a call once. A call that asked for approval was counted when it asked and again when it ran after approval and resume, so the meter ran ahead of the calls made by the number of approved writes (44 against 41 in one run, 40 against 37 in another). A call is now counted when it is allowed to run: one that is waiting, denied or refused by a rule is not counted, and a call the budget cannot afford still does not run.
  • An approved call the budget then refuses is not asked about again. When an approved call resumed and the budget had nothing left, the person’s approval was already marked used, so after a grant the same call asked a second time. The approval is now kept until the call runs.
  • After a crash and resume, the dead segment no longer shows running. When a run resumes, its earlier trace segments that never ended are closed with the new trace status interrupted and an end time equal to their last event. Listings filtered by status=running no longer return them, and the run’s own record is unchanged. interrupted is a new value of the trace status filter.
  • A worker whose lead crashed no longer stays running. When a lead run ends (finished after a resume, cancelled or abandoned with abandon_run), the workers it started with spawn_subagents that are still running with a lapsed lease are marked with the new run status abandoned, and the error names the lead’s run. A worker whose lease is alive is never touched. abandoned is also a value of the status filter of list_runs and GET /runs.
  • A very small budget no longer pauses on every call, once a grant names its amount. Nothing in the ledger is per call: a grant raises the run’s limit and later calls draw on it (a limit of one tool call plus a grant of 3 runs four calls). The pause came from the default: a grant with no amount is the shortfall of the one call that stopped, so a run that stops one call at a time asked again at each. That default is unchanged (a grant never gives more than was asked), and the budget docs now say to name an amount for the work still to do.

0.5.0

September 2026 · Python 3.12+ Give your agent real work. Keep control. Every action checked before it runs. Every run survives a crash without silently redoing anything. Every step on the record. From this release an agent is governed unless it says otherwise.

Changed: governance is on by default

An agent with no governance settings is governed by the permissive-dev profile. Its own tools, workspace, memory, code mode, skills, sub-agents and background runs are allowed, and nothing pauses for a person unless the agent is given a sandbox with network (then a person is asked first). Refused unasked: reading raw secrets, unrestricted host files and network, installing packages, and shell commands on the host. The MCP servers you configure are trusted: starting them and calling their tools is allowed by a named rule. There is no budget until you set one. governance_config={"enabled": False} turns it off. Enabling governance yourself without naming a profile still means interactive-dev, as before, so no one’s policy loosens. See Upgrading, “From 0.4 to 0.5.0”.

New: worker profiles

List the kinds of worker a lead may spawn in worker_profiles, each with its own model, reasoning effort, tools, MCP servers, steps and deny or ask rules, and the lead picks one for each worker (Sub-agents). A profile only narrows what the lead has. Every worker spends the lead’s budgets, and each spawn records the profile, model and effort it ran with.

Changed: a simple policy, the sandbox is the boundary

  • Command rules are prefix rules. A rule names a command by its first words ({"prefix": ["git", "push"]}, or {"program": "rm"}). Only a plain chain of commands is split; anything else (a redirect, $(...), a variable, eval, sh -c…) is one unreadable command, never allowed by a rule and, where the policy has command rules, asked about in every mode (refused in strict). args_any, redirect, env and globs in rules are refused with the new form. This replaces 0.4.4’s parser, which was never published.
  • A sandbox setting the provider does not enforce is refused when the agent is built, naming the provider and the setting. Examples are a workspace mount on a hosted provider, a sandbox lifetime, or a host allowlist on Daytona. Each provider is listed in What stops harm.
  • filesystem_policy is removed. Its path lists were approved and then applied by no provider; a manifest naming it is refused with what to use.
  • Modal and Vercel check a sandbox’s network from inside, as E2B and Daytona do, and refuse one that should be offline but is not.

Changed: OmniServe defaults

OMNICOREAGENT_SERVE_CORS_CREDENTIALS is false by default and never applies with * origins. X-Forwarded-For is read only from the proxies in OMNICOREAGENT_SERVE_TRUSTED_PROXIES. /prometheus counts per route, not per path.

Changed: what a trace keeps by default

Under telemetry_config={"capture": "default"}, a governed agent’s tool and delegation arguments are recorded by name, not value. Every agent is now governed, so under that preset this applies to all of them. The capture you get without setting one, full, keeps the values (personal data redacted).

Fixed

  • A folder delete, move or clear is checked for each file in it: a deny on deleting prod/* did not stop delete_file("prod"). A workspace path through a link is refused.
  • Spawned workers spend the lead’s budgets; they each had their own.
  • OmniServe: error bodies never carry a registered credential, and the token is compared in constant time.
  • A governed run’s evidence under full capture was reported incomplete. Every delegation’s policy request masked its task as [REDACTED], which the trajectory counted as a capture gap. The task text is left out of that record instead (the delegation call keeps it); task_length stays.

Docs

  • A new landing page that leads with what is different, and a comparison with an Agno column and a row for what happens to an interrupted tool call after a crash, each cell from the project’s own docs, read 2026-09-29. Pydantic AI’s cost_limit cell is corrected: it is checked after each response.

Known issues

These come from the 0.5.0 release testing. None breaks a core promise on a normal path. Each is planned for 0.5.1.
  • The first run in a fresh process starts slowly. Loading the model client can take 15–60 s on a cold or loaded machine. Measured in 0.5.1, it is import litellm (about 3,000 modules; about 9 s of CPU on a warm bytecode cache and about 20 s on a cold one, on the test machine), which OmniCoreAgent does not control. It is already off the event loop, on a worker thread, and shows in the trace as a model.client.load span; the headless --timeout starts after it. What is ours is light: import omnicoreagent takes about 10 ms, building an agent and initializing it well under a second, with litellm and openai not yet loaded. OmniServe and background workers load it while they start, so only a short script or a headless run pays for it on its first run.
  • A very small budget pauses on every call. With a limit as low as one tool call, each call pauses on its own, so a person grants each one in turn.
  • read_skill_file can read a sibling skill’s files. Its path check is by prefix, so a sibling whose name starts with the skill’s own name (../my-skill-extra/file from my-skill) is not refused. Paths outside the skills folder are. Fixed in 0.5.1.

0.4.4

Not published: tagged and tested, never uploaded. Its fixes ship in 0.5.0, and its command rules were replaced there by prefix rules. Filming OmniCoreAgent as it really runs (an agent killed mid-payment, a command denied, a budget topped up from another process) found two gaps. This release closes them.

New

  • Rules on shell commands. Every command reached the policy as sh, so no rule could name rm -rf. A process.exec rule can now carry a command matcher (program, prefix, args_any). The command is parsed first, so quoting, paths, $(...), sudo, sh -c '...', xargs and find -exec do not hide what runs. Any matching deny or ask applies; an allow needs every command proven; a command whose effect cannot be proven (eval, ... | sh) is never allowed by one. Rules carry examples, checked when the policy loads. Tested against the bypasses that defeated text denylists elsewhere, and against a real shell.

Fixed

  • An approval for a command now shows the command, and is bound to it. A person asked to approve a command on the host saw sh and an argument count, and the approval did not name the command: approving ls looked the same as approving rm -rf ~. It now lists the commands it would run and applies to that exact text.
  • budget_status and abandon_run work from a process that never ran a query. budget_status answered [], and abandon_run left the run’s budget holds on the day’s counter, when called from a second process.
  • An unknown key in a policy rule is a clear PolicyLoadError naming the rule and the key, not a bare TypeError.

Changed

  • New dependencies: tree-sitter and tree-sitter-bash (MIT, binary wheels), to parse commands.

0.4.3

September 2026 · Python 3.12+ An outside review read these docs and asked the question every durable runtime must answer exactly: what happens when the process dies halfway through a tool that has side effects? Answering it, we found one case where the answer was wrong, and fixed it.

Fixed

  • A call that finished while another call of the same turn was still running is no longer run again after a crash. Results of a turn’s calls are saved once every call of the turn has finished. A call that completed while a sibling was still running, when the process died, had its state saved but not its result, and recovery ran it again whatever its idempotency: a card could be charged twice. Now it runs again only if its tool is idempotent; otherwise the model is told “This call finished (it reported success), but its result was lost … do not call it again without checking.”

New

  • On Python 3.10 or 3.11, installing fails loudly. A guard release, 0.3.10, exists for exactly those Pythons: pip and uv choose it over 0.3.9 (which cannot build an agent) and stop with a message saying how to get Python 3.12. It never shows as the latest release.

Docs

  • What is guaranteed: what a recovered run repeats and what it never does, per call and model-call state (exactly once, at most once, at least once), one owner, how budgets count across a crash, sub-agents and background retries. Four places that claimed more than the code are corrected, and the recovery sample now also catches RunStateConflict.
  • The navigation leads with what is different: Make It Safe, Run It and See and Improve now come before Build.
  • Short paths work: /docs/quickstart, /docs/install and others redirect instead of returning 404.
  • One tagline: The governed runtime for Python agents you can let act.

0.4.2

September 2026 · Python 3.12+ Every section of these docs was rewritten by running its examples, and then tested by agents that had only the docs and a build of this release: they built a refunds desk, an ops service, and an evaluation loop, and every feature they used worked. What they and the writers found in the code is fixed here. Most of it is fixes; five settings now do what their names say, and those are listed first.

Before you upgrade

  • A durable memory store needs its URL. MemoryRouter("sql") without DATABASE_URL (redis without REDIS_URL, mongodb without MONGODB_URI) raises a ValueError naming the variable. It fell back to memory, so runs silently did not survive a restart or resume elsewhere. MemoryRouter("in_memory"), the default, is unchanged.
  • Sandbox settings no provider keeps are refused when the agent is built: denied_hosts (except with http), environment.secret_refs, and a host allowlist on Daytona. They were accepted and silently not enforced.
  • cancel_previous interrupts the run it replaces, within a heartbeat, instead of letting it finish its attempt and then discarding it.
  • misfire_policy: the default is now run_once, which behaves as the old default did. skip_missed now skips: nothing runs for a time missed by more than a minute.
  • Trace retention removes abandoned traces: a trace still running with no activity for longer than retention_days, whose process died. A trace waiting for a person is never removed.
  • OmniServe: a timed-out /run/sync or resume answers 504 with a detail object that names the run (message, run_id, …), where detail was a string. A 403 for a refused request adds capability, reason_code and matched_rule_ids.

New

  • workspace_config["files_dir"]: the file tools work in a directory you already have; the runtime’s own files stay in workspace_dir. Workspace files
  • Harbor trials use the task’s own directory for the agent’s files, whatever the task set (they pointed at the agent’s workspace, and a model spent steps on “Path not found”). harbor doctor takes --ak as run does. Harbor
  • A denial is a decision in the evidence: a deny with denied_by (the person, or system for an approval nobody decided in time), as an approval is an allow with approved_by.
  • The run header says which privacy boundaries redact (harness.privacy).
  • await memory_router.close() closes MongoDB, Redis and SQL connections.

Fixed

Privacy and governance
  • privacy_config["redact_model_io"] now keeps personal data from the model provider. It was accepted but never applied.
  • An approval past its expiry no longer strands its run: one nobody decided in time is a refusal on resume, and a decision made in time stands however late the resume. A person’s denial reaches the model as theirs, note included; with approval_mode="fail", the model is told the call was refused.
  • ls("."), and glob or grep with no path, list the whole workspace; they were refused as invalid arguments.
  • Clearer messages: a strict-policy denial says no rule matched; a YAML policy says YAML is not supported.
Budgets
  • A call that crosses a model_tokens budget is recorded in full, and the next call is stopped before it is made. The crossing call’s tokens were lost and a resume made the call again.
Evidence and telemetry
  • A finished trace read back by a new process was marked incomplete; it reads back complete.
  • Pruned traces stay pruned (they came back on the next start), and prune() counts archived traces.
  • A resumed run’s totals count a step once; a run with nothing priced says None for its cost everywhere.
  • training_records() carries final_answer, and every step has resumed.
  • An exporter’s failure keeps its reason in the trace.
Runs, stores and serving
  • A timed-out run’s record says timeout, as its trace does.
  • list_runs() approvals carry the call’s arguments.
  • The MongoDB store closes its connections; a store built per request leaked a pool each time. SQL stores release their shared pool.
  • omniserve run with a Redis or MongoDB task store and no URL is a clean config error.
  • The prompt-injection guardrail logs through the omnicoreagent.guardrails logger instead of printing to stdout.
CLI
  • A headless run’s result.json counts the whole run’s usage.
  • harbor doctor fails a development runtime nothing can build into a wheel.

Docs

Build, Make It Safe, Run It and See and Improve rewritten page by page, with new pages for Policies, Approvals, Budgets, Privacy and credentials, Stores and scale, Read a run, and Outcomes and training records; a comparison with other agent frameworks; and the generated reference now shows each method’s whole description.

0.4.1

September 2026 · Python 3.12+

Changed: finished run records are removed after 30 days

Every run keeps a record in the memory store; until now they were kept forever. From 0.4.1 a finished run’s record (completed, failed, blocked, cancelled, timed out) is removed 30 days after the run started. A run waiting for a person or a resume is never removed. To keep every record, set agent_config["run_retention_days"] = None. Any other number of days works too; await agent.prune_runs() prunes on demand and reports what it removed, and the retention status has a runs section. This is the run’s record (status, approvals, usage), not its traces: the step-by-step trajectory lives in the traces, still kept 7 days by telemetry_config["retention_days"]. When a run’s traces are gone and its record is not, get_run_trajectory now says so (trace_kept: False, traces_missing) instead of reporting zero tokens and no cost. How long a run’s evidence is kept

New

  • base_url in model_config. Point any provider at another endpoint that speaks its API: vLLM, LM Studio, a gateway, a proxy. Sent with each of the agent’s calls; two agents can use two endpoints. Models

Fixed

  • Agents sharing one ToolRegistry rebound each other’s file tools, so a lead could read a child’s workspace (“File not found” on a file it had listed). Each agent adds its tools to its own copy; your registry is not changed.
  • A file tool that failed (“File not found”, “File already exists”) was recorded as a success; it is the tool’s error now, with the same text.
  • get_latest_trace(session) returned a delegated child’s trace; it returns the agent’s own.
  • A delegate_<name> tool offered the model the child’s tags, provenance and private parameters; it offers the task only.
  • An approval for a call made inside a run_code program carried no arguments; the run’s approval record has them.
  • A partial memory_config ({"mode": ...}) failed with KeyError; it keeps the other defaults. switch_memory_store keeps the window and summary settings.
  • A tool parameter named like a secret (api_key) had its schema redacted, which marked the run’s evidence partial.

Docs

  • The Build section rewritten page by page, every example run; navigation by what a reader is trying to do.
  • How OmniCoreAgent compares, each cell sourced.
  • Skills: allowed-tools is read and grants nothing; the policy decides.
  • Settings explained one by one: the telemetry record_* switches (which are master switches and which narrow them), the guardrail, archive, offload and timeout settings that shared one description, memory_config’s default window, and which limit to use among max_steps, request_limit, total_tokens_limit and budgets.
  • How long a run’s evidence is kept: traces (7 days) and run records (30 days), side by side.

0.4.0

September 2026 · Python 3.12+ 0.4 rebuilds OmniCoreAgent as a runtime for agents that act: an agent that can run commands, spend money and change things needs a policy over what it may do, a person to ask, a sandbox, a budget, a record of everything it did, and a run that survives a crash. 0.4 has each of them, on by default where that is safe and one setting away where it is not. Coming from 0.3? Upgrading to 0.4 lists, as measured, what runs unchanged, what was removed and what took its place, and the defaults that changed.

New

  • Governance. A policy — allow, ask, deny — over every capability: each tool, each MCP server, the sandbox, the network, delegation, background runs. Built-in profiles; ask pauses the run for a person. Security model
  • Durable runs. Every run has a record. It pauses for an approval or a budget top-up and resumes where it stopped (resolve_approval, resume); interrupt, steer and abandon_run; a run whose process died continues from its checkpoint, and a call interrupted mid-way is reported, never repeated. Durable runs
  • Budgets in dollars, tokens, calls and sandbox seconds, per request, session, agent or application.
  • Sandboxed execution. An execute tool whose commands run in Docker, E2B, Modal, Daytona or Vercel — or local where a container is already the boundary — with no network unless the policy allows it, and without your process’s keys and tokens. Sandbox providers
  • Code mode. A run_code tool: the model writes a short Python program that calls your tools, run in Monty, every call inside it governed and traced. Code mode
  • Telemetry. One trace per run, readable end to end, kept on local disk by default; trajectories, get_run_trajectory across a run’s segments, export to OTLP, LangSmith, Opik or JSONL. Outcomes that arrive later (record_outcome) and training_records() for evaluators and trainers. Observability
  • Privacy. Emails, phone numbers, SSNs and card numbers are redacted from the record; the run itself works with the real data.
  • Native tool calls for every provider, streamed (agent.stream(), on_event=), with provider continuation data kept across turns.
  • Sub-agents spawned under the same policy and budgets, each with a linked trace; sub_agents=[...] exposes agents as delegate_* tools.
  • Skills (.agents/skills), AGENTS.md files, workspace files on local disk, S3 or R2.
  • Background agents rebuilt on a durable task store (memory, SQL, Redis, MongoDB): cron, intervals, leases, retries, recovery after a restart.
  • OmniServe gains approvals, budgets, runs, background tasks and traces over REST and SSE. OmniServe
  • The omnicoreagent command: run for one headless run (CI, evaluations); harbor for Terminal-Bench through Harbor, with harbor doctor and harbor results. CLI reference
  • The reference, generated from the code: every method, setting, command and HTTP route. Reference

Changed

  • Python 3.12 or later.
  • Defaults: max_steps 50 (was 15), tool_call_timeout 180 s (was 30), context management and tool offload on, the prompt-injection guardrail on (guardrail_mode="full"), workspace files on, telemetry on.
  • get_trace also takes trace_id= and run_id=.
  • An agent_config key that does not exist is refused with the settings that do, instead of a bare TypeError.

Removed

  • SequentialAgent, ParallelAgent, RouterAgent, DeepAgent, OmniAgent.
  • EventRouter, event_router= and the event-store methods, replaced by telemetry.
  • BackgroundOmniCoreAgent, BackgroundTaskScheduler, APSchedulerBackend, TaskRegistry, and the background extra.
  • agent_config["memory_tool_backend"] and the memory_* tools, replaced by workspace files.
  • The top-level retry and circuit-breaker helpers; omniserve run --reload.

Fixed

Found by four agents that built applications from these docs alone, before the release:
  • A resumed run’s evidence is whole: the call a person approved appears in its own step, as approved by them; the run’s totals and its training record include every segment; get_run() shows what each approval is for.
  • An ask on delegation (delegate_<name>) pauses the lead, as documented; it was a tool error.
  • Full capture keeps governed tool arguments, and the calls a run_code program makes have theirs.
  • budget_status counts a person’s grant, as enforcement does.
  • A run’s deadline holds when the work does not stop when told.
  • A host allowlist a sandbox provider cannot enforce is refused when the agent is built.
  • 0.3.9 was published without omnicoreagent.core.workspace, so OmniCoreAgent(...) failed on import. A test now fails if any source file of the package is left out of the wheel.

0.3.x and earlier

See GitHub Releases.