Skip to main content

Outcomes and training records

A run ends with an answer, but whether the answer was any good is known later: the customer accepts it the next day, the pull request is merged, the tests pass in CI. By the end of this page you will have a run, an outcome recorded for it by a different process, and the run exported as one JSONL line an evaluator or a fine-tuning job can read.
Every output on this page is what the code printed when it was run. The model’s wording, the ids and the token counts will differ on your run.

A run, an outcome, a record

Three processes share one agent definition. The run record lives in SQLite so the other two can find it, and the traces live in the default place, ./workspace/telemetry/.
1

One definition, imported by every process

SQLite needs pip install "omnicoreagent[postgres]" (it brings SQLAlchemy).
2

The run

3

The next day, another process records the outcome

This process ran with no model key in its environment: recording an outcome, and reading runs back, never calls the model.
4

Export the runs worth learning from

The same question was asked once more in the session, and nobody rated that answer. Export only the runs with a reward:
support.jsonl has one line, 23 KB: every message the model was sent at both steps, in full, with the 15 tool definitions (get_order and the built-in workspace tools).

How it works

1

An outcome goes on the record and into the trace

record_outcome appends to the run’s record in the memory store (get_run(run_id)["outcomes"]) and adds a run_outcome event to the run’s latest trace, so the trajectory lists it under outcomes. It needs the run’s record: the memory store must be one both processes reach (SQLite, PostgreSQL, Redis, MongoDB), and the traces must be in the same workspace for the trace half. A run may gather any number of outcomes, from different sources; none replaces another.
2

A training record is one finished run

training_records reads traces from the telemetry store, keeps the runs that finished (completed, failed, blocked, cancelled), and builds one record per run. A run that paused for an approval or a budget and was resumed is one record: trace_ids lists its segments, and each step says which segment it came from. A run still running or waiting is left out until it finishes.
3

Built from what the trace recorded

The record holds only what the trace holds. A run recorded without model prompts (the default capture preset) has nothing to learn from, and is left out.

What a record holds

Each step: Every step has resumed: False for a model turn; True for a step that only ran calls a person approved, on resume, with no messages and those calls in tool_calls.

An outcome

Each outcome gets an outcome_id and recorded_at. Over HTTP, OmniServe takes the same body at POST /runs/{run_id}/outcome:

Choosing and exporting records

training_records takes one filter: With none, it reads every run in the telemetry store. Anything else — by reward, by label, by date, by model — is a Python filter over the records, as rated was above. Each record is plain JSON: json.dumps(record) per line is a JSONL file. The messages are already in the chat format a fine-tuning job reads; which records and which fields to keep for a given trainer is yours to choose. To score a batch, an evaluator reads request, the last step’s response, and outcomes.

Token details

A trainer that reuses a run (rather than only imitating it) needs the tokens the model chose and their probabilities. Ask the provider for them, and record them:
They are large, so they are off by default. Whether they come back is the provider’s decision. With this model, the provider returned them for a plain chat call made directly, but not for the agent’s calls (the model call’s request_settings show both settings were sent), so token_details was null. A hosted reasoning model often returns none. A provider that refuses the setting outright names it, and the call is retried without it (the model call’s retries in the trajectory say so).

What privacy settings change

The record is built from the trace, so it holds what the trace was allowed to keep (privacy). The same run through two agents, one with capture: "default":
A redacted value is redacted in the record too: the api_key above is [REDACTED] in the tool call’s observation and in the tool message of step 2’s messages. A trainer never sees it. Every setting is in the telemetry reference.

How long you have

An outcome can arrive only while the run’s record exists, and a record can be built only while the run’s traces exist: If outcomes arrive weeks later, or you export monthly, raise both (how long a run’s evidence is kept). Once the record is gone but the traces remain, a training record still carries the outcomes that were recorded, read from the trace.

Options

The methods are in the OmniCoreAgent reference.

When things go wrong

The memory store has no record of that run. Either it was pruned (run_retention_days), or the process that ran it kept records in memory, the default, and they went with it. Give every process the same durable store, as agent_def.py does. Over HTTP it is a 404.
source is required and must not be blank. Over HTTP it is a 422.
One of: the runs have not finished (a run waiting for an approval is left out until it ends); the agent records with capture: "default"; the traces are past retention_days; or this process reads a different telemetry store than the one that ran them (another workspace, or storage: "memory" in another process).
Either record_token_details is off, or the provider returned no logprobs for that call (above).

Next

Read a run

The trajectory a training record is built from, step by step.

Observability

Traces, stores and exporters.

Headless runs in CI

Run the agent in a CI job, then record whether the tests passed.

Harbor and Terminal-Bench

Runs on benchmark tasks, each with a verifier’s reward.