Outcomes and training records
A run ends with an answer, but whether the answer was any good is known later: the customer accepts it the next day, the pull request is merged, the tests pass in CI. By the end of this page you will have a run, an outcome recorded for it by a different process, and the run exported as one JSONL line an evaluator or a fine-tuning job can read.Every output on this page is what the code printed when it was run. The
model’s wording, the ids and the token counts will differ on your run.
A run, an outcome, a record
Three processes share one agent definition. The run record lives in SQLite so the other two can find it, and the traces live in the default place,./workspace/telemetry/.
1
One definition, imported by every process
pip install "omnicoreagent[postgres]" (it brings SQLAlchemy).2
The run
3
The next day, another process records the outcome
4
Export the runs worth learning from
The same question was asked once more in the session, and nobody rated
that answer. Export only the runs with a reward:
support.jsonl has one line, 23 KB: every message the model was sent at
both steps, in full, with the 15 tool definitions (get_order and the
built-in workspace tools).How it works
1
An outcome goes on the record and into the trace
record_outcome appends to the run’s record in the memory store
(get_run(run_id)["outcomes"]) and adds a run_outcome event to the
run’s latest trace, so the trajectory lists it under outcomes. It needs
the run’s record: the memory store must be one both processes reach
(SQLite, PostgreSQL, Redis, MongoDB), and the traces must be in the same
workspace for the trace half. A run may gather any number of outcomes, from
different sources; none replaces another.2
A training record is one finished run
training_records reads traces from the telemetry store, keeps the runs
that finished (completed, failed, blocked, cancelled), and builds
one record per run. A run that paused for an approval or a budget and was
resumed is one record: trace_ids lists its segments, and each step says
which segment it came from. A run still running or waiting is left out
until it finishes.3
Built from what the trace recorded
The record holds only what the trace holds. A run recorded without model
prompts (the
default capture preset) has nothing to learn from, and is
left out.What a record holds
Each step:
Every step has
resumed: False for a model turn; True for a step that
only ran calls a person approved, on resume, with no messages and those calls
in tool_calls.
An outcome
Each outcome gets an
outcome_id and recorded_at. Over HTTP,
OmniServe takes the same body at
POST /runs/{run_id}/outcome:
Choosing and exporting records
training_records takes one filter:
With none, it reads every run in the telemetry store. Anything else — by
reward, by label, by date, by model — is a Python filter over the records, as
rated was above. Each record is plain JSON: json.dumps(record) per line is
a JSONL file. The messages are already in the chat format a fine-tuning job
reads; which records and which fields to keep for a given trainer is yours to
choose. To score a batch, an evaluator reads request, the last step’s
response, and outcomes.
Token details
A trainer that reuses a run (rather than only imitating it) needs the tokens the model chose and their probabilities. Ask the provider for them, and record them:request_settings show both settings were sent), so token_details was
null. A hosted reasoning model often returns none. A provider that refuses the setting outright names
it, and the call is retried without it (the model call’s retries in the
trajectory say so).
What privacy settings change
The record is built from the trace, so it holds what the trace was allowed to keep (privacy). The same run through two agents, one withcapture: "default":
A redacted value is redacted in the record too: the
api_key above is
[REDACTED] in the tool call’s observation and in the tool message of step
2’s messages. A trainer never sees it. Every
setting is in the telemetry reference.
How long you have
An outcome can arrive only while the run’s record exists, and a record can be built only while the run’s traces exist:
If outcomes arrive weeks later, or you export monthly, raise both
(how long a run’s evidence is kept).
Once the record is gone but the traces remain, a training record still carries
the outcomes that were recorded, read from the trace.
Options
The methods are in the OmniCoreAgent reference.
When things go wrong
LookupError: No run …
LookupError: No run …
run_retention_days), or the process that ran it kept records in
memory, the default, and they went with it. Give every process the same
durable store, as agent_def.py does. Over HTTP it is a 404.ValueError: record_outcome needs a source
ValueError: record_outcome needs a source
source is required and must not be blank. Over HTTP it is a 422.training_records returns []
training_records returns []
One of: the runs have not finished (a run waiting for an approval is left
out until it ends); the agent records with
capture: "default"; the
traces are past retention_days; or this process reads a different
telemetry store than the one that ran them (another workspace, or
storage: "memory" in another process).token_details is null
token_details is null
Either
record_token_details is off, or the provider returned no
logprobs for that call (above).Next
Read a run
The trajectory a training record is built from, step by step.
Observability
Traces, stores and exporters.
Headless runs in CI
Run the agent in a CI job, then record whether the tests passed.
Harbor and Terminal-Bench
Runs on benchmark tasks, each with a verifier’s reward.