Telemetry and exporters
Every run records a trace: a span for each piece of work and an event for each thing that happened, kept on local disk with no service to run. By the end of this page you will have followed a run live, event by event, sent its trace to an OpenTelemetry collector and a JSONL file, and know what is recorded, where, and for how long. To read a run back step by step, see Read a run.gpt-5.4-mini. The model’s wording, token counts, times and IDs will differ
on your run.Watch a run, then send it on
The collector here is a small stand-in for an OpenTelemetry collector: it accepts OTLP/HTTP on port 4318 and prints what arrives. It needs theotel
extra (pip install "omnicoreagent[otel]").
exports/traces.jsonl has one line: the whole trace, as JSON.
What happened:
- The watcher saw every event as it was recorded: the request, the run’s
header (
run_configuration), each step’s context, model call, the policy’s decision on the tool call (governance is on by default), the tool call, and the answer. It passedrun_id=to see only this run. result["metric"]counted two model requests: this run’s two model calls.- The trace was stored locally (
jsonl, in./workspace/telemetry/) and exported when it ended: 14 spans to the collector, the same trace to the file.
What a trace records
telemetry_config decides what goes into a trace. The facts are always
kept: IDs, the links between records, tokens, cost, latency, outcomes, policy
decisions and the run’s header are event metadata, recorded under every
setting. The payloads (what the model was sent and answered, tool
arguments and results) follow the capture policy:
record_* switch you leave unset; one you set always
wins:
redact_keys (values of
keys like api_key, token, password become [REDACTED] at any depth)
and the privacy filter (personal data, in the
trace only, never in the run). A payload over max_payload_bytes (64,000)
is cut to a marker with its size and checksum, unless it is
offloaded. Each payload carries its capture state, and
Read a run
shows how a trajectory reports what is missing.
Where traces are kept
telemetry/traces-archive/): one body per trace and a SQLite
index, so reading one trace does not read them all. Several server processes
can share one archive: archive_index_url makes the index a shared database,
and archive_target: "object_storage" keeps the bodies in the workspace’s
bucket (Stores and scale).
Recording never fails a run. A write that fails or takes longer than
persistence_timeout_seconds (5) is reported, and the trace is marked
incomplete: True with evidence_status partial. With strict: True the
error is raised instead. A trace’s metadata["telemetry_storage"] says where
it was kept (jsonl or memory).
Read the raw trace
The trajectory is the readable view; the trace is the record as stored, spans and events, as a dict. With the run above:The live event stream
Every event gets astream_cursor, an opaque string that increases with
every event. Keep the last one you saw to carry on from there.
trace_id, run_id, session_id, task_id and event_types
to narrow it; cursor=None starts from the first stored event. A reader
that falls 1,000 events behind is stopped with Telemetry stream queue overflow; reconnect from a cursor rather than silently skipping events.
Over HTTP, OmniServe streams a run while
it runs (POST /run), and serves the stored record under /telemetry. Add
?run_id= to isolate one run in a shared session:
id:, so a client
that reconnects with Last-Event-ID (or ?cursor=) carries on where it
stopped. The same events are behind /events/{session_id}?run_id=....
Usage metrics
For a quick count without reading traces: eachrun() returns a metric
(requests, request_tokens, response_tokens, total_tokens,
total_time), and await agent.get_metrics() adds them up for this agent
object since it was built. requests counts model calls, not runs, and
average_time is per model call. The counts live in the process and start
again with it; OmniServe serves them at GET /metrics. For cost, and for
anything that must survive a restart, read the run’s totals
(Read a run).
Exporters
The trace store needs no service. Exporters send a copy elsewhere: to any OTLP/HTTP endpoint (an OpenTelemetry collector, or a vendor that takes OTLP), LangSmith, Opik, or a JSONL file. Configure them on the agent to export each trace when it ends, or export one when you choose:error. With strict=True, export_trace raises
instead. When an exporter configured on the agent fails at the end of a run,
the trace records a telemetry_error event naming the exporter.
export_trace exports one trace. A run that paused has one per segment, and
a run that delegated has one per child: list_telemetry_traces(run_id=...)
and get_trace_family(trace_id=...) list them to export each.
Each destination
telemetry_exporters takes dicts (destination and the arguments below),
names, or exporter objects; build_telemetry_exporter(destination, **arguments)
builds one.
x-api-key and Langsmith-Project; Authorization,
Comet-Workspace and projectName):
omnicoreagent.run_id, omnicoreagent.session_id,
omnicoreagent.evidence_status, gen_ai.system, gen_ai.request.model,
and, on model calls, gen_ai.usage.input_tokens and
gen_ai.usage.output_tokens. An export at the end of a run is given up
after export_timeout_seconds (5). What is exported is what was recorded:
the same capture policy and redaction.
Large payloads
A payload overmax_payload_bytes is cut short. To keep it whole, offload
it: the trace keeps a reference, and the value is stored beside the traces:
telemetry/payloads/ in the workspace, or in the workspace’s S3 or R2
bucket with offload_target: "object_storage". This is the trace’s copy;
what the model saw is set by tool offload,
separately.
How long traces are kept
None keeps everything. Retention runs once per agent, when it starts its
first run, and also when a process first reads its traces (such as a first
get_run_trajectory). To run it now, and see what it did:
last_prune counts the traces removed and how many of them were abandoned;
skipped_records counts damaged lines found in the log (their traces are
marked incomplete). payload_store is None because this agent does not
offload payloads. OmniServe serves the status at GET /telemetry/retention.
Nothing you exported is touched.
A run’s record is kept longer than its traces (30 days): see how long a
run’s evidence is
kept.
How it works
The loop records as it goes
redact_keys, the privacy filter, and the payload size limit.Writes are batched and never block a run
Subscribers get each event once
At the end, the trace is exported
Options
Intelemetry_config:
telemetry_exporters=[...]. Methods, in the OmniCoreAgent
reference: get_trace, get_latest_trace,
get_trace_family, list_telemetry_traces, get_telemetry_stream_cursor,
stream_telemetry_after, get_telemetry_events_after, export_trace,
read_telemetry_payload, prune_telemetry, telemetry_retention_status,
get_metrics.
When things go wrong
ValueError: telemetry capture must be one of: default, full
ValueError: telemetry capture must be one of: default, full
"default", turn off a
record_* switch.OTLP export failed: [Errno 111] Connection refused
OTLP export failed: [Errno 111] Connection refused
telemetry_error event, and it is still in the local store, so
export it again with export_trace when the collector is up.OpenTelemetry export requires the 'otel' extra
OpenTelemetry export requires the 'otel' extra
telemetry_error event
(export_trace returns it as the result’s error):ValueError: Unknown telemetry exporter destination
ValueError: Unknown telemetry exporter destination
otlp, langsmith, opik, jsonl and memory.
For another backend that takes OTLP, use otlp with its endpoint and
headers.ValueError: No telemetry payload store is configured
ValueError: No telemetry payload store is configured
read_telemetry_payload needs offload_large_payloads: True; without
it, large payloads are truncated and there is nothing to read back.ValueError: Use only one of trace_id, run_id, or session_id
ValueError: Use only one of trace_id, run_id, or session_id
get_trace takes one way of finding the trace.Telemetry stream queue overflow; reconnect from a cursor
Telemetry stream queue overflow; reconnect from a cursor
stream_telemetry_after fell 1,000 events behind. Start
again from the last stream_cursor it got; nothing is lost.The traces are gone after a restart
The traces are gone after a restart
storage: "memory" they live in the process. The default keeps
them on disk, in ./workspace/telemetry/ relative to where the process
runs: a process started elsewhere, or in a container without a volume
there, has its own.