Skip to main content

Basic Usage

The patterns you will use in almost every application: run a request and check how it ended, keep a conversation going per user, stream the answer as it is written, read back what a run did, and handle the ways a run can fail.
This page assumes an agent like the one in the quickstart, and LLM_API_KEY set. Every output shown is what the code printed when it was run; the model’s wording, the ids and the timings will differ on yours.

Run, and keep a conversation

  • Check status before using response. "success" means the model answered. Anything else means the run stopped: response then says why, and termination_reason names the cause (see below).
  • A session_id is a conversation. Use your user or ticket id. Runs with the same one see each other’s messages; a run without one gets a fresh session, whose id is in result["session_id"].
  • clear_session_history(session_id) forgets one conversation; clear_session_history() with no argument forgets every conversation of this agent.
History lives in process memory unless you choose a store, and is gone when the process ends. To keep it:
Each store needs its extra and its URL (Configuration, Memory).

Stream the answer

stream() runs the same loop as run() and yields the text as the model writes it, then one final event with the whole result.
  • The complete event carries everything run() returns (status, response, the ids, metric). Use it, not the joined deltas, as the answer: text the model writes alongside a tool call is intermediate.
  • An exception inside the run arrives as an error event instead of being raised.
  • Leaving the loop early, with aclosing, cancels the run.
  • Deltas are live only; the run’s trace keeps the final answer, not the pieces.
To stream without changing how you call the agent, pass a callback to run(). A plain function or a coroutine both work; each event has type text_delta and the text:
More on streaming and the events around it: Events.

Read back what a run did

Every run leaves a trace. result["trace_id"] finds it; result["run_id"] finds the run’s events.
The trajectory is the readable one; the quickstart prints it, and Observability has every field.

Tools

Tools are covered on their own pages: your Python functions in Local tools (the quickstart registers one), MCP servers in MCP. For MCP, call await agent.connect_mcp_servers() before the first run.

When things go wrong

A run that fails does not raise. It returns a result whose status is "error", and termination_reason says why. This one script meets three of them: an injection attempt, a request limit, and a model that does not exist.
Other statuses are not failures: awaiting_approval and awaiting_budget wait for a person, and interrupted was stopped on request. Each is continued with resume(run_id) (Durable runs). What does raise is a problem with the setup, before the run can start: a bad setting (ValueError from the constructor), no LLM_API_KEY (ValueError from the first run), a store that cannot be reached. So a production call looks like:
When the guardrail blocks an input it logs a warning such as THREAT DETECTED: CRITICAL (score: 32, confidence: 0.95) on the omnicoreagent.guardrails logger. Inside an agent you see it only if your application configures logging (for example logging.basicConfig()); the refusal itself is in the result (termination_reason safety_guard) and the trace. It is the guardrail working, not an error in your code.

Next

Configuration

Every layer of settings: model, agent, memory, workspace, telemetry.

Memory

Durable stores, history windows and summaries.

Every run is evidence

Trajectories, outcomes, training records, exporters.

Durable runs

Pauses, approvals, crashes and resume.