Skip to main content

Models

The model is two lines of model_config — a provider and a model name — and the key is always LLM_API_KEY. By the end of this page your agent will run on the provider you choose, with the settings you choose, and you will have seen what was actually sent.

Pick a provider

Only the OpenAI configuration was run for this page. The others are shown as configuration; check the model name against your provider’s current list.

Set the model’s parameters

Every model call’s trace records the model the provider says served it, the settings that were sent, the tokens, and an estimated cost (from LiteLLM’s price table, not your invoice).

How it works

  • One route: LiteLLM. provider and model become LiteLLM’s "<provider>/<model>" (openai/gpt-5.4-mini above). Every provider returns the same turns to the runtime, so tools, streaming, memory and traces work the same on all of them.
  • One key. LLM_API_KEY is read when the agent starts (its first run) and handed to the provider you named — whose own key variable you never set. An agent can carry its own instead — "api_key": "..." in its model_config — for two providers in one process. Keys are never written into a trace. Ollama needs none.
  • Settings are sent as given. Nothing is added for you: with no max_tokens, the provider’s own limit applies (with a budget, a call is held to 4,096 output tokens). If a model refuses a setting by name — for example temperature on a reasoning model — the call is retried once without it, the setting is not sent again to that model, and the retry is recorded on the model call (facts["retries"], with the setting under dropped).
  • Transient failures are retried — rate limits, timeouts, connection errors — up to three times with backoff. A wrong key, an empty account or an unknown model is not retried: the run ends at once with the reason. A streamed call is not retried, since its text was already delivered.

Thinking and reasoning models

Thinking models return data with each turn that must be sent back unchanged on the next request, or a multi-step tool loop fails on its next step. OmniCoreAgent keeps it with the turn and sends it back to the provider that issued it: This works with run() and stream(), across runs in the same session (the data is kept in stored history), and after context compression. Other providers receive exactly the fields they always did. Enable thinking with the provider’s usual setting, for example:
Memory privacy redaction never alters signatures or encrypted reasoning. If the readable thinking text contains personal data that redaction must change, the stored copy of that turn leaves the signed data out (and records continuation_dropped) instead of storing a signature the provider would reject; the running agent keeps its own copy. Traces record that continuation data was present (counts and a digest on each model call) and never store signatures or encrypted values; the thinking text itself is recorded only with capture="full".

Azure OpenAI and Ollama

The key is your Azure OpenAI key, in LLM_API_KEY.

Your own endpoint (base_url)

A server that speaks a provider’s API — vLLM, LM Studio, llama.cpp’s server, a company gateway or a proxy — is reached by naming the provider whose API it speaks and giving its URL:
  • The URL is sent with each of this agent’s calls, streamed or not. It does not change the process environment, so two agents in one process can use two endpoints.
  • The key is still LLM_API_KEY. A local server that checks no key accepts any value: export LLM_API_KEY=local.
  • The URL is not recorded with the run’s model settings (a URL can carry a token).
  • A dollar budget needs a published price for the model name. A model with no price is governed by the token and call limits: set a model_tokens or model_calls budget, or total_tokens_limit, for it.
  • It works with any provider: ollama with base_url reaches Ollama on another machine.

Options

Other keys are accepted and ignored. See also the OmniCoreAgent reference.

When things go wrong

The first model call loads the model client (LiteLLM), about 2 seconds on a typical server and longer on a slow machine or a cold start. Later calls in the same process do not pay it. OmniServe and the background manager load it while they start, so no request pays it there.
Raised by the first run() when the key is not set in the shell that runs the script: export LLM_API_KEY=.... The runtime does not read .env files itself; call load_dotenv() from python-dotenv at the top of your script if your key is in one.
The provider name is lowercase and one of these nine. For another OpenAI-compatible service, use openrouter if it is listed there.
Name the model as well as the provider. (model_config.provider is required is the same message for the provider.)
The run ends with status "error" and this response. The key is wrong, or it belongs to a different provider than model_config["provider"]. An account out of credits reads the provider account has no credits left (insufficient_quota) in the same place.
The model name is misspelled, retired, or not available to your account. Check it against the provider’s model list, and that provider is the one that serves it.

Next

Quickstart

A tool, a memory and a trace, in five minutes.

Budgets and durable runs

Put a price on runs, checked before every model call.

Streaming and events

Show the answer as the model writes it.

Configuration

Every setting of an agent in one place.