Models
The model is two lines ofmodel_config — a provider and a model name — and
the key is always LLM_API_KEY. By the end of this page your agent will run on
the provider you choose, with the settings you choose, and you will have seen
what was actually sent.
Pick a provider
Only the OpenAI configuration was run for this page. The others are shown as
configuration; check the model name against your provider’s current list.
Set the model’s parameters
How it works
- One route: LiteLLM.
providerandmodelbecome LiteLLM’s"<provider>/<model>"(openai/gpt-5.4-miniabove). Every provider returns the same turns to the runtime, so tools, streaming, memory and traces work the same on all of them. - One key.
LLM_API_KEYis read when the agent starts (its first run) and handed to the provider you named — whose own key variable you never set. An agent can carry its own instead —"api_key": "..."in itsmodel_config— for two providers in one process. Keys are never written into a trace. Ollama needs none. - Settings are sent as given. Nothing is added for you: with no
max_tokens, the provider’s own limit applies (with a budget, a call is held to 4,096 output tokens). If a model refuses a setting by name — for exampletemperatureon a reasoning model — the call is retried once without it, the setting is not sent again to that model, and the retry is recorded on the model call (facts["retries"], with the setting underdropped). - Transient failures are retried — rate limits, timeouts, connection errors — up to three times with backoff. A wrong key, an empty account or an unknown model is not retried: the run ends at once with the reason. A streamed call is not retried, since its text was already delivered.
Thinking and reasoning models
Thinking models return data with each turn that must be sent back unchanged on the next request, or a multi-step tool loop fails on its next step. OmniCoreAgent keeps it with the turn and sends it back to the provider that issued it:
This works with
run() and stream(), across runs in the same session (the
data is kept in stored history), and after context compression. Other
providers receive exactly the fields they always did. Enable thinking with the
provider’s usual setting, for example:
continuation_dropped) instead of storing a signature the provider would
reject; the running agent keeps its own copy.
Traces record that continuation data was present (counts and a digest on each
model call) and never store signatures or encrypted values; the thinking text
itself is recorded only with capture="full".
Azure OpenAI and Ollama
- Azure OpenAI
- Ollama
LLM_API_KEY.Your own endpoint (base_url)
A server that speaks a provider’s API — vLLM, LM Studio, llama.cpp’s server, a
company gateway or a proxy — is reached by naming the provider whose API it
speaks and giving its URL:
- The URL is sent with each of this agent’s calls, streamed or not. It does not change the process environment, so two agents in one process can use two endpoints.
- The key is still
LLM_API_KEY. A local server that checks no key accepts any value:export LLM_API_KEY=local. - The URL is not recorded with the run’s model settings (a URL can carry a token).
- A dollar budget needs a published price for the model name. A model with no
price is governed by the token and call limits: set a
model_tokensormodel_callsbudget, ortotal_tokens_limit, for it. - It works with any provider:
ollamawithbase_urlreaches Ollama on another machine.
Options
Other keys are accepted and ignored. See also the
OmniCoreAgent reference.
When things go wrong
The first run in a process is slower than the rest
The first run in a process is slower than the rest
The first model call loads the model client (LiteLLM), about 2 seconds on
a typical server and longer on a slow machine or a cold start. Later calls
in the same process do not pay it. OmniServe and the background manager
load it while they start, so no request pays it there.
ValueError: LLM_API_KEY not found in environment variables
ValueError: LLM_API_KEY not found in environment variables
Raised by the first
run() when the key is not set in the shell that runs
the script: export LLM_API_KEY=.... The runtime does not read .env
files itself; call load_dotenv() from python-dotenv at the top of your
script if your key is in one.ValueError: Unsupported provider: OpenAI. Supported: openai, anthropic, groq, ollama, azure, gemini, deepseek, mistral, openrouter
ValueError: Unsupported provider: OpenAI. Supported: openai, anthropic, groq, ollama, azure, gemini, deepseek, mistral, openrouter
The provider name is lowercase and one of these nine. For another
OpenAI-compatible service, use
openrouter if it is listed there.ValueError: model_config.model is required
ValueError: model_config.model is required
Name the model as well as the provider. (
model_config.provider is required
is the same message for the provider.)The model call was refused: the provider rejected the API key (invalid_api_key). Fix the account; retrying will not help.
The model call was refused: the provider rejected the API key (invalid_api_key). Fix the account; retrying will not help.
The run ends with
status "error" and this response. The key is
wrong, or it belongs to a different provider than model_config["provider"].
An account out of credits reads the provider account has no credits left (insufficient_quota) in the same place.The model call was refused: the provider does not serve the model … to this account
The model call was refused: the provider does not serve the model … to this account
The model name is misspelled, retired, or not available to your account.
Check it against the provider’s model list, and that
provider is the one
that serves it.Next
Quickstart
A tool, a memory and a trace, in five minutes.
Budgets and durable runs
Put a price on runs, checked before every model call.
Streaming and events
Show the answer as the model writes it.
Configuration
Every setting of an agent in one place.