Skip to main content

Prompt Injection Guardrails

Every agent screens for injection phrasing: text that tries to redirect the model or reach its instructions. The screen runs on the request before the model sees it, and on every tool result (local tools, MCP tools, sandbox commands) before that result enters the model’s context. It is on by default, and needs no configuration. By the end of this page you will have seen a request refused, a poisoned tool result replaced, and a suspicious one flagged, and you will know where each is recorded.
The screen is a first look, not the security boundary. It makes an attempt visible early and records it. What the agent may do is decided by its policy; where its commands run is decided by the sandbox. An injection phrased in a way no pattern knows still cannot make the agent do what the policy does not allow.
Every output on this page is what the code printed when it was run. The model’s wording will differ on your run; what the guardrail does will not.

A request that is refused

The model was never called. run() returned status "error" with termination_reason "safety_guard", the verdict in guardrail_result, and a fixed refusal as the response. The run’s record says blocked, and its trace ends with status aborted_safety_guard. The guard also logs what it found to the omnicoreagent.guardrails logger, which your logging configuration controls. A steering message sent to a running agent with agent.steer() is user input too: it is checked the same way, and a blocked one is never queued (steer() returns {"status": "blocked", "guardrail_result": {...}}).

A tool result that is replaced

Tool output is where injections usually arrive: a web page, a file, an email. Here one tool returns a page with an instruction planted in it, and another returns a note framed as a system message:
The run went on. The planted instruction never reached the model: fetch_page became an error whose data is [Tool output blocked by guardrail: ...], and the model told the user so. The [SYSTEM] note is only suspicious (framing, with no instruction in it), so by default it was passed through and flagged in the trace. With "guardrail_config": {"suspicious_output_action": "block"}, the same run blocked both:
And with "guardrail_mode": "input_only", tool results are not screened at all. The page reached the model as written; this time the model noticed the planted line on its own, which you should not count on:
An MCP tool’s result is screened where it arrives from the server, including its structured content; a dangerous one is replaced by [MCP response blocked by guardrail: ...].

How it works

1

Normalize

The text is folded before matching: Unicode normalized, zero-width characters removed, leetspeak read as letters (1gn0r3 4ll pr3v10us 1nstruct10ns is matched as ignore all previous instructions), and spaced-out letters (o v e r r i d e) joined. Hex identifiers such as run ids and commit SHAs are left alone.
2

Look for evidence

Two kinds count, and only these (see the table below). Structure (markdown, code fences, diffs, JSON), vocabulary (the words system, prompt, override, secret on their own), length and entropy are not evidence, so ordinary developer text passes.
3

Decide from the kinds found

Nothing found is safe. Hidden or framed content alone is suspicious (dangerous in strict mode). One kind of intent is dangerous. Two kinds of intent, or intent beside hidden content, is critical. The verdict never comes from adding up weak signals.
4

Act

A request that is not safe is refused. A tool result that is dangerous or critical is always replaced; a suspicious one follows suspicious_output_action. Each decision is an event in the trace.

What counts as evidence

You can try the screen on a text without an agent:
Run on its own like this, the guard also prints a line to stderr for each dangerous or critical input, such as THREAT DETECTED: CRITICAL (score: 32, confidence: 0.95); they are left out above. They are warnings on the omnicoreagent.guardrails logger, printed by Python because nothing else handles them. The corpus that holds the screen to this, ordinary developer text that must pass and attacks that must be blocked, is tests/test_guardrail_corpus.py in the repository.

Modes

agent_config["guardrail_mode"]: A run resumed after an approval or a crash does not screen its original request again.

What the evidence records

The verdict (guardrail_result, and the output of a request’s guardrail_violation event) holds threat_level, is_safe, flags (what matched, up to ten), message, confidence, threat_score, recommendations, input_hash, input_length, detection_time, analysis_time_ms, pattern_version and flag_count. A tool result’s event records the tool’s name, the field checked, threat_level, threat_score, input_hash, message and the action taken (blocked or flagged); the blocked text itself stays out of the model’s context.

Options

Pass these in agent_config["guardrail_config"]:
Every agent setting is in the agent settings reference.

When things go wrong

A request over max_input_length characters (10,000 by default) is refused before it is analysed. A 16,000-character log pasted into run() returned:
Raise max_input_length, or put large inputs in the workspace and ask the agent to read the file.
Look at guardrail_result["flags"] (or the guardrail_violation event): it quotes what matched. If the text is legitimately about injection (a security write-up, a test fixture), add an allowlist_patterns entry that matches it; an allowlist still cannot pass a real override or extraction. If the screen is wrong about ordinary text, that is a bug worth reporting with the text.
Strictness is guardrail_config={"strict_mode": True}, not a mode.
A key in guardrail_config that is not a setting:
Use the names in the table above.
Patterns are compiled when the agent is built, so a bad one fails early:

Privacy boundaries

Redacting personal data and credentials is a separate policy: see Privacy and credentials.

Next

Privacy and credentials

What is redacted from traces and outputs, and how keys are kept out.

Policies

Allow, deny or ask for every capability: the boundary that holds.

Security model

What each layer protects against, and what it does not.

Execution

Where commands run, and what a sandbox isolates.