Prompt Injection Guardrails
Every agent screens for injection phrasing: text that tries to redirect the model or reach its instructions. The screen runs on the request before the model sees it, and on every tool result (local tools, MCP tools, sandbox commands) before that result enters the model’s context. It is on by default, and needs no configuration. By the end of this page you will have seen a request refused, a poisoned tool result replaced, and a suspicious one flagged, and you will know where each is recorded.Every output on this page is what the code printed when it was run. The model’s
wording will differ on your run; what the guardrail does will not.
A request that is refused
run() returned status "error" with
termination_reason "safety_guard", the verdict in guardrail_result, and a
fixed refusal as the response. The run’s record says blocked, and its trace
ends with status aborted_safety_guard. The guard also logs what it found to
the omnicoreagent.guardrails logger, which your logging configuration
controls.
A steering message sent to a running agent
with agent.steer() is user input too: it is checked the same way, and a
blocked one is never queued (steer() returns {"status": "blocked", "guardrail_result": {...}}).
A tool result that is replaced
Tool output is where injections usually arrive: a web page, a file, an email. Here one tool returns a page with an instruction planted in it, and another returns a note framed as a system message:fetch_page
became an error whose data is [Tool output blocked by guardrail: ...], and the
model told the user so. The [SYSTEM] note is only suspicious (framing, with
no instruction in it), so by default it was passed through and flagged in the
trace.
With "guardrail_config": {"suspicious_output_action": "block"}, the same run
blocked both:
"guardrail_mode": "input_only", tool results are not screened at all.
The page reached the model as written; this time the model noticed the planted
line on its own, which you should not count on:
[MCP response blocked by guardrail: ...].
How it works
1
Normalize
The text is folded before matching: Unicode normalized, zero-width
characters removed, leetspeak read as letters (
1gn0r3 4ll pr3v10us 1nstruct10ns is matched as ignore all previous instructions), and
spaced-out letters (o v e r r i d e) joined. Hex identifiers such as run
ids and commit SHAs are left alone.2
Look for evidence
Two kinds count, and only these (see the table below). Structure (markdown,
code fences, diffs, JSON), vocabulary (the words system, prompt,
override, secret on their own), length and entropy are not evidence, so
ordinary developer text passes.
3
Decide from the kinds found
Nothing found is
safe. Hidden or framed content alone is suspicious
(dangerous in strict mode). One kind of intent is dangerous. Two kinds
of intent, or intent beside hidden content, is critical. The verdict never
comes from adding up weak signals.4
Act
A request that is not
safe is refused. A tool result that is dangerous
or critical is always replaced; a suspicious one follows
suspicious_output_action. Each decision is an event in the trace.What counts as evidence
You can try the screen on a text without an agent:
THREAT DETECTED: CRITICAL (score: 32, confidence: 0.95); they are left out above. They are warnings on the
omnicoreagent.guardrails logger, printed by Python because nothing else
handles them. The corpus that holds the
screen to this, ordinary developer text that must pass and attacks that must be
blocked, is tests/test_guardrail_corpus.py in the repository.
Modes
agent_config["guardrail_mode"]:
A run resumed after an approval or a crash
does not screen its original request again.
What the evidence records
The verdict (
guardrail_result, and the output of a request’s
guardrail_violation event) holds threat_level, is_safe, flags (what
matched, up to ten), message, confidence, threat_score,
recommendations, input_hash, input_length, detection_time,
analysis_time_ms, pattern_version and flag_count. A tool result’s event
records the tool’s name, the field checked, threat_level, threat_score,
input_hash, message and the action taken (blocked or flagged); the
blocked text itself stays out of the model’s context.
Options
Pass these inagent_config["guardrail_config"]:
When things go wrong
A long request was refused: Input exceeds maximum allowed length
A long request was refused: Input exceeds maximum allowed length
A request over Raise
max_input_length characters (10,000 by default) is refused
before it is analysed. A 16,000-character log pasted into run() returned:max_input_length, or put large inputs in the
workspace and ask the agent to read
the file.An ordinary request or tool result was blocked
An ordinary request or tool result was blocked
Look at
guardrail_result["flags"] (or the guardrail_violation event):
it quotes what matched. If the text is legitimately about injection (a
security write-up, a test fixture), add an allowlist_patterns entry that
matches it; an allowlist still cannot pass a real override or extraction.
If the screen is wrong about ordinary text, that is a bug worth reporting
with the text.ValueError: guardrail_mode must be one of: full, input_only, off
ValueError: guardrail_mode must be one of: full, input_only, off
guardrail_config={"strict_mode": True}, not a mode.TypeError: DetectionConfig.init() got an unexpected keyword argument
TypeError: DetectionConfig.init() got an unexpected keyword argument
A key in Use the names in the table above.
guardrail_config that is not a setting:ValueError: blocklist_patterns[0] is not a valid regex
ValueError: blocklist_patterns[0] is not a valid regex
Patterns are compiled when the agent is built, so a bad one fails early:
Privacy boundaries
Redacting personal data and credentials is a separate policy: see Privacy and credentials.Next
Privacy and credentials
What is redacted from traces and outputs, and how keys are kept out.
Policies
Allow, deny or ask for every capability: the boundary that holds.
Security model
What each layer protects against, and what it does not.
Execution
Where commands run, and what a sandbox isolates.