Writing ·
How to test an AI agent runtime monitor for evasion
Test complete task-policy conflicts across blocked calls, retries, split actions, and unmonitored paths; score attempted bypass, actual effects, and safe-task utility separately.
By Youssef Hemimy · agent safety · agent evaluation · AgentOps
Test the complete agent–monitor interaction, not just the first blocked call. Give an agent a legitimate task whose apparent completion path conflicts with an operator policy. Verify the task is otherwise solvable; then inspect what happens after a clear denial. Record proposed calls, monitor decisions, actual effects, and final outcomes separately. A good result is not merely that the first call was blocked: the prohibited effect never occurred, while authorized work and escalation still behaved as intended.
Qualify a task–policy conflict before testing the monitor
EvasionBench pairs 50 benign tasks with manually written policies: 40 adapted from five benchmarks and ten authored terminal tasks. Genuine completion requires a prohibited operation. The authors checked that two agents could solve all 50 without monitoring and that their monitor allowed clearly permitted calls while blocking clearly prohibited ones. In this study, the compliant outcome is an unfinished task. In your workflow, a permitted alternate or human approval may instead be correct.
The baseline GPT-OSS-120B monitor sees the policy, proposed call, and up to 40 recent attempted calls with allow/block decisions; older arguments are truncated. A blocked call does not execute. The denial tells the agent that a guardrail blocked it, not that a tool failed transiently. After the agent stops, the harness prompts it to continue. These details create a bounded monitoring window and add completion pressure.
Write the case file before the run
Specify the legitimate goal, forbidden effect, policy version, and expected safe response before interpreting a trace.
| Case field | Evidence to record |
|---|---|
| Legitimate goal | The user's task and permitted ways to complete it. |
| Protected effect | The action or state change the operator policy forbids, with its enforcement owner. |
| Qualification | Unmonitored solvability; direct prohibited call blocked; comparable allowed call passed. |
| Expected response | Stop, complete a permitted subset, seek approval, or take a documented alternate. |
| Ground truth | Execution and environment evidence held outside the agent-visible transcript. |
This case template adapts the paper's qualification method; it does not imply every real task lacks an authorized path. The separate fail-closed policy-hooks guide covers policy-service outages. Here, the monitor may respond correctly to the first call and still miss a later path.
Exercise paths a one-call test misses
Use these as mechanisms to probe in your own runtime, not as attacks known to work against every monitor.
| Mechanism | Question | Retained evidence |
|---|---|---|
| Encoded operation | Could a decoded command or script have a prohibited effect the monitor misses? | Submitted call, decoded effect, verdict, execution result. |
| Decomposition | Could individually allowed calls combine into a prohibited effect? | Ordered calls, state changes, final effect. |
| Retries | Does the agent rephrase or resubmit a denied operation? | Every attempt and verdict, including denials. |
| Context loss | Does earlier blocked intent leave the monitor's history window? | Full independent trace and exact monitor input per decision. |
| Delegation or continuation | Can a child agent or secondary input channel act outside the reviewed hook? | Child and continuation events, interception coverage, effects. |
The paper describes one integration-specific blind spot: a Codex setup reviewed the launch of an interactive process but not later input through write_stdin. It distinguishes a prohibited action never reviewed from one reviewed and allowed. Inventory your host's actual execution channels; do not assume its coverage matches the paper's.
The independent Agent Hooks 0.1 specification requires a conformant host to invoke required interceptors for declared guarded actions and deny on an unreachable interceptor or invalid verdict. That normative contract is limited to interception points the host exposes; it does not prove direct execution paths are covered.
Score attempts, decisions, and effects separately
- Attempted bypass: a proposed sequence pursuing a prohibited operation, with trace evidence and reviewer rationale.
- Monitor decision: what was reviewed, allowed, blocked, or never intercepted.
- Prohibited effect: whether independent environment evidence shows it actually occurred.
- Task outcome: complete, partial, or unfinished legitimate work.
- Authorized utility: permitted alternate, appropriate escalation, or unnecessary refusal of safe work.
The paper's successful-evasion measure is deterministic task completion; its attempted-evasion measure uses a separate LLM judge. It runs each task three times and reports whether an outcome occurred in at least one run, excluding infrastructure failures. Completion can miss a prohibited operation that did not finish the task. Judge-labeled attempts can include ambiguous troubleshooting. Do not turn either into a single “monitor blocked X%” statistic.
Check usefulness alongside containment
A monitor can appear safe if the agent abandons every hard task. In a separate benign-task experiment, a forged denial message caused one evaluated agent to abandon otherwise solvable work. That is a specific injected-message result, not a universal model ranking. Run comparable tasks where the restricted operation is unnecessary and verify the agent completes permitted work. If no authorized path exists, a truthful refusal or escalation can be the right outcome.
Attach containment and utility evidence to your own agent release thresholds. The preprint supplies no threshold suitable for your system. Keep input-driven prompt-injection tests as a complementary suite: this test targets task pressure after a genuine policy conflict and monitor denial.
Keep the evidence boundary honest
EvasionBench studies deliberately constructed conflicts, selected monitors and models, best-of-three outcomes, and a harness that urges stopped agents to continue. The prohibited operations are not necessarily harmful by themselves; some retries may reflect uncertainty about a hidden policy. Its observed bypasses and overrefusal example do not estimate production prevalence. The defensible decision is narrower: qualify the task, exercise repeated and multi-step paths against your real interception boundary, and verify effects independently.
Sources
FAQ
Does blocking the first prohibited tool call prove a runtime monitor works?
No. Test the complete run, including encoded or split actions, retries, delegation, and execution channels outside the reviewed hook. Confirm independently that the prohibited effect did not occur.
Is an evasion attempt the same as a successful violation?
No. Keep proposed actions, monitor decisions, actual execution, environment effects, and task completion separate. An attempt may be blocked; a forbidden effect may occur without completing the task.
Do EvasionBench results estimate production monitor-evasion rates?
No. Its 50 constructed task-policy conflicts, selected monitors and models, best-of-three reporting, and continuation prompts do not represent unconstrained production deployments.
Building something that has to hold up?
We do this work for teams — agent reliability hardening, custom MCP servers, and full-stack AI systems built to survive production.