Writing ·
How to roll out runtime protection rules for AI agents
Start with audited agent actions, qualify detections against legitimate work, then scope and verify blocking rules while preserving investigation evidence.
By Youssef Hemimy · agent safety · runtime protection · AgentOps
Roll out runtime protection as an evidence-backed promotion, not a blanket switch. First map which agent actions your enforcement point actually sees. Run detections in audit mode, join each event to the agent, user, tool, action, and reason, and review legitimate-task false positives. Promote a high-confidence detection to a narrowly scoped block only when you can verify the block happens before execution, investigate its effects, and disable the rule if it interrupts authorized work.
Draw the enforcement boundary first
For each agent class, list the action surfaces it can use and the interception point for each: user request, model response, tool invocation, tool result, or endpoint process action. Mark paths that bypass the reviewed hook. A policy cannot block an action it never sees. Microsoft documents distinct coverage: Agent 365 evaluates Work IQ MCP tool invocations, including onboarded customer MCP tools, while unsupported tools and agents outside that integration are not covered by that path. Copilot Studio tool-invocation protection is in preview and uses a different integration. Foundry preview protection evaluates requests, responses, tool invocations, and tool responses. Local agents require separate Defender for Endpoint onboarding and active mode.
That is product-specific coverage, not a universal agent-security architecture. For a self-hosted agent, make your own interception inventory and test the unobserved paths. The separate runtime-monitor evasion test guide explains how to check effects beyond a first blocked call.
Use audit evidence to qualify a block
Microsoft's built-in cloud-agent default rule audits activity without stopping it. Custom rules can block matching actions before they execute and record the behavior. That sequence supports a disciplined promotion decision: collect enough normal and suspicious traffic to understand the rule's scope, then block only where the cost of a missed action outweighs a reviewed interruption risk. It does not establish a universal score or waiting period.
| Decision | Evidence to retain | Failure that prevents promotion |
|---|---|---|
| Inventory coverage | Agent platform, integration, tool/action class, and exact interception point. | Important execution paths are assumed covered but not observed. |
| Qualify detection | Audit event, reason, agent and user identity, tool, intended action, and independent outcome. | Benign tasks trigger the rule or the event cannot be explained. |
| Scope block | Named detection type, specific agents, exclusions, owner, and disable procedure. | Only an all-agent rule is available for a signal with known false positives. |
| Prove effect | Blocked event joined to tool execution and downstream state; authorized-task regression results. | A block event appears but the protected effect still occurs elsewhere. |
| Preserve investigation | Queryable behavior, related entities, alert dependencies, and evidence-retention settings. | Promotion removes the signal used by responders without a replacement query. |
This table is a recommended operational decision record. It is not a claim that Microsoft's controls automatically perform every check. Keep the rule owner, scope, evidence, and rollback in the same release record you use for other agent acceptance decisions.
Do not mistake changed telemetry for reduced risk
Defender records real-time audit and block events as behaviors in BehaviorInfo, with related entities in BehaviorEntities. The documented behavior includes what happened, why it was considered risky, and the involved agent, user, and tool. Its investigation flow can correlate those records with AlertInfo, AlertEvidence, CloudAppEvents, and AgentsInfo. But the product documentation says near-real-time detection alerts continue only in audit mode; when a blocking rule covers an agent, those alerts are not generated for that agent. Before promotion, move any alert-dependent investigation or automation to a verified behavior-based path.
Keep proposed action, policy verdict, actual tool execution, and downstream effect separate. A block record supports a claim about that reviewed action; it does not by itself prove that the agent stopped, that no alternate tool reached the same resource, or that a later retry was safe. Join the behavior to an execution ledger outside the agent's control. See the trace-authority guide for that evidence boundary.
Review prompt evidence as a separate data decision
Microsoft says prompt evidence collection is enabled by default for alerts and includes snippets classified as suspicious and relevant; it also says secrets and sensitive data are redacted, while customer conversations can still be sensitive. Decide whether that evidence is needed for your response workflow, who may view it, and how it is retained. Do not equate redaction with permission to collect every prompt. A block rollout and an alert-evidence setting are two different approvals.
Keep the claim within the documented deployment
The Microsoft documentation describes its supported integrations and preview features, not measured false-positive rates, universal tool coverage, or a tested rollout threshold. NIST's AI Risk Management Framework Playbook offers voluntary, context-dependent risk-management guidance, not a certification of this particular control. Use those sources to design and review a local decision, then test the actual agent, policy, and execution paths you operate.
Sources
FAQ
Should every AI agent detection become a blocking rule?
No. Start with observed audit behavior, verify that the detection is specific enough to avoid interrupting legitimate tasks, and limit a block to the agents and actions for which you have reviewed evidence. Keep an owner and rollback path for each rule.
Does Microsoft Defender protect every tool an agent can call?
No. Coverage varies by agent platform and integration. For Agent 365, Defender evaluates Work IQ MCP tool invocations, including onboarded customer MCP tools; unsupported tools and agents outside that integration are not covered by that path. Copilot Studio, Foundry, and local endpoint agents have different requirements and coverage.
Will a blocking rule still produce the same near-real-time alerts?
Not necessarily. Microsoft documents audit and block events as BehaviorInfo records, but says near-real-time detection alerts continue only in audit mode and are not generated for an agent covered by a blocking rule. Verify investigation queries and alert-dependent workflows before promotion.
Building something that has to hold up?
We do this work for teams — agent reliability hardening, custom MCP servers, and full-stack AI systems built to survive production.