Abstract
Tool-using language-model agents fail when lower-integrity or untrusted information gains causal control over privileged action fields, including recipients, URLs, paths, repositories, payments, schedules, and accounts. Existing defenses commonly mediate whole tool calls, suspicious text, global provenance, or broad approvals, whereas the security decision at a protected sink depends on which source is authorized to cause each field value. We present CAGE, a runtime reference monitor for field-level authorization over dependency graphs. Before a protected action executes, CAGE computes an ancestor-closed per-field dependency closure, canonicalizes the field value, and checks that every provenance source in the closure is authorized for that field, sink, value scope, and trace.
Critical fields also require supporting evidence from an authorized source, scoped approval, declassification, or trusted delegation. Under an explicit adapter-soundness contract and fail-closed mediation, we prove a field-level integrity property, implement CAGE with typed dependency graphs, canonicalizers, scoped approvals and delegations, memory and denial-feedback containment, and audit certificates, and evaluate it on recorded model-backed suites plus scoped AgentDojo evidence. In these suites, CAGE blocks the observed targeted attacks while preserving utility and reducing approval burden relative to approval-only mediation.