We study how agents behave when they use tools and change things around them. Where does a review help, what does it miss, and how can we measure the difference?
These are preliminary results on converted examples, not proof of safety in live deployments.
DARC
Review the action in its context.
Action → Context → Decision
What we're trying to understand.
Our agenda covers individual tool calls and the larger systems they belong to: permissions, memory, other agents and the people responsible for the work.
Runtime safety
How should a system review a tool call before an agent changes a file, sends a message or touches a live service?
Tools and permissions
A tool's reach depends on its permissions. We study how private context, memory and access to sensitive resources change the review.
Multi-agent coordination
One agent's action can become another's input. We examine delegation, shared tools and failures that spread across a workflow.
Interpreting risk
We investigate what internal signals, action history and policy matches can tell us about the risk of a proposed action.
Identity and provenance
Who created an agent? What could it access, and who approved its actions? We study how to keep those connections traceable.
Evidence for deployment
Evaluations, traces and human reviews can inform deployment decisions. We examine what that evidence supports and where it falls short.
Work you can inspect.
We want the research to be useful outside the paper. Our work is designed to produce better evaluations, practical review interfaces and evidence that teams can examine.