Skip to content

Research for AIthat acts.

We study how agents behave when they use tools and change things around them. Where does a review help, what does it miss, and how can we measure the difference?

Dipesh Tharu Mahato

Our preliminary DARC study evaluates runtime controls on 34,734 action-level examples converted from four public agent-safety benchmarks.

Read the research note

These are preliminary results on converted examples, not proof of safety in live deployments.

What we're trying
to understand.

Our agenda covers individual tool calls and the larger systems they belong to: permissions, memory, other agents and the people responsible for the work.

Runtime safety

How should a system review a tool call before an agent changes a file, sends a message or touches a live service?

Tools and permissions

A tool's reach depends on its permissions. We study how private context, memory and access to sensitive resources change the review.

Multi-agent coordination

One agent's action can become another's input. We examine delegation, shared tools and failures that spread across a workflow.

Interpreting risk

We investigate what internal signals, action history and policy matches can tell us about the risk of a proposed action.

Identity and provenance

Who created an agent? What could it access, and who approved its actions? We study how to keep those connections traceable.

Evidence for deployment

Evaluations, traces and human reviews can inform deployment decisions. We examine what that evidence supports and where it falls short.

Work you
can inspect.

We want the research to be useful outside the paper. Our work is designed to produce better evaluations, practical review interfaces and evidence that teams can examine.

Explore our open source

Action-level benchmarks

Evaluations that test what an agent does, including the tools and environment around it.

Control protocols

Policy and review interfaces that developers can place before execution.

Risk taxonomies

Shared ways to describe risky actions and the evidence collected during a review.

Technical reports

Methods, results and limitations that other researchers can inspect.

Working on
the same questions?

If you build agent systems, design evaluations or study AI safety, we'd like to hear about your work.