Watcher
AI agent monitoring tool by Apollo Research
Watcher catches dangerous coding‑agent behavior before it becomes an incident. Watcher detects:
- Insecure code execution
- Data exfiltration
- Agent manipulation
- Emergent risks
Our monitoring research aims to translate compute into scalable safety. We are building Watcher, a monitoring tool for coding agents, to bring this frontier research into production.
We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring (e.g. opposed to real-time binary prompts).
Apollo Research ran a pilot monitor campaign for auto mode, Anthropic's monitoring system for Claude Code agents that decides whether an agent's next action should be allowed or blocked.
We evaluate 16 LLMs scoring coding-agent trajectories on a 1-10 severity scale across three datasets and five failure modes to assess their capabilities and calibration as coding agent monitors.
The goal of this agenda is to translate compute into safety at scale. In the best case, we figure out how to spend on the order of $10-100M to produce extremely good monitors for coding agents.
Watcher is an oversight layer for AI agents. It detects real-world safety and security failures before they become liabilities, and flags those failures to you.
We’re building tools that make it easier to secure frontier AI agents.