From Research to Product
AI Agent Monitoring
Our monitoring research aims to translate compute into scalable security. We are building Watcher, a monitoring tool for coding agents, to bring this frontier research into production.
highlights
Latest updates
July 23, 2026
What makes a good monitoring prompt?
We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring, as opposed to real-time binary prompts.
Product
July 13, 2026
Red-teaming auto mode: lessons from our first external monitor campaign with Anthropic
Apollo Research ran a pilot monitor campaign for auto mode, Anthropic's monitoring system for Claude Code agents that decides whether an agent's next action should be allowed or blocked.
Research
July 7, 2026
Evaluating LLM Calibration for Coding-Agent Monitoring
We evaluate 16 LLMs scoring coding-agent trajectories on a 1-10 severity scale across three datasets and five failure modes to assess their capabilities and calibration as coding agent monitors.
Research
May 8, 2026
A scalable monitoring research agenda
The goal of this agenda is to translate compute into security at scale. In the best case, we figure out how to spend on the order of $10-100M to produce extremely good monitors for coding agents.
Research
our findings
All research
Apollo’s product vision
Product


watcher
Your AI agent monitoring tool by Apollo Research
Watcher catches dangerous coding agent behavior before it becomes an incident.

