From Research to Product

AI Agent Monitoring

Our monitoring research aims to translate compute into scalable security. We are building Watcher, a monitoring tool for coding agents, to bring this frontier research into production.

highlights

Latest updates

July 23, 2026

What makes a good monitoring prompt?

We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring, as opposed to real-time binary prompts.

Read more

Product

July 13, 2026

Red-teaming auto mode: lessons from our first external monitor campaign with Anthropic

Apollo Research ran a pilot monitor campaign for auto mode, Anthropic's monitoring system for Claude Code agents that decides whether an agent's next action should be allowed or blocked.

Read more

Research

July 7, 2026

Evaluating LLM Calibration for Coding-Agent Monitoring

We evaluate 16 LLMs scoring coding-agent trajectories on a 1-10 severity scale across three datasets and five failure modes to assess their capabilities and calibration as coding agent monitors.

Read more

Research

May 8, 2026

A scalable monitoring research agenda

The goal of this agenda is to translate compute into security at scale. In the best case, we figure out how to spend on the order of $10-100M to produce extremely good monitors for coding agents.

Read more

Research

watcher

Your AI agent monitoring tool by Apollo Research

Watcher catches dangerous coding agent behavior before it becomes an incident.

Meet Watcher