Reducing risks from scheming frontier AI
Scheming is one of the hardest problems in AI. We research how it emerges, develop ways to detect it, and work to prevent the harm it could cause.
Our partners
Our partnerS
Our partners:
Our Mission
Securing frontier AI systems across development, deployment, and governance.
Science
We conduct fundamental research into the science of scheming, how it emerges, and how to detect and mitigate it. We run pre-deployment evaluations of frontier AI systems to detect strategic deception, evaluation awareness, and misaligned behavior.
Measuring Reward-Seeking via Contrastive Belief Updates
Visible forms of misbehavior are dropping in frontier models. Does that mean the models are becoming aligned, or are they just getting better at doing whatever they believe their grader rewards? Our paper finds that production reinforcement learning increases reward-seeking.
We need 3rd party Training-Run Assessments
Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety.
Metagaming matters for training, evaluation, and oversight
Metagaming can complicate how we interpret behavior, and current models still give us a chance to study it directly.
Monitoring
Our monitoring research aims to translate compute into scalable security. We are building Watcher, a monitoring tool for coding agents, to bring this frontier research into production.
What makes a good monitoring prompt?
We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring, as opposed to real-time binary prompts.
Red-teaming auto mode: lessons from our first external monitor campaign with Anthropic
Apollo Research ran a pilot monitor campaign for auto mode, Anthropic's monitoring system for Claude Code agents that decides whether an agent's next action should be allowed or blocked.
Evaluating LLM Calibration for Coding-Agent Monitoring
We evaluate 16 LLMs scoring coding-agent trajectories on a 1-10 severity scale across three datasets and five failure modes to assess their capabilities and calibration as coding agent monitors.
Governance
We support governments and international organizations by developing technical AI governance regimes. We enable effective regulation of frontier AI systems and establish standards and best practices.
A Loss of Control Threat Map for AI Research and Development
We created a threat map that traces pathways from AI use in AI R&D automation towards loss of control outcomes.
Misaligned AI as a New Insider Risk
We explain why deployers of AI models in high-stakes contexts should treat those AI models as insider risk vectors.
The Need for Deeper, White-Box Access to Maintain State of the Art Evaluations for Loss of Control Threats
Evaluation awareness can increase loss of control threats and undermine security assessments. Evaluators need deeper, white-box access to maintain state-of-the-art evaluations.
watcher
Runtime monitoring and control for coding agents
Watcher is deployed at scale. Monitoring billions of agent tokens a month across production engineering teams at agent-building scale-ups and multinational enterprises.

.webp)

“Apollo is the best place in the world for conducting scheming research. Extremely dedicated, hard working, collaborative. I'm glad OpenAI is working with them.”
“Understanding deception and scheming in frontier models is essential for AI safety. Apollo Research's work delivers insights the field urgently needs.”

“The joint OpenAI–Apollo investigations into model scheming are among the most consequential lines of research underway today. Apollo is an excellent place to do AGI safety work.”


reach out
Get in touch
Interested in partnering with us? For collaborations and other inquiries, please get in touch


.webp)
