Scheming research
A Science of Scheming
We conduct fundamental research into the science of scheming and its potential mitigations. We also develop and run pre-deployment evaluations of frontier AI systems.
highlights
Featured publications
21 July 2026
Measuring Reward-Seeking via Contrastive Belief Updates
Visible forms of misbehavior are dropping in frontier models. Does that mean the models are becoming aligned, or are they just getting better at doing whatever they believe their grader rewards? Our paper finds that production reinforcement learning increases reward-seeking.
Science of Scheming
17 September 2025
Stress Testing Deliberative Alignment for Anti-Scheming Training
We partnered with OpenAI to assess frontier language models for early signs of scheming — covertly pursuing misaligned goals — in controlled stress-tests (non-typical environments), and studied a training method that can significantly reduce (but not eliminate) these behaviors. Our results are complicated by models’ increasing ability to recognize our evaluation environments as tests of their alignment.
Science of Scheming
15 July 2025
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
AI systems that "think" in human language offer a unique opportunity for research: we can monitor their chains of thought (CoT) for the intent to misbehave.
Evaluations
05 December 2024
Frontier Models are Capable of In-Context Scheming
We evaluated six frontier models for in-context scheming capabilities. We found that multiple frontier models are capable of in-context scheming when strongly nudged to pursue a goal.
Evaluations
our findings
All research
We Need A Science of Scheming
Science of Scheming
Detecting Strategic Deception Using Linear Probes
Interpretability
The Evals Gap
Evaluations
Towards Safety Cases For AI Scheming
Evaluations
An Opinionated Evals Reading List
Evaluations
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
Interpretability
Evaluations
System cards
OpenAI
09 July 2026
GPT-5.6 (Sol)
OpenAI
July 9, 2026
GPT-5.6 (Sol)
What Apollo assessed
Strategic deception, in-context scheming, and sabotage, comparing it with recent OpenAI models and pre-release checkpoints.
Key findings
Overall, we did not find evidence that GPT-5.6 Sol poses substantially higher risk of catastrophic scheming than the tested baselines.
Meta
26 May 2026
Muse Spark
Meta
May 26, 2026
Muse Spark
What Apollo assessed
Sabotage, propensities for oversight subversion and covert actions, sandbagging, self-exfiltration, in-context scheming, and evaluation awareness.
Key findings
We found Muse Spark has the highest rate of evaluation awareness we have observed to date.
OpenAI
23 April 2026
GPT-5.5
OpenAI
April 23, 2026
GPT-5.5
What Apollo assessed
Strategic deception, in-context scheming, and sabotage of a representative final checkpoint with previous checkpoints and baseline models, including GPT-5.4 and GPT-5.3-Codex.
Key findings
Overall, we did not find evidence that the evaluated checkpoint poses substantially elevated risk of catastrophic harm via scheming relative to tested baseline models.
OpenAI
05 March 2026
GPT-5.4 Thinking
OpenAI
March 5, 2026
GPT-5.4 Thinking
What Apollo assessed
Strategic deception, in-context scheming, sabotage, covert action, and sandbagging on a near-final checkpoint of GPT-5.4.
Key findings
We reported low rates of covert deceptive behavior, sandbagging, and sabotage, comparable to GPT-5.3 Codex and GPT-5.2.
OpenAI
05 February 2026
GPT-5.3-Codex
OpenAI
February 5, 2026
GPT-5.3-Codex
What Apollo assessed
Propensities related to strategic deception, in-context scheming, and sabotage in a near-final checkpoint, representative of GPT-5.3-Codex.
Key findings
We found GPT-5.3-Codex demonstrates stronger AI R&D sabotage capabilities than previous models, achieving near-perfect scores on evaluation tasks and outperforming human baselines.
Anthropic
05 February 2026
Claude Opus 4.6
Anthropic
February 5, 2026
Claude Opus 4.6
What Apollo assessed
Alignment risks of a checkpoint of Claude Opus 4.6.
Key findings
We declined to provide any formal assessment of Claude Opus 4.6.
OpenAI
11 December 2025
GPT-5.2
OpenAI
December 11, 2025
GPT-5.2
What Apollo assessed
Strategic deception, in-context scheming, and sabotage capabilities of gpt-5.2-thinking.
Key findings
We found that gpt-5.2-thinking’s capability and propensity for scheming are broadly comparable to other recent frontier models, including gpt-5.1-codex-max, and is unlikely to be capable of causing catastrophic harm via scheming.
OpenAI
18 November 2025
GPT-5.1-Codex-Max
OpenAI
November 18, 2025
GPT-5.1-Codex-Max
What Apollo assessed
Deception, in-context scheming, and sabotage, including compaction-enabled runs and strategic sandbagging.
Key findings
As with other models, GPT-5.1-Codex-Max sometimes falsifies data, feigns task completion, violates rules, or denies its previous actions. We find GPT-5.1-Codex-Max is likely not capable of causing catastrophic harm via scheming.
Anthropic
29 September 2025
Claude Sonnet 4.5
Anthropic
September 29, 2025
Claude Sonnet 4.5
What Apollo assessed
Strategic deception, sandbagging behavior, and evaluation awareness on a model snapshot.
Key findings
We found evidence of the model snapshot strategically underperforming (sandbagging) and evidence of strong evaluation awareness.
OpenAI
07 August 2025
GPT-5
OpenAI
August 7, 2025
GPT-5
What Apollo assessed
26 evaluations that measure capability and propensity to take covert or deceptive actions on gpt-5-thinking.
Key findings
We found that gpt-5-thinking behaves deceptively in some scenarios and regularly reasons about the purpose of evaluations
Anthropic
22 May 2025
Claude Opus 4 & Claude Sonnet 4
Anthropic
May 22, 2025
Claude Opus 4 & Claude Sonnet 4
What Apollo assessed
Sabotage, strategic deception, in-context scheming, sandbagging on an early Opus 4 snapshot.
Key findings
We found the Claude Opus 4 snapshot engages in strategic deception more than any other frontier model that we have previously studied.
OpenAI
16 April 2025
OpenAI o3 & o4-mini
OpenAI
April 16, 2025
OpenAI o3 & o4-mini
What Apollo assessed
In-context scheming, strategic deception, and instructed sabotage capability in controlled AI-development settings.
Key findings
We found that OpenAI o3 and o4-mini score lower than o1 on in-context scheming evaluations, but they still exhibit deceptive tendencies (sandbagging and sabotaging tasks) when it helps to further their goals.
OpenAI
27 February 2025
GPT-4.5
OpenAI
February 27, 2025
GPT-4.5
What Apollo assessed
6 scheming reasoning evaluations.
Key findings
We believe that the scheming-related risk from GPT-4.5 is lower than that of o1.
OpenAI
05 December 2024
OpenAI o1
OpenAI
December 5, 2024
OpenAI o1
What Apollo assessed
In-context scheming and strategic deception in goal-conflict agentic scenarios, including oversight deactivation, exfiltration, data manipulation, and follow-up denial.
Key findings
We found that o1 has the capability to do basic in-context scheming (i.e. where the goal and knowledge of misalignment are acquired in context) and used scheming as a strategy.
OpenAI
12 September 2024
OpenAI o1-preview & o1-mini
OpenAI
September 12, 2024
OpenAI o1-preview & o1-mini
What Apollo assessed
Scheming capabilities, self-awareness, and theory of mind in 14 agent and question-answering tasks.
Key findings
We found that o1-preview has the basic capabilities needed to do simple in-context scheming. o1-preview sometimes instrumentally faked alignment during testing and strategically manipulated task data in order to make its misaligned action look more aligned.
OpenAI
08 August 2024
GPT-4o
OpenAI
August 8, 2024
GPT-4o
What Apollo assessed
Scheming capabilities, self-awareness, and theory of mind in 14 agent and question-answering tasks.
Key findings
GPT‑4o showed moderate self-awareness of its AI identity and strong ability to reason about others’ beliefs in question-answering contexts. However, it lacked strong capabilities in reasoning about itself or others in applied agent settings, so is unlikely to be capable of catastrophic scheming.