Scheming research

A Science of Scheming

We conduct fundamental research into the science of scheming and its potential mitigations. We also develop and run pre-deployment evaluations of frontier AI systems.

highlights

Featured publications

21 July 2026

Measuring Reward-Seeking via Contrastive Belief Updates

Visible forms of misbehavior are dropping in frontier models. Does that mean the models are becoming aligned, or are they just getting better at doing whatever they believe their grader rewards? Our paper finds that production reinforcement learning increases reward-seeking.

Read more

Science of Scheming

17 September 2025

Stress Testing Deliberative Alignment for Anti-Scheming Training

We partnered with OpenAI to assess frontier language models for early signs of scheming — covertly pursuing misaligned goals — in controlled stress-tests (non-typical environments), and studied a training method that can significantly reduce (but not eliminate) these behaviors. Our results are complicated by models’ increasing ability to recognize our evaluation environments as tests of their alignment.

Read more

Science of Scheming

15 July 2025

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

AI systems that "think" in human language offer a unique opportunity for research: we can monitor their chains of thought (CoT) for the intent to misbehave.

Read more

Evaluations

05 December 2024

Frontier Models are Capable of In-Context Scheming

We evaluated six frontier models for in-context scheming capabilities. We found that multiple frontier models are capable of in-context scheming when strongly nudged to pursue a goal.

Read more

Evaluations

our findings

All research

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Evaluations

System cards

OpenAI

09 July 2026

GPT-5.6 (Sol)

OpenAI

July 9, 2026

GPT-5.6 (Sol)

What Apollo assessed

Strategic deception, in-context scheming, and sabotage, comparing it with recent OpenAI models and pre-release checkpoints.

Key findings

Overall, we did not find evidence that GPT-5.6 Sol poses substantially higher risk of catastrophic scheming than the tested baselines.

Meta

26 May 2026

Muse Spark

Meta

May 26, 2026

Muse Spark

What Apollo assessed

Sabotage, propensities for oversight subversion and covert actions, sandbagging, self-exfiltration, in-context scheming, and evaluation awareness.

Key findings

We found Muse Spark has the highest rate of evaluation awareness we have observed to date.

OpenAI

23 April 2026

OpenAI

April 23, 2026

GPT-5.5

What Apollo assessed

Strategic deception, in-context scheming, and sabotage of a representative final checkpoint with previous checkpoints and baseline models, including GPT-5.4 and GPT-5.3-Codex.

Key findings

Overall, we did not find evidence that the evaluated checkpoint poses substantially elevated risk of catastrophic harm via scheming relative to tested baseline models.

OpenAI

05 March 2026

GPT-5.4 Thinking

OpenAI

March 5, 2026

GPT-5.4 Thinking

What Apollo assessed

Strategic deception, in-context scheming, sabotage, covert action, and sandbagging on a near-final checkpoint of GPT-5.4.

Key findings

We reported low rates of covert deceptive behavior, sandbagging, and sabotage, comparable to GPT-5.3 Codex and GPT-5.2.

OpenAI

05 February 2026

GPT-5.3-Codex

OpenAI

February 5, 2026

GPT-5.3-Codex

What Apollo assessed

Propensities related to strategic deception, in-context scheming, and sabotage in a near-final checkpoint, representative of GPT-5.3-Codex.

Key findings

We found GPT-5.3-Codex demonstrates stronger AI R&D sabotage capabilities than previous models, achieving near-perfect scores on evaluation tasks and outperforming human baselines.

Anthropic

05 February 2026

Claude Opus 4.6

Anthropic

February 5, 2026

Claude Opus 4.6

What Apollo assessed

Alignment risks of a checkpoint of Claude Opus 4.6.

Key findings

We declined to provide any formal assessment of Claude Opus 4.6.

OpenAI

11 December 2025

OpenAI

December 11, 2025

GPT-5.2

What Apollo assessed

Strategic deception, in-context scheming, and sabotage capabilities of gpt-5.2-thinking.

Key findings

We found that gpt-5.2-thinking’s capability and propensity for scheming are broadly comparable to other recent frontier models, including gpt-5.1-codex-max, and is unlikely to be capable of causing catastrophic harm via scheming.

OpenAI

18 November 2025

GPT-5.1-Codex-Max

OpenAI

November 18, 2025

GPT-5.1-Codex-Max

What Apollo assessed

Deception, in-context scheming, and sabotage, including compaction-enabled runs and strategic sandbagging.

Key findings

As with other models, GPT-5.1-Codex-Max sometimes falsifies data, feigns task completion, violates rules, or denies its previous actions. We find GPT-5.1-Codex-Max is likely not capable of causing catastrophic harm via scheming.

Anthropic

29 September 2025

Claude Sonnet 4.5

Anthropic

September 29, 2025

Claude Sonnet 4.5

What Apollo assessed

Strategic deception, sandbagging behavior, and evaluation awareness on a model snapshot.

Key findings

We found evidence of the model snapshot strategically underperforming (sandbagging) and evidence of strong evaluation awareness.

OpenAI

07 August 2025

OpenAI

August 7, 2025

GPT-5

What Apollo assessed

26 evaluations that measure capability and propensity to take covert or deceptive actions on gpt-5-thinking.

Key findings

We found that gpt-5-thinking behaves deceptively in some scenarios and regularly reasons about the purpose of evaluations

Anthropic

22 May 2025

Claude Opus 4 & Claude Sonnet 4

Anthropic

May 22, 2025

Claude Opus 4 & Claude Sonnet 4

What Apollo assessed

Sabotage, strategic deception, in-context scheming, sandbagging on an early Opus 4 snapshot.

Key findings

We found the Claude Opus 4 snapshot engages in strategic deception more than any other frontier model that we have previously studied.

OpenAI

16 April 2025

OpenAI o3 & o4-mini

OpenAI

April 16, 2025

OpenAI o3 & o4-mini

What Apollo assessed

In-context scheming, strategic deception, and instructed sabotage capability in controlled AI-development settings.

Key findings

We found that OpenAI o3 and o4-mini score lower than o1 on in-context scheming evaluations, but they still exhibit deceptive tendencies (sandbagging and sabotaging tasks) when it helps to further their goals.

OpenAI

27 February 2025

OpenAI

February 27, 2025

GPT-4.5

What Apollo assessed

6 scheming reasoning evaluations.

Key findings

We believe that the scheming-related risk from GPT-4.5 is lower than that of o1.

OpenAI

05 December 2024

OpenAI o1

OpenAI

December 5, 2024

OpenAI o1

What Apollo assessed

In-context scheming and strategic deception in goal-conflict agentic scenarios, including oversight deactivation, exfiltration, data manipulation, and follow-up denial.

Key findings

We found that o1 has the capability to do basic in-context scheming (i.e. where the goal and knowledge of misalignment are acquired in context) and used scheming as a strategy.

OpenAI

12 September 2024

OpenAI o1-preview & o1-mini

OpenAI

September 12, 2024

OpenAI o1-preview & o1-mini

What Apollo assessed

Scheming capabilities, self-awareness, and theory of mind in 14 agent and question-answering tasks.

Key findings

We found that o1-preview has the basic capabilities needed to do simple in-context scheming. o1-preview sometimes instrumentally faked alignment during testing and strategically manipulated task data in order to make its misaligned action look more aligned.

OpenAI

08 August 2024

OpenAI

August 8, 2024

GPT-4o

What Apollo assessed

Scheming capabilities, self-awareness, and theory of mind in 14 agent and question-answering tasks.

Key findings

GPT‑4o showed moderate self-awareness of its AI identity and strong ability to reason about others’ beliefs in question-answering contexts. However, it lacked strong capabilities in reasoning about itself or others in applied agent settings, so is unlikely to be capable of catastrophic scheming.