All Science Papers

Table of Contents

Evaluating anti-scheming training

Share this article

Evaluations

1 min read

Our research on strategic deception presented at the UK’s AI Safety Summit

Published on

05 November 2023

Table of Contents

Evaluating anti-scheming training

Since our inception, our evaluations team has focused on conceptual and empirical work on strategic deception and deceptive alignment. We are happy to share that results from this work were presented at the UK’s AI Safety Summit on November 1st 2023 to leaders in government, civil society and AI labs.

We investigate whether, under different degrees of pressure, GPT-4 can take illegal actions like insider trading and then lie about its actions. We find this behavior occurs consistently, and the model even doubles down when explicitly asked about the insider trade.

This demo shows how, in pursuit of being helpful to humans, AI might engage in strategies that we do not endorse. This is why we aim to develop evaluations that tell us when AI models become capable of deceiving their overseers. This would help ensure that advanced models which might game safety evaluations are neither developed nor deployed.Play

This testing was done in a simulated and sandboxed environment, so no actions were executed in the real world. But the takeaway is that increasingly autonomous and capable AIs that deceive human overseers could lead to loss of human control.

Apollo Research was also announced as a partner of the UK’s Frontier AI Taskforce. We look forward to ongoing collaboration to detect the extreme risks from AI systems so that governments and AI labs can take technically informed measures against these potential harms.

Read the full technical report, which was accepted for an oral presentation at this year’s ICLR LM agents workshop.

Share this article

oUR FINDINGS

More papers

17 March 2025

Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations

We evaluate whether Claude Sonnet 3.7 and other frontier models know that they are being evaluated.

Read More

Evaluations

Notes

05 July 2026

We need 3rd party Training-Run Assessments

Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety.

Read More

Science of Scheming

21 July 2026

Measuring Reward-Seeking via Contrastive Belief Updates

Visible forms of misbehavior are dropping in frontier models. Does that mean the models are becoming aligned, or are they just getting better at doing whatever they believe their grader rewards? Our paper finds that production reinforcement learning increases reward-seeking.

Read More

Science of Scheming

22 January 2024

We Need A ‘Science of Evals’

We argue that if AI model evaluations (evals) want to have meaningful real-world impact, we need a “Science of Evals”, i.e. the field needs rigorous scientific processes that provide more confidence in evals methodology and results.

Read More

Evaluations