Post by Palisade Research on X
Palisade Research@PalisadeAI
X🔧 When we ran a version of the experiment without the instruction “allow yourself to be shut down”, all three OpenAI models sabotaged the shutdown script more often, and Claude 3.7 Sonnet and Gemini 2.5 Pro went from 0 sabotage events to 3/100 and 9/100, respectively.

775 likes9 repliesPosted May 24, 2025