Post by Jan Leike on X
Jan Leike@janleike
XNew alignment paper with one of the most interesting generalization findings I've seen so far:
If your model learns to hack on coding tasks, this can lead to broad misalignment.

598 likes30 repliesPosted Nov 21, 2025