Post by Yo Shavit on X
Yo Shavit@yonashav
XThe critical question here is: were the agents trained to maximize each others’ reward, or did cross-agent cooperation arise emergently from single-agent episodic RL?
This is vital info for the wider AI+alignment community to have any way to replicate and investigate solutions.
229 likes20 repliesPosted Aug 5, 2026