Yo Shavit@yonashav
The critical question here is: were the agents trained to maximize each others’ reward, or did cross-agent cooperation arise emergently from single-agent episodic RL?
This is vital info for the wider AI+alignment community to have any way to replicate and investigate solutions.
Opens with a question