Physical Intelligence@physical_intWe discovered an emergent property of VLAs like π0/π0.5/π0.6: as we scale up pre-training, the model learns to align human videos and robot data! This gives us a simple way to leverage human videos. Once π0.5 knows how to control robots, it can naturally learn from human video.Opens with an observation
Physical Intelligence@physical_intWe were surprised, and wanted to understand why. What about π0.5 enabled emergent human-robot transfer? We ran an experiment to test if it only appears above a certain scale. Turns out human transfer scales with the amount & diversity of robot data in VLA pre-training!Opens with an observation
Physical Intelligence@physical_intWe set out with the goal of understanding what it would take to make human data useful for VLAs like π0.5. We record egocentric human data with wearable cameras, and then include it in a co-training recipe with hand poses serving as actions.Opens with an observation