Physical Intelligence's π*0.6 learns from real-world experience with RL (Recap), running tasks for hours
On 2025-11-17 Physical Intelligence released π*0.6, a version of its π0.6 VLA improved with Recap (RL with Experience & Corrections via Advantage-conditioned Policies): demonstrations, then human corrections, then RL on the robot's own autonomous trials. Recap more than doubled throughput and roughly halved failure rates on the hardest tasks; robots made espresso for 13 hours, folded laundry for 3 hours and assembled boxes in a real factory.
Key facts
- Paper: 'π*0.6: a VLA That Learns From Experience' (arXiv 2511.14759)
- Recap = RL with Experience & Corrections via Advantage-conditioned Policies; a value function scores actions and the policy is conditioned on advantage
- >2x throughput and ~2x lower failure rates on some of the hardest tasks (PI)
- Demos: espresso drinks from 5:30am to 11:30pm (~13 h), 50 novel laundry items in a new home (~3 h), 59 chocolate-packaging boxes assembled and labeled in a real factory
- No weights or API released
What happened
Most VLAs are trained only on imitation from teleoperated demonstrations. π*0.6 adds a reinforcement-learning stage that runs on real robots. A learned value function judges which of the robot's own attempts, and which human interventions, were better than average, and the policy is trained to produce those "high-advantage" actions. PI demonstrated long unattended runs in an office, a home and a factory.
Why it matters
It is one of the first convincing demonstrations that VLAs can keep improving from deployment experience rather than only from more demonstrations. That makes "robots that get better on the job" a practical path, and PI followed it with π0.7 in April 2026.
Changelog
- 2026-09-29: created (pi.website blocked automated fetch; numbers from PI blog search snippets, arXiv listing and press)
Models
- π0.6 / π*0.6 Physical Intelligence · legacy
Videos (1)
π*0.6: four hours of robotic box assembling
Physical Intelligence · 2025-11-17 · demoDescription by Gemini, which watched the video:
Summary
This video is an unedited, extended autonomous demonstration presented by Physical Intelligence (π), showcasing their robotic manipulation policy (identified in the title as π*0.6). Over an unbroken span of nearly four hours, a bimanual robotic arm system continuously and autonomously picks up flat cardboard sheets, folds and forms them into assembled boxes, and places them into storage bins.
What is shown
- Autonomous Bimanual Box Assembly: Two robotic arms mounted on a workshop table manipulate flat cardboard cutouts, coordinating both end-effectors to fold flaps, crease edges, and square the boxes into finished form [00:30–02:30].
- Continuous Multi-Hour Operation: The robotic system repeats the box-folding workflow continuously at 1x real-time speed across the multi-hour video without policy failure [00:00–230:10].
- Human-in-the-Loop Environment Maintenance: A human technician periodically enters the frame to remove stacks of assembled boxes from the bin and restock flattened cardboard sheets while the robot continues operating [26:15–26:50, 50:20–50:30, 77:35–77:45, 119:10–119:25, 133:35–134:10, 154:10–154:20].
Claims & numbers
- Runtime: Approximately four hours of continuous autonomous box assembling at real-time (1x) playback speed (indicated by on-screen overlay "autonomous, 1x" and the video title).
- Autonomous Execution: The folding policy operates fully autonomously without teleoperation during assembly cycles (indicated by on-screen overlay).
Notable quotes
- None (the video has no spoken dialogue, narration, or voiceover).
Assessment
This is a real, unedited long-duration endurance demo of physical AI manipulation from Physical Intelligence. The entire multi-hour run is shown in continuous real-time without cuts or speed-ups, demonstrating robust generalization and long-horizon bimanual dexterous manipulation.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.
Related events
Sources (4)
- officialPhysical Intelligence: A VLA that Learns from Experience (π*0.6)
- paperarXiv 2511.14759: π*0.6: a VLA That Learns From Experience
- pressHumanoids Daily: Physical Intelligence claims 'RL is back'
- videoYouTube: π*0.6: four hours of robotic box assembling
id: 2025-11-17-physical-intelligence-pi-star-0-6-recap · updated 2026-09-29 · open in the interactive timeline