Xiaomi-Robotics-0 (4.7B VLA)
4.7B parameters, Qwen3-VL-4B-Instruct backbone; pretrained on cross-embodiment robot trajectories plus vision-language data. Real-robot evals: Lego disassembly and towel folding (bimanual). Checkpoints: -Pretrain, -LIBERO, -Calvin-ABC_D, -Calvin-ABCD_D, -SimplerEnv-WidowX, -SimplerEnv-Google-Robot (HF, 2026-02-10). Paper arXiv 2602.12684 (2026-02-13). Superseded by xiaomi-robotics-1 (July 2026).
- Input
- image, text
- Output
- action
- License
- apache-2.0
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | XiaomiRobotics/Xiaomi-Robotics-0-Pretrain | huggingface.co/XiaomiRobotics/Xiaomi-Robotics-0-Pretrain | — |
| Project page | — | xiaomi-robotics-0.github.io | — |
Notable capabilities (2)
- Real-time asynchronous execution on a consumer GPU: Post-trained for asynchronous execution with aligned timesteps between consecutive action chunks, so rollouts stay smooth despite inference latency; runs on a consumer-grade GPU (per paper). source
- Strong open sim-benchmark results: LIBERO 98.7% avg; SimplerEnv Visual Matching 85.5%, Visual Aggregation 74.7%, WidowX 79.2%; CALVIN avg length 4.75 (ABC-D) / 4.80 (ABCD-D) (authors). source
Xiaomi's first open VLA. Use xiaomi-robotics-1 for new work.
Sources: project page, arXiv 2602.12684, HF.
Other Xiaomi models
Xiaomi-Robotics-1 (XR-1, 5B) · Xiaomi-Robotics-U0 (38B) / U0-4B