Xiaomi-Robotics-U0 (38B) / U0-4B
'first' flag is Xiaomi's claim (first model with high-quality multi-view scene generation across multiple robot embodiments). Paper says 38B params; the HF README table says 34B. U0 and U0-FlashAR weights 2026-07-13; U0-4B, U0-Sequence and U0-4B-Sequence weights plus FSDP training code 2026-09-08. U0-Video announced as coming soon. Authors report beating GPT-Image-2.0 in human evals of embodied scene generation/transfer and #1 on World Arena for embodied video. Not an action model: it generates observations/data, not motor commands.
- Input
- text, image
- Output
- image, video, text
- License
- apache-2.0
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | XiaomiRobotics/Xiaomi-Robotics-U0 | huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0 | — |
| Hugging Face | XiaomiRobotics/Xiaomi-Robotics-U0-4B | huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B | — |
| GitHub | — | github.com/XiaomiRobotics/Xiaomi-Robotics-U0 | — |
| ModelScope | — | modelscope.cn/collections/XiaomiRobotics/Xiaomi-Robotics-U0 | — |
Notable capabilities (3)
- FIRST Unified embodied synthesis: One autoregressive model (shared discrete visual tokenizer, next-token objective, initialized from Emu3.5) does text-to-image, image editing, multi-view robot scene generation, embodied transfer (editing scenes while keeping multi-view consistency) and embodied video rollout. source
- Data engine for VLAs: Synthetic data from U0 raised π0.5's out-of-distribution success on hard real-world manipulation tasks from 36.9% to 63.2% (authors). source
- FlashAR fast decoding: Anti-diagonal grouped visual-token decoding plus vLLM batching: 5.44 s per 1024x1024 image on one H20, 82.86x faster than eager AR. source
Xiaomi's open embodied world model / synthetic-data engine, the generation-side companion of its VLAs (xiaomi-robotics-1).
Sources: arXiv 2607.11643, HF U0-4B card, project page.
Timeline entry
- Xiaomi open-sources Xiaomi-Robotics-U0, a 38B unified world model that generates multi-view robot scenes and training data ★★
On 2026-07-13 Xiaomi released Xiaomi-Robotics-U0 (arXiv 2607.11643, Apache-2.0), a 38B autoregressive model initialized from Emu3.5 that handles text-to-image, image editing, multi-view embodied scene generation, embodied transfer and embodied video in one next-token framework; its synthetic data…
Other Xiaomi models
MiMo-V2.6-Flash · MiMo-V2.6-Pro · Xiaomi-Robotics-1 (XR-1, 5B) · Xiaomi-Robotics-0 (4.7B VLA)