RDT2 (and RDT-1B)
'First' claim is the authors' own hedged wording. HF RDT2-VQ repo created 2025-09-22.
- Input
- image, text
- Output
- action
- License
- apache-2.0
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | robotics-diffusion-transformer/RDT2-VQ | huggingface.co/robotics-diffusion-transformer/RDT2-VQ | — |
| Hugging Face (RDT-1B, MIT) | robotics-diffusion-transformer/rdt-1b | huggingface.co/robotics-diffusion-transformer/rdt-1b | — |
| GitHub | — | github.com/thu-ml/RDT2 | docs |
Notable capabilities (2)
- FIRST Zero-shot deployment on unseen embodiments: RDT2 (8B, Qwen2.5-VL-7B based, residual-VQ action tokens; RDT2-FM flow-matching variant) trained on 10k+ h of UMI-gripper human manipulation from 100+ scenes; authors call it possibly the first foundation model to deploy zero-shot on unseen embodiments (UR5e, Franka FR3) for simple open-vocabulary tasks. source
- Large diffusion foundation model for bimanual manipulation (RDT-1B): RDT-1B (Oct 2024, 1.2B) was billed as the largest diffusion-based foundation model for bimanual manipulation, pretrained on 46 datasets (1M+ episodes) and fine-tuned on a 6K+ episode ALOHA dataset. source