OpenVLA (7B) and OpenVLA-OFT
The most-downloaded open VLA checkpoint on HF (500k+ downloads at check time); widely used as a research baseline. Superseded in capability by pi0-family and newer open VLAs but still a standard reference. Release day: arXiv 2406.09246 v1 dated 2024-06-13 (HF repo created 2024-06-10).
- Input
- image, text
- Output
- action
- License
- mit
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | openvla/openvla-7b | huggingface.co/openvla/openvla-7b | — |
| Hugging Face (OFT fine-tunes) | moojink/openvla-7b-oft-finetuned-libero-spatial | huggingface.co/moojink/openvla-7b-oft-finetuned-libero-spatial | — |
| GitHub | — | github.com/openvla/openvla | docs |
Notable capabilities (2)
- Open 7B generalist VLA beating a 55B closed model: Llama 2 7B backbone with fused DINOv2 + SigLIP vision, trained on ~970k Open X-Embodiment episodes; outperformed RT-2-X (55B) by 16.5% absolute success over 29 tasks with 7x fewer parameters, and fine-tunes with LoRA on consumer GPUs. source
- OFT fine-tuning recipe (Feb 2025) (found after launch): OpenVLA-OFT (parallel decoding, action chunking, continuous actions, L1 loss) raised LIBERO average success from 76.5% to 97.1% and action throughput 26x; on bimanual ALOHA it beat pi0 and RDT-1B by up to 15% absolute. source