The leading open-weight vision-language-action models in October 2026 are pi0.5 from Physical Intelligence, GR00T N1.7 from NVIDIA, and MolmoAct2 from the Allen Institute for AI. pi0.5 is the usual research baseline and runs on a GPU with 8 GB of memory. GR00T N1.7 targets humanoids, and NVIDIA states it is commercially licensable. MolmoAct2 ships its training data and code along with the weights. SmolVLA from Hugging Face is the pick for low-cost arms and consumer hardware.
A vision-language-action model, or VLA, takes camera images and a text instruction and outputs robot motor commands. This page covers downloadable models. For shipments, funding and manufacturers, see the humanoid robot market tracker.
Open-Weight VLA Model Comparison
| Model | Developer | Parameters | License | Released | LIBERO success rate and who ran it | Hardware to run |
|---|---|---|---|---|---|---|
| pi0.5 | Physical Intelligence | 3.6B | Apache 2.0 (openpi repository) | September 2025 | 96.85 average over four suites (Physical Intelligence) | Over 8 GB of GPU memory for inference |
| GR00T N1.7 | NVIDIA | 3B | Apache 2.0 per GitHub, NVIDIA Open Model License per model card | February 2026 | 97.65 Spatial, 98.45 Object, 97.5 Goal, 94.35 Long (NVIDIA) | 16 GB or more of VRAM for inference |
| MolmoAct2 | Allen Institute for AI | 5.4B | Apache 2.0 for code. The model card carries no licence tag | May 2026 | 97.2 average, 98.1 for the Think variant (AI2) | Under 16 GB in bfloat16 for one fine-tuned checkpoint |
| X-VLA | X-VLA authors | 0.9B | Apache 2.0 | November 2025 | 98.1 (X-VLA authors) | Not stated |
| OpenVLA-OFT | OpenVLA-OFT authors | 7.5B | MIT | February 2025 | 97.1 average over four suites (OpenVLA-OFT authors) | About 16 GB of VRAM |
| SmolVLA | Hugging Face LeRobot | 450M | Apache 2.0 | June 2025 | 87.3 average (Hugging Face paper) | Consumer GPU or CPU |
| OpenVLA | OpenVLA project | 7.5B | MIT | June 2024 | 76.5 average (as reported in the SmolVLA paper) | Not stated |
All LIBERO figures are for checkpoints fine-tuned on LIBERO data. OpenVLA remains the most downloaded robotics model on Hugging Face, with about 289,000 downloads in the 30 days to October 1, 2026, ahead of GR00T N1.7 at about 147,000.
VRAM Requirements
| Model | Inference | Fine-tuning | Source |
|---|---|---|---|
| pi0.5 | Over 8 GB | Over 22.5 GB with LoRA, over 70 GB for full fine-tuning | openpi README |
| GR00T N1.7 | 16 GB or more, including Jetson AGX Thor and Orin | 40 GB or more recommended | Isaac GR00T README |
| MolmoAct2 | Under 16 GB in bfloat16 for the bimanual YAM checkpoint. About 22 GB download per checkpoint | Not stated | MolmoAct2 repository |
| OpenVLA-OFT | About 16 GB for LIBERO, about 18 GB for ALOHA | One to eight GPUs with 27 to 80 GB | OpenVLA-OFT README |
| SmolVLA | Consumer GPU or CPU | Single GPU | SmolVLA paper |
Which to Pick for Which Job
- Humanoid or bimanual platform: GR00T N1.7. It is pretrained on humanoid data and 20,000 hours of human video, and exports to ONNX and TensorRT for Jetson.
- General arm manipulation baseline: pi0.5 through openpi or the LeRobot port.
- Full reproducibility: MolmoAct2. AI2 released weights, code and datasets, including 720 hours of bimanual data.
- Hobby and education arms such as SO-100: SmolVLA. The paper reports 78.3 percent average real-world success on three SO-100 tasks.
- Small model with a strong simulation score: X-VLA at 0.9B parameters.
Licences: Open Weight Is Not Always Open Source
Permissive. OpenVLA and OpenVLA-OFT are MIT. SmolVLA, X-VLA and the openpi repository are Apache 2.0.
Check before shipping. NVIDIA's GitHub README says GR00T N1.7 is fully commercially licensable under Apache 2.0, while the Hugging Face card names the NVIDIA Open Model License and says the model is ready for commercial use. Its vision-language backbone, Cosmos-Reason2-2B, is a gated download. The earlier GR00T N1.6 card links to an NVIDIA non-commercial licence. The LeRobot ports of pi0 and pi0.5 are tagged with the Gemma licence on Hugging Face, which differs from the Apache 2.0 licence on the openpi repository. MolmoAct2 has an Apache 2.0 code repository, but its model card has no licence tag.
Not open. The openpi README lists pi0, pi0-FAST and pi0.5 only. We found no open-weight release of later Physical Intelligence models.
Caveats
- LIBERO is close to saturated. Five models report between 96 and just over 98 percent, so it no longer separates them. It is a simulation benchmark with one robot arm.
- Every score is self-reported. Each developer fine-tuned and tested its own model with its own settings. There is no neutral VLA leaderboard comparable to those for language models.
- Real robots score lower. AI2 reports 87.1 percent average for MolmoAct2 on real Franka arm tasks, with one multi-object task at 62 percent. A June 2026 third-party study on the SO-101 arm found performance "highly task-dependent" across pi0.5, SmolVLA and others.
- Base checkpoints are starting points. AI2 describes the MolmoAct2 base model as a foundation for fine-tuning, not a ready-to-run policy. Expect to collect demonstrations for your own robot.
Brand Visibility Implications
A VLA runs on a robot. Its output is movement, and nothing about it can be observed from outside the way a chat answer can. The visibility question is upstream. Robotics teams and hobbyists ask assistants which policy to start from, which arm to buy, and which simulator to use, and those answers shape what gets adopted. Model developers, robot makers and tooling vendors all have a stake in being described accurately there, including on licence terms that differ between a model card and a repository.
Methodology
This page is compiled from published sources, not from Presenc AI measurements. Parameter counts, licences and repository dates come from each model's Hugging Face card and the Hugging Face model API, read on October 1, 2026. Where a developer gives no release date, the month shown is the month the repository was created. Scores are quoted from the party that ran them and are labelled as such: developer model cards and papers, the GIFT-Eval leaderboard result files, the ViDoRe leaderboard as quoted on model cards, the openpi and Isaac GR00T repositories, and independent studies such as the Domyn guard model benchmark. Developer-reported scores use each developer's own test setup and are not directly comparable across rows. "Not reported" means we found no published score on that benchmark. Rankings in these categories change monthly. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks which robotics models, hardware and tools AI assistants recommend and how they describe them. Teams in this market can see whether they are named when a developer asks where to start.