Research

Best Open-Weight Robotics VLA Models 2026

Open-weight vision-language-action model comparison for October 2026: pi0.5, GR00T N1.7, MolmoAct2, X-VLA, OpenVLA-OFT, SmolVLA. LIBERO scores, licences, and the VRAM each one needs.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: October 2026

The leading open-weight vision-language-action models in October 2026 are pi0.5 from Physical Intelligence, GR00T N1.7 from NVIDIA, and MolmoAct2 from the Allen Institute for AI. pi0.5 is the usual research baseline and runs on a GPU with 8 GB of memory. GR00T N1.7 targets humanoids, and NVIDIA states it is commercially licensable. MolmoAct2 ships its training data and code along with the weights. SmolVLA from Hugging Face is the pick for low-cost arms and consumer hardware.

A vision-language-action model, or VLA, takes camera images and a text instruction and outputs robot motor commands. This page covers downloadable models. For shipments, funding and manufacturers, see the humanoid robot market tracker.

Open-Weight VLA Model Comparison

ModelDeveloperParametersLicenseReleasedLIBERO success rate and who ran itHardware to run
pi0.5Physical Intelligence3.6BApache 2.0 (openpi repository)September 202596.85 average over four suites (Physical Intelligence)Over 8 GB of GPU memory for inference
GR00T N1.7NVIDIA3BApache 2.0 per GitHub, NVIDIA Open Model License per model cardFebruary 202697.65 Spatial, 98.45 Object, 97.5 Goal, 94.35 Long (NVIDIA)16 GB or more of VRAM for inference
MolmoAct2Allen Institute for AI5.4BApache 2.0 for code. The model card carries no licence tagMay 202697.2 average, 98.1 for the Think variant (AI2)Under 16 GB in bfloat16 for one fine-tuned checkpoint
X-VLAX-VLA authors0.9BApache 2.0November 202598.1 (X-VLA authors)Not stated
OpenVLA-OFTOpenVLA-OFT authors7.5BMITFebruary 202597.1 average over four suites (OpenVLA-OFT authors)About 16 GB of VRAM
SmolVLAHugging Face LeRobot450MApache 2.0June 202587.3 average (Hugging Face paper)Consumer GPU or CPU
OpenVLAOpenVLA project7.5BMITJune 202476.5 average (as reported in the SmolVLA paper)Not stated

All LIBERO figures are for checkpoints fine-tuned on LIBERO data. OpenVLA remains the most downloaded robotics model on Hugging Face, with about 289,000 downloads in the 30 days to October 1, 2026, ahead of GR00T N1.7 at about 147,000.

VRAM Requirements

ModelInferenceFine-tuningSource
pi0.5Over 8 GBOver 22.5 GB with LoRA, over 70 GB for full fine-tuningopenpi README
GR00T N1.716 GB or more, including Jetson AGX Thor and Orin40 GB or more recommendedIsaac GR00T README
MolmoAct2Under 16 GB in bfloat16 for the bimanual YAM checkpoint. About 22 GB download per checkpointNot statedMolmoAct2 repository
OpenVLA-OFTAbout 16 GB for LIBERO, about 18 GB for ALOHAOne to eight GPUs with 27 to 80 GBOpenVLA-OFT README
SmolVLAConsumer GPU or CPUSingle GPUSmolVLA paper

Which to Pick for Which Job

  • Humanoid or bimanual platform: GR00T N1.7. It is pretrained on humanoid data and 20,000 hours of human video, and exports to ONNX and TensorRT for Jetson.
  • General arm manipulation baseline: pi0.5 through openpi or the LeRobot port.
  • Full reproducibility: MolmoAct2. AI2 released weights, code and datasets, including 720 hours of bimanual data.
  • Hobby and education arms such as SO-100: SmolVLA. The paper reports 78.3 percent average real-world success on three SO-100 tasks.
  • Small model with a strong simulation score: X-VLA at 0.9B parameters.

Licences: Open Weight Is Not Always Open Source

Permissive. OpenVLA and OpenVLA-OFT are MIT. SmolVLA, X-VLA and the openpi repository are Apache 2.0.

Check before shipping. NVIDIA's GitHub README says GR00T N1.7 is fully commercially licensable under Apache 2.0, while the Hugging Face card names the NVIDIA Open Model License and says the model is ready for commercial use. Its vision-language backbone, Cosmos-Reason2-2B, is a gated download. The earlier GR00T N1.6 card links to an NVIDIA non-commercial licence. The LeRobot ports of pi0 and pi0.5 are tagged with the Gemma licence on Hugging Face, which differs from the Apache 2.0 licence on the openpi repository. MolmoAct2 has an Apache 2.0 code repository, but its model card has no licence tag.

Not open. The openpi README lists pi0, pi0-FAST and pi0.5 only. We found no open-weight release of later Physical Intelligence models.

Caveats

  • LIBERO is close to saturated. Five models report between 96 and just over 98 percent, so it no longer separates them. It is a simulation benchmark with one robot arm.
  • Every score is self-reported. Each developer fine-tuned and tested its own model with its own settings. There is no neutral VLA leaderboard comparable to those for language models.
  • Real robots score lower. AI2 reports 87.1 percent average for MolmoAct2 on real Franka arm tasks, with one multi-object task at 62 percent. A June 2026 third-party study on the SO-101 arm found performance "highly task-dependent" across pi0.5, SmolVLA and others.
  • Base checkpoints are starting points. AI2 describes the MolmoAct2 base model as a foundation for fine-tuning, not a ready-to-run policy. Expect to collect demonstrations for your own robot.

Brand Visibility Implications

A VLA runs on a robot. Its output is movement, and nothing about it can be observed from outside the way a chat answer can. The visibility question is upstream. Robotics teams and hobbyists ask assistants which policy to start from, which arm to buy, and which simulator to use, and those answers shape what gets adopted. Model developers, robot makers and tooling vendors all have a stake in being described accurately there, including on licence terms that differ between a model card and a repository.

Methodology

This page is compiled from published sources, not from Presenc AI measurements. Parameter counts, licences and repository dates come from each model's Hugging Face card and the Hugging Face model API, read on October 1, 2026. Where a developer gives no release date, the month shown is the month the repository was created. Scores are quoted from the party that ran them and are labelled as such: developer model cards and papers, the GIFT-Eval leaderboard result files, the ViDoRe leaderboard as quoted on model cards, the openpi and Isaac GR00T repositories, and independent studies such as the Domyn guard model benchmark. Developer-reported scores use each developer's own test setup and are not directly comparable across rows. "Not reported" means we found no published score on that benchmark. Rankings in these categories change monthly. Status as of October 1, 2026.

How Presenc AI Helps

Presenc AI tracks which robotics models, hardware and tools AI assistants recommend and how they describe them. Teams in this market can see whether they are named when a developer asks where to start.

Frequently Asked Questions

There is no single winner. pi0.5 from Physical Intelligence is the common research baseline and reports 96.85 percent on LIBERO. GR00T N1.7 from NVIDIA is the choice for humanoids. MolmoAct2 from the Allen Institute for AI reports 97.2 percent and releases its training data. SmolVLA is the best fit for low-cost arms.
Not a neutral one. LIBERO is the most quoted benchmark, but each developer runs it on its own fine-tuned checkpoint and the top models all score about 96 to 98 percent. Treat published LIBERO numbers as a sanity check, not a ranking, and test on your own robot.
OpenVLA and OpenVLA-OFT are MIT. SmolVLA, X-VLA and the openpi repository that hosts pi0.5 are Apache 2.0. NVIDIA states that GR00T N1.7 is commercially licensable, though its GitHub page and model card name different licences. Read the licence file in the repository you download from.
For inference, pi0.5 needs over 8 GB, GR00T N1.7 needs 16 GB or more, and OpenVLA-OFT needs about 16 GB. SmolVLA runs on a consumer GPU or a CPU. Fine-tuning needs more: over 22.5 GB for pi0.5 with LoRA and 40 GB or more recommended for GR00T N1.7.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Join the waitlist for early access.