Research

Best Open-Weight Guardrail and Safety Models 2026

Open-weight guardrail model comparison for October 2026: Qwen3Guard, Shieldstral, gpt-oss-safeguard, Granite Guardian 4.1, Llama Guard 4, ShieldGemma 2. Licences, benchmark scores, and hardware.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: October 2026

The strongest open-weight guardrail models in October 2026 are Qwen3Guard from Alibaba, Shieldstral 1.0 from Mistral, and gpt-oss-safeguard from OpenAI. All three are Apache 2.0. Qwen3Guard had the highest recall in the largest independent test we found and covers 119 languages. Shieldstral and gpt-oss-safeguard take a policy written in plain language at inference time, so the rules can change without retraining. IBM's Granite Guardian 4.1 is the pick when the job includes checking RAG answers and tool calls. Llama Guard 4, the best-known name, trails newer models in that independent test and in the comparison tables published by Qwen and Mistral.

Open-Weight Guardrail Model Comparison

ModelDeveloperParametersLicenseReleasedKey benchmark and scoreHardware to run
Qwen3Guard-Gen-8BAlibaba Qwen8B (also 0.6B and 4B)Apache 2.0September 2025Average F1 of 90.0 on prompts and 83.9 on responses across English benchmarks (Qwen technical report)Not stated on card
Shieldstral 1.0Mistral AI3BApache 2.0August 2026WildGuardTest prompt F1 88.1, ToxicChat F1 84.1 (Mistral model card)16 GB VRAM in BF16
gpt-oss-safeguard-20bOpenAI with ROOST21B (3.6B active). A 120b version has 117B (5.1B active)Apache 2.0October 2025WildGuardTest prompt F1 87.3, ToxicChat F1 79.8 (as run by Mistral). OpenAI's card reports none16 GB VRAM for 20b, one H100 for 120b
Granite Guardian 4.1 8BIBM8BApache 2.0April 2026Aggregate F1 0.79 on out-of-distribution safety sets, balanced accuracy 0.76 on RAG hallucination (IBM model card)Not stated on card
Nemotron 3.5 Content SafetyNVIDIA4BOpenMDW 1.1 plus Gemma termsMay 2026WildGuardTest prompt F1 84.4, Aegis v2 F1 86.3 (as run by Mistral)Not stated on card
Llama Guard 4Meta12BLlama 4 Community LicenseApril 2025English F1 61 percent, recall 69 percent, false positive rate 11 percent (Meta model card)Single GPU
ShieldGemma 2Google4BGemma Terms of UseMarch 2025Image F1 88.6 sexually explicit, 93.7 dangerous content, 85.0 violence (Google internal benchmark)Not stated on card
Llama Prompt Guard 2Meta86M and 22MLlama 4 Community LicenseApril 202597.5 percent recall at 1 percent false positive rate on English jailbreaks for 86M (Meta model card)19.3 ms per input on an A100 for 22M

Llama Prompt Guard 2 detects prompt injection and jailbreaks, not harmful content. ShieldGemma 2 classifies images only.

What Independent Tests Show

The largest independent comparison we found is a Domyn paper on arXiv, dated April 2026. It ran 14 open guard models on 79,331 samples from HarmBench, StrongREJECT, RealToxicityPrompts and BeaverTails.

Model as tested by DomynSizeRecallF1
Qwen3Guard4B0.8400.756
Nemotron Safety Guard 8B v38B0.7730.761
Granite Guardian 3.38B0.6880.726
ShieldGemma2B0.4550.586
Llama Guard 412B0.3330.468
gpt-oss-safeguard20B0.2490.380

The authors conclude that model size does not predict detection quality. Two cautions apply. The study predates Granite Guardian 4.1, Nemotron 3.5 and Shieldstral. And gpt-oss-safeguard follows whatever policy it is given, so its score depends on the policy text used in the test. Mistral's card, with its own policies, puts the same model within one point of the best on WildGuardTest.

Which to Pick for Which Job

JobPickReason
General moderation, many languagesQwen3Guard-Gen 4B or 8BHighest recall in the Domyn test, 119 languages. Granite Guardian 4.1 is English only
Custom policy that changes oftenShieldstral 1.0 or gpt-oss-safeguardPolicy is supplied as text at inference time
Text and image moderation in one modelShieldstral 1.0 or Nemotron 3.5 Content SafetyBoth accept images. Shieldstral reports F1 97.7 on VLGuard
RAG groundedness and tool-call checksGranite Guardian 4.1Built-in detectors for context relevance, groundedness and function-call hallucination
Streaming output moderationQwen3Guard-StreamToken-level classification while the answer is generated
Prompt injection screening on CPU or edgeLlama Prompt Guard 2 22MSmall enough to run in front of every request

Licences: Open Weight Is Not Always Open Source

Permissive. Qwen3Guard, Shieldstral, gpt-oss-safeguard and Granite Guardian are Apache 2.0, which allows commercial use and modification.

Restricted. Llama Guard 4 and Prompt Guard 2 use the Llama 4 Community License, and ShieldGemma uses Google's Gemma terms. Both allow commercial use but add an acceptable use policy and other conditions. Nemotron 3.5 Content Safety is under OpenMDW 1.1 and also inherits Gemma terms from its base model.

For the wider picture see the open-weight licence landscape.

Caveats

  • Scores in the first table come from different test sets and thresholds. Compare within a source, not across sources.
  • Recall and false positives trade off. Llama Guard 4 and gpt-oss-safeguard are conservative in the Domyn test, which means few false alarms and many misses.
  • Artificial Analysis and NVIDIA published a guardrail benchmark in June 2026 on WildGuardTest, ToxicChat and XSTest that also measures latency. Its figures are in interactive charts, so we do not quote them.

Brand Visibility Implications

Guard models run inside other products. Their decisions are not visible from outside, so nobody can track a brand's presence in them the way they can in a chat assistant. The visibility question sits one step earlier. Engineers increasingly choose a guard model by asking an assistant which one to use, and the answer often still names Llama Guard out of habit. Developers of newer models, and vendors of safety tooling built on them, depend on whether assistants have caught up with the current evidence. Related context is in the AI safety incident tracker.

Methodology

This page is compiled from published sources, not from Presenc AI measurements. Parameter counts, licences and repository dates come from each model's Hugging Face card and the Hugging Face model API, read on October 1, 2026. Where a developer gives no release date, the month shown is the month the repository was created. Scores are quoted from the party that ran them and are labelled as such: developer model cards and papers, the GIFT-Eval leaderboard result files, the ViDoRe leaderboard as quoted on model cards, the openpi and Isaac GR00T repositories, and independent studies such as the Domyn guard model benchmark. Developer-reported scores use each developer's own test setup and are not directly comparable across rows. "Not reported" means we found no published score on that benchmark. Rankings in these categories change monthly. Status as of October 1, 2026.

How Presenc AI Helps

Presenc AI tracks which models and tools ChatGPT, Claude, Gemini and Perplexity recommend when developers ask for a guardrail or safety classifier. Vendors in this category can see whether assistants name them and which sources those answers draw on.

Frequently Asked Questions

On published evidence, Qwen3Guard is the strongest fixed-taxonomy option. It had the highest recall (0.840) in Domyn's independent test of 14 guard models and is Apache 2.0. Shieldstral 1.0 from Mistral and gpt-oss-safeguard from OpenAI are the leading choices when you need to supply your own policy as text. Granite Guardian 4.1 is the pick for RAG and tool-call checks.
There is no single official leaderboard. The closest things are Domyn's arXiv benchmark of 14 open guard models (April 2026), the Artificial Analysis and NVIDIA guardrail benchmark (June 2026), and the comparison tables in the Qwen3Guard report and the Shieldstral model card. They use different datasets, so rankings differ between them.
Some are. Qwen3Guard, Shieldstral, gpt-oss-safeguard and Granite Guardian use Apache 2.0. Llama Guard 4 uses the Llama 4 Community License and ShieldGemma uses Gemma terms, which permit commercial use with added conditions. Check the licence on the model card before deploying.
It is widely supported and handles text and images, but it scores below newer models in the tests we found. Domyn measured recall of 0.333 and F1 of 0.468. The Qwen3Guard report gives it an average prompt F1 of 75.9 against 90.0 for Qwen3Guard-Gen-8B. It produces few false positives, which suits some products.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.