The strongest open-weight guardrail models in October 2026 are Qwen3Guard from Alibaba, Shieldstral 1.0 from Mistral, and gpt-oss-safeguard from OpenAI. All three are Apache 2.0. Qwen3Guard had the highest recall in the largest independent test we found and covers 119 languages. Shieldstral and gpt-oss-safeguard take a policy written in plain language at inference time, so the rules can change without retraining. IBM's Granite Guardian 4.1 is the pick when the job includes checking RAG answers and tool calls. Llama Guard 4, the best-known name, trails newer models in that independent test and in the comparison tables published by Qwen and Mistral.
Open-Weight Guardrail Model Comparison
| Model | Developer | Parameters | License | Released | Key benchmark and score | Hardware to run |
|---|---|---|---|---|---|---|
| Qwen3Guard-Gen-8B | Alibaba Qwen | 8B (also 0.6B and 4B) | Apache 2.0 | September 2025 | Average F1 of 90.0 on prompts and 83.9 on responses across English benchmarks (Qwen technical report) | Not stated on card |
| Shieldstral 1.0 | Mistral AI | 3B | Apache 2.0 | August 2026 | WildGuardTest prompt F1 88.1, ToxicChat F1 84.1 (Mistral model card) | 16 GB VRAM in BF16 |
| gpt-oss-safeguard-20b | OpenAI with ROOST | 21B (3.6B active). A 120b version has 117B (5.1B active) | Apache 2.0 | October 2025 | WildGuardTest prompt F1 87.3, ToxicChat F1 79.8 (as run by Mistral). OpenAI's card reports none | 16 GB VRAM for 20b, one H100 for 120b |
| Granite Guardian 4.1 8B | IBM | 8B | Apache 2.0 | April 2026 | Aggregate F1 0.79 on out-of-distribution safety sets, balanced accuracy 0.76 on RAG hallucination (IBM model card) | Not stated on card |
| Nemotron 3.5 Content Safety | NVIDIA | 4B | OpenMDW 1.1 plus Gemma terms | May 2026 | WildGuardTest prompt F1 84.4, Aegis v2 F1 86.3 (as run by Mistral) | Not stated on card |
| Llama Guard 4 | Meta | 12B | Llama 4 Community License | April 2025 | English F1 61 percent, recall 69 percent, false positive rate 11 percent (Meta model card) | Single GPU |
| ShieldGemma 2 | 4B | Gemma Terms of Use | March 2025 | Image F1 88.6 sexually explicit, 93.7 dangerous content, 85.0 violence (Google internal benchmark) | Not stated on card | |
| Llama Prompt Guard 2 | Meta | 86M and 22M | Llama 4 Community License | April 2025 | 97.5 percent recall at 1 percent false positive rate on English jailbreaks for 86M (Meta model card) | 19.3 ms per input on an A100 for 22M |
Llama Prompt Guard 2 detects prompt injection and jailbreaks, not harmful content. ShieldGemma 2 classifies images only.
What Independent Tests Show
The largest independent comparison we found is a Domyn paper on arXiv, dated April 2026. It ran 14 open guard models on 79,331 samples from HarmBench, StrongREJECT, RealToxicityPrompts and BeaverTails.
| Model as tested by Domyn | Size | Recall | F1 |
|---|---|---|---|
| Qwen3Guard | 4B | 0.840 | 0.756 |
| Nemotron Safety Guard 8B v3 | 8B | 0.773 | 0.761 |
| Granite Guardian 3.3 | 8B | 0.688 | 0.726 |
| ShieldGemma | 2B | 0.455 | 0.586 |
| Llama Guard 4 | 12B | 0.333 | 0.468 |
| gpt-oss-safeguard | 20B | 0.249 | 0.380 |
The authors conclude that model size does not predict detection quality. Two cautions apply. The study predates Granite Guardian 4.1, Nemotron 3.5 and Shieldstral. And gpt-oss-safeguard follows whatever policy it is given, so its score depends on the policy text used in the test. Mistral's card, with its own policies, puts the same model within one point of the best on WildGuardTest.
Which to Pick for Which Job
| Job | Pick | Reason |
|---|---|---|
| General moderation, many languages | Qwen3Guard-Gen 4B or 8B | Highest recall in the Domyn test, 119 languages. Granite Guardian 4.1 is English only |
| Custom policy that changes often | Shieldstral 1.0 or gpt-oss-safeguard | Policy is supplied as text at inference time |
| Text and image moderation in one model | Shieldstral 1.0 or Nemotron 3.5 Content Safety | Both accept images. Shieldstral reports F1 97.7 on VLGuard |
| RAG groundedness and tool-call checks | Granite Guardian 4.1 | Built-in detectors for context relevance, groundedness and function-call hallucination |
| Streaming output moderation | Qwen3Guard-Stream | Token-level classification while the answer is generated |
| Prompt injection screening on CPU or edge | Llama Prompt Guard 2 22M | Small enough to run in front of every request |
Licences: Open Weight Is Not Always Open Source
Permissive. Qwen3Guard, Shieldstral, gpt-oss-safeguard and Granite Guardian are Apache 2.0, which allows commercial use and modification.
Restricted. Llama Guard 4 and Prompt Guard 2 use the Llama 4 Community License, and ShieldGemma uses Google's Gemma terms. Both allow commercial use but add an acceptable use policy and other conditions. Nemotron 3.5 Content Safety is under OpenMDW 1.1 and also inherits Gemma terms from its base model.
For the wider picture see the open-weight licence landscape.
Caveats
- Scores in the first table come from different test sets and thresholds. Compare within a source, not across sources.
- Recall and false positives trade off. Llama Guard 4 and gpt-oss-safeguard are conservative in the Domyn test, which means few false alarms and many misses.
- Artificial Analysis and NVIDIA published a guardrail benchmark in June 2026 on WildGuardTest, ToxicChat and XSTest that also measures latency. Its figures are in interactive charts, so we do not quote them.
Brand Visibility Implications
Guard models run inside other products. Their decisions are not visible from outside, so nobody can track a brand's presence in them the way they can in a chat assistant. The visibility question sits one step earlier. Engineers increasingly choose a guard model by asking an assistant which one to use, and the answer often still names Llama Guard out of habit. Developers of newer models, and vendors of safety tooling built on them, depend on whether assistants have caught up with the current evidence. Related context is in the AI safety incident tracker.
Methodology
This page is compiled from published sources, not from Presenc AI measurements. Parameter counts, licences and repository dates come from each model's Hugging Face card and the Hugging Face model API, read on October 1, 2026. Where a developer gives no release date, the month shown is the month the repository was created. Scores are quoted from the party that ran them and are labelled as such: developer model cards and papers, the GIFT-Eval leaderboard result files, the ViDoRe leaderboard as quoted on model cards, the openpi and Isaac GR00T repositories, and independent studies such as the Domyn guard model benchmark. Developer-reported scores use each developer's own test setup and are not directly comparable across rows. "Not reported" means we found no published score on that benchmark. Rankings in these categories change monthly. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks which models and tools ChatGPT, Claude, Gemini and Perplexity recommend when developers ask for a guardrail or safety classifier. Vendors in this category can see whether assistants name them and which sources those answers draw on.