As of October 1, 2026, the highest LiveCodeBench score we could trace to a primary source is 93.5 percent pass@1 for DeepSeek V4 Pro at its Max reasoning setting, reported by DeepSeek on its own model card. That number is not on the official leaderboard. The official board still runs on release v6, its data file was last changed on August 1, 2025, and its top entry is OpenAI's o4-mini at high effort with 80.2 percent on the default date window. Anyone quoting a LiveCodeBench score in 2026 is almost always quoting a vendor's own run.
Official LiveCodeBench Leaderboard
These rows are calculated from the leaderboard's public data file for the default view: 454 problems released between August 1, 2024 and May 1, 2025. The board holds 28 models and none was released in 2026.
| Rank | Model | Pass@1 | Hard problems | Reported by | Date |
|---|---|---|---|---|---|
| 1 | o4-mini (high) | 80.2% | 63.5% | Official leaderboard | Data file dated August 1, 2025 |
| 2 | o3 (high) | 75.8% | 57.1% | Official leaderboard | Same |
| 3 | o4-mini (medium) | 74.2% | 52.7% | Official leaderboard | Same |
| 4 | Gemini 2.5 Pro (06-05) | 73.6% | 50.2% | Official leaderboard | Same |
| 5 | DeepSeek R1 (0528) | 73.1% | 50.7% | Official leaderboard | Same |
| 6 | Gemini 2.5 Pro (05-06) | 71.8% | 50.2% | Official leaderboard | Same |
On the full v6 set of 1,055 problems, o4-mini (high) scores 87.3 percent. The date window changes the number by seven points for the same model, which is why the window has to be stated.
2026 Scores Reported Outside the Official Board
This table is not a ranking. The rows come from different runs, and several do not state which release or window they used.
| Model | Score | Problem set | Reported by | Date |
|---|---|---|---|---|
| DeepSeek V4 Pro (Max) | 93.5% | Not stated | DeepSeek model card | V4 previewed April 2026 |
| Sakana Fugu-Ultra | 93.2% | v6 | Provider figure, listed by BenchLM | Viewed October 1, 2026 |
| Gemini 3 Pro Preview (high) | 91.7% | Artificial Analysis run | Artificial Analysis, independent | Viewed October 1, 2026 |
| DeepSeek V4 Flash (Max) | 91.6% | Not stated | DeepSeek model card | V4 previewed April 2026 |
| Qwen3.7 Max | 91.6% | Not stated | Listed by BenchLM | Viewed October 1, 2026 |
| Gemini 3 Flash Preview (reasoning) | 90.8% | Artificial Analysis run | Artificial Analysis, independent | Viewed October 1, 2026 |
| DeepSeek V4 Pro (High) | 89.8% | Not stated | DeepSeek model card | V4 previewed April 2026 |
| Claude Opus 4.6 (Max) | 88.8% | Not stated | DeepSeek model card, as a comparison | V4 previewed April 2026 |
We could not find a published LiveCodeBench score for DeepSeek V4.1 Flash. Trackers that cite its model card list a Codeforces rating of 3471 instead. The aggregator LLM Stats lists 75 LiveCodeBench results and marks all 75 as self-reported, with none verified.
Benchmark Facts
| Item | Detail |
|---|---|
| What it tests | Competitive programming problems. Four scenarios: code generation, self-repair, test output prediction, and code execution. Leaderboards quote code generation |
| Problem sources | LeetCode, AtCoder, and Codeforces. In v6: 602 AtCoder, 444 LeetCode, 9 Codeforces |
| Problem count | 1,055 in release v6, covering May 2023 to April 2025. Release v1 had 400 |
| Difficulty mix in v6 | 322 easy, 383 medium, 350 hard |
| Scoring | Pass@1: the share of problems where generated code passes hidden tests. Default settings are 10 samples at temperature 0.2 |
| Human baseline | None published |
| Versions | v1 (400), v2 (511), v3 (612), v4 (713), v5 (880), v6 (1,055). LiveCodeBench Pro is a separately listed benchmark and its scores are not interchangeable |
Why It Resists Contamination
Every problem carries the date it was published on the contest site. A model can then be scored only on problems that appeared after its training data ends, so it cannot have seen the solutions. The official board has a date slider for this and highlights a model when its listed date falls after the start of the chosen window. This protection depends on new problems being added. The newest problem in v6 is from April 2025, and models released in 2026 generally have training data that runs past that date, so the protection is weak for current models on v6. For the general problem see benchmark saturation and contamination.
How to Read LiveCodeBench Scores
- Ask which release and window. A score on v6 and a score on an earlier set of about 500 problems are different tests.
- Check the reasoning setting. DeepSeek reports V4 Pro at 93.5 percent on Max, 89.8 percent on High, and 56.8 percent with thinking off. That is one model with a spread of about 37 points.
- Look at hard problems. On v6, every model in the official top ten solves 98 percent or more of easy problems. The differences are in the hard tier.
- Treat vendor figures as claims. Scores above 90 percent in 2026 come from model cards, not from the benchmark's maintainers.
- It is not an agent benchmark. LiveCodeBench tests single self-contained problems. For repository and terminal work see the Terminal-Bench leaderboard, the SWE-bench Pro leaderboard, and coding agent benchmarks.
Brand Visibility Implications
Open-weight vendors use LiveCodeBench as a headline coding number, and strong scores help a model get adopted inside coding assistants and self-hosted developer tools. Those tools then answer questions such as which library, API, or hosting provider to use, and increasingly install the choice themselves. A developer-tool brand is chosen or skipped by whichever model sits behind the assistant. When that model is open-weight and runs locally, the recommendation comes from training data with no search step, as described in the local LLM visibility blind spot. See also how AI agents choose brands.
Methodology
This page is built from official leaderboards, benchmark papers, and vendor announcements, not from Presenc AI measurements. Primary sources are the Terminal-Bench leaderboard and its 4.0 release note, the OSWorld 2.0 site, paper and repository, the OSWorld-Verified leaderboard, the LiveCodeBench leaderboard and repository, Anthropic's Claude Opus 5.5 announcement, and DeepSeek's V4 model card. Independent runs come from Artificial Analysis for Terminal-Bench 4.0 and LiveCodeBench. OpenAI's own pages could not be read by our fetcher, so OpenAI figures that are not on a public board come from launch coverage and are labelled that way. Every score is shown with who reported it and which benchmark version it belongs to. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks how coding models and assistants describe and recommend developer-tool brands across the prompts buyers use. That shows whether the models winning coding benchmarks name your product when a developer asks what to use.