Research

LiveCodeBench Leaderboard 2026

LiveCodeBench scores as of October 1, 2026: the official leaderboard, vendor-reported 2026 results led by DeepSeek V4 Pro at 93.5 percent, what the benchmark measures, and how it limits contamination.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: October 2026

As of October 1, 2026, the highest LiveCodeBench score we could trace to a primary source is 93.5 percent pass@1 for DeepSeek V4 Pro at its Max reasoning setting, reported by DeepSeek on its own model card. That number is not on the official leaderboard. The official board still runs on release v6, its data file was last changed on August 1, 2025, and its top entry is OpenAI's o4-mini at high effort with 80.2 percent on the default date window. Anyone quoting a LiveCodeBench score in 2026 is almost always quoting a vendor's own run.

Official LiveCodeBench Leaderboard

These rows are calculated from the leaderboard's public data file for the default view: 454 problems released between August 1, 2024 and May 1, 2025. The board holds 28 models and none was released in 2026.

RankModelPass@1Hard problemsReported byDate
1o4-mini (high)80.2%63.5%Official leaderboardData file dated August 1, 2025
2o3 (high)75.8%57.1%Official leaderboardSame
3o4-mini (medium)74.2%52.7%Official leaderboardSame
4Gemini 2.5 Pro (06-05)73.6%50.2%Official leaderboardSame
5DeepSeek R1 (0528)73.1%50.7%Official leaderboardSame
6Gemini 2.5 Pro (05-06)71.8%50.2%Official leaderboardSame

On the full v6 set of 1,055 problems, o4-mini (high) scores 87.3 percent. The date window changes the number by seven points for the same model, which is why the window has to be stated.

2026 Scores Reported Outside the Official Board

This table is not a ranking. The rows come from different runs, and several do not state which release or window they used.

ModelScoreProblem setReported byDate
DeepSeek V4 Pro (Max)93.5%Not statedDeepSeek model cardV4 previewed April 2026
Sakana Fugu-Ultra93.2%v6Provider figure, listed by BenchLMViewed October 1, 2026
Gemini 3 Pro Preview (high)91.7%Artificial Analysis runArtificial Analysis, independentViewed October 1, 2026
DeepSeek V4 Flash (Max)91.6%Not statedDeepSeek model cardV4 previewed April 2026
Qwen3.7 Max91.6%Not statedListed by BenchLMViewed October 1, 2026
Gemini 3 Flash Preview (reasoning)90.8%Artificial Analysis runArtificial Analysis, independentViewed October 1, 2026
DeepSeek V4 Pro (High)89.8%Not statedDeepSeek model cardV4 previewed April 2026
Claude Opus 4.6 (Max)88.8%Not statedDeepSeek model card, as a comparisonV4 previewed April 2026

We could not find a published LiveCodeBench score for DeepSeek V4.1 Flash. Trackers that cite its model card list a Codeforces rating of 3471 instead. The aggregator LLM Stats lists 75 LiveCodeBench results and marks all 75 as self-reported, with none verified.

Benchmark Facts

ItemDetail
What it testsCompetitive programming problems. Four scenarios: code generation, self-repair, test output prediction, and code execution. Leaderboards quote code generation
Problem sourcesLeetCode, AtCoder, and Codeforces. In v6: 602 AtCoder, 444 LeetCode, 9 Codeforces
Problem count1,055 in release v6, covering May 2023 to April 2025. Release v1 had 400
Difficulty mix in v6322 easy, 383 medium, 350 hard
ScoringPass@1: the share of problems where generated code passes hidden tests. Default settings are 10 samples at temperature 0.2
Human baselineNone published
Versionsv1 (400), v2 (511), v3 (612), v4 (713), v5 (880), v6 (1,055). LiveCodeBench Pro is a separately listed benchmark and its scores are not interchangeable

Why It Resists Contamination

Every problem carries the date it was published on the contest site. A model can then be scored only on problems that appeared after its training data ends, so it cannot have seen the solutions. The official board has a date slider for this and highlights a model when its listed date falls after the start of the chosen window. This protection depends on new problems being added. The newest problem in v6 is from April 2025, and models released in 2026 generally have training data that runs past that date, so the protection is weak for current models on v6. For the general problem see benchmark saturation and contamination.

How to Read LiveCodeBench Scores

  • Ask which release and window. A score on v6 and a score on an earlier set of about 500 problems are different tests.
  • Check the reasoning setting. DeepSeek reports V4 Pro at 93.5 percent on Max, 89.8 percent on High, and 56.8 percent with thinking off. That is one model with a spread of about 37 points.
  • Look at hard problems. On v6, every model in the official top ten solves 98 percent or more of easy problems. The differences are in the hard tier.
  • Treat vendor figures as claims. Scores above 90 percent in 2026 come from model cards, not from the benchmark's maintainers.
  • It is not an agent benchmark. LiveCodeBench tests single self-contained problems. For repository and terminal work see the Terminal-Bench leaderboard, the SWE-bench Pro leaderboard, and coding agent benchmarks.

Brand Visibility Implications

Open-weight vendors use LiveCodeBench as a headline coding number, and strong scores help a model get adopted inside coding assistants and self-hosted developer tools. Those tools then answer questions such as which library, API, or hosting provider to use, and increasingly install the choice themselves. A developer-tool brand is chosen or skipped by whichever model sits behind the assistant. When that model is open-weight and runs locally, the recommendation comes from training data with no search step, as described in the local LLM visibility blind spot. See also how AI agents choose brands.

Methodology

This page is built from official leaderboards, benchmark papers, and vendor announcements, not from Presenc AI measurements. Primary sources are the Terminal-Bench leaderboard and its 4.0 release note, the OSWorld 2.0 site, paper and repository, the OSWorld-Verified leaderboard, the LiveCodeBench leaderboard and repository, Anthropic's Claude Opus 5.5 announcement, and DeepSeek's V4 model card. Independent runs come from Artificial Analysis for Terminal-Bench 4.0 and LiveCodeBench. OpenAI's own pages could not be read by our fetcher, so OpenAI figures that are not on a public board come from launch coverage and are labelled that way. Every score is shown with who reported it and which benchmark version it belongs to. Status as of October 1, 2026.

How Presenc AI Helps

Presenc AI tracks how coding models and assistants describe and recommend developer-tool brands across the prompts buyers use. That shows whether the models winning coding benchmarks name your product when a developer asks what to use.

Frequently Asked Questions

LiveCodeBench is a coding benchmark built from competitive programming problems published on LeetCode, AtCoder, and Codeforces. Each problem keeps its publication date, so a model can be scored only on problems released after its training data ends. The current release, v6, has 1,055 problems from May 2023 to April 2025, and the headline metric is pass@1 on code generation.
It depends on the source. As of October 1, 2026, the official leaderboard has not added 2026 models and is led by o4-mini (high) at 80.2 percent on its default window. The highest figure we could trace to a primary source is DeepSeek V4 Pro (Max) at 93.5 percent, which DeepSeek reports on its own model card without stating the release or date window.
We could not find a published LiveCodeBench score for DeepSeek V4.1 Flash as of October 1, 2026. Trackers that cite its model card list a Codeforces rating of 3471. The earlier DeepSeek V4 Flash is reported by DeepSeek at 91.6 percent on Max, 88.4 percent on High, and 55.2 percent with thinking off.
They are useful with caveats. The contamination protection works only for problems newer than a model's training data, and the newest problem in release v6 dates from April 2025. Most 2026 scores are self-reported by vendors, often without the release or window. Compare scores only when the release, window, and reasoning setting match.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.