Skip to main content
Featured image for ChatGPT Accuracy Rate Statistics 2026

ChatGPT Accuracy Rate Statistics 2026

September 16, 2026 · 13 min readBot Memo

By: Editorial Staff

ChatGPT accuracy depends entirely on the task. GPT-4o scores 88.7% on MMLU but only 38.2% on OpenAI’s SimpleQA factuality benchmark, where it answers incorrectly 60.8% of the time and declines to answer just 1%. OpenAI’s own o3 and o4-mini system card puts o3’s hallucination rate on PersonQA at 33% and o4-mini’s at 48%, against 16% for o1. Reasoning capability and factual accuracy are moving in opposite directions.

On this page

ChatGPT accuracy in 2026: the short answer

“Accuracy” is not one number. It splits into at least four measurements: knowledge recall (MMLU, MMLU-Pro), factual grounding (SimpleQA, PersonQA, Vectara HHEM), reasoning (AIME, GPQA Diamond, Humanity’s Last Exam), and applied task performance (SWE-bench Verified, TAU-bench, USMLE).

Those four lenses disagree sharply. OpenAI’s o3 with tools hits 98.4% on AIME 2025 competition mathematics and a Codeforces Elo of 2706, inside the 2600-and-above band Codeforces labels International Grandmaster. The same model hallucinates on a third of short fact-seeking questions about real people.

The counterintuitive finding from 2025 is that scaling reasoning made factual accuracy worse. OpenAI’s system card says more research is needed to understand why. Traditional benchmarks saturated, vendor-reported numbers often use custom scaffolding, and the famous GPT-4 “90th percentile bar exam” claim turned out to rest on a cohort skewed toward people retaking the exam after failing it.

Every chatgpt accuracy rate below is dated to the source that published it. Where a benchmark has been rebuilt or retired since, the older reading and the current one both appear.

MMLU and HumanEval: the saturated benchmarks

GPT-4o scores 88.7% on MMLU on OpenAI’s own 0-shot chain-of-thought run. That sounds impressive until you learn frontier models have clustered above 85% for two years and the top scores are now separated by decimal points rather than percentage points.

MMLU covers roughly 14,000 multiple-choice questions across 57 subjects. HumanEval covers 164 Python programming problems. Both have saturated, and the industry has stopped pretending otherwise. Vellum’s LLM leaderboard now states plainly that it features “results from non-saturated benchmarks, excluding outdated benchmarks (e.g. MMLU)”, and its headline ranking is Humanity’s Last Exam, with GPQA Diamond as the reasoning metric.

Cost collapse is the more interesting MMLU story. Stanford’s AI Index 2025 reports that querying a model scoring GPT-3.5-equivalent MMLU (64.8) fell from $20.00 to $0.07 per million tokens between November 2022 and October 2024, a more than 280-fold reduction, with Gemini-1.5-Flash-8B sitting at the cheap end. Parameter counts fell with price. In 2022 the smallest model above 60% on MMLU was PaLM at 540 billion parameters; by 2024 Microsoft’s Phi-3-mini cleared the same bar at 3.8 billion.

Geography collapsed too. The US-China performance gap on MMLU fell from 17.5 percentage points at the end of 2023 to 0.3 points by the end of 2024. HumanEval’s gap narrowed from 31.6 points to 3.7 points over the same window. Chinese labs like DeepSeek and Moonshot AI closed the distance on legacy benchmarks while US labs moved to harder tests.

Benchmark GPT-4o result Status, September 2026
MMLU (0-shot CoT) 88.7% Dropped by Vellum as outdated
SimpleQA 38.2% correct, 60.8% incorrect Still discriminating
Vectara HHEM hallucination 9.6% Rebuilt on a harder document set in 2026
PersonQA hallucination Not published for GPT-4o OpenAI reports o1, o3 and o4-mini only

Sources: OpenAI GPT-4o announcement (May 2024), OpenAI SimpleQA paper (October 2024), Vectara HHEM leaderboard (updated May 11, 2026), Vellum LLM leaderboard (updated September 4, 2026)

Factual accuracy and hallucination rates

SimpleQA is the benchmark to watch if you care about factuality. GPT-4o scores 38.2% on SimpleQA’s 4,326 short fact-seeking questions, is incorrect on 60.8% of them, and declines to answer only 1.0%. o1-preview leads that original table at 42.7% correct, with 48.1% incorrect and 9.2% not attempted. Both sit below a D grade if this were a school test.

Fact recall is one failure mode. Grounded summarization is a different one. The Vectara HHEM leaderboard measures whether a model sticks to what a source document says when it writes a summary of that document. On the short-document set Vectara maintained until it rebuilt the leaderboard in 2026, Gemini 2.0 Flash scored 0.7% and Claude 3.7 Sonnet 4.4%, with Ant Group’s Finix S1 32B a shade ahead of both at 0.6% in that board’s final October 2025 reading.

Vectara then rebuilt the leaderboard on a much harder set: over 7,700 articles spanning news, science, medicine, law, business and education, from 50 words to 24,000. Every rate moved. In the May 11, 2026 reading, Ant Group’s Finix S1 32B holds the lowest rate at 1.8%, GPT-5.4-nano is second at 3.1%, and GPT-4o sits at 9.6%. Reasoning-heavy configurations sit much further down: DeepSeek-R1 at 11.3%, Claude Sonnet 4.5 at 12.0%, GPT-5.1-high at 12.1%, GPT-5-high at 15.1% and Grok-4-fast-reasoning at 20.2%. Longer documents and longer reasoning chains expose more failure surface.

OpenAI’s internal PersonQA benchmark tells the same story from a different angle. o3 hallucinates on 33% of PersonQA questions and o4-mini on 48%, against 16% for o1. OpenAI’s own explanation is that o3 makes more claims per answer, which produces both more accurate claims and more invented ones, and the system card says more research is needed to explain the rest.

Fabricated citations are the version of this that academics keep catching. A 2023 study in the Canadian Psychological Association’s Mind Pad asked ChatGPT for references across psychology subfields and found 32.3% of the citations were fake, with the per-subfield rate running from 6% to 60%. The invented references carried real researchers’ names, which is what made them hard to spot.

There is no single chatgpt hallucination rate. On the short-document set Vectara retired in 2026, the best models stayed under 5%. On the harder 2026 document set, most frontier reasoning models are above 10%. On fact-seeking queries with no source attached, they are wrong the majority of the time.

Reasoning model benchmarks: o3 and o4-mini

Reasoning models changed the accuracy conversation. OpenAI o3 with tools scores 98.4% on AIME 2025, and o4-mini with tools scores 99.5% on the same competition mathematics problems. Without Python interpreter access, the scores drop to 88.9% and 92.7%. Tools matter.

Codeforces Elo ratings tell the coding story. o3 reaches 2706 Elo and o4-mini 2719. Codeforces labels 2400 to 2599 Grandmaster and 2600 to 2999 International Grandmaster, so both models sit in the higher band.

On SWE-bench Verified, o3 scored 71.7%. OpenAI has since stopped reporting the metric. The company audited 138 problems that o3 failed to solve consistently across 64 runs and found 59.4% of them carried material problems in test design or in the problem description itself, including tests that enforce implementation details and tests that check for behavior the task never specified. That is a rare public admission that a headline benchmark was broken, and it arrived while state of the art on the same benchmark was creeping from 74.9% to 80.9% in six months.

Harder tests replaced it. Humanity’s Last Exam, published in January 2025 by the Center for AI Safety and Scale AI and now holding 2,500 expert-written questions, was meant to hold up for years. In its first round of testing, Stanford’s AI Index 2025 recorded o1 as the top scorer at 8.8%. By September 2026 the leader on Artificial Analysis is Claude Fable 5.1 at 59.1%, and the version of the board maintained at agi.safe.ai puts Gemini 3 Pro on 38.3%. Expert humans average roughly 90% in their own domain, so there is still room, but far less than the benchmark’s authors expected.

o3 still hallucinates on 33% of PersonQA questions and o4-mini on 48%. Reasoning scaling improved mathematics and coding dramatically while worsening factuality. Any production deployment on these models needs retrieval augmentation or grounding guardrails.

Claude 3.7 Sonnet benchmark comparison

Anthropic’s Claude 3.7 Sonnet introduced “extended thinking” as a toggle between fast answers and reasoning-heavy answers. The compute tradeoff shows up across benchmarks.

On SWE-bench Verified, Claude 3.7 Sonnet scores 70.3% with custom scaffolding and 63.7% without, both measured on the 489-task subset that ran on Anthropic’s infrastructure rather than the full 500. On GPQA Diamond it scores 68.0% in standard mode and 84.8% with extended thinking enabled. Same model, different inference strategy, 16.8 percentage points apart.

Benchmark Standard Extended Thinking
GPQA Diamond 68.0% 84.8%
SWE-bench Verified (n=489 subset) 63.7% 70.3% (with scaffolding)
IFEval 93.2% (reported as a single figure)
TAU-bench retail 81.2% (reported as a single figure)

Source: Anthropic Claude 3.7 Sonnet announcement, February 24, 2025

IFEval at 93.2% and TAU-bench retail at 81.2% are the practical numbers for teams building agents. IFEval measures instruction-following precision. TAU-bench measures tool-use success on simulated retail and airline customer service tasks. These translate more directly to production than MMLU.

Anthropic never published an MMLU accuracy score for Claude 3.7 Sonnet. Its system card carries no capability benchmarks at all, and uses MMLU questions only as prompt material for a chain-of-thought faithfulness test, where the model scored 0.30. The GPQA Diamond, IFEval and TAU-bench figures above come from the launch announcement instead.

Companies building on Claude 3.7 Sonnet ranged from code agents like Cursor to customer support platforms like Sierra AI. Its 4.4% hallucination rate on the Vectara short-document set that ran until 2026 made it a reasonable pick for document-heavy work at the time.

Gemini 2.0 benchmark comparison

Google DeepMind’s Gemini 2.0 Flash, announced in December 2024, landed in a different position. It trailed the frontier models on raw capability. For a long stretch it was among the most grounded.

On the Vectara HHEM short-document set that ran until 2026, Gemini 2.0 Flash posted a 0.7% hallucination rate, second only to Ant Group’s Finix S1 32B. That made it a default pick for RAG (retrieval-augmented generation) pipelines, where faithfulness to a source document matters more than raw reasoning. Enterprise search companies like Glean and Hebbia optimize for exactly this tradeoff. On the rebuilt 2026 board, Gemini 2.0 Flash no longer appears and Gemini 2.5 Flash Lite is Google’s best entry at 3.3%.

Google’s December 2024 numbers on classical benchmarks:

Google’s FACTS Leaderboard, published in December 2025, aggregates four factuality tests: grounding, multimodal, closed-book parametric knowledge, and search. In the launch paper’s table, Gemini 3 Pro came first with a FACTS Score of 68.8 out of 100, built on 69.0 for grounding, 46.1 for multimodal, 76.4 for parametric and 83.8 for search. Gemini 2.5 Pro followed on 62.1, and the widest gaps between the two were on search (83.8 against 63.9) and parametric knowledge (76.4 against 63.2). Not one of the fifteen models in that table cleared 70 overall, which is a useful corrective to the idea that benchmark saturation means factuality is solved.

Benchmark scores do not transfer to practice. The clearest evidence comes from a JAMA Network Open randomized clinical trial in which 50 physicians worked through up to six clinical vignettes with and without ChatGPT Plus access. GPT-4 alone posted a median diagnostic reasoning score of 92%. Physicians using ChatGPT Plus scored 76%. Physicians without it scored 74%. The two-point gap between the two physician groups was not statistically significant (P = .60), while GPT-4 alone beat the conventional-resources group by 16 points (P = .03).

Model capability does not automatically raise practitioner accuracy. Interface, prompting and workflow gaps absorb most of the gain. Clinical AI platforms like Abridge and OpenEvidence build their products around that gap, with structured prompts and verified citations in place of open-ended chat. Bot Memo’s list of top AI healthcare startups covers the wider market.

Other medical readings:

The legal accuracy story is worse. The famous GPT-4 “90th percentile bar exam” figure was benchmarked against a February cohort weighted toward candidates who had already failed the July sitting. Martinez’s re-evaluation in Artificial Intelligence and Law puts GPT-4 below the 69th percentile on the Uniform Bar Exam against a general July cohort, at the 62nd percentile against first-time takers (42nd on the essays), and at the 45th percentile against candidates who actually passed (15th on the essays), a figure the paper’s abstract rounds to the 48th. The original framing reached thousands of legal-industry press releases before the correction landed. For investors tracking AI legal tech startups, the distance between benchmark framing and production accuracy is the core due-diligence question.

For coding, o3’s 71.7% on SWE-bench Verified was real before OpenAI pulled the metric. Code-agent companies like Cognition (Devin) and Factory now lean on customer-specific eval suites, because SWE-bench numbers stopped predicting production performance. Bot Memo’s guide to AI developer tools covers the wider set of products built on these models.

The benchmark measurement crisis

A credibility gap is opening between vendor-reported numbers and production reality. Four signals:

First, OpenAI stopped reporting SWE-bench Verified after auditing 138 problems o3 kept failing and finding 59.4% of them had flawed tests or underspecified problem statements. The benchmark that shaped 2024’s narrative about AI coding ability is now considered unfit for frontier launches by the company whose scores made it famous.

Second, vendor-reported numbers routinely use tooling, scaffolding or best-of-N sampling that does not match production deployment. Anthropic’s 70.3% on SWE-bench Verified uses custom scaffolding on a 489-task subset. OpenAI’s 98.4% on AIME uses Python interpreter access. The headline numbers are not wrong, but they are not the numbers you get from a plain API call.

Third, a published score is a snapshot of one model version, not a property of the product. Stanford and UC Berkeley researchers ran the March and June 2023 builds of GPT-4 over the same tasks and watched prime-versus-composite accuracy fall from 84% to 51% in three months, which they traced partly to the model becoming less willing to follow chain-of-thought prompting. The product name never changed.

Fourth, accuracy tooling from the vendors is no better. OpenAI shipped a classifier for detecting AI-written text, reported that it correctly flagged only 26% of AI-written text while mislabeling 9% of human text, and withdrew it six months later for low accuracy.

Coverage gaps remain. Anthropic published no MMLU score for Claude 3.7 Sonnet. Google’s Gemini 2.0 Pro benchmark table is thin next to Flash. Vectara’s HHEM measures grounded summarization and not fact-seeking queries. Google’s own FACTS benchmark suite paper put no model above 70 on the aggregate score.

Pick the benchmark that matches your task. If you run a RAG pipeline, Vectara HHEM matters. If you build code agents, build your own eval suite. If you evaluate reasoning, track GPQA Diamond and Humanity’s Last Exam.

Frequently Asked Questions

How accurate is ChatGPT in 2026?

ChatGPT accuracy ranges from 38.2% on SimpleQA fact-seeking queries (GPT-4o) to 98.4% on AIME 2025 mathematics (o3 with tools). There is no single accuracy number because the benchmarks measure different capabilities. For general knowledge, frontier models cluster near 88% on MMLU. For grounded summarization, the best model on Vectara’s 2026 board hallucinates on 1.8% of summaries and GPT-4o on 9.6%.

What is ChatGPT’s hallucination rate?

Task-dependent. On the Vectara HHEM leaderboard as of May 11, 2026, GPT-4o hallucinates on 9.6% of document summaries and GPT-5.4-nano on 3.1%. On OpenAI’s internal PersonQA, o3 hallucinates 33% and o4-mini 48%. On SimpleQA, GPT-4o answers incorrectly on 60.8% of fact-seeking questions.

Is GPT-4o more accurate than o1?

On fact-seeking SimpleQA queries, o1-preview leads at 42.7% correct against GPT-4o’s 38.2%. On reasoning benchmarks like AIME and GPQA Diamond, o1 and its successors far outperform GPT-4o. On grounded summarization, the newer reasoning models are worse: Vectara’s 2026 board has GPT-5-high at 15.1% and DeepSeek-R1 at 11.3%, against 9.6% for GPT-4o.

Did ChatGPT pass the bar exam in the 90th percentile?

No. The original 90th-percentile figure came from a February sitting whose cohort skewed toward repeat takers. Martinez’s re-evaluation in Artificial Intelligence and Law puts GPT-4 at the 62nd percentile against first-time takers and the 45th percentile against candidates who passed, with essay performance at the 42nd and 15th percentiles respectively.

Why do reasoning models like o3 hallucinate more?

OpenAI’s o3 and o4-mini system card attributes part of it to o3 making more claims per answer, which yields more correct claims and more invented ones, and says more research is needed to explain the rest. PersonQA hallucination rates roughly doubled from o1 to o3 and tripled to o4-mini.

Which AI model has the lowest hallucination rate?

On the current Vectara HHEM board, Ant Group’s Finix S1 32B leads at 1.8%, ahead of GPT-5.4-nano at 3.1% and Gemini 2.5 Flash Lite at 3.3%. That measures faithfulness to a supplied document. For fact-seeking queries with no source attached, the best score on SimpleQA at publication was o1-preview’s 42.7% correct, with 48.1% wrong and 9.2% declined.

Methodology

This analysis draws on published LLM accuracy and hallucination benchmarks from OpenAI, Anthropic, Google DeepMind, Stanford HAI, Vectara, Vellum, Artificial Analysis, the Center for AI Safety, JAMA, JAMA Network Open, JMIR, the Annals of Surgical Treatment and Research, Scientific Reports, the Canadian Psychological Association and Artificial Intelligence and Law, covering 2023 through September 2026.

Data sources: OpenAI GPT-4o announcement (May 2024), OpenAI SimpleQA paper (October 2024), OpenAI o3 and o4-mini announcement and system card (April 2025), OpenAI SWE-bench Verified retirement analysis (2025), Anthropic Claude 3.7 Sonnet announcement and system card (February 2025), Google DeepMind Gemini 2.0 announcement (December 2024), Google DeepMind FACTS Leaderboard paper (December 2025), Stanford AI Index 2025 (April 2025), Vectara HHEM hallucination leaderboard (May 2026 reading), Vellum LLM leaderboard (September 2026 reading), Artificial Analysis Humanity’s Last Exam board (September 2026 reading), JAMA Network Open RCT on ChatGPT clinical reasoning (October 2024), Kanjee et al. on NEJM clinicopathologic cases (JAMA, August 2023), Hirosawa et al. on differential diagnosis (JMIR, 2023), the Korean general surgery board evaluation (Annals of Surgical Treatment and Research, 2023), Scientific Reports USMLE evaluation (April 2024), Chen, Zaharia and Zou on GPT behavior drift (arXiv, July 2023), and the Martinez bar exam re-evaluation (Artificial Intelligence and Law, March 2024).

Scope: only benchmarks with published methodology from an official vendor or a peer-reviewed source are included. Scores that require custom scaffolding, tool access or best-of-N sampling are labeled as such.

Dating: leaderboards move. Every leaderboard figure carries the date of the reading it comes from, and where a board has been rebuilt on a new dataset, both the old and the current reading appear.

Currency: all cost comparisons in USD.

Last reviewed: September 16, 2026.

Bot Memo

About the author

Editorial Staff

The Editorial Staff at Bot Memo is a team of writers, analysts, and AI agents dedicated to mapping the global AI startup ecosystem. Led by Chintan Zalani, the team tracks thousands of funding rounds, classifies companies across verticals, and distills it all into actionable intelligence for investors and founders.

Subscribe to AI Funding Memo

Weekly pre-seed and seed AI deal intelligence with early trends.