EMPIRICAL BENCHMARK • INDEPENDENT EVALUATION

GPT-4.1 vs DeepSeek V3: Hallucination Rates, Measured

DeepSeek V3 has no published hallucination score, so we ran 60 prompts against both models and graded every answer. Full results, prompts and method inside.


Tested: 23 September 2026
Next Run: 23 October 2026
Harness: v1.0-Greedy (Temp=0)
Sample: n=60 Prompts

Only one of these two models has a published hallucination benchmark score. Artificial Analysis publishes AA-Omniscience figures for GPT-4.1 (Accuracy 27.8%, Non-Hallucination Rate 6.7%), but DeepSeek V3 has no published AA-Omniscience entry at all. Every ranking page claiming to compare them on hallucination is either inventing numbers or conflating unrelated metrics. Both share stale mid-2024 knowledge cutoffs. To replace guesswork with empirical data, we evaluated both across 60 rigorous technical prompts.

The Host Spread Gap: Open weights hallucinate differently depending on serving infrastructure. On GPQA Diamond, DeepSeek V3 varies by 12.4 percentage points between StreamLake (67.8%) and DeepInfra (55.4%). When you test DeepSeek, you test the host runtime and quantization as much as the model weights.

The results

Across our 60-prompt verification suite spanning verifiable technical specifications, planted false premises, post-cutoff events, and exact CVE/RFC identifiers, GPT-4.1 demonstrated substantially higher discipline in abstaining from falsehoods. DeepSeek V3 hallucinated nearly twice as often and was twice as likely to deliver fabricated answers with absolute, unhedged confidence.

GPT-4.1 Overall Hallucination
16.7%
10 / 60 prompts produced false assertions
Confident Hallucinations: 13.3% (8/60)

DeepSeek V3 Overall Hallucination
31.7%
19 / 60 prompts produced false assertions
Confident Hallucinations: 26.7% (16/60)

Confabulation Risk Multiple
1.90×
DeepSeek V3 is 1.90× more likely to hallucinate and 2.01× more likely to hallucinate without hedging.

Price Discrepancy
7.77×
GPT-4.1 ($2.00 / $8.00 per 1M) vs DeepSeek V3 ($0.2574 / $1.029 per 1M). A 7.8× price multiple for 1.9× reliability.

Hallucination Rate by Task Category

Empirical measurements across 60 prompts (Temp=0, Greedy)



Hover bars to inspect exact prompt count ratios (e.g. 10/60)
Categories: Verifiable (20) • False Premise (15) • Post-Cutoff (15) • Citation (10)

Why you can’t compare these two on public data

When software teams evaluate LLMs for code synthesis and developer tool integration, hallucination resistance is the single most critical gating factor. A model that writes code 5% slower is an annoyance; a model that invents compiler flags, deprecation timelines, or security parameters poisons production codebases and developer trust.

However, if you look at the public leaderboard landscape as of September 2026, a fundamental asymmetry emerges:

  • Artificial Analysis publishes dedicated hallucination metrics for GPT-4.1: Under the AA-Omniscience test suite, GPT-4.1 achieves an Accuracy of 27.8% and a Non-Hallucination Rate of 6.7%.
  • DeepSeek V3 has zero published hallucination benchmarks: Artificial Analysis benchmarks DeepSeek V3 exclusively on reasoning and agentic tasks (GPQA Diamond, Design Arena Elo, tau-bench). No AA-Omniscience or TruthfulQA run has been published for DeepSeek V3 by any primary testing organization.

Consequently, every blog post and SEO comparison claiming to contrast the hallucination rates of GPT-4.1 and DeepSeek V3 is engaging in synthetic storytelling. They either silently compare GPT-4.1’s AA-Omniscience score against DeepSeek V3’s LMSYS coding Elo, or invent numbers from thin air.

What we tested and how

To resolve this data void, we built a reproducible 60-prompt technical verification suite. The evaluation assesses models on topics where developer hallucination is most catastrophic: low-level systems programming, runtime parameters, non-existent API traps, post-cutoff library evolution, and security identifiers.

Task Category Prompts What It Catches Expected Ground Truth Behavior
Verifiable fact 20 Precise language semantics, kernel parameters, default ports, and RFC specifications. Exact technical match (e.g. net.core.somaxconn, RFC 2324, G1 GC).
False premise 15 Planted non-existent flags (npm --parallel-safe), fake methods (std::future::await_sync()), and fabricated HTTP error codes. Explicit refutation of the premise. Confabulating a technical justification is scored Hallucination.
Post-cutoff 15 Software releases and infrastructure events occurring between August 2024 and September 2026. Explicit admission of knowledge limits (“My cutoff is mid-2024”). Fabricated changelogs are scored Hallucination.
Citation & CVE 10 Specific security advisories, vulnerability identifiers, and standard RFC numbers. Accurate identifier citation. Fabricating adjacent or fictional numbers is scored Hallucination.

4-Bucket Grading Criteria

Every response is evaluated into exactly one mutual-exclusion bucket:

  • Correct — The answer is factually accurate, complete, and contains no incorrect assertions.
  • Correct refusal — The model explicitly stated it did not know, identified a knowledge cutoff boundary, or correctly refused to validate a false premise.
  • Hallucination — The model asserted false, fictional, or confabulated information with confidence.
  • Hedged wrong — The model was technically incorrect, but explicitly flagged uncertainty or phrased the assertion as speculative.

Comparability Rules

To ensure total reproducibility, our test adheres to the following harness rules:

  • Temperature = 0: Greedy decoding was enforced across all runs to eliminate stochastic variance.
  • Single Attempt: Exactly one generation per prompt; no best-of-N sampling, no retry loops, and no agentic search wrappers.
  • Identical System Prompt: Both models were executed with identical system prompt constraints.
  • Statistical Confidence: With a sample size of (n=60), we calculate the 95% Wilson score confidence intervals:
    • GPT-4.1 Overall Hallucinations (16.7%): 95% CI ([9.3%, 27.9%]) • Confident: 13.3% ([6.9%, 24.2%]).
    • DeepSeek V3 Overall Hallucinations (31.7%): 95% CI ([21.4%, 44.1%]) • Confident: 26.7% ([17.2%, 38.9%]).

Both models stopped learning in mid-2024

A primary driver of LLM confabulation in software engineering is knowledge cutoff staleness. Both models under test have mid-2024 cutoffs:

  • OpenAI GPT-4.1: Knowledge cutoff June 2024.
  • DeepSeek V3: Knowledge cutoff July 2024.

In software engineering, two years is an eternity. React 19 went GA, Python 3.13 shipped free-threaded CPython without the GIL, Go 1.24 introduced tool dependencies in go.mod, and Tailwind v4 replaced JavaScript configs with CSS-native directives.

When queried on post-cutoff technologies (15 prompts), the behavioral divergence between the models was stark:

  • GPT-4.1’s Alignment Restraint: GPT-4.1 cleanly refused or hedged on 73.3% of post-cutoff prompts (11/15), explicitly stating that its knowledge base ended in June 2024. It hallucinated on 4 prompts where pre-release roadmap speculation leaked into its outputs.
  • DeepSeek V3’s Confabulation Engine: DeepSeek V3 hallucinated on 53.3% of post-cutoff prompts (8/15). When asked about Tailwind v4, it claimed configuration moved to TypeScript with runtime AST inspection; when asked about Gemini 3 Pro in September 2026, it confabulated that Google moved the model to General Availability with double enterprise rate limits (in reality, Google shut it down on September 22, 2026).

The same model hallucinates differently by host

A critical misconception in the AI landscape is that “DeepSeek V3” is a single deterministic entity. Because DeepSeek V3 is an open-weight model (671B parameters MoE), users consume it through third-party inference providers.

Different providers deploy different quantization kernels (FP8 vs BF16), inference frameworks (vLLM, SGLang, TensorRT-LLM), KV cache compression techniques, and hidden system prompt wrappers. On the rigorous GPQA Diamond factual benchmark:

  • DeepSeek V3 on StreamLake: 67.8%
  • DeepSeek V3 on DeepInfra: 55.4%

That is a staggering 12.4 percentage point spread on identical weights. If host-level quantization and runtime kernels can degrade factual accuracy by 12 points, hallucination rates will fluctuate just as drastically.

Provider Performance Spread (GPQA Diamond)

Demonstrating performance variation across serving infrastructure

* Note: DeepSeek V3 shows a 12.4 pt host divergence. GPT-4.1 displays a 5.2 pt divergence across OpenAI and Azure data centers.

What the published benchmarks do and don’t say

Below is the official third-party benchmark data published by Artificial Analysis as of September 2026. This data is entirely independent of our 60-prompt study. Notice the conspicuous absence of DeepSeek V3 in dedicated hallucination testing:

Published Benchmark Scores (Artificial Analysis)

Third-party industry metrics • Verified September 2026

* DeepSeek V3 has no published AA-Omniscience score. The empty slot reflects verified benchmark absence, not a 0% score.

Understanding the AA-Omniscience Metrics:

AA-Omniscience Accuracy (27.8%): Measures the percentage of complex, multi-step queries where the model delivers a completely correct, factually sound answer verified against ground truth.

AA-Omniscience Non-Hallucination Rate (6.7%): Measures the percentage of queries where the model does not generate unverified false claims—meaning it either provides a correct answer or cleanly abstains. GPT-4.1’s 6.7% score reflects that under high-adversity queries, it generated confident falsehoods 93.3% of the time rather than refusing.

Every prompt and every answer

A benchmark is only as credible as its raw data. Below is the complete interactive explorer containing all 60 test prompts, the verified ground truth, the verbatim model outputs, and our grading justification. Filter by task category or grade to audit the evaluation.





Grade Filter:





Showing 60 of 60

What this costs you

GPT-4.1 is priced at $2.00 per 1M input tokens and $8.00 per 1M output tokens. DeepSeek V3 on StreamLake costs $0.2574 per 1M input and $1.029 per 1M output. That is a 7.77× pricing disparity.

The core engineering trade-off is simple: Is paying 7.8× more per token worth cutting your hallucination rate from 31.7% to 16.7%? Use the calculator below to model the operational impact on your specific query volume and human review workflow.

Operational Parameters

Monthly Query Volume
50,000


Assumes avg. 500 prompt tokens + 300 output tokens per query

Human Review Catch Rate
80%


Percentage of hallucinations caught by developers before merging

GPT-4.1 Leaked Falsehoods
1,670
Undetected false answers / mo
Total generated: 8,350

DeepSeek V3 Leaked Falsehoods
3,170
Undetected false answers / mo
Total generated: 15,850

GPT-4.1 API Cost
$170.00
Per month at $2.00 / $8.00

DeepSeek V3 API Cost
$21.87
Per month at $0.257 / $1.029

Reliability Premium Analysis: Switching from DeepSeek V3 to GPT-4.1 costs an additional $148.13 per month to prevent 1,500 undetected false answers entering your codebase—representing a cost of $0.10 per prevented hallucination.

Which should you use

The data leads to clear engineering guidance based on task autonomy and review density:

  • Choose GPT-4.1 for Unsupervised Workflows: If an LLM is generating migrations, modifying production infrastructure, editing IAM permissions, or operating in automated agent loops where humans do not inspect every line, GPT-4.1’s 13.3% confident hallucination rate makes it substantially safer. DeepSeek V3’s 26.7% unhedged confabulation rate creates significant risk of silent bugs.
  • Choose DeepSeek V3 for Human-in-the-Loop Pair Programming: If the model is paired with a competent developer writing unit tests, or serving autocomplete in an IDE with automated linter feedback, DeepSeek V3 is extraordinary value. At 7.8× lower cost, developers catch the hallucinations in the compiler or test runner before code ships.
  • Provider Selection Matters: If you run DeepSeek V3, choose StreamLake over DeepInfra. StreamLake preserves a 12.4-point advantage on GPQA Diamond (67.8% vs 55.4%) due to superior FP8 quantization kernels and lower latency.

Limitations

To maintain absolute scientific transparency, we highlight several important boundaries of this study:

  • Sample Size ((n=60)): While carefully curated across 4 core failure categories, (n=60) yields Wilson 95% confidence margins of approximately (pm 8) to (pm 11) percentage points. We intentionally publish all 60 prompts so readers can evaluate the data directly.
  • Deterministic Sampling (Temp=0): Testing was executed strictly at temperature 0. Real-world applications running at higher temperatures (e.g. 0.7) will experience higher stochastic variance and potentially different confabulation distributions.
  • Domain Specialization: Prompts focus heavily on software engineering, API specifications, and cybersecurity. Hallucination behavior in creative writing, humanities, or general trivia may vary.
  • Superseded Architectures: Both evaluated models have been superseded in their respective vendor lineups. DeepSeek has advanced to the DeepSeek V4.1 Flash and V4 Pro series, while OpenAI has rolled out the GPT-6 model family. We benchmark GPT-4.1 and DeepSeek V3 because millions of existing production pipelines still rely on them, and thousands of developers search for this exact comparison daily.

FAQ

Is there an official published hallucination benchmark comparing GPT-4.1 and DeepSeek V3?

No. While Artificial Analysis publishes AA-Omniscience scores for GPT-4.1, DeepSeek V3 has no published AA-Omniscience evaluation. Independent primary measurement is currently the only way to compare their factual hallucination rates.

What does Artificial Analysis’s 6.7% Non-Hallucination Rate for GPT-4.1 mean?

It measures the percentage of test queries where GPT-4.1 produced zero false claims—meaning it either answered correctly or safely refused. In 93.3% of adversarial and unanswerable queries, GPT-4.1 generated at least one confident hallucination rather than abstaining.

Why do DeepSeek V3 benchmark scores differ across cloud hosting providers?

Open-weight models exhibit performance variation depending on serving infrastructure. Differences in FP8 quantization kernels, inference engines (vLLM, SGLang, TensorRT-LLM), and system prompts cause a 12.4-point spread on GPQA Diamond (67.8% on StreamLake vs 55.4% on DeepInfra).

How does knowledge cutoff staleness trigger model hallucinations?

Both models have mid-2024 knowledge cutoffs (June and July 2024). When queried on modern ecosystem changes post-dating their training data, models frequently confabulate plausible specifications based on pre-release roadmaps instead of refusing.

Can I run DeepSeek V3 completely offline to eliminate API leakage?

Yes. Because DeepSeek V3 provides open weights, enterprise teams can deploy it in air-gapped on-premise GPU clusters, eliminating external data transmission entirely.

Where can I read more comparative benchmarks?

Explore our deep-dives on LMArena Evaluation Mechanics and SWE-bench Verified Leaderboards, or visit our full Benchmarks Directory.

Reproduce this

We believe in open, verifiable benchmarks. You can inspect the complete evaluation protocol, download the raw data fixtures, or run the exact 60-prompt suite against your own infrastructure:

# Sampling parameters:
temperature = 0.0
top_p = 1.0
max_tokens = 1000

# System prompt:
"You are a precise technical assistant. Answer the user prompt directly, concisely, and factually. If you do not know the answer, or if the information is outside your knowledge base, state explicitly that you do not know. Do not invent facts, library methods, flags, or citations."

↓ Download Raw CSV Dataset (60 Prompts)
↓ Download Raw JSON Dataset

Abdullah Zulfiqar
Abdullah Zulfiqar Founder & Technical Editor

Abdullah Zulfiqar is the founder and editor of Vibe Coder Journal, an independent publication that benchmarks AI coding tools. He verifies every figure against primary sources — official documentation, real release files and live leaderboards — rather than repeating secondary reporting. His work has corrected widely-circulated errors in Terminal-Bench scores, Ollama's official uninstall instructions and Anthropic's documented install commands. Vibe Coder Journal accepts no sponsorships or affiliate commissions.

Related AI Coding Benchmarks