Advanced Language Models Still Guess—and Benchmarks Reward It

More capable language models can answer more questions correctly while still producing plausible errors instead of admitting uncertainty. That warning, demonstrated across several model families in 2024, remains relevant—but it does not mean every newer model is less factual than every older one.
The important update is that researchers now have a clearer explanation for the pattern. Evidence published through 2026 indicates that hallucinations have declined in newer systems, yet common benchmarks can still reward a model for guessing when a cautious refusal would be more reliable.
The original study measured reliability, not deliberate deception
Calling these errors “lies” is rhetorically powerful but technically misleading. A lie normally implies an intention to deceive. The research examined whether language models returned correct answers, incorrect answers or avoidant responses; it did not establish that the systems understood the truth and consciously chose to conceal it.
The September 2024 Nature study compared generations of GPT, LLaMA and BLOOM models across five areas: addition, anagrams, geographical knowledge, science and information transformations. It also tested sensitivity to different natural phrasings and used human studies to estimate perceived difficulty and people’s ability to recognize model errors.
Scaling and post-training generally improved the proportion of correct responses and made answers more stable across prompt variations. The troubling trade-off was lower prudence: shaped, instruction-following models avoided fewer questions and more often supplied an apparently reasonable but incorrect answer. Their failures also did not settle into a dependable boundary where users could safely assume that easy-looking tasks would always work.
This distinction matters. A model can become more accurate overall and less cautious at the same time. If it moves from answering 50 questions to answering 90, it may produce more correct answers and more incorrect answers even when its accuracy rate improves. “Smarter” and “reliable” therefore describe different properties.
Why accuracy scores encourage confident guesses
Most conventional benchmarks give credit for a correct response and none for an incorrect response or an abstention. Under that rule, guessing has an upside and saying “I don’t know” does not. A model optimized around such evaluations can improve its headline accuracy by attempting questions beyond its dependable knowledge.
OpenAI illustrated the conflict in 2025 with a SimpleQA comparison. Its published hallucination analysis reported 22% accuracy, 26% errors and 52% abstentions for gpt-5-thinking-mini, compared with 24% accuracy, 75% errors and 1% abstention for o4-mini. The older model was two percentage points more accurate under the benchmark’s conventional measure, but it returned a wrong answer far more often.
Those figures describe two particular model configurations on one factual benchmark; they are not a universal ranking of the systems. They demonstrate why accuracy alone can hide a meaningful difference in user risk. For a creator checking an obscure date, attribution or quotation, a refusal leaves a visible research gap. A fluent invention can pass directly into a script, post or client document.
The 2026 evidence adds nuance, not a reversal
A later peer-reviewed analysis strengthened the incentives explanation while rejecting the idea that progress has stopped. The April 2026 Nature paper says hallucination rates have fallen substantially since early language models, but argues that next-token training creates statistical pressure toward errors on sparsely represented facts and that accuracy-led evaluations continue to favor guessing.
The researchers tested a simple consistency check in which a model generated two responses, compared them and abstained when they conflicted. Across 4,326 questions per model, the method reduced incorrect responses but also reduced the number of correct responses. That apparent loss of accuracy can discourage adoption even when the resulting system is safer for situations where a false answer costs more than no answer.
The paper proposes open rubrics that tell a model how correct answers, errors and abstentions will be scored. This allows behavior to change with the stakes: guessing may be acceptable in a low-consequence brainstorming exercise, while uncertainty should trigger abstention in factual or safety-critical work. The central issue is therefore not simply model size. Training objectives, evaluation rules, prompting, access to evidence and the cost of a mistake all affect observed reliability.
What creators should change in their workflow
A polished response should be treated as generated copy, not as verification. Fluency gives users few reliable clues about whether a claim came from strong evidence, a weak association or an unsupported completion. Asking the model whether it is confident is also insufficient unless the answer can be checked against external material.
For research-heavy creative work, the practical response is to separate ideation from factual approval:
- Use the model freely for outlines, alternative wording and questions to investigate, while marking factual claims as unverified.
- Ask for explicit uncertainty and permission to abstain when evidence is missing, rather than demanding an answer to every prompt.
- Require a retrievable primary document for quotations, dates, statistics, product claims and scientific conclusions.
- Open each cited page and confirm that it supports the precise sentence; fabricated or mismatched citations are themselves a form of hallucination.
- Apply stricter review to claims that could affect health, money, reputation, contracts or publication corrections.
The durable lesson is narrower—and more useful—than “the smartest AI lies the most.” Capability gains do not automatically produce dependable judgment about when to answer. Until models and their evaluations consistently reward calibrated uncertainty, the safest system is one that combines stronger generation with visible abstention and independent verification.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.