AI can sound equally certain when it knows a lot and when it is guessing. Tone is generated separately from factual correctness, so confidence should not be treated as evidence.
OpenAI argues that standard training and evaluation can reward guessing over admitting uncertainty. If a system is scored only on whether it produces the right answer, guessing can sometimes outperform saying “I don’t know” — even when that creates many more confident errors. OpenAI: Why language models hallucinate.
Why better models can still be overconfident
Research from MIT CSAIL has shown that reinforcement-learning approaches used in reasoning systems can worsen calibration if they reward correct final answers without rewarding honest uncertainty. Their work found that models could become more capable and more overconfident at the same time. MIT News: Teaching AI models to say “I’m not sure”.
A broader evaluation across multiple models also found persistent overconfidence: predicted confidence often did not match actual correctness. Chhikara et al. — Confidence calibration in language models.
Calibration matters more than tone
Calibration asks whether a system’s stated confidence matches its real accuracy. A model that is right roughly 70% of the time when it says it is 70% confident is well calibrated. A model that sounds 95% certain while being right only half the time is not.
What to do with this
- Judge claims on source quality and verifiability, not tone.
- Ask what the model is uncertain about.
- Give it permission to say “I don’t know”.
- Ask the same factual question in a separate session and compare the specifics.
- Prefer systems that expose uncertainty over systems that always sound certain.
Rule to remember: certainty is a presentation style unless the evidence behind it is independently checkable.




