Monday, 10 Aug 2026
  • Contact
  • Privacy Policy
  • Terms & Conditions
  • DMCA
logo logo
  • World
  • Politics
  • Crime
  • Economy
  • Tech & Science
  • Sports
  • Entertainment
  • More
    • Education
    • Celebrities
    • Culture and Arts
    • Environment
    • Health and Wellness
    • Lifestyle
  • πŸ”₯
  • Trump
  • House
  • White
  • ScienceAlert
  • VIDEO
  • man
  • Trumps
  • Season
  • star
  • Years
Font ResizerAa
American FocusAmerican Focus
Search
  • World
  • Politics
  • Crime
  • Economy
  • Tech & Science
  • Sports
  • Entertainment
  • More
    • Education
    • Celebrities
    • Culture and Arts
    • Environment
    • Health and Wellness
    • Lifestyle
Follow US
Β© 2024 americanfocus.online – All Rights Reserved.
American Focus > Blog > Tech and Science > OpenAI’s new reasoning AI models hallucinate more
Tech and Science

OpenAI’s new reasoning AI models hallucinate more

Last updated: April 18, 2025 11:30 pm
Share
OpenAI’s new reasoning AI models hallucinate more
SHARE

OpenAI’s Latest AI Models Still Struggle with Hallucinations

OpenAI recently introduced its o3 and o4-mini AI models, which are considered state-of-the-art in many aspects. However, these new models still face a significant challenge – they tend to hallucinate, or make up information, even more than some of OpenAI’s older models.

Hallucinations have long been a tough nut to crack in the field of AI, affecting even the most advanced systems available today. Traditionally, each new model has shown slight improvements in reducing hallucinations compared to its predecessor. However, this trend seems to have taken a step back with the o3 and o4-mini models.

According to OpenAI’s internal evaluations, the reasoning models o3 and o4-mini exhibit a higher rate of hallucinations compared to the company’s previous reasoning models like o1, o1-mini, and o3-mini, as well as the non-reasoning models such as GPT-4o.

One concerning aspect is that OpenAI is still uncertain about the root cause of this increased hallucination phenomenon. In their technical report for o3 and o4-mini, OpenAI states that further research is required to understand why these models are experiencing more hallucinations as they scale up reasoning capabilities. While these models excel in certain tasks related to coding and math, the increased number of claims they make leads to both accurate and inaccurate/hallucinated claims.

OpenAI’s findings reveal that o3 hallucinates in response to 33% of questions on PersonQA, which is used to gauge a model’s knowledge accuracy about people. This rate is double that of previous reasoning models like o1 and o3-mini. Surprisingly, o4-mini performs even worse on PersonQA, hallucinating 48% of the time.

See also  Analyst Says His AI Stock Models β€˜Don’t Like’ Palantir Technologies (PLTR); Valuation Unjustifiable

Third-party testing conducted by Transluce, a nonprofit AI research lab, also highlighted o3’s tendency to fabricate actions it supposedly took to arrive at answers. This behavior raises concerns about the model’s reliability and accuracy in real-world applications.

Experts like Neil Chowdhury and Sarah Schwettmann from Transluce suggest that the reinforcement learning techniques used in o-series models might be amplifying these issues, leading to an increased rate of hallucinations. While o3 shows promise in coding workflows, it still struggles with hallucinating broken website links, which could impact its usability.

Although hallucinations can sometimes lead to creative ideas, they pose a significant challenge for businesses that require high accuracy, such as law firms reviewing contracts. One potential solution to improve model accuracy is by incorporating web search capabilities, as demonstrated by OpenAI’s GPT-4o with web search achieving 90% accuracy on SimpleQA.

As the AI industry shifts towards reasoning models for better performance on various tasks, the issue of hallucinations remains a critical area of concern. OpenAI acknowledges the need to address hallucinations across all models and continues to focus on enhancing accuracy and reliability.

In conclusion, while reasoning models offer significant benefits, they also bring about new challenges such as increased hallucinations. Finding a balance between performance and accuracy will be crucial for the future development of AI models.

TAGGED:hallucinatemodelsOpenAIsreasoning
Share This Article
Twitter Email Copy Link Print
Previous Article Crash levels post at Oamaru crossing Crash levels post at Oamaru crossing
Next Article Beauty Marks: The Best Beauty Looks of The Week Beauty Marks: The Best Beauty Looks of The Week
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *


The reCAPTCHA verification period has expired. Please reload the page.

Popular Posts

Oprah Set to Land Huge TV Exclusive Thanks to Old Pal Meghan Markle

Meghan Markle Reaches Out to Brooklyn Beckham Amid Backlash A source close to the situation…

February 13, 2026

All About Savannah Guthrie’s Family Including Siblings Annie and Camron

Savannah Guthrie's mother, Nancy Guthrie, has always been a source of inspiration and support for…

February 27, 2026

Bravo Boss Frances Berwick on BravoCon’s Profits, RHOSLC, Karen Huger

BravoCon 2025: An Inside Look at the Ultimate Bravo Fan Experience As the curtains close…

November 20, 2025

Style, Power, and Modern Glamour

Henson stuns in a black-and-white embellished gown by Sebastian Gunawan at the Ebony Power 100…

November 10, 2025

Indiana mom busted after trying to sell baby daughter for sex

Indiana Mom Accused of Attempted Child Sex Trafficking A shocking case has emerged in Indiana,…

July 19, 2025

You Might Also Like

The Shai-Hulud npm worm didn't fake its security check β€” it earned a legitimate one
Tech and Science

The Shai-Hulud npm worm didn't fake its security check β€” it earned a legitimate one

August 10, 2026
Just 30 Minutes of One Type of Exercise Improves Sleep The Most : ScienceAlert
Tech and Science

Just 30 Minutes of One Type of Exercise Improves Sleep The Most : ScienceAlert

August 10, 2026
Dune Colour, Pixel Watch 5 & Celebrity Hosts – Tech Advisor
Tech and Science

Dune Colour, Pixel Watch 5 & Celebrity Hosts – Tech Advisor

August 10, 2026
Europe’s wildfires are igniting forgotten WWII and Cold War bombs
Tech and Science

Europe’s wildfires are igniting forgotten WWII and Cold War bombs

August 9, 2026
logo logo
Facebook Twitter Youtube

About US


Explore global affairs, political insights, and linguistic origins. Stay informed with our comprehensive coverage of world news, politics, and Lifestyle.

Top Categories
  • Crime
  • Environment
  • Sports
  • Tech and Science
Usefull Links
  • Contact
  • Privacy Policy
  • Terms & Conditions
  • DMCA

Β© 2024 americanfocus.online –Β  All Rights Reserved.

Welcome Back!

Sign in to your account

Lost your password?