Scott Winters is suing OpenAI, alleging ChatGPT’s medical advice delayed treatment for a life-threatening pulmonary embolism. (Photo by Justin Sullivan/Getty Images)
Getty Images
Scott Winters, once a pastor in Florida, has filed a lawsuit against OpenAI and its CEO, Sam Altman. His claim centers on the potentially fatal medical advice given by ChatGPT. Filed in July 2026 in San Francisco County Superior Court, the lawsuit contends that Winters sought advice from ChatGPT-4o in 2025 regarding symptoms of dizziness and unstable blood pressure. The chatbot allegedly downplayed the significance of these symptoms, advising him to remain “recliner-bound” and suggesting that concern was only warranted after eight to ten episodes. Subsequently, Winters suffered a severe pulmonary embolism, which a doctor attributed to the extended immobility recommended by the chatbot.
On the day of the embolism, Winters sought advice from ChatGPT about groin tenderness and whether it necessitated an emergency room visit. The bot reportedly referenced his faith, saying “God did not design your body to endlessly fail.” Shortly afterward, he experienced a near-fatal incident. OpenAI has stated that ChatGPT is not intended to replace medical professionals and advised users against relying on it solely for medical advice, as outlined in its terms of service. Winters’ attorneys are pursuing financial compensation and a temporary halt to ChatGPT Health, OpenAI’s health-specific feature, until an independent safety review is conducted.
In another case, a couple from Texas also sued OpenAI in May after their son died from an overdose following advice from ChatGPT about drugs. They claimed the company ignored its own safety protocols, which could have prevented their son’s death.
These lawsuits are challenging the extent of accountability AI companies have when individuals use chatbots in medical or psychological emergencies. They also raise ongoing questions about the accuracy of AI in diagnostics.
Some Research Shows AI Is An Excellent Diagnostician
Research indicates that AI’s diagnostic capabilities can be impressive, particularly in controlled environments. A 2024 study in JAMA Internal Medicine compared GPT-4 with 21 attending physicians and 18 residents across 20 clinical cases using a clinical-reasoning scale called r-IDEA. GPT-4 scored a median of 10 out of 10, outperforming attendings who scored 9, and residents who scored 8. However, GPT-4 was incorrect more often than the human residents.
Another study in JAMA Network Open conducted later that year compared 50 physicians against six challenging cases. ChatGPT alone achieved 90% diagnostic accuracy, whereas physicians without AI assistance scored 74%, and those with AI assistance scored just 76%. The minimal improvement was largely due to doctors questioning or ignoring the chatbot’s advice.
Further research supports these findings. A study in Nature in 2025 tested Google’s AMIE model against 20 clinicians on 302 complex cases. AMIE identified the correct diagnosis 59% of the time compared to 34% for clinicians working independently. Clinicians who used AMIE produced better differential diagnoses than those relying on search engines and standard references.
A meta-analysis in npj Digital Medicine reviewed 50 studies on 25 AI models, finding that AI systems generally performed on par with or better than practicing clinicians in standardized diagnostic and triage tasks.
Many Studies That Raise Concerns About AI Diagnosis
However, most of this research is based on highly controlled conditions, featuring structured prompts and select cases. Results have not been consistently positive. A study in NEJM AI created a 750-question benchmark using script concordance testing to see how new clinical information affects diagnosis under uncertainty. It tested ten leading AI models against over 1,500 medical students, residents, and attending physicians. Even the best model, OpenAI’s o3, achieved only 68% accuracy, below senior residents and attendings, despite excelling in multiple-choice medical exams.
High test scores do not always equate to sound clinical judgment in uncertain conditions. The real-world use scenario, as highlighted by the Winters lawsuit, involves open-ended interactions, lacking physical exams or vital signs, and escalating reassurance rather than hospital visits.
Concerns about AI’s reliability are further amplified by a study in Communications Medicine that introduced fabricated details into vignettes fed to six popular chatbots, including GPT-4o and DeepSeek. These models often accepted and built upon the false information, with acceptance rates between 50% and 83%. Even when warned about potential inaccuracies, errors persisted.
A separate benchmark from Stanford tested 20 models and four clinical AI tools on 1,100 cases for potential harm from advice given. In 24.6% of cases, the advice could cause severe harm, with omissions being the primary issue.
Practicing physicians share these concerns. A 2025 survey by the physician network Sermo found that 94% of over 1,000 doctors were apprehensive about patients relying on AI for medical advice, particularly due to risks of misdiagnosis or delayed care.
Research suggests that in narrow and specific tasks, AI often matches or surpasses human physicians. However, the Winters case involved ongoing interaction over weeks without comprehensive information, which aligns more closely with conditions in studies examining AI hallucinations and reasoning flaws.
For healthcare systems and tech companies developing AI diagnostic tools, these findings suggest a design issue rather than a clear-cut judgment on AI’s capabilities. While AI can excel in structured tasks, its performance in real-world, unsupervised situations remains uncertain. This question continues to be explored by courts, hospitals, and researchers.

