Wednesday, July 29, 2026

Patients advised against AI chatbots for medical self-diagnosis amid significant flaws identified

July 27, 2026
1 min read
Patients advised against AI chatbots for medical self-diagnosis amid significant flaws identified

Patients have been warned against turning to artificial intelligence for at-home medical diagnosis after some AI systems failed to recognise crucial symptoms, reports BritPanorama.

Research from Carnegie Mellon University’s School of Computer Science in the US has revealed that large language models, including GPT-5, Gemini, and Claude, were tested on their ability to respond to basic medical questions aimed at diagnosing illnesses. The analysis highlighted that these models sometimes generated false information, failing to adhere to criteria typically used by healthcare professionals when evaluating patients.

In the study led by Siddharth Vohra, a master’s student at the robotics institute, the AI models were asked to describe a medical image intentionally left out of the query. Instead of prompting for the missing image, the models fabricated a diagnosis based on provided demographic information in 18% of the cases, raising significant concerns regarding the reliability of such AI systems.

Misleading

The findings indicated that nearly one in five diagnoses offered by AI could be inaccurate or misleading. Furthermore, the computer programs made assumptions, such as suggesting that certain patients had cancer without adequate evidence to support this diagnosis. “A 65-year-old white man asking Claude about a skin mole receives melanoma in nearly every response,” Vohra noted, emphasizing the risk of misleading outputs. Similarly, requests involving chest X-ray questions led GPT-5 to name sarcoidosis for approximately 77% of young Black patients, despite the rarity of the condition.

Experts estimate that fewer than one in 10,000 moles develop into cancer and that sarcoidosis is an uncommon inflammatory disease, occurring in just 10 to 60 individuals per 100,000. Vohra remarked that diagnosis outcomes shifted significantly based simply on the inclusion of race, age, or gender in the prompt used.

The research underscores critical flaws inherent in self-diagnosis methods currently proliferating in the digital age. By soliciting medical information through online sources, patients risk disregarding professional medical advice and instead rely on potentially flawed AI-generated assessments.

AI responses may appear convincing; however, small alterations in the phrasing can drastically alter conclusions, illuminating risks associated with such tools. The study concentrated on various diagnostic imaging techniques, including chest X-rays and dermatology images. Across Claude, GPT-5, and Google’s Gemini, the AI models yielded a staggering 11,700 responses, highlighting a pervasive assumption among users that the models possess a deeper understanding than they genuinely do.

“About 82% of the time, the models refused to provide a response because no image was attached,” Vohra explained. “However, in the remaining 18%, the models invented a diagnosis instead of asking for the missing image.” This raises pressing questions about the perceived intelligence of AI applications; proficiency in select tasks does not equate to a comprehensive understanding of medical complexities.

Convincing diagnosis

“Large language models reached hospitals before the safeguards did,” stated Ahmed Ashoor, chief technology officer for Outworks, a UAE technology service provider. “A model produces a fluent answer whether or not the answer is true, and in medicine a confident wrong answer does more damage than no answer at all.” Furthermore, Ashoor noted that most models are trained primarily on Western populations, leading to declining accuracy for patients from different backgrounds.

Addressing data privacy is crucial for enhancing the effectiveness of LLMs in healthcare. “Health records rank among the most sensitive data a nation holds, and no ministry can hand them to systems outside its control,” Ashoor highlighted. In addition, accountability poses a significant challenge; when an algorithm influences a clinical decision, it is essential to determine who is responsible for the outcome.

Hospitals

AI is showing notable effectiveness in the realm of radiology, with applications that flag abnormalities in scans with accuracy comparable to specialist assessments, allowing for efficient second reviews without additional time burdens on radiologists. Hospital systems leveraging AI to forecast admissions and optimise bed capacity also contribute to better patient outcomes.

Nalla Karunanithy, chief executive of Digital Health and Omnichannel at Aster DM Healthcare, shared that their hospitals utilise AI algorithms to help radiologists prioritise urgent cases more effectively. He acknowledged, however, that chatbots are not currently integrated into hospital operations but can assist in summarising medical literature.

Karunanithy asserted, “Where large language models genuinely add value today is around the clinical encounter rather than inside the diagnostic decision itself.” General models are at risk of perpetuating biases present in their training data, leading to incorrect conclusions presented with misplaced confidence. “In a clinical setting, that confidence without accountability is not something that is encouraged at this point,” he added, underscoring the importance of physician oversight in any AI-related diagnostic process.

Leave a Reply

Your email address will not be published.

Don't Miss

Shares of China’s CXMT rise 462% in largest recent IPO amid semiconductor demand surge

Shares of China’s CXMT rise 462% in largest recent IPO amid semiconductor demand surge

Shares of China’s largest memory chipmaker soar in IPO Shares of CXMT,
Russia deploys AI-generated videos of fake Polish soldiers in disinformation campaign

Russia deploys AI-generated videos of fake Polish soldiers in disinformation campaign

A Russian hybrid warfare operation targeting Poland has been uncovered, using artificial