How doctors are using AI

Doctors’ use of artificial intelligence has increased rapidly in the last few years.  According to a 2026 American Medical Association survey: “Over 80% of physician respondents currently use AI in a professional context – double the share reported in 2023.”  That percentage is expected to continue to grow.  A similar survey of 3151 US physicians by Doximity  reported that “94% of physicians surveyed are currently using AI or are interested in doing so.”

Physicians’ current uses of AI can be classified in three main categories: administration, diagnosis, and treatment.  When Doximity asked: “What excites you most about the potential of AI in your Practice?” the top answers were:

  1. Less administrative workload (69%)
  2. Better work-life balance (67%)
  3. Greater job satisfaction (50%)
  4. More time with patients (43%)
  5. Better patient care, outcomes (41%)

All five of these are related to the time doctors spend filling out forms and related administrative tasks.  A decade ago, a classic survey of 4,720 doctors “found the average doctor spends 8.7 hours per week, or about 16.6% of working hours, on administrative tasks, and that this burden was directly associated with lower career satisfaction.”  That percentage has been reduced slightly since then, but the amount of time doctors must spend on administrative tasks still remains a significant problem that affects not just doctors but also their patients. 

When the Doximity survey asked “Do you believe that AI can help increase time for patient care by reducing administrative workload?” over 90% of doctors said yes, including 27% who said it already has.

How?  One leading approach is the use of “ambient AI scribes,” programs which record doctor-patient conversations and generate medical notes, minus any chit chat.  Others include programs that can reduce paperwork by analyzing existing medical records to generate answers to patient portal questions, generate insurance codes, treatment orders, bills, and more.

There is a fair amount of controversy over exactly how much administrative time these tools saves the average doctor today.  Vendors often claim much higher time savings than peer-reviewed research articles.  But systematic academic research has shown that even when the time savings are modest, the psychological effects on self-reported burnout and exhaustion scores are both real and significant.  For example, when six health care systems introduced ambient scribes, they found that within 30 days “burnout among those working in ambulatory clinics decreased significantly from 51.9% to 38.8%. There were also significant improvements in the cognitive task load, time spent documenting after hours, focused attention on patients, and urgent access to care.”

From a patient perspective, this is more significant than it might sound.  Physician burnout roughly doubles the odds of safety incidents and of patient dissatisfaction, according to a systematic review of 170 studies involving 239,246 physicians.

No matter how significant this may be, the more interesting question is: can AI directly improve medical diagnosis and treatment?  Thousands of studies have been published on this question, leading to a number of breakthroughs and a fair amount of controversy.  For example, consider the contradictory titles from two studies published in April 2026: “AI Is Starting to Beat Doctors at Making Correct Diagnoses” and “AI Fails at Primary Patient Diagnosis More Than 80% of the Time, Study Finds.”  Why the discrepancy?  Because the real answer is it depends.  It depends on the medical condition, the exact details of how each study was conducted, the way success was defined and much more.

This type of conclusion is familiar to scientists and is one reason replication of results is such a critical part of the scientific method.  It’s also the reason that the best way to determine the status of this type of research question is a statistical technique called meta-analysis, which compares and combines the results of all available independent studies. 

When it comes to medical diagnosis using LLMs (the large language models behind ChatGPT and related products; see Part 3), the most definitive meta-analysis to date was published in 2025 in the Journal of Medical Internet Research (JMIR).  This extensive review of several databases in English and Chinese identified 30 studies that met their demanding criteria.  When combined, these studies represented a total of 4,762 medical cases, primarily in internal medicine, radiology and ophthalmology.  24 of the 30 studies used versions of ChatGPT, and six used other LLMs.

They concluded that the best summary of research to date is that “clinical professionals generally outperformed LLMs in diagnostic accuracy… [but] each medical field shows different ways and effects of LLMs’ application.” 

When it comes to using AI in medical treatment, some of the most promising research to date has been performed at the Mayo Clinic.  For example, one set of studies focused on sepsis, which is one of the leading causes of hospital deaths, killing more than 250,000 Americans every year.  Sepsis occurs when the immune system overreacts to an infection, leading to organ failure.  One problem with treating sepsis is that traditional methods of identifying the problem often detect the condition hours after optimal treatment windows have passed. 

Mayo clinic researchers developed an AI machine learning system (see Part 2) which combines “real-time analysis of over 100 clinical variables per patient… [with] machine learning models trained on millions of patient encounters [and]… continuous learning algorithms [which improve] accuracy based on patient outcomes.”  Ultimately this tool provided an “average 3-hour advance warning before traditional clinical identification.”  This enabled earlier intervention which in turn led to a “sepsis mortality reduction of 18%.”  In addition, after patients were released from the hospital “wireless sensor technology [provided] 24/7 patient monitoring… which enabled a 40% reduction in hospital readmissions.”

In another study, OpenAI and Penda (a health care provider in Kenya) built a tool called AI Consult, “to provide clinicians with LLM-written recommendations at key points during a patient visit. AI Consult acts as a real-time safety net that activates only when there might be an error, keeping clinicians fully in control.”

The result: “In a study of 39,849 patient visits across 15 clinics, clinicians with AI Consult had a 16% relative reduction in diagnostic errors and a 13% reduction in treatment errors compared to those without.”  Programs like this could have a major impact on “expanding access to safe, high-quality care” in developing countries.

While studies like the ones quoted above have led to great optimism, AI is still a new technology, and many doctors are concerned with problems it may raise.  In the Doximity survey “Accuracy and reliability of AI outputs… was the top concern… [for] 71% of surveyed physicians.”  In addition, “47% reported the AI decision‑making process at their institution is ‘still evolving.’”

Despite these challenges, when Elsevier surveyed 2,757 doctors and nurses in 118 countries for its Clinician of the Future 2026 report, they found that “80% said AI will become a critical assistant within the next decade.”  To maximize its benefits, 92% [of doctors] said they want more education and training on AI (according to the latest AMA survey).   

Interestingly, the AMA named that report the “2026 Physician Survey on Augmented Intelligence.” Not artificial intelligence, but augmented intelligence.  According to Gartner,  “Augmented intelligence is all about people taking advantage of AI. As AI technology evolves, the combined human and AI capabilities that augmented intelligence allows will deliver the greatest benefits.  The AMA House of Delegates prefers the word augmented to artificial because it focuses “on AI’s assistive role, emphasizing that its design enhances human intelligence rather than replaces it.”

One of the best summaries of AI’s proven capabilities to date appeared in the JMIR meta-analysis described above: “While LLMs still have a long way to go in accurately diagnosing real-world clinical scenarios… they undeniably possess significant potential as health care assistants… With… advancements… in technology, it is anticipated that LLMs will play an increasingly important role in future clinical diagnostics.” The same could be said of other AI products which are based on expert systems and other non-LLM approaches.