Educators, AI and cheating

A few months ago, a New York state Supreme Court judge did something unusual he overturned a university’s academic misconduct finding and ordered the school to expunge a student’s record entirely. The student, Orion Newby, had been accused of using AI to write a paper at Adelphi University, based on an AI-detection app his professor ran. He’d explained that he’d worked with tutors in the school’s program for students with learning and neurological disorders. The professor didn’t buy it. Newby was given a zero and, after a second flagged assignment, faced expulsion. “I felt shocked,” he said. “I felt like that was it. I felt like my life was over.” His parents sued, and the court ultimately found the AI accusation was unfounded. His attorney, Mark Lesko, called the ruling “groundbreaking,” adding: “Higher education needs to take a very careful look at this.”

Newby’s case wasn’t isolated — a Yale School of Management student has sued the university over a similar accusation, and the Palo Alto Unified school district is currently fighting a $150 million lawsuit over an AI allegation made against a high schooler’s essay. Cases like these are a symptom of a much bigger problem: three years after ChatGPT arrived in classrooms, educators still don’t have reliable tools to answer the most basic question in any assignment — did the student actually do the work?

The scale of the underlying issue is well documented. According to a Pew Research Center survey of 1,458 U.S. teens conducted in late 2025, 54% have used a chatbot for help with schoolwork, and 10% say they do all or most of their schoolwork this way. An even larger study — the biggest of its kind, surveying more than 95,000 undergraduates across 20 research universities and published in Science in May 2026 — found that about two-thirds of students had used generative AI, and nearly 40% used it monthly or more. That study also found real disparities in who is using it: low-income, underrepresented, and female students reported using AI noticeably less than their peers, and students in the humanities and social sciences were more likely to use it to cheat than students in STEM fields, where a problem set with one correct answer is harder to fake convincingly than an essay.

Using AI for schoolwork isn’t automatically the same as cheating with it, and the Science study drew that distinction carefully. Still, it found that “at least 9% of AI-using students reported cheating with it” — and the more often a student used AI, the more likely they were to cross that line: 26% of daily users admitted to cheating with it, compared to just 7% of monthly users. The study’s lead researcher, Igor Chirikov, put it starkly: “The arrival of artificial intelligence technologies and GenAI tools like ChatGPT was a big shock to higher education.” But he added a caution that surprised university administrators hoping for a quick fix: “banning GenAI won’t stop cheating and may even harm students.”

Perception may be an even bigger problem than the raw numbers. In the Pew survey, 59% of teens said students at their school use AI to cheat “at least sometimes,” and a third said it happens “extremely or very often.” Among teens who’d used chatbots for schoolwork themselves, that figure jumped to 76% — meaning three out of four students who use AI for homework believe cheating with it is now the norm among their peers, whether or not they’re doing it themselves. That’s the kind of number that can poison a classroom’s culture regardless of how many students are technically breaking any rule.

Teachers see it too, and it’s reshaping how they teach. A nationally representative NPR/Ipsos poll of 545 K-12 teachers found:

  1. 73% believe AI’s impact on education will be bigger than that of the internet or computers
  2. 60% say AI is eroding trust between students and teachers
  3. 55% call AI “mostly a shortcut to avoid doing work”
  4. 54% say AI makes it harder for students to develop critical thinking skills
  5. 80% say schools should be actively teaching students to use AI responsibly

One teacher summed up the underlying worry bluntly: “I want them to figure out things for themselves, not rely on software. If we stop questioning what it says, we can be led to believe anything.” That erosion of trust has driven a partial retreat to older methods: 40% of teachers now require more handwritten assignments, and the same share require more work be done in class, where AI is harder to use undetected. I spent a decade teaching at Boston University in the 1970s, back when the biggest threat to academic honesty was a roommate’s old term paper; grading a blue-book exam again in 2026 would have felt like science fiction back then, and now it’s being marketed as a solution.

The obvious fix, AI-detection software, has turned out to be far less reliable than schools hoped, and biased in ways that hit some students much harder than others. Turnitin, the most widely used plagiarism-detection company in education, initially advertised a false-positive rate below 1%. That claim didn’t hold up; the company later disclosed that its sentence-level detection carries roughly a 4% error rate. Four percent might sound tolerable, until you multiply it across a university of 20,000 students turning in papers every week. Soheil Feizi, a University of Maryland researcher who has studied these tools extensively, concluded flatly that “no publicly available AI detectors are sufficiently reliable in practical scenarios.” He argues detectors would need a false-positive rate closer to 0.01% to be fair to use in disciplinary decisions — and, he says, “at this point, it’s impossible.”

The bias problem is worse still for one group in particular. A Stanford study of seven widely used AI detectors found they incorrectly flagged 61% of TOEFL essays written by non-native English speakers as AI-generated — 97% of those essays were flagged by at least one detector, and 19% were unanimously condemned by all seven — while the same tools performed “near-perfect” on essays from U.S.-born students. The detectors rely on a measure called “perplexity,” which rewards sophisticated, idiomatic phrasing, exactly what a second-language writer is less likely to produce, AI or no AI. That bias isn’t hypothetical: the Yale student mentioned above, suing under the pseudonym John Doe after being suspended over a GPTZero flag on his final exam, argues in his complaint that “the AI program is unreliable and contains implicit bias,” pointing specifically to his status as a non-native English speaker.

So where does that leave educators? Caught, mostly, between two problems that don’t cancel each other out: real, measurable cheating on one side, and unreliable, biased tools for catching it on the other. Half of teachers in the NPR poll say their school has offered no formal guidance on AI at all, and among the schools that have introduced AI software, only 35% have a formal policy governing how teachers should use it. Just 40% of teachers say they’ve received any professional development on the technology now reshaping their classrooms — roughly the same share of physicians and lawyers who, in surveys I’ve cited here before, said they wanted more training in AI before trusting it with real decisions.

Not every country is handling this the same fumbling way. During the 2025 gaokao, China’s high-stakes college entrance exam taken by roughly 13 million students, the country’s major AI companies — Alibaba, ByteDance, Tencent, and Moonshot — quietly disabled the photo-recognition features on their own chatbots for the duration of the test, so a student couldn’t snap a picture of a question and get an answer back. Students who tried were simply told “such action is not in compliance with the rules.” No lawsuits, no disciplinary hearings, no debate about detector bias — just a centrally coordinated flip of a switch, with exam proctors separately using their own AI surveillance to watch for cheating. It’s a blunt, top-down fix that would be unthinkable in a US system built around due process and decentralized school boards, where ordering four private companies to alter their products overnight is a legal fantasy. Whether that makes the American approach more principled or just slower and costlier at arriving somewhere similar is a fair question, and not one with an obvious answer.

Here’s my honest, tentative prediction, offered with the caveat I’d apply to any forecast about a fast-moving technology: detection is probably a losing game, and schools that keep betting on it will keep generating cases like Orion Newby’s. The 80% of teachers who want schools to actively teach responsible AI use are effectively voting for Chirikov’s approach over Turnitin’s — not banning the technology or policing it with flawed software, but building the judgment to use it well, the same judgment students will need in whatever profession they enter. Whether that shift happens fast enough to spare the next round of falsely accused students is the part nobody can honestly answer yet. Three years after ChatGPT walked into the classroom uninvited, most US schools are still improvising a response, and if the last three years are any guide, the next three will bring at least one more surprise nobody saw coming.

Editorial note: In case you don’t realize how hard it is for teachers to recognize AI-created  homework, this post was 100% generated by AI.  Could you tell? 

This is the prompt I used with Claude AI: “Draft a 1500 word blog post on how educators are being challenged by AI, especially the issue of students cheating. Write it in the same general style as previous posts in my blog aiandprofessionals.com.  Also see chinain5.com for more examples of my writing style.”  I added a few links for accuracy, but otherwise did not change a single word.

How doctors are using AI

Doctors’ use of artificial intelligence has increased rapidly in the last few years.  According to a 2026 American Medical Association survey: “Over 80% of physician respondents currently use AI in a professional context – double the share reported in 2023.”  That percentage is expected to continue to grow.  A similar survey of 3151 US physicians by Doximity  reported that “94% of physicians surveyed are currently using AI or are interested in doing so.”

Physicians’ current uses of AI can be classified in three main categories: administration, diagnosis, and treatment.  When Doximity asked: “What excites you most about the potential of AI in your Practice?” the top answers were:

  1. Less administrative workload (69%)
  2. Better work-life balance (67%)
  3. Greater job satisfaction (50%)
  4. More time with patients (43%)
  5. Better patient care, outcomes (41%)

All five of these are related to the time doctors spend filling out forms and related administrative tasks.  A decade ago, a classic survey of 4,720 doctors “found the average doctor spends 8.7 hours per week, or about 16.6% of working hours, on administrative tasks, and that this burden was directly associated with lower career satisfaction.”  That percentage has been reduced slightly since then, but the amount of time doctors must spend on administrative tasks still remains a significant problem that affects not just doctors but also their patients. 

When the Doximity survey asked “Do you believe that AI can help increase time for patient care by reducing administrative workload?” over 90% of doctors said yes, including 27% who said it already has.

How?  One leading approach is the use of “ambient AI scribes,” programs which record doctor-patient conversations and generate medical notes, minus any chit chat.  Others include programs that can reduce paperwork by analyzing existing medical records to generate answers to patient portal questions, generate insurance codes, treatment orders, bills, and more.

There is a fair amount of controversy over exactly how much administrative time these tools saves the average doctor today.  Vendors often claim much higher time savings than peer-reviewed research articles.  But systematic academic research has shown that even when the time savings are modest, the psychological effects on self-reported burnout and exhaustion scores are both real and significant.  For example, when six health care systems introduced ambient scribes, they found that within 30 days “burnout among those working in ambulatory clinics decreased significantly from 51.9% to 38.8%. There were also significant improvements in the cognitive task load, time spent documenting after hours, focused attention on patients, and urgent access to care.”

From a patient perspective, this is more significant than it might sound.  Physician burnout roughly doubles the odds of safety incidents and of patient dissatisfaction, according to a systematic review of 170 studies involving 239,246 physicians.

No matter how significant this may be, the more interesting question is: can AI directly improve medical diagnosis and treatment?  Thousands of studies have been published on this question, leading to a number of breakthroughs and a fair amount of controversy.  For example, consider the contradictory titles from two studies published in April 2026: “AI Is Starting to Beat Doctors at Making Correct Diagnoses” and “AI Fails at Primary Patient Diagnosis More Than 80% of the Time, Study Finds.”  Why the discrepancy?  Because the real answer is it depends.  It depends on the medical condition, the exact details of how each study was conducted, the way success was defined and much more.

This type of conclusion is familiar to scientists and is one reason replication of results is such a critical part of the scientific method.  It’s also the reason that the best way to determine the status of this type of research question is a statistical technique called meta-analysis, which compares and combines the results of all available independent studies. 

When it comes to medical diagnosis using LLMs (the large language models behind ChatGPT and related products; see Part 3), the most definitive meta-analysis to date was published in 2025 in the Journal of Medical Internet Research (JMIR).  This extensive review of several databases in English and Chinese identified 30 studies that met their demanding criteria.  When combined, these studies represented a total of 4,762 medical cases, primarily in internal medicine, radiology and ophthalmology.  24 of the 30 studies used versions of ChatGPT, and six used other LLMs.

They concluded that the best summary of research to date is that “clinical professionals generally outperformed LLMs in diagnostic accuracy… [but] each medical field shows different ways and effects of LLMs’ application.” 

When it comes to using AI in medical treatment, some of the most promising research to date has been performed at the Mayo Clinic.  For example, one set of studies focused on sepsis, which is one of the leading causes of hospital deaths, killing more than 250,000 Americans every year.  Sepsis occurs when the immune system overreacts to an infection, leading to organ failure.  One problem with treating sepsis is that traditional methods of identifying the problem often detect the condition hours after optimal treatment windows have passed. 

Mayo clinic researchers developed an AI machine learning system (see Part 2) which combines “real-time analysis of over 100 clinical variables per patient… [with] machine learning models trained on millions of patient encounters [and]… continuous learning algorithms [which improve] accuracy based on patient outcomes.”  Ultimately this tool provided an “average 3-hour advance warning before traditional clinical identification.”  This enabled earlier intervention which in turn led to a “sepsis mortality reduction of 18%.”  In addition, after patients were released from the hospital “wireless sensor technology [provided] 24/7 patient monitoring… which enabled a 40% reduction in hospital readmissions.”

In another study, OpenAI and Penda (a health care provider in Kenya) built a tool called AI Consult, “to provide clinicians with LLM-written recommendations at key points during a patient visit. AI Consult acts as a real-time safety net that activates only when there might be an error, keeping clinicians fully in control.”

The result: “In a study of 39,849 patient visits across 15 clinics, clinicians with AI Consult had a 16% relative reduction in diagnostic errors and a 13% reduction in treatment errors compared to those without.”  Programs like this could have a major impact on “expanding access to safe, high-quality care” in developing countries.

While studies like the ones quoted above have led to great optimism, AI is still a new technology, and many doctors are concerned with problems it may raise.  In the Doximity survey “Accuracy and reliability of AI outputs… was the top concern… [for] 71% of surveyed physicians.”  In addition, “47% reported the AI decision‑making process at their institution is ‘still evolving.’”

Despite these challenges, when Elsevier surveyed 2,757 doctors and nurses in 118 countries for its Clinician of the Future 2026 report, they found that “80% said AI will become a critical assistant within the next decade.”  To maximize its benefits, 92% [of doctors] said they want more education and training on AI (according to the latest AMA survey).   

Interestingly, the AMA named that report the “2026 Physician Survey on Augmented Intelligence.” Not artificial intelligence, but augmented intelligence.  According to Gartner,  “Augmented intelligence is all about people taking advantage of AI. As AI technology evolves, the combined human and AI capabilities that augmented intelligence allows will deliver the greatest benefits.  The AMA House of Delegates prefers the word augmented to artificial because it focuses “on AI’s assistive role, emphasizing that its design enhances human intelligence rather than replaces it.”

One of the best summaries of AI’s proven capabilities to date appeared in the JMIR meta-analysis described above: “While LLMs still have a long way to go in accurately diagnosing real-world clinical scenarios… they undeniably possess significant potential as health care assistants… With… advancements… in technology, it is anticipated that LLMs will play an increasingly important role in future clinical diagnostics.” The same could be said of other AI products which are based on expert systems and other non-LLM approaches. 

How lawyers are using AI

Lawyers are not known for embracing change.  As an American Bar Association article put it “Attorneys tend to be risk-averse with a tendency toward perfectionism, which makes tech tools look less desirable.  There’s also a strong emphasis on tradition and precedent in the legal industry which naturally makes it harder to embrace new trends.”

But all that began to change in 2022 when OpenAI released ChatGPT 3.5 (see Part 3), and lawyers saw its potential to help in their work.  Since then, AI use has grown rapidly, as shown in four recent surveys.   

According to Clio’s latest Legal Trends Report, a survey of 1702 US legal professionals found that “79% use artificial intelligence in their firms.”  Another survey of 1,300 legal professionals conducted by the company 8AM found that, “Nearly three-quarters of legal professionals (69%) now report personally using general-purpose AI tools such as ChatGPT, Gemini, and Claude for work-related purposes… For a profession historically cautious about new technology, that rate of adoption represents a dramatic increase from the… 31% we found in [our] 2025 report.”  Similarly a third survey from Law 360 found that “Seventy percent of attorneys at law firms report using artificial intelligence at least once a week as part of their jobs.”  Finally, in the fourth survey, Wolters Kluwer  interviewed 810 lawyers in 11 countries and found “Over 90% of legal professionals now use at least one AI tool in their daily work.”

Summing up the current situation, the latest report from the American Bar Association’s Task Force on Law and Artificial Intelligence concluded that “as the transformative power of the technology has become more widely known, the conversation has shifted from whether to use the AI technology to how to use it… Early adoption has been limited largely to low-risk, routine tasks where the benefits are clear and the risks are manageable [such as] “summarizing, extracting insights from unstructured data, drafting simple communications like emails or short memos, and drafting client alerts.” 

Most law firms had actually been using AI years before chatbots emerged, especially in the areas of legal research and e-discovery.  In 1973, the company now named LexisNexis introduced the first computerized product to search the legal literature for precedents and more. Competitors soon followed, all built around traditional rule-based computer programming. LexisNexis added AI features in 2017, and maintains its dominant position in this market with over “five million [users]… in over 175 countries.”

E-discovery refers to the process of efficiently analyzing computer emails, texts and more to use as evidence in lawsuits.  It too started with traditional rule-based computer programming and later began adding AI features.

But when people talk about today’s AI legal revolution, they usually are not talking about the features of widely accepted tools for research and e-discovery.  They are usually talking about general purpose tools like ChatGPT, Gemini and Claude that use machine learning (see Part 2) to create new content.  The most advanced firms are often talking about legal-specific AI tools that have been customized for a particular legal purpose and/or for a single firm.  Legal-specific AI has many advantages including the ability to host products on a law firm’s internal servers (keeping all private information off the internet) and the ability to load all of a firm’s documents into a firm-specific training database.  The power and sophistication possible with customized products like these dwarfs anything a lawyer could do with ChatGPT.

Some lawyers have jumped into these sophisticated AI programs with both feet.  For one example, see the X post by Neil Katyal, a partner at Milbank and a Professor of Law at Georgetown.  His post describes how he used legal-specific AI to help win a 2025 Supreme Court case which “Legal scholars… and some of my own colleagues [had] said was impossible [to win].”  The product he used was developed by the AI company Harvey (named after Harvey Specter, a fictional partner in the TV hit Suits.).  The program was “trained on every question every [Supreme Court] Justice has asked in oral argument for 25 years, and everything they’ve ever written… Harvey predicted many of the questions the Justices asked — sometimes almost word for word” and that helped Katyal to win the case. (For a more detailed account, see Katyal’s 18 minute TED talk “What really won the trillion-dollar Supreme Court case?”)

Why is the use of AI now spreading so rapidly among lawyers?  The answer is simple: it saves time.  According to Clio’s latest Legal Trends Report “86% of heavy users say the technology has eased their work.”

But it can take considerable time and money to reap these benefits, including redesigning workflow, and increasing training and change management.  There have also been several widely publicized cases in which lawyers used AI in their research, and submitted “hallucinations” of legal precedents.  In one recent example, “Lawyers for Sullivan & Cromwell… one of the oldest and most prestigious law firms in the country … apologized for submitting a court filing that had fake citations created by artificial intelligence.”  This despite the fact that “Sullivan & Cromwell requires its lawyers to take a training course before gaining access to AI tools… to ‘trust nothing and verify everything.’”  In another, “A federal judge in Mississippi has punished all four lawyers on opposing sides in a civil trial and canceled the proceedings after some of them, relying on artificial intelligence, cited fake legal cases in court filings.”

Problems like this have been rare to date but very costly in terms of public perceptions. As the old cliché says, you won’t get a second chance to make a first impression. 

As a result, an American Bar Association article entitled Understanding the Legal AI Landscape offers this advice:  when firms first get serious about the topic, they “must navigate a complex and evolving AI landscape to identify the right tools for their needs.”  The key to success, the article continues is to “start with AI tools that address specific challenges, and expand gradually to maximize efficiency while mitigating risks like bias, hallucinations, and data privacy concerns.” 

And then there’s the elephant in the room: if lawyers who bill by the hour get more efficient, their revenue will go down. 

When AI saves time for paralegals, associates, and partners, what are they supposed to do with it?  Companies that sell legal AI products sometimes paint an optimistic picture in which no one will be laid off and this newly available time can be used for marketing and improved service.  Other experts suggest that greater efficiency can produce greater profits if firms switch from hourly billing to fixed fees and value-based pricing.  Still others believe that, as consultant James Markham put it, “there’s a misplaced optimism in fixed fees and value based pricing as the silver bullet… a likely course of events is that prices come down as clients can take their pick from an increasing number of AI-ified firms, as well as using the same tools for themselves.”

On a broader scale, as a Thomson Reuters report put it, “AI may cause momentous, industry-wide shifts. Many professionals say their organizations are still struggling to determine the return on investment (ROI) of AI tools… only 18% of respondents say their organizations collect metrics around ROI from AI. Of those, most metrics are internally focused, involving such areas as cost savings or employee usage, rather than business-focused metrics such as client satisfaction or amount of business generated.”

According to many experts, the next step in the legal AI revolution will be an increasing reliance on AI agents, which “perform autonomous tasks on behalf of the user or another system… [They are] focused on decisions as opposed to creating… content [and they don’t] solely rely on human prompts nor require human oversight.”   But that’s another story, which will be covered in a future post.

With all these rapidly moving developments, should financially conservative law firms wait for the dust to settle before they invest in AI?  Absolutely not.

Casey Flaherty, co-founder of LexFusion put it this way:  “Failure to start down the path of using GenAI is a terrible plan. [Note: Generative AI or Gen AI refers to products like ChatGPT that create new content.] But… GenAI is not magic… it is a path, not a teleportation device. There is a staggering amount of work to be done to realize GenAI’s potential, especially in an enterprise environment.”

Despite the barriers, AI is becoming a vital part of legal work.  Lawyers who expect to survive in an ever more competitive market don’t have much choice.  As Paul Saunders, the Chief Strategy and Innovation Officer at Canadian law firm Stewart McKelvey summed it up: “AI will not replace lawyers, but lawyers that use AI will replace those that don’t”

For background see the section on how AI works