A few months ago, a New York state Supreme Court judge did something unusual he overturned a university’s academic misconduct finding and ordered the school to expunge a student’s record entirely. The student, Orion Newby, had been accused of using AI to write a paper at Adelphi University, based on an AI-detection app his professor ran. He’d explained that he’d worked with tutors in the school’s program for students with learning and neurological disorders. The professor didn’t buy it. Newby was given a zero and, after a second flagged assignment, faced expulsion. “I felt shocked,” he said. “I felt like that was it. I felt like my life was over.” His parents sued, and the court ultimately found the AI accusation was unfounded. His attorney, Mark Lesko, called the ruling “groundbreaking,” adding: “Higher education needs to take a very careful look at this.”
Newby’s case wasn’t isolated — a Yale School of Management student has sued the university over a similar accusation, and the Palo Alto Unified school district is currently fighting a $150 million lawsuit over an AI allegation made against a high schooler’s essay. Cases like these are a symptom of a much bigger problem: three years after ChatGPT arrived in classrooms, educators still don’t have reliable tools to answer the most basic question in any assignment — did the student actually do the work?
The scale of the underlying issue is well documented. According to a Pew Research Center survey of 1,458 U.S. teens conducted in late 2025, 54% have used a chatbot for help with schoolwork, and 10% say they do all or most of their schoolwork this way. An even larger study — the biggest of its kind, surveying more than 95,000 undergraduates across 20 research universities and published in Science in May 2026 — found that about two-thirds of students had used generative AI, and nearly 40% used it monthly or more. That study also found real disparities in who is using it: low-income, underrepresented, and female students reported using AI noticeably less than their peers, and students in the humanities and social sciences were more likely to use it to cheat than students in STEM fields, where a problem set with one correct answer is harder to fake convincingly than an essay.
Using AI for schoolwork isn’t automatically the same as cheating with it, and the Science study drew that distinction carefully. Still, it found that “at least 9% of AI-using students reported cheating with it” — and the more often a student used AI, the more likely they were to cross that line: 26% of daily users admitted to cheating with it, compared to just 7% of monthly users. The study’s lead researcher, Igor Chirikov, put it starkly: “The arrival of artificial intelligence technologies and GenAI tools like ChatGPT was a big shock to higher education.” But he added a caution that surprised university administrators hoping for a quick fix: “banning GenAI won’t stop cheating and may even harm students.”
Perception may be an even bigger problem than the raw numbers. In the Pew survey, 59% of teens said students at their school use AI to cheat “at least sometimes,” and a third said it happens “extremely or very often.” Among teens who’d used chatbots for schoolwork themselves, that figure jumped to 76% — meaning three out of four students who use AI for homework believe cheating with it is now the norm among their peers, whether or not they’re doing it themselves. That’s the kind of number that can poison a classroom’s culture regardless of how many students are technically breaking any rule.
Teachers see it too, and it’s reshaping how they teach. A nationally representative NPR/Ipsos poll of 545 K-12 teachers found:
- 73% believe AI’s impact on education will be bigger than that of the internet or computers
- 60% say AI is eroding trust between students and teachers
- 55% call AI “mostly a shortcut to avoid doing work”
- 54% say AI makes it harder for students to develop critical thinking skills
- 80% say schools should be actively teaching students to use AI responsibly
One teacher summed up the underlying worry bluntly: “I want them to figure out things for themselves, not rely on software. If we stop questioning what it says, we can be led to believe anything.” That erosion of trust has driven a partial retreat to older methods: 40% of teachers now require more handwritten assignments, and the same share require more work be done in class, where AI is harder to use undetected. I spent a decade teaching at Boston University in the 1970s, back when the biggest threat to academic honesty was a roommate’s old term paper; grading a blue-book exam again in 2026 would have felt like science fiction back then, and now it’s being marketed as a solution.
The obvious fix, AI-detection software, has turned out to be far less reliable than schools hoped, and biased in ways that hit some students much harder than others. Turnitin, the most widely used plagiarism-detection company in education, initially advertised a false-positive rate below 1%. That claim didn’t hold up; the company later disclosed that its sentence-level detection carries roughly a 4% error rate. Four percent might sound tolerable, until you multiply it across a university of 20,000 students turning in papers every week. Soheil Feizi, a University of Maryland researcher who has studied these tools extensively, concluded flatly that “no publicly available AI detectors are sufficiently reliable in practical scenarios.” He argues detectors would need a false-positive rate closer to 0.01% to be fair to use in disciplinary decisions — and, he says, “at this point, it’s impossible.”
The bias problem is worse still for one group in particular. A Stanford study of seven widely used AI detectors found they incorrectly flagged 61% of TOEFL essays written by non-native English speakers as AI-generated — 97% of those essays were flagged by at least one detector, and 19% were unanimously condemned by all seven — while the same tools performed “near-perfect” on essays from U.S.-born students. The detectors rely on a measure called “perplexity,” which rewards sophisticated, idiomatic phrasing, exactly what a second-language writer is less likely to produce, AI or no AI. That bias isn’t hypothetical: the Yale student mentioned above, suing under the pseudonym John Doe after being suspended over a GPTZero flag on his final exam, argues in his complaint that “the AI program is unreliable and contains implicit bias,” pointing specifically to his status as a non-native English speaker.
So where does that leave educators? Caught, mostly, between two problems that don’t cancel each other out: real, measurable cheating on one side, and unreliable, biased tools for catching it on the other. Half of teachers in the NPR poll say their school has offered no formal guidance on AI at all, and among the schools that have introduced AI software, only 35% have a formal policy governing how teachers should use it. Just 40% of teachers say they’ve received any professional development on the technology now reshaping their classrooms — roughly the same share of physicians and lawyers who, in surveys I’ve cited here before, said they wanted more training in AI before trusting it with real decisions.
Not every country is handling this the same fumbling way. During the 2025 gaokao, China’s high-stakes college entrance exam taken by roughly 13 million students, the country’s major AI companies — Alibaba, ByteDance, Tencent, and Moonshot — quietly disabled the photo-recognition features on their own chatbots for the duration of the test, so a student couldn’t snap a picture of a question and get an answer back. Students who tried were simply told “such action is not in compliance with the rules.” No lawsuits, no disciplinary hearings, no debate about detector bias — just a centrally coordinated flip of a switch, with exam proctors separately using their own AI surveillance to watch for cheating. It’s a blunt, top-down fix that would be unthinkable in a US system built around due process and decentralized school boards, where ordering four private companies to alter their products overnight is a legal fantasy. Whether that makes the American approach more principled or just slower and costlier at arriving somewhere similar is a fair question, and not one with an obvious answer.
Here’s my honest, tentative prediction, offered with the caveat I’d apply to any forecast about a fast-moving technology: detection is probably a losing game, and schools that keep betting on it will keep generating cases like Orion Newby’s. The 80% of teachers who want schools to actively teach responsible AI use are effectively voting for Chirikov’s approach over Turnitin’s — not banning the technology or policing it with flawed software, but building the judgment to use it well, the same judgment students will need in whatever profession they enter. Whether that shift happens fast enough to spare the next round of falsely accused students is the part nobody can honestly answer yet. Three years after ChatGPT walked into the classroom uninvited, most US schools are still improvising a response, and if the last three years are any guide, the next three will bring at least one more surprise nobody saw coming.
Editorial note: In case you don’t realize how hard it is for teachers to recognize AI-created homework, this post was 100% generated by AI. Could you tell?
This is the prompt I used with Claude AI: “Draft a 1500 word blog post on how educators are being challenged by AI, especially the issue of students cheating. Write it in the same general style as previous posts in my blog aiandprofessionals.com. Also see chinain5.com for more examples of my writing style.” I added a few links for accuracy, but otherwise did not change a single word.