Picture someone rolling a forklift into a gym. The weights go up. The plates clink. The log sheet says 400 pounds, five sets, clean form. And the person holding the controls is exactly as strong as they were an hour ago. That is homework in 2025. The output looks right. The rep never happened.
I spend most of my time here testing AI tools and telling you which ones hold up. So I read the teacher accounts of AI-completed homework the way I read a product bug report: not as a moral panic, but as a spec failure. Homework was a tool. It had a job. Someone found a way to produce the artifact without doing the work, and the tool stopped measuring the thing it claimed to measure.
What one instructor actually changed
The version of this story that stuck with me came from an instructor who described AI handling essentially all of the assigned homework. His response was not a detection arms race. It was a redesign.
Written reflections used to be attached to nearly every assignment. Those are gone, replaced by in-person conversation. After each assignment, students book a one-on-one with a TA and talk through what they did. Oral and written exams carry more weight. Most assignments were rebuilt from scratch.
He was also blunt about the cost: the shift meant walking away from evidence-based practices he had used for years. That admission is the most credible part of the whole account. Anyone selling you a clean, painless adaptation to AI in the classroom is selling something.
Why detection was always the wrong product
Every few months a new tool promises to flag AI-written student work. I have tested enough of these to have a standing position: treating this as a detection problem puts you in a race you cannot win, against a system that improves monthly, with a false-positive rate you have to defend to a nineteen-year-old and their parents.
The teachers moving to oral exams and live discussion figured out something reviewers learn the hard way. When you cannot verify the artifact, stop grading the artifact. Verify the person. A student who understands their own submission can explain a choice they made, defend it, and revise it out loud. A student who does not will stall in about ninety seconds. No subscription required.
Homework as process, not product
The deeper change in these accounts is a reframe. Homework used to be a deliverable you handed in. Now it is treated as a learning process that happens to leave a paper trail. The paper is evidence, not the product.
That opens up assignment designs that were never practical before. Two that show up repeatedly in the research discussion:
- Error analysis. Hand students AI-generated solutions and ask them to find and fix what is wrong. The model becomes the thing being graded, and the student has to be the one who knows better.
- Self-generated problems. Ask students to write their own question modeled on the class material, then solve it. Writing a good problem requires understanding the shape of the concept, which is harder to outsource than an answer.
Both designs share a property I look for in any tool: they degrade gracefully. If a student uses AI on them, the AI use is either visible or actually part of the learning.
The part nobody wants to review honestly
Now the tradeoff column, because that is the job. One-on-one meetings after every assignment do not scale. They are fine in a course with TA support and a manageable roster. They are a fantasy in a 300-student lecture or a high school teacher’s 150-student load. Oral exams also carry real equity problems. Students with anxiety, students speaking a second language, and students who process slowly under pressure are all penalized by a format that rewards quick verbal recall. That is not a reason to skip the format. It is a reason to build accommodations into it deliberately rather than discovering the gap in week ten.
There is also the quiet loss the instructor named. Some of the practices he dropped were supported by actual evidence. Replacing them with in-person interviews may be the right call under the circumstances, but it is a downgrade forced by conditions, not an upgrade. Calling it progress would be dishonest.
What I would take from this
If you teach, the useful question is not how to catch AI. It is which of your assessments still measure something when the artifact is free to produce. Most graded writing that exists mainly to prove effort is already dead. Anything requiring live reasoning, critique, or original problem construction is not.
If you build tools, the market signal here is loud. The demand is not for better detectors. It is for infrastructure that makes small-group verification cheap enough to survive a 300-person roster. That product barely exists yet. Whoever ships it well is going to matter more than every AI-checker on the market combined.
🕒 Published: