AI in Education: Why Student Cheating Is a Design Problem
In spring, one of the most prominent game theorists within the realm of American economics experienced firsthand as his students went into a standard ‘prisoner’s dilemma’ as outlined in textbooks; they ultimately failed.
Dr. Roberto Serrano, the Harrison S. Kravis University Professor of Economics at Brown University and Editor of ‘Games and Economic Behavior,’ had demonstrated a level of compassionate integrity toward his students following the occurrence of a mass shooting that created a highly negative emotional environment on campus by converting his mid-term exams for the 86 students enrolled in his advanced mathematical economics class into take-home exams to minimize the impact of trauma and address their mental health problems following the traumatic event.
It was reported that the mid-term examination had an average score of 96 and that 40 out of 86 students had achieved a perfect score on the mid-term, whereas the final exam, which was a traditional in-classroom exam versus being take-home, had an average score of 48%, and 27 students had withdrawn from the course before the mid-term, 22 of whom completed the mid-term with a perfect score.
The forensic evidence was clear. The graders discovered what Serrano later described as “unusual passages from students’ work that matched results produced when the questions were input into the ChatGPT modeling system.” On one problem involving a short but elegant proof, the ChatGPT model produced an absurdly complicated argument; that same absurdly complicated argument could be found, word-for-word, in over a dozen exams.
Serrano reported his findings to his dean and the provost; initially, he was ignored. Only after being escalated to Brown University’s Academic Code Committee was the incident recognized as the largest ever documented case of AI-enhanced cheating in all the Ivy League institutions. At a recent class meeting, Serrano put an important question to his students—one that countless faculty members have been asking themselves (and perhaps questioning their purpose): “Why have you chosen to attend a university if you are unwilling to learn, work hard, and commit the required time to the development of critical thinking?”
The question is legitimate. But it may not be the right first question.
The Policies We Currently Have May Not Be the Correct Strategy
Generative AI has received a similar three-part response from higher education: (1) bans, (2) detection measures, and (3) creating a moral framework around its use. None of these responses has produced solutions.
Bans don’t work because the technology is undetectable by traditional proctoring methods. The changes made at Princeton University, which require students to take tests in the classroom with a proctor, reflect a realization that the only reliable form of verification available to educational institutions may be through human observation.
Detection measures are not judicial instruments, and their high error rates render them useless for imposing penalties against students. Furthermore, with every improvement made in detection, an equivalent improvement will quickly be made to avoid detection. Furthermore, most students are already aware that using AI to cheat is wrong; they are not confused about the rules; they simply do not have adequate incentives to follow them. Therefore, those who argue that prohibiting AI in schools will not be effective at “protecting students” are making a design argument disguised as a policy argument.
The Nash Equilibrium in the Classroom
Our modern-day assessments are an unassessed take-home exam governed by ambiguous AI guidelines without a verification stage, and represent a hidden action by being a one-time action game. Therefore, the position of using AI dominates the position of not using it for any student who values grades more than they value the abstract moral principle of completing work on their own.
This is not cynicism; it is the outcome matrix. If no one cheats, then all students will learn. If everyone cheats, grades will rise equally for all students, with no single student being disadvantaged. However, if some students cheat and others do not cheat, then the honest students will suffer twice, once from the effort put into an assignment and once from a grade curve that is affected negatively by students who have cheated. Therefore, current bans and tools to detect cheating result in a very small chance that an individual student will suffer; consequently, as they do not change the outcome of the situation, they only change the appearance or perception of the situation.
The fictitious play model is not evidence that our students today are more corrupt than previous generations; rather, it is an indication that the assessments given to students today have become games that we would reject if we could observe them written out on a whiteboard.
Redesign the Game
When a game theorist comes across an unsatisfactory equilibrium, they do not try to work with the players to find a new equilibrium; instead, they simply modify the game itself. A faculty member like Serrano who successfully modifies their game can pull on any of three levers (or all three) to make changes.
Lever #1: Change the payoffs from the artifact to the defenses of the artifact. Generative models can produce artifacts; they cannot defend a student’s artifact in real time (to a human interlocutor). After the cheating scandal, Serrano permanently banned take-home exams and set weekly homework to have zero grading weight. This change denotes the maximus of making payoffs from the artifact to the defense of the artifact, as now, homework is simply for practice and can be done however a student wants, without concern for the tools available.
Lever #2: Change the information structure. Unclear, ambiguous use of AI guesses can be an advantage to the student who is taking the strategic approach. It is important to be very clear (in writing) about which form of AI assistance is permitted for each assignment and to have each student submit a statement of the AI tools used, their purpose, and the prompts for the tools used. In addition to the basics, it is also important to include version histories and draft trails as part of the delivered assignment. The purpose of these components is not to surveil but rather to create an enforceable separate equilibrium in which the honest students can clearly differentiate themselves from the dishonest students at a low cost.
Lever #3: Change the order of all the movements in the activity. There are ways for students to hide how they are completing single-submission assignments, but there are no ways for students to hide how they are completing multi-stage assignments. If the same activity as assigning a research paper is split into multiple sub-activities, such as a research memo defended by way of oral presentation, an annotated bibliography with an oral presentation, a workshopped draft, a second draft in response to a peer review, and a final assessment with an oral defense, this makes it easier for students to cooperate since they can observe one another completing and producing the same task.
Instruments that can be used to measure and evaluate this concept include adversarial prompting assignments in which students critique or improve an AI’s answer, changing the AI from being used as a replacement for a student to being used for academic research about AIs. Calibration scoring is a system used to score based on forecasts that offer rewards for accurately self-assessing and penalties for being overconfident because of receiving an unearned grade. Two envelopes to respond with take-home assignments at a low weighting, with the same questions assessed in-class, as a different assessment will be weighted differently on the final grade.
What Institutions Owe the Faculty Doing This Work
Redesign is expensive and time-consuming. Oral defenses are lengthy. Proctored examinations require physical space within which to take the examinations. To achieve an honest assessment, universities must provide financial support for those rigorous forms of assessment instead of encouraging a “friction-free” take-home model of learning that has turned educational institutions (classrooms) into gambling establishments using an “honor” system.
When a faculty member reports their experiences to the institution, the administration must take those reports seriously upon receipt. The fact that Serrano was denied by his administration initially will impact how all faculty members will view the experience and whether they will escalate issues to avoid the possible political consequences.
AI detectors should not be part of the academic assessment process. AI detectors are tools used only as indicators of whether a more in-depth investigation is necessary; they do not provide sufficient evidence to determine an outcome.
The Guardrails We Need
Serrano states that his students could benefit from “guardrails” and that “we need to put the right guardrails in place, and if they fail, we need to implement consequences.” He is right, but the guardrails he is talking about are not about surveillance. They are about structure, i.e., making sure we have structured assessments that AI can pass vs. assessments they cannot pass. So, the point he is making is hierarchical for any provost or dean reading this.
He told Fortune, “If we do not defend truth and decency and honesty, then what kind of credibility do we have as academics?” The answer is contingent on the design we implement, rather than the policies we create. Generative AI will not get un-invented; bans are only going to continue to be announced, continue to be circumvented, and fail to protect anyone from anything.
The faculty who will flourish in the next 10 years of higher education will not be those with the strictest rules; they will be those who, like Serrano, have examined the incentive structure fairly and redesigned their game. Therefore, we should stop asking how we can prevent students from cheating and start asking the more significant question: Given what these tools allow for, what is it we want these students to do?