The perfect scores arrived first, then the questions. That is the order of things now in assessment, and the Universidad Nacional Autónoma de México is living through the consequences 1. When UNAM rolled out its first fully online admission test, it saw a surge in perfect results—anomalous enough that the university ordered a second, supervised control exam for thousands of applicants, with results due in late August 1. The cheating involved AI tools and hidden devices, which is to say, the test was not measuring aptitude so much as access to better cheating software.
The classroom problem here is not that students cheat. It is that our assessment instruments have become unmoored from the learning they claim to measure. The UNAM case is the acute version of a chronic condition. A large-scale study in China, tracking 26,811 students, found that AI homework assistance raised assignment grades by 18% while exam scores fell by 20% 3. That gap is the real scandal. It tells us that homework, as currently designed and graded, has drifted so far from the knowledge it is supposed to build that it now rewards the opposite: the efficient outsourcing of thought.
The incentive producing this is structural. Teachers are evaluated on completion and compliance, not on the integrity of the learning process. Institutions are rewarded for throughput and retention metrics, not for the durability of knowledge. When homework grades rise and exam scores fall, the system has an incentive to celebrate the homework and ignore the exam. The Chinese study is useful precisely because it quantifies the divergence: an 18% gain in one column, a 20% loss in the other 3. That is not a tradeoff; it is a warning.
The conventional reform, of course, is to ban the tools. Mexico is now weighing nationwide cellphone restrictions in basic-education classrooms, with a consultation scheduled for August 24–25 across the Technical School Councils, involving more than a million teachers 4. Sixteen states have already acted 4. Banning phones is satisfying and photographable, but it mistakes the instrument for the incentive. The phone is not the problem; the assignment that a phone can complete is the problem. If a homework task can be finished by a chatbot in thirty seconds, the task was never assessing what we thought it was assessing.
The same logic applies to the Dominican Republic, where the 2026-2027 school year begins Monday amid a dispute over readiness 2. The Ministry of Education insists it has 100% of teachers and enough classrooms; the ADP union protests infrastructure problems and unpaid incentives 2. Both claims may be true. The ministry is measuring coverage; the union is measuring conditions. Neither is measuring learning. That is the recurring failure of institutional incentives: they measure what is countable, not what matters.
A realistic alternative is not to fight AI but to redesign assessment around it. The UNAM control exam is a start—a supervised, in-person test that closes the gap between what students can do alone and what they can do with tools 1. But it should not be a one-time correction. It should be the model. If we accept that AI is now part of the working environment, then assessment must shift toward supervised performance, in-class writing, and oral defense. The Chinese data suggests the cost of not doing this is measurable: students who look successful on paper are actually learning less 3.
The tradeoff is uncomfortable. Supervised assessment is more expensive, more labor-intensive, and less scalable than the online model that just failed. UNAM will spend considerable resources re-testing thousands of applicants 1. The Dominican Republic cannot even agree on whether classrooms are ready 2. But the alternative is a system where the perfect score is the surest sign of failure. The question that matters is not whether students will cheat. It is whether institutions will keep building tests that make cheating the rational choice.
