The Setup: High Stakes, First Deployment
UNAM is not a minor institution. More than 100,000 applicants compete each year for roughly 22,000 spots. Rejection rates run between 80% and 90%. For many Mexican students, UNAM represents the only realistic path into higher education, given the prohibitive cost of private universities.
The university contracted Territorium Life to administer the online exam across May and June, with AI surveillance described as capable of “capturing and recording any behavior and action carried out by each applicant,” supported by human monitors.
What the system encountered was a motivated, resourceful population with strong financial incentives to find workarounds — and a secondary market that apparently moved quickly to supply them.
The Signal That Broke the System
The anomaly that triggered the investigation was statistical, not behavioral. The number of perfect scores tripled compared to previous years. Simultaneously, the number of near-zero scores also rose sharply — a pattern consistent with a split population: those who used external tools effectively, and those who may have attempted to and failed, or who were simply underprepared.
Alma Maldonado, a member of the appointed technical commission, described the pattern as “really shocking.” The distribution did not look like a normal exam cohort. It looked like contamination.
This is a critical point: the AI monitoring system did not catch the cheating in real time. The anomaly was detected afterward, through score analysis. The behavioral surveillance — the layer meant to prevent cheating as it happened — appears to have been insufficient to stop it at scale.
What the Cheating Actually Looked Like
The investigation found that some students used phones, brought companions into the exam environment, and in some cases committed identity fraud. Viral TikTok videos allegedly demonstrated methods for defeating the monitoring software. One student reported being offered a program that would solve the exam automatically for approximately $150 USD — and knew of peers who purchased and used it.
This points to a structural problem with remote proctoring in high-stakes contexts: the attack surface is wide, and the incentive to exploit it is proportional to what is at stake.
AI proctoring systems are generally trained to detect known behavioral signals — unusual eye movement, audio anomalies, screen switching, multiple faces in frame. They are less effective against:
- Purpose-built software designed to mimic compliant behavior
- Off-screen assistance that stays outside the camera’s field of view
- Identity substitution, where a different person sits the exam entirely
- Pre-leaked question sets, which require no in-session cheating at all
The UNAM case appears to have involved several of these simultaneously.
The Vendor’s Position and Its Limits
Territorium Life stated that its platform “operated as planned” and that integrity “does not depend solely on technological infrastructure” — requiring instead the coordination of both technological and academic security components.
That framing is technically defensible. No proctoring vendor can fully compensate for gaps in question security, identity verification at enrollment, or the absence of sufficient human oversight during the exam itself.
But it also illustrates a recurring dynamic in AI tool deployment: when something fails, responsibility distributes across the system. The institution points to the vendor. The vendor points to the broader process. Students caught in the middle — including those with no involvement in cheating — absorb the consequences.
The Collateral Damage Problem
This is where the UNAM case becomes particularly instructive for anyone evaluating AI-assisted processes in high-stakes environments.
Students like the young man from Veracruz who scored 94 out of 120 on his fourth attempt — after three previous trips to Mexico City — now face an uncertain wait and a mandatory in-person retake. He is not suspected of cheating. He is simply caught in the remediation net cast wide enough to cover 58,000 applicants.
When an AI system fails to prevent fraud at scale, the corrective action often falls on the entire population, not just the offenders. That is a significant cost that rarely appears in vendor capability assessments or institutional planning documents.
What This Means for AI Proctoring as a Category
AI proctoring tools occupy a genuine and growing market. Remote exams are more accessible, cheaper to administer, and increasingly expected in a post-pandemic education environment. The tools themselves have improved considerably in terms of behavioral detection, anomaly flagging, and integration with learning management systems.
But the UNAM case surfaces a set of conditions under which these tools face structural limits:
- Extreme incentive environments. When the outcome of an exam determines access to a scarce, high-value resource, the motivation to cheat scales accordingly. Most proctoring systems are calibrated for lower-stakes contexts.
- First-time remote deployment at scale. UNAM had no prior baseline for online exam behavior. Anomaly detection depends on having a reliable reference point.
- Insufficient identity verification upstream. Behavioral monitoring during an exam cannot compensate for weak identity checks at registration.
- Question security as a separate failure mode. If questions are leaked before the exam, no amount of in-session surveillance is relevant.
None of these are arguments against AI proctoring as a tool. They are arguments for understanding where it fits in a security architecture — and where it does not substitute for other controls in higher education.
The Practical Takeaway
UNAM’s experience is a useful reference point for any institution, certification body, or organization considering AI proctoring for high-stakes assessment.
The question to ask is not “does this tool detect cheating?” Most do, to varying degrees, under standard conditions. The more precise questions are:
- What is the incentive level of the population being tested, and how does that affect adversarial behavior?
- What happens when the system fails — who bears the cost, and how is it distributed?
- Is the proctoring layer part of a layered security model, or is it being asked to carry the entire integrity burden?
- Has the institution stress-tested the system before deploying it at scale in a consequential context?
AI proctoring tools are not inherently unreliable. But they are being deployed in environments that were not always part of their design assumptions. The gap between what a tool is capable of and what an institution expects it to do is where scandals like UNAM’s tend to originate.
Comments (0) No comments yet
Want to join this discussion? Login or Register.
No comments yet. Be the first to share your thoughts!