The short answer: AI proctoring used by an educational or vocational institution to monitor and detect prohibited behavior during an exam can fall within the EU AI Act’s high-risk education category. But classification depends on the system’s intended purpose, how its output is used, and whether the AI performs a consequential decision or a narrow supporting task.

The European Commission’s AI Act Service Desk now gives proctoring-specific examples. That makes 2026 the right time for universities, certification providers, and technology vendors to document their use cases instead of relying on a generic label such as ‘AI monitoring.’ This article is a practical planning guide, not legal advice.

Why the intended purpose matters more than the product name

The AI Act classifies systems by what they are intended to do in a real deployment. Two products may use similar models but fall into different risk analyses because one detects prohibited behavior during a certification exam while another only assists with a procedural identity check.

Examples the Commission lists as potentially in scope

  • AI-enabled proctoring used during a certification exam to monitor access to unauthorized material or communication with another person.

  • Real-time behavior analysis used to detect patterns that may indicate cheating during an online exam.

  • AI exam monitoring used to detect prohibited objects or behavior during an in-person assessment.

Examples that may fall outside that specific high-risk use case

  • A plagiarism checker that analyzes work produced outside a live or supervised test.

  • Identity verification that processes documents or facial data but leaves access and sanction decisions to a human proctor may qualify as a narrow procedural task.

  • AI used by a human proctor to confirm or challenge an observation already made by that proctor may qualify as improving a previously completed human activity.

These examples are useful, but they are not a shortcut. An organization must document the actual workflow, model outputs, decision points, affected people, and consequences. A feature name or a contractual disclaimer will not override how a system is designed and used.

Does human review make AI proctoring low risk?

Not automatically. Human oversight is important, but a reviewer must have the authority, information, time, and training to challenge an AI output. A rubber-stamp review does not change the practical impact of the system. The key questions are whether the AI initiates or determines a consequential outcome, whether its result materially shapes the decision, and whether the reviewer can independently assess the evidence.

For exam integrity, a safer operating model is to treat AI outputs as risk indicators. The system should show the underlying event timeline and relevant evidence, while a trained reviewer applies the institution’s rules. Candidates should have access to a documented challenge or appeal process.

The current EU AI Act timeline

The European Commission’s current implementation timeline says transparency rules apply from August 2026. Following the 2026 simplification agreement, the high-risk rules for standalone systems in sensitive areas, including education, are scheduled to apply from 2 December 2027. Organizations should verify the latest official timeline when making decisions because implementation dates and supporting guidance can change.

The additional preparation time should not be treated as permission to wait. Evidence governance, human oversight, data quality, accuracy testing, and candidate redress are operational systems that take time to design and validate.

A practical compliance-readiness checklist

1. Write a precise intended-purpose statement

Describe the users, exam type, monitored behavior, model outputs, and decisions the output may influence. Separate identity verification, behavior detection, evidence prioritization, and final misconduct decisions instead of describing everything as one AI feature.

2. Map every human and automated decision

Document who sees each flag, what evidence they receive, whether they can override it, and what happens next. Identify any situation where an automated result blocks entry, invalidates an attempt, changes a score, or triggers a sanction.

3. Build evidence logging and traceability

Keep a reliable record of the rule, event, timestamp, model or control involved, reviewer action, and final outcome. Logging should support quality assurance and appeals without collecting unrelated candidate data.

4. Test accuracy in the real candidate population

Measure false positives and false negatives by exam type, device, network condition, lighting, language, disability or accommodation path, and other relevant operating conditions. A global model metric is not enough to understand deployment risk.

5. Define meaningful human oversight

Train reviewers, give them access to the underlying evidence, prevent automation bias, and monitor reviewer agreement. Set escalation thresholds for ambiguous or high-impact cases.

6. Minimize and govern personal data

Define the lawful basis and purpose for each data type, limit access, publish retention and deletion rules, and assess biometric and video processing separately. Procurement teams should understand where recordings and logs are stored and whether the customer can control data residency.

7. Inform candidates and provide redress

Use plain language to explain monitoring, prohibited behavior, AI involvement, evidence use, retention, accommodations, technical support, and appeals. A candidate should know how to request human review of a contested result.

8. Control vendor and model changes

Require documentation for material model, threshold, feature, and data-flow changes. Re-test the workflow when a change could affect accuracy, classification, or the candidate experience.

Questions to ask an AI proctoring vendor

  • What exact decisions does the system make, recommend, or prioritize?

  • Can the customer configure AI flags as advisory evidence only?

  • What evidence is shown to a human reviewer for each flag?

  • How are false positives measured and investigated?

  • How does the system handle low bandwidth, poor lighting, assistive technology, and accommodations?

  • Where are video, biometric data, and event logs stored, and who controls retention?

  • How are virtual cameras, replayed media, remote access, and device-level bypasses addressed?

  • What documentation and change notices are available for governance and audit?

Frequently asked questions

Is every AI proctoring system high-risk under the EU AI Act?

No. The analysis depends on intended purpose and use. Real-time AI monitoring for prohibited behavior during educational or certification exams is explicitly identified as an in-scope example, while narrow procedural identity verification or AI that only supports a prior human assessment may qualify for an exception in some circumstances.

Does keeping a human in the loop guarantee an exception?

No. Human involvement must be meaningful. Reviewers need authority and sufficient evidence to challenge the output, and the AI must not effectively determine the result before the review occurs.

Is identity verification treated the same as cheating detection?

Not necessarily. The Commission gives an example where identity verification that streamlines document or facial checks, without deciding exam access or sanctions, may be treated as a procedural task. The full workflow still needs to be assessed.

What should organizations do first in 2026?

Create a use-case inventory and decision map. Those two documents reveal which systems require deeper classification analysis and where evidence, human oversight, privacy, or redress controls are missing.

Sources and further reading

Orken Rakhmatulla

Head of Education

Share