Skip to content
All resourcesTRUSTEXAM RESOURCES

How to evaluate an AI proctoring pilot

A practical pilot plan: representative participants, completion rates, useful signals, false alarms and human review.

A useful proctoring pilot answers a decision: can this configuration support this examination programme under its actual operating conditions? A demonstration can explain features. A pilot should test readiness, participant experience and the quality of human review.

1. Define the assessment before the technology

Record the exam format, stakes, permitted references, devices, languages and locations. Identify the programme owner, technical lead, support contact and people authorised to make integrity decisions. Agree how participants request accommodations and how technical interruptions are handled.

2. Choose a representative cohort

Include the setups that matter in production, not only the newest devices or the most experienced staff. Where the programme uses several languages or venues, include those variations. Record which conditions the pilot did not cover; those remain untested, rather than implicitly approved.

3. Set measurable criteria in advance

Define a completed session before calculating completion rate. Divide completed eligible sessions by all eligible started sessions and report both counts. Record cancellations, practice runs and authorised retakes separately. This prevents a changing denominator from making the pilot look better than it was.

For operations, measure setup failures, support requests, interruptions and review time. For review quality, examine useful signals, signals with an innocent explanation and disagreements between reviewers. Set thresholds appropriate to this programme; there is no universal acceptable number of flags.

4. Review more than flagged sessions

Have authorised reviewers examine a sample of sessions with no automated flags as well as flagged sessions. Without that comparison, the pilot cannot show what the system may have missed. Use agreed rules and, where practical, an independent second review of disputed examples.

5. Make a documented rollout decision

Summarise what worked, what failed, the conditions tested and the remaining limitations. A higher flag count does not by itself prove more cheating or better detection. Decide whether to proceed, change the configuration or repeat a defined part of the pilot. Keep the previous configuration and a practical return path until the new process is accepted.

What to bring to a TrustExam discussion

Bring the exam format, concurrent session estimate, platform, languages, devices and the review team’s responsibilities. These inputs help scope an appropriate demonstration and a pilot whose results can support an informed decision.

Practical guides

YOUR PROGRAMME. YOUR NEXT STEP.

Let’s explore how this works for you.

Book a demo