The short answer: PISA 2025 does not show that every use of AI harms learning. It shows a more useful pattern. Students’ outcomes differ by why, how often, and how critically they use AI. For exam owners, the implication is clear: assessment integrity in 2026 cannot be reduced to banning a tool or running an AI detector. Institutions need assessments that make valid evidence of learning visible.

The OECD published PISA 2025 results in September 2026, providing a timely view of how students use AI chatbots for schoolwork and how those patterns relate to performance. The findings should be read as associations, not proof that AI use causes a particular result. They nevertheless give education leaders a strong basis for redesigning policy, assessment, and supervision.

What PISA 2025 says about student AI use

  • AI use for schoolwork is widespread across participating countries and economies.

  • On average across OECD countries, 14% of students reported never or almost never using AI for any of the schoolwork purposes examined.

  • Students who did not use AI for specific tasks such as summarizing texts or preliminary research generally outperformed users in many countries and economies.

  • For some learning-oriented uses, moderate or weekly use was associated with outcomes similar to or slightly better than non-use, while very frequent use was often associated with lower performance.

  • Students who learned to evaluate the quality of AI-generated information showed more positive patterns, but access to that instruction was unequal.

These findings do not support a simple ‘AI is good’ or ‘AI is bad’ conclusion. They support a distinction between AI that helps a student think and AI that replaces the thinking an assessment is meant to measure.

What the findings mean for assessment integrity

1. A polished output is weaker evidence than it used to be

Generative AI can produce fluent essays, summaries, code, images, and explanations. When an assessment grades only the final artifact, it may be difficult to distinguish demonstrated competence from delegated production. The answer is not to assume that every strong response is suspicious. It is to collect better evidence.

2. AI policy must be tied to the learning outcome

If the objective is unaided recall or independent reasoning, AI access may invalidate the task and should be restricted. If the objective is to evaluate evidence, critique an AI output, or use professional tools responsibly, supervised AI use may be appropriate. The policy should state the allowed level, required disclosure, and evidence expected from the student.

3. Detection alone cannot preserve validity

AI-text detectors cannot reliably reconstruct how a response was produced, and device monitoring cannot observe every off-platform path. Detection can identify risk signals, but valid assessment requires design choices that make the student’s process, judgment, and identity observable.

Five principles for AI-resilient assessment design

1. Start with the claim the result must support

Write one sentence describing what a passing result should prove. Then ask what evidence would remain persuasive if the candidate had access to a powerful AI assistant. This shifts the design conversation from policing a tool to protecting the meaning of the credential.

2. Use specific context and original inputs

Ofqual’s 2026 advice notes that tasks tied to a particular context, data set, or problem may be less susceptible to generic AI-generated responses. Use local cases, original observations, unique data, staged information, or role-specific constraints. Specificity is not a guarantee, but it makes generic delegation less effective and improves the quality of evidence.

3. Assess process as well as product

Collect plans, intermediate calculations, version history, source notes, short reflections, demonstrations, or oral defenses where these are relevant to the learning outcome. Do not add process artifacts as bureaucracy. Choose the smallest set that helps a reviewer understand how the candidate reached the result.

4. Add supervised checkpoints where consequences are high

A short live explanation, controlled practical task, identity-verified exam, or proctored validation exercise can confirm whether the candidate can reproduce or defend the competence shown in unsupervised work. The higher the consequence of a false result, the stronger this confirmation should be.

5. Teach critical AI literacy

PISA 2025 highlights the value of learning to assess AI-generated information. Students should practice checking evidence, identifying uncertainty, comparing sources, documenting prompts or assistance where required, and taking responsibility for the final work. Integrity improves when policy distinguishes transparent assistance from misrepresentation.

A simple three-zone AI policy

Red zone: independent performance required

No generative AI or unauthorized external assistance. Use this zone for skills that must be demonstrated without support, high-stakes knowledge checks, licensing decisions, or controlled verification of earlier work. State the permitted tools precisely and apply proportionate supervision.

Amber zone: limited AI assistance with disclosure

Allow defined activities such as brainstorming, grammar support, translation, or feedback, while requiring the candidate to document use and remain responsible for accuracy. Grade the underlying reasoning and evidence, not only the polish of the output.

Green zone: AI use is part of the assessed competence

Let candidates use AI openly and assess prompt strategy, verification, judgment, domain expertise, and the ability to improve or reject an AI output. The assessment should still reveal what the student understands and can defend.

When proctoring is appropriate

Proctoring is most useful when an institution needs controlled evidence of independent performance: admissions, professional certification, licensing, final validation of course outcomes, or a supervised checkpoint in a broader assessment. It should not be the automatic answer for every assignment.

Where proctoring is justified, combine clear rules, identity assurance, device and environment controls, multi-signal monitoring, human review, privacy safeguards, accommodations, and appeals. Configure the level of control to the stakes instead of applying maximum restrictions to every learner.

An implementation sequence for institutions

  • Inventory assessments by learning outcome and consequence of an invalid result.

  • Assign each task to a red, amber, or green AI-use zone.

  • Redesign high-risk tasks to collect process, context, or supervised evidence.

  • Publish candidate-facing examples of permitted and prohibited AI use.

  • Choose proportionate identity and integrity controls for supervised checkpoints.

  • Train educators and reviewers to evaluate evidence without over-relying on automated flags.

  • Measure candidate completion, support needs, accessibility exceptions, reviewer agreement, appeals, and learning outcomes.

  • Review the policy as AI capabilities and official guidance change.

Frequently asked questions

Does PISA 2025 prove that AI lowers student performance?

No. PISA reports associations between self-reported AI-use patterns and performance. The relationship varies by purpose and frequency, and the results do not establish simple causation.

What is an AI-resilient assessment?

It is an assessment designed to produce valid evidence of competence even when generative AI is widely available. It clarifies permitted use, collects the right evidence, and makes it difficult for unauthorized delegation to substitute for the skill being assessed.

Are oral exams the only AI-resilient option?

No. Useful approaches include specific contextual tasks, original data, staged problems, process evidence, practical demonstrations, controlled checkpoints, oral follow-ups, and assessments where responsible AI use is itself part of the competence.

Should universities ban AI in all assessments?

A blanket ban ignores the difference between learning outcomes. Some tasks require unaided performance; others should teach and assess responsible AI use. A task-level policy is clearer and more defensible than one rule for every context.

Sources and further reading

Orken Rakhmatulla

Head of Education

Share