MODERATION · EXTERNAL EXAMINING

AI in Moderation and Second Marking: What External Examiners Need to See

UK higher education already accepts that markers disagree. That is why moderation, second marking and external examining exist. AI does not remove the need for any of them. Used well, it gives every script a consistent first reading, shows moderators where to look, and gives external examiners better evidence than a random sample ever could.

By Eduface · September 2026 · 12 min read

The exam board meets in June. The external examiner has read a sample of twelve scripts from a module of 340. In two of them, she thinks the first-class mark is generous. The module leader explains that the second marker agreed. The external asks how the sample was chosen. Nobody is quite sure. Somewhere in the other 328 scripts there are probably more marks like those two, and there is no way to find them before the board signs off.

How can AI be used in moderation and second marking?

AI can give every script the same criterion-level first reading, act as an independent comparison point when lecturers mark in blind mode, and flag the scripts and markers where human and AI judgements diverge most. Moderators then look first where disagreement is highest, instead of at a random sample. The lecturer approves every mark, and external examiners see the same evidence as before, plus an audit trail and disagreement data that make the sample far more informative.

How does moderation work in UK higher education today?

Most UK institutions use a mix of four practices, often in the same programme.

Practice

How it works

Typical use

Double-blind marking

Two markers mark independently, without seeing each other’s marks, then agree a final mark

Dissertations, final-year projects, high-stakes work

Second marking

A second marker reviews the first marker’s marks, usually on a sample

Most modules

Moderation

A moderator checks a sample to confirm standards have been applied consistently, without re-marking every script

Large modules with many markers

External examining

An examiner from another institution reviews samples, advises on standards and on whether regulations were followed

Every programme leading to an award

Table 1: The four main moderation practices in UK higher education. Local policies vary.

External examining has its own sector principles. The External Examining Principles, published in 2022 on behalf of the UK Standing Committee for Quality Assessment, say that external examiners should “take part in calibration activities within their discipline” and “advise the institution on whether assessment regulations are followed in determining student marks, outcomes, classifications and awards”.¹

The QAA’s accompanying guidance is realistic about what an examiner can see. Because of the volume of assessment, “an external examiner is unlikely to be able to view all the assessed work… so they are usually viewing samples”. Those samples “need to be of sufficient size” and “should include examples from different degree classifications, including fails”. And external examiners “should not expect or encourage an examination board to raise or lower marks for individual students, because this would be unfair to those candidates whose work is not part of the sample”.²

That last point is the heart of the problem. If a sample reveals a generous or harsh marker, the fair response is to review that marker’s whole allocation, not to adjust the sampled scripts. But without a way to find the other affected scripts, that review is slow, expensive and often skipped.

Why do markers disagree in the first place?

Because academic judgement is hard to write down. The Higher Education Academy concluded in 2018 that “although the sector has produced vast quantities of agreed written standards, achieving a shared interpretation of their meaning is a very different matter”, and that repeated studies “demonstrate considerable inconsistency in academics’ judgements about student performance”. It added that studies of external examiners have found similar inconsistency.³

Bloxham, Boyd and Orr showed how this happens in practice. Twelve UK lecturers thought aloud while marking. Most formed a holistic judgement first, then used the written criteria to refine it, drawing on tacit standards built from experience and on comparisons between scripts.⁴ That is how experienced marking works, and it is also why two experienced markers can land in different places.

None of this is a criticism of markers. It is the reason moderation exists. The question is whether AI can make moderation better at its job.

Where does AI fit into moderation?

AI adds one thing that no human process has: a reading of every script against the same rubric, in the same way, with the reasoning written down. That consistent reference point can be used in four ways.

Model

How it works

What it adds

What stays human

1. AI first reading, human approval

AI drafts criterion scores and reasoning for every script; the marker reviews, edits and approves

Every marker starts from the same reading of the rubric

Every mark is the marker’s decision

2. AI as independent comparison (blind mode)

The marker marks first without seeing the AI; the AI draft is revealed afterwards for comparison

A second reading of every script, not only a sample

The marker decides whether to change anything

3. Targeted moderation

Scripts and markers where human and AI marks diverge most are flagged for the moderator

The moderator looks where the risk is, not at random

The moderator’s judgement

4. Calibration

Exemplar scripts are run through the AI before marking starts; the team compares their marks and the AI’s criterion by criterion

Makes disagreements about descriptors visible early

The team agrees the standard

Table 2: Four ways to use AI in moderation, from lightest to most integrated.

Model 2 deserves attention. In Jisc’s year-long pilot of AI in marking and feedback, involving 38 UK colleges and universities, one finding was that keeping the human in the loop “can be hard to maintain in practice”, and that methods such as parallel marking helped build it in.⁵ Blind mode is parallel marking by design: the marker’s independent judgement is recorded before the AI’s is visible, so it cannot anchor on the AI. For summative work in the first years of adoption, we recommend institutions make it compulsory.

1

Marker marks script independently (blind mode)

2

AI draft revealed for comparison

3

Small gap: marker confirms · Large gap: marker reviews, then confirms or changes

4

Grader comparison dashboard aggregates gaps per marker

5

Moderator reviews flagged markers and scripts first

6

Exam board sees approved marks, disagreement summary and audit trail

Every script gets two independent readings. Moderation starts from the largest disagreements.

What does targeted moderation look like in practice?

Take a module of 340 scripts and five markers. A conventional moderation sample might be ten per cent, chosen to cover grade bands. With an AI reading of every script, you can choose the sample differently.

Sample component

How it is chosen

Why

Largest disagreements

The scripts where the marker’s mark and the AI draft differ most

Most likely to contain a marking error, in either direction

Marker drift

Scripts from any marker whose marks sit consistently above or below both colleagues and the AI

Finds a generous or harsh marker across their whole allocation

Boundary scripts

Scripts near classification boundaries

Where a small error changes an outcome

Fails and firsts

A spread across classifications, including fails

As the QAA guidance on external samples expects

Random control

A small random selection

Checks that low-disagreement scripts are also sound

Table 3: A risk-based moderation sample. Illustrative; adapt to your institution’s policy.

Eduface’s grader comparison dashboard shows this at a glance: for each marker, how their marks compare with the rest of the team and with the AI draft. A marker whose marks sit consistently above both is visible early in the marking period, not at the exam board.

Marker

Mean difference from the AI draft

Status

Marker A

Around zero

Within agreed tolerance

Marker B

Around zero

Within agreed tolerance

Marker C

Consistently about +4, narrow range

Flagged for moderation

Marker D

Around zero

Within agreed tolerance

Marker E

Around zero

Within agreed tolerance

Figure 1: Marker drift becomes visible early when every script has an AI reading to compare against. Illustrative.

When a drifting marker is found, the fair response the QAA guidance implies becomes practical: review that marker’s whole allocation, with the AI’s reasoning as a starting point for each script, rather than adjusting only the scripts that happened to be sampled.

What do external examiners need to see?

External examiners do not need to understand how an AI model works. They need to be confident that marks were reached properly and consistently, and that regulations were followed. AI-assisted marking should give them more evidence, not less.

What stays the same. The examiner reviews a sample of scripts with the approved marks and feedback, across classifications including fails, and comments on standards, consistency and process.

What you can add:

The marking process in one page. Which modules used AI assistance, in which mode (blind or AI-visible), and the rule that no mark was released without a qualified marker’s approval.

Disagreement data. For each module: how often and by how much marker and AI differed, and how large disagreements were resolved.

The moderation sample and why. A risk-based sample as in Table 3, with the reason each script is in it.

The audit trail. For any sampled script, the AI’s criterion scores and reasoning, the marker’s changes and the final approved mark.

Calibration evidence. If the team calibrated on exemplar scripts before marking, what was compared and what changed in the rubric as a result.

A checklist for the external examiner’s pack

Process summary with the approval rule · list of modules and modes used · disagreement summary per module · risk-based sample with reasons · audit trail for each sampled script · calibration notes · any rubric changes made during the cycle.

This fits the External Examining Principles well. Examiners are asked to take part in calibration within their discipline and to advise on whether regulations were followed.¹ An audit trail per criterion and a documented approval rule make both easier to do properly.

How should you write this into policy?

Moderation and assessment regulations should say explicitly how AI assistance is used. Vague policy is where problems at appeal start. Here is illustrative wording to adapt with your quality team.

Illustrative policy wording

“Where AI assistance is used in marking, the AI produces a draft assessment against the approved rubric. No mark or feedback is released to a student until it has been reviewed and approved by a qualified marker, who remains responsible for the mark. For summative assessment, markers mark independently before viewing the AI draft. Moderation samples include scripts where the marker’s mark and the AI draft differ by more than [agreed tolerance], alongside samples required by this policy. A record of the AI draft, the marker’s decision and any changes is retained for [period] to support moderation, external examining and appeals.”

Four decisions sit behind that wording, and each belongs to your institution:

Which modes are allowed for which assessments (for example, blind mode compulsory for summative work).

The disagreement tolerance that triggers a moderator’s review.

The retention period for the audit trail, set in your Data Processing Agreement.

How students are told, in module handbooks and assessment briefs, where AI assistance is used and that a qualified marker approves every mark.

What are the risks, and how do you manage them?

Risk

What it looks like

How to manage it

Anchoring

Markers accept the AI draft instead of judging

Blind mode for summative work; the marker’s own mark is recorded first

Rubber-stamping

Review becomes a click

Monitor time spent and the rate of changes; sample approved scripts where nothing changed

Shared blind spots

Markers and AI agree, but both are wrong

Keep a random element in every sample; external examiner review

A weak rubric

AI and markers disagree because descriptors are vague

Calibrate before marking; see how to write a rubric AI can mark against

Regulatory exposure

AI used without effective oversight

Approval rule in policy; audit trail; transparency to students

Table 4: Risks of AI in moderation and how to manage them.

On regulation: AI that evaluates learning outcomes is classified as high-risk under Annex III, point 3(b), of the EU AI Act, which requires effective human oversight and record-keeping. Following the Digital Omnibus on AI, those obligations apply to Annex III systems from 2 December 2027.⁶ A process in which a qualified marker approves every mark, with the reasoning logged, is designed for that. For more, see human in the loop: AI assessment and the EU AI Act.

How accurate is the AI reading you are moderating against?

A comparison point is only useful if it is close to good human judgement. Two measurements from UK pilots:

Lecturers changed an average of 5% of each final grade the AI drafted.

In a pilot at one UK university (435 submissions, six modules, 13 markers), the AI’s suggested grade came within 94% of the marker’s grade on average, and within 98% for markers who had tuned the model to their standards. Accuracy means how small the gap is between suggestion and marker.

For open exam answers, Eduface’s Exam Grader uses three independent AI agents per answer and a fourth that reconciles them, flagging disagreement for the lecturer. That internal disagreement is itself a moderation signal: an answer the AI agents cannot agree on is often one human markers would disagree on too.

Frequently asked questions

Can AI be used as a second marker?

AI can provide an independent second reading of every script, but it should not be the second marker of record. In blind mode, the marker marks first and then compares with the AI draft; large differences prompt review. A qualified human still approves every mark, and moderation and external examining continue as normal.

Do external examiners need to approve the use of AI in marking?

External examiners advise on standards and on whether regulations are followed. They do not usually approve tools. But they should be told how AI assistance was used, and it helps to give them the process summary, disagreement data and audit trail described above.

Does AI-assisted marking reduce the need for moderation?

It changes moderation rather than reducing it. Instead of a random sample, moderators can review the scripts and markers where human and AI judgements diverge most, which finds problems a random sample would miss.

What is blind mode in AI marking?

The marker marks the script without seeing the AI’s suggestion. Only afterwards is the AI draft revealed for comparison. It prevents anchoring and records the marker’s independent judgement, which is what moderation and external examining rely on.

How do we explain AI-assisted marking to students?

Say it in the module handbook and assessment brief: where AI assistance is used, that a qualified marker reviews and approves every mark, and how to raise concerns. Eduface labels reviewed feedback “Lecturer + AI”.

References

1. UK Standing Committee for Quality Assessment. (2022). External Examining Principles. Published by QAA, Universities UK and GuildHE. [Principle 1: examiners take part in calibration activities within their discipline; Principle 3: advise whether assessment regulations are followed.]

2. Quality Assurance Agency for Higher Education. (2022). External Examining: Putting the Principles into Practice. [Examiners usually view samples, which should include different classifications including fails; examiners should not encourage boards to change individual sampled marks.]

3. Higher Education Academy. (2018), quoted in Lancaster University, We need to talk about marking, CEDA briefing paper. [Considerable inconsistency in academics’ judgements; similar inconsistency among external examiners.]

4. Bloxham, S., Boyd, P., & Orr, S. (2011). Mark my words: the role of assessment criteria in UK higher education grading practices. Studies in Higher Education, 36(6), 655–670.

5. FE News. (2026, 21 May). New findings from Jisc highlight the benefits of a collaborative approach to AI in assessment. [38 UK colleges and universities; human in the loop hard to maintain in practice; parallel marking helped.]

6. European Parliament and Council of the EU. (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act), Annex III, point 3(b); as amended by the Digital Omnibus on AI (in force 27 July 2026). [High-risk obligations for Annex III systems apply from 2 December 2027.]

See it on your own assignments

Eduface drafts criterion-level grades and feedback for your lecturers to approve. Book a demo or start free.