COMPLIANCE · HUMAN OVERSIGHT
Human-in-the-Loop Marking: How to Use AI for Grading Without Breaking Ofqual Rules
Everyone says their tool has one. Fewer can tell you what theirs does, or why it satisfies a regulator rather than just sounding reassuring in a sales deck.
By Eduface · July 2026 · 9 min read
Human in the loop is one of those phrases that has been used so often, by so many vendors, that it has started to lose its edges. Everyone says their tool has one. Fewer can tell you exactly what theirs actually does, or why it would satisfy a regulator rather than just sound reassuring in a sales deck. That gap matters, because a vague version of human oversight will not get you through an Ofqual audit, and a genuine version isn’t actually that hard to build a workflow around once you know what you’re looking for.
What the phrase needs to actually guarantee
Our companion piece on Ofqual’s marking rules covers the regulatory baseline: AI cannot be the sole mechanism for awarding a grade, and a human judgement has to be part of the decision. What that means in practice is where things get interesting, because there’s a useful philosophical framework for thinking about this that comes from a very different field entirely.
Filippo Santoni de Sio and Jeroen van den Hoven, writing about autonomous weapons systems rather than exam marking, proposed that meaningful human control over an automated system needs two conditions to hold. The first, which they call the tracking condition, is that the system’s outputs need to respond to the actual reasoning and judgement of the humans responsible for it, not just to statistical patterns. The second, the tracing condition, is that there has to be a real human being somewhere in the chain who can be identified as responsible for the outcome (Santoni de Sio & van den Hoven, 2018). It’s an odd source to borrow from for an education blog, but the logic transfers cleanly. A marking tool passes the tracking test if an assessor’s actual judgement shapes the final grade, not just the AI’s pattern matching. It passes the tracing test if there’s a specific person who reviewed that specific submission and can be asked why.
A simpler version of the same test: would an assessor notice, and be able to act on it, if the AI got something wrong? If the grade sits in draft form, the assessor can see the reasoning behind it, and changing the mark is a normal and expected part of the process, the answer is yes. If the grade goes out automatically and a correction only happens when someone stumbles across a problem nobody flagged, the answer is no, whatever the product page says.
Four things a genuinely compliant workflow needs
A draft state, not a released one. The AI’s output should sit as a proposal until an assessor takes a deliberate action to approve it. If a learner can see a grade before that approval happens, the human step has already been skipped in substance.
Visible reasoning, not just a number. Nobody can meaningfully review a decision they can’t see the logic behind. Criterion-by-criterion reasoning is what makes a review real rather than symbolic, and it echoes a much older point from the formative assessment literature: feedback only changes anything when the reviewer can actually see the gap between where the work is and where it needs to be, not just the final verdict (Hattie & Timperley, 2007).
A normal path to disagreement. If overriding the AI is technically possible but awkward, buried three menus deep, or quietly discouraged by the interface, assessors will do it rarely even when they should do it often.
A record of what happened. If a grade is ever appealed, you want to show what the AI proposed, what the assessor changed, and why. That’s what turns a human was technically involved into actual evidence that a human decision took place.
The failure mode nobody designs for on purpose
Here’s the uncomfortable part. The most common way human oversight quietly stops being real isn’t fraud, and it isn’t laziness in any simple sense. It’s habituation, and there’s decades of research on exactly this pattern from human factors science. Raja Parasuraman and Dietrich Manzey’s well known review of automation complacency found that people monitoring a reliable automated system become measurably worse at catching its failures over time, and crucially, that this isn’t something you can train your way out of with more practice or more experience (Parasuraman & Manzey, 2010). One of the more striking findings in that literature: operators working with a consistently high reliability system were around half as likely to catch a genuine failure as operators working with a system they knew to be less reliable. Trust, in other words, is precisely what erodes vigilance.
Translate that into a marking context and the risk is obvious. An assessor who reviews a hundred AI drafted grades and finds the first ninety accurate will, entirely predictably, start clicking approve without genuinely re-reading the ninety-first. That’s not a character flaw. It’s what the research says happens to essentially everyone under these conditions, which is exactly why Ofqual’s concern about bias and inaccuracy in AI marking isn’t a hypothetical worry, it’s a documented human tendency waiting for the right conditions to show up.
What actually pushes back against that drift
A tool that presents every output with identical, uniform confidence makes the habituation problem worse, because there’s no signal telling the assessor when to slow down. A tool that occasionally flags genuine uncertainty, a criterion where the reasoning is thinner, a submission that sits close to a grade boundary, gives the assessor something to actually pay attention to, rather than asking them to maintain uniform vigilance across every single case indefinitely, which the automation literature is fairly clear humans are bad at doing.
The second common failure is structural rather than behavioural. Some tools technically allow an assessor to change a grade, but release it to the learner automatically after a set period if nobody acts. That is not human in the loop marking. That’s automated marking with an opt-out window, and it doesn’t clear the bar Ofqual has set, however the marketing describes it.
A short checklist before you commit to a tool
Does the grade require an explicit approval action, with no automatic release path under any setting? Can the assessor see reasoning per criterion, not just a final number? Is overriding a grade a normal, easy action, not a workaround? Is there an audit trail pairing the AI’s original proposal with the assessor’s final decision? Does the tool ever signal its own uncertainty, or does every output look equally confident? If a vendor can’t answer each of those specifically, that’s useful information in itself.
Frequently asked questions
Is human in the loop just a marketing phrase, or does it mean something specific?
It needs to mean something specific: a genuine review that can change the outcome, with a real record that it happened. The phrase alone guarantees nothing.
What’s the biggest practical risk with human-in-the-loop marking?
Habituation. Assessors who see a tool get things right most of the time predictably start approving without properly re-checking, which quietly turns real oversight into a formality, exactly as the automation complacency research would predict.
Does a high agreement rate between AI and assessors mean the review step is unnecessary?
No. It shows the tool is a strong starting point. The minority of cases where the assessor disagrees are exactly why the review step has to stay genuine.
What’s the one question that cuts through vendor marketing on this?
Whether a grade can reach a learner without explicit assessor approval, under any configuration or timeout. If there’s any path where that can happen, the oversight isn’t structurally guaranteed, whatever it’s called.
Sources
Santoni de Sio, F., & van den Hoven, J. (2018). Meaningful Human Control over Autonomous Systems: A Philosophical Account. Frontiers in Robotics and AI, 5, Article 15.
Parasuraman, R., & Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, 52(3), 381-410.
Hattie, J., & Timperley, H. (2007). The Power of Feedback. Review of Educational Research, 77(1), 81-112.
Ofqual (2026). Using AI in Marking: Why Technical Capability, Fairness, and Transparency All Matter. Ofqual blog, gov.uk.
Human control that is structural, not optional
Eduface holds every grade as a draft, shows reasoning per criterion, and logs every assessor decision. Book a demo or start free.