COMPARISON · UK AI MARKING

Graide vs KEATH vs Eduface: Which AI Marking Tool Fits a UK University? (2026)

Three AI marking tools come up again and again in UK conversations: Graide, KEATH and Eduface. They are built for different kinds of assessment and different kinds of institution. Here is what each one does, what the Jisc pilot found about AI marking in practice, and how to choose.

By Eduface · September 2026 · 16 min read

Jisc’s AI in marking and feedback pilot has put AI marking on the agenda of UK colleges and universities that had not looked at it before. Two of the three purpose-built tools in that pilot were Graide and KEATH. Eduface was not in the pilot, but it is an approved supplier on the Jisc/CHEST framework, and it is the third name that comes up when institutions compare options. So the question we hear most in UK conversations is simple: what is the difference between them?

What is the difference between Graide, KEATH and Eduface?

Graide, now part of Inspera, learns from how an educator marks one response and reuses that feedback on similar responses, which makes it strongest for structured, step-by-step STEM work, including handwritten maths. KEATH marks written assignments against templated or customised rubrics with teachers able to change any score or feedback, and puts particular emphasis on the Extended Project Qualification. Eduface drafts criterion-level grades and feedback on university essays, reports and open exam answers with six discipline models, holds every grade until a lecturer approves it, and adds oral examination.

We are Eduface, so read this with that in mind. We have described Graide and KEATH only in the terms each company uses about itself, or as reported by Jisc and the University of Birmingham, with the date of each source. Both products change, so check the current state with each vendor before you decide.

What did the Jisc pilot find about AI marking?

Before comparing tools, it helps to know what the largest UK trial of AI marking has found so far, because its findings apply to every tool in this article.

Jisc’s National Centre for AI ran a year-long pilot of AI in marking and feedback from 2025. It tested three purpose-built tools, Graide, KEATH and TeacherMatic, alongside general-purpose AI tools such as ChatGPT, Gemini and Copilot, guided by a best-practice toolkit. According to a May 2026 write-up of Jisc’s findings, 38 UK colleges and universities took part.¹

Two findings from that write-up stand out:

Keeping the human in the loop is harder in practice than on paper. The pilot found that “keeping the human in the loop when using AI can be hard to maintain in practice”, and that methods such as parallel marking helped to build it into the workflow.¹

Formative feedback is the safest place to start. Across the pilots, formative assessment was identified as “the most effective and low risk use case for AI in marking and feedback”.¹

Both points matter when you compare tools. The first is a question about workflow design: does the tool make human review the default, or something a busy marker can skip? The second is a question about scope: does the tool support draft feedback, not only final marks?

An earlier Jisc pilot focused on Graide alone. It ran from March 2022 to January 2023, interviewed nine universities and had five test Graide hands-on, with particular confidence reported for mathematical topics.²

What is Graide?

Graide was developed by 6 Bit Education and first evaluated at the University of Birmingham. In June 2026 it became part of Inspera, the digital assessment company.³ Graide describes itself as an explainable AI model, and its website positions it explicitly against general-purpose chatbots.

How it works: Replay Grading. Graide’s core idea, as described in its University of Birmingham research report, is that the educator grades a response once and the system learns from it. When a new submission arrives, Graide checks whether it has seen the same approach before. If it has, it can grade it automatically or suggest the feedback the educator gave last time. If the approach is new, it goes to the educator, whose feedback is then learned. The goal, in Graide’s words, is that educators “do not have to grade the same answer twice”.⁴

What it is built for. The Birmingham report describes students entering work through a visual editor with an inline maths editor, a markdown editor that supports LaTeX, stylus handwriting, or optical character recognition on handwritten paper work. That tells you where Graide is strongest: structured, step-by-step responses in STEM, where many students take the same path through a problem and a few take unusual ones.

What the research found. In Graide’s University of Birmingham study, published in February 2022, researchers compared paper grading with grading in Graide for 10 questions with 172 submissions per question. Median grading time fell by 74%, and the number of words of feedback per script rose by a factor of 7.2, from 23 to 166.⁴ Graide’s other pages quote different time-saving figures for the same study, so if you use a number, cite the report itself.

Where it fits less well. Replay Grading depends on many students giving similar responses. Long essays, where every student’s argument is different, give it less to reuse. That is not a weakness for STEM problem sets. It is a sign of what the tool was designed for.

What is KEATH?

KEATH is a UK AI assessment and marking platform. On its own website it describes grading “for all forms of assignments, through user selected, KEATH templated rubrics”, with teachers able to “select recommended rubrics and tweak as needed”.⁵

Human in the loop. KEATH describes a “Human in the Loop system” that “allows teachers to directly modify scores and feedback, and to customize grading rubrics”, and says it emphasises “the pivotal role of teachers”.⁵

A focus on the Extended Project Qualification. KEATH’s site gives the EPQ particular prominence, describing it as “the most difficult assessment” it has mastered. The EPQ is a Level 3 qualification, typically taken alongside A-levels, which makes KEATH a natural fit for schools and colleges as well as universities.

Integrations. KEATH says it integrates directly with learning management systems and can build custom integrations with other systems.

Recognition. KEATH lists a number of sector awards and shortlistings on its website, and it was one of the three purpose-built tools in Jisc’s AI in marking and feedback pilot.¹

KEATH’s public materials also include accuracy claims. We have not been able to verify them against a primary source, so we have not repeated them here. Ask KEATH directly how its accuracy was measured, on what kind of work, and against whose marks.

What is Eduface?

Eduface is an AI assessment platform built for written and oral assessment in higher education.

How it works. The lecturer sets up an assignment with the brief, the rubric and instructions for how feedback should be written; if there is no rubric, Eduface generates one from the brief. Every submission that arrives through the LMS is read against the rubric by the Paper Grader, which annotates the text, scores every criterion and writes the reasoning. Three independent AI agents evaluate each submission without seeing each other’s results, and a fourth reconciles them into one draft. The lecturer reviews, edits and approves. Nothing reaches the student before that.

What it is built for. University essays, reports, dissertations and open exam answers, across disciplines. The Paper Grader runs on six discipline models (Law, Economics, Social Sciences, STEM, Humanities and Health Sciences), each trained on the conventions of its field. Oral Examination (early access) runs structured spoken assessment, and Academic Integrity (beta) asks students to defend their own submitted work.

How it keeps the human in the loop. This is where the Jisc finding matters. In Eduface, lecturer approval is not a setting. It is the only route by which a grade reaches a student. Institutions can also make blind mode compulsory: the lecturer marks first, and only then sees the AI draft for comparison. That is, in effect, the parallel marking the Jisc pilot found effective, built into the product.

Evidence. In UK pilots, lecturers changed an average of 5% of each final grade the AI drafted. In a pilot at one UK university (435 submissions, six modules, 13 markers), the AI’s suggested grade came within 94% of the marker’s grade on average, and within 98% for markers who had tuned the model to their standards; accuracy means how small the gap is between suggestion and marker. In an independent student test of eight tools, Eduface averaged ±0.15 deviation from lecturer grades across five psychology papers.⁶

What it does not do. Eduface does not read handwritten scripts or maths notation. For handwritten STEM work, a tool built for it is the better choice.

How do Graide, KEATH and Eduface compare side by side?


Graide (part of Inspera)

KEATH

Eduface

Core approach

Learns from an educator’s grading and reuses it on similar responses (Replay Grading)

Grades against templated or customised rubrics; teachers modify scores and feedback

Drafts criterion-level grades and feedback against the lecturer’s rubric; three AI agents plus reconciliation

Strongest for

Structured, step-by-step STEM work, including handwritten maths

Written assignments, with emphasis on the EPQ

University essays, reports, dissertations and open exam answers across disciplines

Handwriting and maths input

Yes: maths editor, LaTeX, stylus and OCR

Not a stated focus

No

Discipline-specific models

Not described

Not described

Six discipline models

Human review

Educator grades new approaches; system reuses them

Teachers can modify any score or feedback

Lecturer approval required for every grade; blind mode can be made compulsory

Formative feedback on drafts

Feedback on submitted work

Feedback on submitted work

Draft-by-draft feedback with progress tracking, four feedback styles

Oral assessment

No

No

Oral Examination (early access), Academic Integrity (beta)

In the Jisc AI marking pilot

Yes

Yes

No

UK procurement framework

Ask the vendor

Ask the vendor

Jisc/CHEST (UK) and HEAnet (Ireland)

Published evidence

University of Birmingham study (2022): median grading time -74%, feedback words x7.2

Ask the vendor for the method behind its accuracy claims

UK pilots: 5% of each grade changed; one-university pilot: 94% average accuracy, 98% after tuning

Table 1: Graide, KEATH and Eduface compared, based on each vendor’s public materials and Jisc and University of Birmingham reports, September 2026.

Graide

Structured, step-by-step responses: maths and problem sets

Spans school through university

KEATH

Centre of the range, lower levels

School and college, with an EPQ focus

Eduface

Open, argued work: essays, reports, dissertations

University levels 4 to 7

Oral examination as well as written

Indicative, based on each vendor’s stated focus.

Figure 1: Indicative positioning of the three tools by type of response and level of study.

How should a UK university choose between them?

Start with the assessment, not the tool. The right choice depends mostly on three questions.

1. What kind of work are you marking?

Handwritten or step-by-step STEM work with many similar answers: Graide’s Replay Grading is built for exactly this.

EPQ projects and written work in a school or college setting: KEATH puts its emphasis here.

University essays, reports and dissertations across several disciplines, or oral assessment: Eduface is built for this.

2. How do you want to keep the lecturer in charge?

The Jisc pilot found that human review is hard to sustain in practice. Ask each vendor what happens when a busy marker does not review: is anything released to students? Can the institution require the marker to mark independently first? In Eduface, nothing is released without approval, and blind mode can be made compulsory.

3. Where do you want to start: formative or summative?

Jisc found formative feedback to be the lowest-risk, most effective starting point. If that is where you want to begin, check how each tool handles feedback on drafts, not only on final submissions.

If your first use case is…

Consider

Faster marking of first-year maths or engineering problem sets

Graide

EPQ marking in a sixth form or FE college

KEATH

Consistent marking and feedback on large essay-based modules

Eduface

Draft feedback before submission on essays and reports

Eduface

Verifying authorship with an oral follow-up

Eduface

A faculty with both handwritten STEM and essay-based programmes

Graide for STEM plus Eduface for written work, or a pilot of both

Table 2: A starting point for choosing, by first use case.

How do you run a fair pilot of more than one tool?

If you are shortlisting two or three tools, the most useful thing you can do is test them on the same work, with the same rubric, against the same agreed marks. Demos show each tool at its best on material the vendor chose. A side-by-side pilot shows each tool on your material.

Choose the material. Pick one assignment type you mark in volume. Take ten scripts from last year that have been marked and moderated, including at least one fail, two borderline scripts and one first. If you are comparing a STEM tool with an essay tool, run two separate pilots: one on problem sets, one on essays. A single pilot on one type of work will favour whichever tool was built for it.

Fix the inputs. Every tool gets the same rubric, the same assignment brief and the same marked scripts. Record any setup help each vendor gives, so you compare tools rather than onboarding.

Score each tool on the same criteria. Here is a scoring template you can adapt.

Criterion

Weight

How to score it

Closeness to agreed marks

30%

Average gap between the tool’s draft and your moderated mark, criterion by criterion as well as overall

Feedback quality

20%

Do comments identify the same strengths and weaknesses your markers did? Would a student know what to do next?

Human oversight

15%

Can anything reach a student without approval? Can independent marking before the AI suggestion be enforced?

Marker time

10%

Minutes per script to review and approve, timed, not estimated

Consistency

10%

Run two scripts twice; how much does the output change?

Data and compliance

10%

Processing location, DPA, training on student data, audit trail

Fit with your LMS and procurement route

5%

LTI integration, grade passback, framework availability

Table 3: A scoring template for comparing AI marking tools in a pilot. Illustrative weights; agree your own before you start.

Let markers score blind where you can. If markers review the tools’ feedback without knowing which tool produced it, you remove brand from the judgement.

Decide the weights before you see results. Otherwise, the weights tend to drift towards whichever tool people liked in the demo.

What about general-purpose AI tools like ChatGPT?

Jisc’s pilot also tested general-purpose AI tools such as ChatGPT, Gemini and Copilot, guided by a best-practice toolkit.¹ Many lecturers already use them informally, so it is worth being clear about how they compare with purpose-built tools.

In our independent student test of eight grading tools on papers with known lecturer grades, general assistants were consistently too generous. On five psychology papers, their average deviation from the lecturer’s grade ranged from ±0.7 (Copilot) to ±1.2 (Gemini), against ±0.15 for Eduface. On a Dutch-language constitutional law essay graded 4.4 by the lecturer, ChatGPT returned 7.1 and Claude 7.2.⁶ They responded to how well the student wrote, not how well the student reasoned. They also have no rubric interface, no approval workflow and, in default settings, process student work outside the EU.

General tools are useful for drafting feedback templates, brainstorming questions or explaining concepts. For marking real student work, a purpose-built tool with enforced human review is the safer choice, whichever one you pick.

What should you ask every vendor?

Whichever tools you shortlist, ask all of them the same questions, in writing.

On accuracy: How was accuracy measured, on what kind of work, against whose marks, and by whom? What is the run-to-run consistency on the same submission?

On human oversight: Can a grade reach a student without a lecturer approving it? Can the institution require independent marking before the AI suggestion is shown? What does the audit trail record?

On data: Where is student work processed? Is it used to train any model? Is there a Data Processing Agreement? Which sub-processors are involved?

On regulation: How does the tool meet the EU AI Act’s requirements for high-risk AI in education, which apply to Annex III systems from 2 December 2027 following the Digital Omnibus on AI?⁷

On procurement: Are you available through Jisc/CHEST or HEAnet, or does this need a separate tender?

On fit: Can we run three of our own marked scripts through the tool before we sign, and compare the result with our agreed marks?

Run the same test on every shortlisted tool

Choose five scripts you have already marked and moderated, including a borderline one and a fail. Run them through each tool with the same rubric. Compare each tool’s draft with your agreed marks, not with how convincing its feedback sounds. It takes an afternoon and tells you more than any demo.

Frequently asked questions

Was Eduface part of the Jisc AI marking pilot?

No. The three purpose-built tools in Jisc’s AI in marking and feedback pilot were Graide, KEATH and TeacherMatic. Eduface is a separate supplier, approved on the Jisc/CHEST framework, which UK institutions use to buy without running their own tender.

Is Graide still an independent company?

Graide became part of Inspera, the digital assessment company, in June 2026. Check with Inspera how Graide is licensed and supported now.

Which AI marking tool is best for essays?

For university essays and reports, look for a tool that reads each submission against your rubric, explains every criterion score and holds the grade until a lecturer approves it. Eduface is built for this. Graide’s Replay Grading is designed for structured responses where many students take similar approaches, which is less common in essays.

Which AI marking tool is best for maths?

For handwritten or step-by-step maths, Graide’s input options (maths editor, LaTeX, stylus and OCR) and Replay Grading are designed for exactly that kind of work. Eduface does not read handwritten scripts or maths notation.

Can we use more than one AI marking tool?

Yes. A faculty with both STEM problem sets and essay-based programmes may use different tools for each. Keep the human oversight rules and data processing terms the same across all of them.

How do we avoid the “human in the loop” slipping in practice?

Make review structural rather than optional: nothing released without approval, and independent marking before the AI suggestion is shown for summative work. The Jisc pilot found parallel marking helped. Eduface’s blind mode does this by design.

Where should a university start with AI marking: formative or summative?

Jisc’s pilot identified formative assessment as the most effective and lowest-risk use case for AI in marking and feedback. A practical route is to start with feedback on drafts, where the stakes are lower and students benefit straight away, then move to AI-assisted summative marking with blind mode once markers trust how the tool reads the rubric.

Which AI marking tool is best for further education colleges?

It depends on the qualifications you mark. KEATH puts particular emphasis on the EPQ, which many colleges and sixth forms deliver. For vocational qualifications regulated by Ofqual, check each vendor against Ofqual’s expectations on human oversight, which we cover in is AI allowed to mark assignments?

References

1. FE News. (2026, 21 May). New findings from Jisc highlight the benefits of a collaborative approach to AI in assessment. [Year-long AI in marking and feedback pilot; Graide, KEATH and TeacherMatic plus general-purpose AI tools; 38 UK colleges and universities; findings on human-in-the-loop and formative use.]

2. Jisc National Centre for AI. (2023, 14 June). National Centre for AI Graide pilot overview. [Pilot March 2022 to January 2023; nine universities interviewed, five tested Graide.]

3. Inspera. (2026, June). Graide acquisition [Press release].

4. 6 Bit Education (Graide). (2022, 21 February). University of Birmingham: Graide efficacy research. [10 questions, 172 submissions per question; median grading time reduced by 74%; feedback words increased by a factor of 7.2, from 23 to 166.]

5. KEATH. (2026). Company website, accessed 28 September 2026. [Templated and customisable rubrics; “Human in the Loop system”; EPQ focus; LMS integration.]

6. Eduface. (2026). The Complete Guide to AI Grading Tools for Higher Education. [Independent student test of eight tools on six papers with known lecturer grades.]

7. European Parliament and Council of the EU. (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act), Annex III, point 3(b); as amended by the Digital Omnibus on AI (in force 27 July 2026).

See it on your own rubric

Bring a rubric and five marked scripts, and compare the draft grades with your markers’. Book a demo or start free.