RUBRICS · AI MARKING
How to Write a Rubric AI Can Mark Against: Templates and Examples for Higher Education
Across studies of AI grading, the most consistent finding is that rubric quality decides how well it works. The same is true for human markers. Here is how to write criteria and descriptors that an AI, a new marker and an external examiner all read the same way, with templates, before-and-after examples and a calibration method.
By Eduface · September 2026 · 15 min read
A module leader uploads a rubric to an AI grading tool and runs ten essays through it. The results are all over the place. One essay she would give a 2:1 comes back as a first; another comes back as a 2:2. She concludes that the AI cannot mark her subject. Then she looks at her rubric. The criterion worth 30% of the mark says “Critical analysis: excellent / good / satisfactory / weak”. Four words, no description. Her two most experienced colleagues, asked separately, would read that criterion differently too.
How do you write a rubric that AI can mark against?
Write a rubric that a new human marker could apply without asking you anything. Use four to six criteria, each measuring one thing. For every criterion, describe what the work looks like at each level in observable terms, not with adjectives like “good” or “excellent”. Add weights, mark any pass/fail criteria, and test the rubric on five to ten scripts you have already marked. A rubric that works for a new marker works for AI.
Why does the rubric matter so much for AI marking?
An AI grader has only two things to go on: the student’s work and your rubric. It does not know what your department means by “critical”, what last year’s cohort produced, or what your external examiner complained about. Everything your experienced markers carry in their heads has to be in the rubric, or the AI cannot use it.
That is not a new problem. It is the same problem human marking has always had, made visible.
Human markers disagree more than we like to admit. A briefing paper from Lancaster University quotes the Higher Education Academy’s 2018 conclusion that “although the sector has produced vast quantities of agreed written standards, achieving a shared interpretation of their meaning is a very different matter. Repeated studies over many years demonstrate considerable inconsistency in academics’ judgements about student performance.”¹
Experienced markers often decide first and check the criteria afterwards. When Bloxham, Boyd and Orr asked twelve UK lecturers to think aloud while marking, they found markers typically made a holistic judgement first and then used the written criteria to refine it, drawing on tacit standards built from experience and on comparisons between scripts.² A human can get away with that. An AI cannot, because it has no tacit standard. It has your rubric.
The evidence on rubrics is clear about what helps. Jonsson and Svingby’s review of 75 studies concluded that rubrics improve the reliability of scoring, especially when they are analytic, topic-specific, and supported by exemplars or rater training.³ Brookhart’s later review of rubric use in higher education adds a definition worth holding on to: a rubric needs both criteria and descriptions of performance at each level. Without the descriptions, it is a rating scale, not a rubric.⁴
So when an AI grading tool gives erratic results, the first place to look is almost always the rubric. And when you fix the rubric for the AI, you fix it for your markers too.
What does a good rubric for AI marking look like?
Seven properties separate a rubric that AI and humans can apply consistently from one they cannot.
Property
What it means
Why it matters for AI
1. One thing per criterion
“Use of evidence” and “structure” are separate criteria, not combined
A combined criterion forces the AI to average two judgements you wanted to see separately
2. Observable descriptors
Each level describes what the work does, not how good it is
“Excellent analysis” means nothing to a model; “compares at least two interpretations and explains why they differ” is checkable
3. Levels that differ in one direction
Each level adds something the level below lacks
The model can place work on a ladder instead of guessing between adjectives
4. A sensible number of levels
Usually four or five
More levels mean finer distinctions, and more room for disagreement between any two markers
5. Explicit weights
Each criterion has a stated weight
The model can combine criterion scores the way you intend
6. Gateway criteria where needed
Criteria that must be met to pass, regardless of the total
Strong writing cannot average away a safety, integrity or core-competence failure
7. Tested on marked work
Checked against five to ten scripts with agreed marks
The only way to find the descriptors that read differently than you intended
Table 1: Seven properties of a rubric that AI and human markers can apply consistently.
On the number of levels, there is one piece of recent evidence worth knowing. A 2026 study of rubric-conditioned grading by a large language model found that its agreement with human graders dropped when the rubric moved from two levels to five: accuracy fell from 76% to 57%.⁵ That study used short science answers and a general-purpose model, so it does not transfer directly to essays. But it points the same way as the human evidence: every extra level is an extra boundary someone has to judge. Use as many levels as your grading scheme needs, and make every boundary between them explicit.
How do you turn a vague criterion into one an AI can mark?
This is the most useful skill in rubric writing, and it is easier to show than describe. Here are three common vague criteria and how to rewrite them.
### Example 1: Critical analysis
Before:
Critical analysis (30%): Excellent / Good / Satisfactory / Weak
After:
Level
Descriptor
First (70+)
Compares at least two competing interpretations or theories, explains why they differ, and uses that comparison to reach and defend a position of the student’s own
2:1 (60-69)
Compares at least two interpretations and reaches a position, but the reasons for preferring one are asserted more than argued
2:2 (50-59)
Presents more than one interpretation, but mostly in sequence, with limited comparison and no clear position
Third (40-49)
Describes one interpretation or summarises sources without evaluating them
Fail (<40)
No analysis; the work is description or opinion without reference to evidence
What changed: every level now describes something you could point to in the text. A marker, human or AI, can find the comparison or not find it.
### Example 2: Use of evidence
Before:
Use of evidence (25%): Uses a wide range of relevant sources.
After:
Level
Descriptor
4
Every main claim is supported by a relevant, credible source; sources are evaluated (method, scope or limitation noted) and applied to the argument, not just cited
3
Most main claims are supported; sources are applied to the argument, with occasional evaluation
2
Sources are cited but mostly summarised rather than applied; some main claims are unsupported
1
Few sources, or sources cited without connection to the argument
What changed: “wide range” is gone. Range is not the point. Whether each claim is supported, and whether sources are used rather than listed, is the point, and both are observable.
### Example 3: Structure and communication
Before:
Structure (15%): Well structured and clearly written.
After:
Level
Descriptor
4
The introduction states the question and the answer; each section advances the argument; the conclusion answers the question and follows from the analysis
3
Clear structure; one or two sections do not obviously advance the argument; the conclusion answers the question
2
Structure is present but the argument is hard to follow; the conclusion summarises rather than answers
1
No discernible structure; the question is not answered
What changed: “well structured” became questions a reader can check. Notice that grammar and spelling are not in this criterion. If they matter, give them their own small criterion, so a strong argument in imperfect English is not penalised twice.
Adjective only
Excellent / Good / Satisfactory / Weak
Wide range of relevant sources
Well structured and clearly written
Observable descriptor
Describes what the work actually does at each level
Draws on sources the marker can name, and applies each one to the argument
Each section opens with the claim it defends and closes by carrying it forward
The single most effective rubric edit: replace adjectives with descriptions of what the work does.
Figure 1: The single most effective rubric edit: replace adjectives with descriptions of what the work does.
Which type of rubric should you use?
There are three main types, and each suits a different job.
Type
What it looks like
Best for
Watch out for
Analytic
Several criteria, each with its own levels and descriptors
Summative marking, moderation, AI marking, and feedback students can act on
Takes longer to write; too many criteria fragment the judgement
Holistic
One overall scale with a description of each grade band
Quick overall judgements, short pieces, experienced teams
Hard to moderate; gives the student little to act on; hard for AI to explain
Single-point
One description of the expected standard per criterion, with space for feedback above and below
Formative feedback on drafts
Not designed to produce a grade on its own
Table 2: Analytic, holistic and single-point rubrics compared.
For AI marking of summative work, use an analytic rubric. It is what lets an AI score each criterion separately and explain each score, and it is what reviews of rubric research associate with more reliable marking.³ For formative feedback on drafts, a single-point rubric can work well, because the goal is guidance, not a grade. Panadero and Jonsson’s review of rubrics used for formative purposes links them to greater transparency for students and better self-regulation of their own work.⁶
A rubric template you can adapt
Here is a complete analytic rubric for a 2,500-word undergraduate essay in the social sciences. Adapt the criteria and wording to your discipline, keep the structure.
Criterion (weight)
70+
60-69
50-59
40-49
<40
Answering the question (15%)
States a clear answer in the introduction and sustains it throughout
Clear answer, occasionally drifts
Answer implied rather than stated
Addresses the topic, not the question
Does not address the question
Knowledge and understanding (20%)
Accurate, precise use of key concepts, including their limits
Accurate use of key concepts
Mostly accurate; some imprecision
Significant errors or gaps
Fundamental misunderstanding
Critical analysis (30%)
Compares competing interpretations, explains differences, defends own position
Compares interpretations; position asserted more than argued
Interpretations presented in sequence, limited comparison
Description or summary without evaluation
No analysis
Use of evidence (20%)
Every main claim supported; sources evaluated and applied
Most claims supported and applied
Sources mostly summarised; some claims unsupported
Few sources or disconnected from argument
No meaningful evidence
Structure and conclusion (10%)
Each section advances the argument; conclusion answers and follows
Clear, with minor digressions
Hard to follow in places; conclusion summarises
Weak structure
No discernible structure
Academic conventions (5%)
Referencing consistent and complete; clear, accurate prose
Minor referencing or language slips
Several slips that do not obscure meaning
Frequent errors that obscure meaning
Referencing absent
Table 3: A template analytic rubric for an undergraduate essay. Illustrative; adapt to your discipline and grading scheme.
Three notes on this template. The heaviest weight sits on the criterion that most distinguishes strong work (analysis), not on the easiest to check (conventions). Every cell describes something observable. And language accuracy has its own small criterion, so it cannot dominate the mark.
When do you need gateway criteria?
Some assessments contain one thing that must be right regardless of everything else: safe clinical practice in nursing, correct application of the law to the facts in a law problem question, ethical handling of data in a research proposal. In a weighted rubric, strong scores elsewhere can average out a failure there.
The fix is a gateway criterion: a criterion that must reach a threshold for the work to pass, whatever the total. Describe the failure specifically (“describes administering a medication without checking the prescription and does not recognise this as unsafe”), so that it can be flagged. An AI grader will then highlight the passage and explain the concern; the lecturer makes the decision. We walk through this for nursing in AI grading software for nursing assignments.
How do you test a rubric before a live cohort?
Never put a new rubric in front of a whole cohort, human or AI, without testing it. This calibration takes about two hours and saves weeks of disputes.
1
Choose 5-10 marked and moderated scripts (include a fail and a borderline)
2
Run them through the AI with the new rubric
3
Compare criterion by criterion with the agreed marks
4
Find the descriptors where the AI and the markers disagree
5
Rewrite those descriptors
6
Run again until the gaps are small
Most disagreements come from two or three descriptors. Fix those and the rest usually follows.
Step 1: choose the scripts. Five to ten scripts from last year, already marked and moderated. Include at least one fail, one borderline, and one first. Jonsson and Svingby’s review found that exemplars support reliable use of rubrics, and these scripts are your exemplars.³
Step 2: run them. With Eduface, set up the assignment with the brief, the rubric and your feedback instructions, and choose the discipline model closest to your module. Then run the scripts through the Paper Grader.
Step 3: compare criterion by criterion. Do not compare only the total. A total that matches can hide one criterion scored too high and another too low.
Step 4: rewrite the descriptors where they disagree. If the AI scores “use of evidence” higher than your markers, read the descriptor again. Usually it rewards citing sources rather than applying them. Make the difference explicit.
Step 5: run again. Two or three rounds is typical. In a pilot at one UK university, markers who spent time tuning the model to their own standards saw the AI’s suggestion come within 98% of their grades on average, against 94% overall. Accuracy there means how small the gap is between suggestion and marker. The difference is almost entirely in the rubric and instructions.
The new-marker test
Give your rubric to a colleague who has never taught the module. Ask them to mark two of your calibration scripts using only the rubric. Every question they ask you (“what counts as a comparison?”, “is this evidence or example?”) is a descriptor that needs rewriting. An AI grader will stumble in exactly the same places.
What about feedback instructions?
A rubric tells the AI what to judge. Feedback instructions tell it how to talk to the student. They are separate, and both matter.
Good feedback instructions say:
What to prioritise: “Comment first on the criterion with the lowest score.”
What to avoid: “Do not comment on referencing format; the library covers it.”
How to phrase it: tone, length, and whether to ask questions or give directions.
What the student should do next: “End each comment with one concrete action for the next draft.”
Eduface offers four feedback styles on top of your instructions: Reflective and Socratic, Constructive and Direct, Went Well and Needs Improvement, and Supportive and Encouraging.
What if you do not have a rubric at all?
Many modules have marking criteria in the handbook but no rubric with levels. You have two options.
Write one from your criteria, using the template above and the before-and-after examples.
Let Eduface generate a first draft from your assignment brief, then edit it. A generated rubric is a starting point, not a finished product. Read every descriptor, apply the new-marker test and calibrate on marked scripts before using it on a live cohort.
What are the most common rubric mistakes?
Mistake
What goes wrong
Fix
Adjectives instead of descriptions
Markers and AI read “good” differently
Describe what the work does at each level
Two things in one criterion
Scores average what you wanted separated
Split the criterion
Language accuracy everywhere
Second-language writers penalised in every criterion
One small criterion for language
Too many criteria (10+)
Judgement fragments; weights become meaningless
Four to six criteria
Levels that overlap
“Some evidence” vs “limited evidence”
Each level adds one observable thing
No gateway where one is needed
Strong writing masks a critical failure
Add a pass/fail criterion for the thing that must be right
Never tested
Problems discovered by students and appeals
Calibrate on 5-10 marked scripts
Table 4: Common rubric mistakes and how to fix them.
For a wider view of fairness to students writing in a second language, see does AI grading disadvantage non-native English speakers?
Frequently asked questions
Does AI grading need a different kind of rubric from human marking?
No. It needs the rubric human marking always should have had: separate criteria, observable descriptors at each level, explicit weights and a test on marked work. A rubric a new marker can apply without asking questions is a rubric an AI can apply too.
How many criteria should a rubric have?
Four to six for most essays and reports. Fewer and you combine things that should be separate; more and the judgement fragments and weights lose meaning.
How many performance levels should a rubric have?
Enough to match your grading scheme, usually four or five. Each extra level adds a boundary that has to be judged, so make every boundary explicit. A 2026 study found a general-purpose model’s agreement with human graders fell as the number of levels rose.
Can AI write the rubric for me?
It can draft one. Eduface generates a rubric from the assignment brief if there is none. Treat the draft as a starting point: rewrite vague descriptors, test it with a colleague and calibrate it on marked scripts before use.
Why does my AI grader score higher than my markers on one criterion?
Usually because the descriptor rewards something easier to spot than what you meant, such as citing sources rather than applying them. Compare criterion by criterion on marked scripts, find the descriptor, and make the difference explicit.
Should grammar and spelling be part of every criterion?
No. Give language accuracy its own small criterion. Otherwise a strong argument in imperfect English is penalised several times, which is unfair to students writing in a second language.
References
1. Lancaster University. We need to talk about marking. Centre for Engagement and Development in Academia (CEDA) briefing paper. [Quotes Higher Education Academy (2018) on inconsistency in academics’ judgements; cites McConlogue (2020).]
2. Bloxham, S., Boyd, P., & Orr, S. (2011). Mark my words: the role of assessment criteria in UK higher education grading practices. Studies in Higher Education, 36(6), 655–670. [Twelve lecturers thinking aloud; holistic judgement first, criteria used to refine it.]
3. Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational Research Review, 2(2), 130–144. [Review of 75 studies; rubrics improve reliability, especially when analytic, topic-specific and supported by exemplars or rater training.]
4. Brookhart, S. M. (2018). Appropriate criteria: key to effective rubrics. Frontiers in Education, 3, 22. [A rubric requires criteria and descriptions of performance levels; rating-scale language is less useful for learning.]
5. Deng, Farber, Lee, & Tang. (2026). Rubric-conditioned LLM grading: alignment, uncertainty, and robustness. arXiv:2601.08843. [Accuracy fell from 76% to 57%, and Cohen’s kappa from 0.51 to 0.34, moving from two-level to five-level grading of short science answers.]
6. Panadero, E., & Jonsson, A. (2013). The use of scoring rubrics for formative assessment purposes revisited: a review. Educational Research Review, 9, 129–144.
Test your rubric on real scripts
Upload a rubric and a handful of marked scripts, and see where the AI and your markers disagree. Book a demo or start free.