RUBRICS · AI MARKING

How to Write a Rubric AI Can Mark Against: Templates and Examples for Higher Education

Across studies of AI grading, the most consistent finding is that rubric quality decides how well it works. The same is true for human markers. Here is how to write criteria and descriptors that an AI, a new marker and an external examiner all read the same way, with templates, before-and-after examples and a calibration method.

By Eduface · September 2026 · 15 min read

A module leader uploads a rubric to an AI grading tool and runs ten essays through it. The results are all over the place. One essay she would give a 2:1 comes back as a first; another comes back as a 2:2. She concludes that the AI cannot mark her subject. Then she looks at her rubric. The criterion worth 30% of the mark says “Critical analysis: excellent / good / satisfactory / weak”. Four words, no description. Her two most experienced colleagues, asked separately, would read that criterion differently too.

How do you write a rubric that AI can mark against?

Write a rubric that a new human marker could apply without asking you anything. Use four to six criteria, each measuring one thing. For every criterion, describe what the work looks like at each level in observable terms, not with adjectives like “good” or “excellent”. Add weights, mark any pass/fail criteria, and test the rubric on five to ten scripts you have already marked. A rubric that works for a new marker works for AI.

Why does the rubric matter so much for AI marking?

An AI grader has only two things to go on: the student’s work and your rubric. It does not know what your department means by “critical”, what last year’s cohort produced, or what your external examiner complained about. Everything your experienced markers carry in their heads has to be in the rubric, or the AI cannot use it.

That is not a new problem. It is the same problem human marking has always had, made visible.

Human markers disagree more than we like to admit. A briefing paper from Lancaster University quotes the Higher Education Academy’s 2018 conclusion that “although the sector has produced vast quantities of agreed written standards, achieving a shared interpretation of their meaning is a very different matter. Repeated studies over many years demonstrate considerable inconsistency in academics’ judgements about student performance.”¹

Experienced markers often decide first and check the criteria afterwards. When Bloxham, Boyd and Orr asked twelve UK lecturers to think aloud while marking, they found markers typically made a holistic judgement first and then used the written criteria to refine it, drawing on tacit standards built from experience and on comparisons between scripts.² A human can get away with that. An AI cannot, because it has no tacit standard. It has your rubric.

The evidence on rubrics is clear about what helps. Jonsson and Svingby’s review of 75 studies concluded that rubrics improve the reliability of scoring, especially when they are analytic, topic-specific, and supported by exemplars or rater training.³ Brookhart’s later review of rubric use in higher education adds a definition worth holding on to: a rubric needs both criteria and descriptions of performance at each level. Without the descriptions, it is a rating scale, not a rubric.⁴

So when an AI grading tool gives erratic results, the first place to look is almost always the rubric. And when you fix the rubric for the AI, you fix it for your markers too.

What does a good rubric for AI marking look like?

Seven properties separate a rubric that AI and humans can apply consistently from one they cannot.

Property

What it means

Why it matters for AI

1. One thing per criterion

“Use of evidence” and “structure” are separate criteria, not combined

A combined criterion forces the AI to average two judgements you wanted to see separately

2. Observable descriptors

Each level describes what the work does, not how good it is

“Excellent analysis” means nothing to a model; “compares at least two interpretations and explains why they differ” is checkable

3. Levels that differ in one direction

Each level adds something the level below lacks

The model can place work on a ladder instead of guessing between adjectives

4. A sensible number of levels

Usually four or five

More levels mean finer distinctions, and more room for disagreement between any two markers

5. Explicit weights

Each criterion has a stated weight

The model can combine criterion scores the way you intend

6. Gateway criteria where needed

Criteria that must be met to pass, regardless of the total

Strong writing cannot average away a safety, integrity or core-competence failure

7. Tested on marked work

Checked against five to ten scripts with agreed marks

The only way to find the descriptors that read differently than you intended

Table 1: Seven properties of a rubric that AI and human markers can apply consistently.

On the number of levels, there is one piece of recent evidence worth knowing. A 2026 study of rubric-conditioned grading by a large language model found that its agreement with human graders dropped when the rubric moved from two levels to five: accuracy fell from 76% to 57%.⁵ That study used short science answers and a general-purpose model, so it does not transfer directly to essays. But it points the same way as the human evidence: every extra level is an extra boundary someone has to judge. Use as many levels as your grading scheme needs, and make every boundary between them explicit.

How do you turn a vague criterion into one an AI can mark?

This is the most useful skill in rubric writing, and it is easier to show than describe. Here are three common vague criteria and how to rewrite them.

### Example 1: Critical analysis

Before:

Critical analysis (30%): Excellent / Good / Satisfactory / Weak


After:

Level

Descriptor

First (70+)

Compares at least two competing interpretations or theories, explains why they differ, and uses that comparison to reach and defend a position of the student’s own

2:1 (60-69)

Compares at least two interpretations and reaches a position, but the reasons for preferring one are asserted more than argued

2:2 (50-59)

Presents more than one interpretation, but mostly in sequence, with limited comparison and no clear position

Third (40-49)

Describes one interpretation or summarises sources without evaluating them

Fail (<40)

No analysis; the work is description or opinion without reference to evidence

What changed: every level now describes something you could point to in the text. A marker, human or AI, can find the comparison or not find it.

### Example 2: Use of evidence

Before:

Use of evidence (25%): Uses a wide range of relevant sources.


After:

Level

Descriptor

4

Every main claim is supported by a relevant, credible source; sources are evaluated (method, scope or limitation noted) and applied to the argument, not just cited

3

Most main claims are supported; sources are applied to the argument, with occasional evaluation

2

Sources are cited but mostly summarised rather than applied; some main claims are unsupported

1

Few sources, or sources cited without connection to the argument

What changed: “wide range” is gone. Range is not the point. Whether each claim is supported, and whether sources are used rather than listed, is the point, and both are observable.

### Example 3: Structure and communication

Before:

Structure (15%): Well structured and clearly written.


After:

Level

Descriptor

4

The introduction states the question and the answer; each section advances the argument; the conclusion answers the question and follows from the analysis

3

Clear structure; one or two sections do not obviously advance the argument; the conclusion answers the question

2

Structure is present but the argument is hard to follow; the conclusion summarises rather than answers

1

No discernible structure; the question is not answered

What changed: “well structured” became questions a reader can check. Notice that grammar and spelling are not in this criterion. If they matter, give them their own small criterion, so a strong argument in imperfect English is not penalised twice.

Adjective only

Excellent / Good / Satisfactory / Weak

Wide range of relevant sources

Well structured and clearly written

Observable descriptor

Describes what the work actually does at each level

Draws on sources the marker can name, and applies each one to the argument

Each section opens with the claim it defends and closes by carrying it forward

The single most effective rubric edit: replace adjectives with descriptions of what the work does.

Figure 1: The single most effective rubric edit: replace adjectives with descriptions of what the work does.

Which type of rubric should you use?

There are three main types, and each suits a different job.

Type

What it looks like

Best for

Watch out for

Analytic

Several criteria, each with its own levels and descriptors

Summative marking, moderation, AI marking, and feedback students can act on

Takes longer to write; too many criteria fragment the judgement

Holistic

One overall scale with a description of each grade band

Quick overall judgements, short pieces, experienced teams

Hard to moderate; gives the student little to act on; hard for AI to explain

Single-point

One description of the expected standard per criterion, with space for feedback above and below

Formative feedback on drafts

Not designed to produce a grade on its own

Table 2: Analytic, holistic and single-point rubrics compared.

For AI marking of summative work, use an analytic rubric. It is what lets an AI score each criterion separately and explain each score, and it is what reviews of rubric research associate with more reliable marking.³ For formative feedback on drafts, a single-point rubric can work well, because the goal is guidance, not a grade. Panadero and Jonsson’s review of rubrics used for formative purposes links them to greater transparency for students and better self-regulation of their own work.⁶

A rubric template you can adapt

Here is a complete analytic rubric for a 2,500-word undergraduate essay in the social sciences. Adapt the criteria and wording to your discipline, keep the structure.

Criterion (weight)

70+

60-69

50-59

40-49

<40

Answering the question (15%)

States a clear answer in the introduction and sustains it throughout

Clear answer, occasionally drifts

Answer implied rather than stated

Addresses the topic, not the question

Does not address the question

Knowledge and understanding (20%)

Accurate, precise use of key concepts, including their limits

Accurate use of key concepts

Mostly accurate; some imprecision

Significant errors or gaps

Fundamental misunderstanding

Critical analysis (30%)

Compares competing interpretations, explains differences, defends own position

Compares interpretations; position asserted more than argued

Interpretations presented in sequence, limited comparison

Description or summary without evaluation

No analysis

Use of evidence (20%)

Every main claim supported; sources evaluated and applied

Most claims supported and applied

Sources mostly summarised; some claims unsupported

Few sources or disconnected from argument

No meaningful evidence

Structure and conclusion (10%)

Each section advances the argument; conclusion answers and follows

Clear, with minor digressions

Hard to follow in places; conclusion summarises

Weak structure

No discernible structure

Academic conventions (5%)

Referencing consistent and complete; clear, accurate prose

Minor referencing or language slips

Several slips that do not obscure meaning

Frequent errors that obscure meaning

Referencing absent

Table 3: A template analytic rubric for an undergraduate essay. Illustrative; adapt to your discipline and grading scheme.

Three notes on this template. The heaviest weight sits on the criterion that most distinguishes strong work (analysis), not on the easiest to check (conventions). Every cell describes something observable. And language accuracy has its own small criterion, so it cannot dominate the mark.

When do you need gateway criteria?

Some assessments contain one thing that must be right regardless of everything else: safe clinical practice in nursing, correct application of the law to the facts in a law problem question, ethical handling of data in a research proposal. In a weighted rubric, strong scores elsewhere can average out a failure there.

The fix is a gateway criterion: a criterion that must reach a threshold for the work to pass, whatever the total. Describe the failure specifically (“describes administering a medication without checking the prescription and does not recognise this as unsafe”), so that it can be flagged. An AI grader will then highlight the passage and explain the concern; the lecturer makes the decision. We walk through this for nursing in AI grading software for nursing assignments.

How do you test a rubric before a live cohort?

Never put a new rubric in front of a whole cohort, human or AI, without testing it. This calibration takes about two hours and saves weeks of disputes.

1

Choose 5-10 marked and moderated scripts (include a fail and a borderline)

2

Run them through the AI with the new rubric

3

Compare criterion by criterion with the agreed marks

4

Find the descriptors where the AI and the markers disagree

5

Rewrite those descriptors

6

Run again until the gaps are small

Most disagreements come from two or three descriptors. Fix those and the rest usually follows.

Step 1: choose the scripts. Five to ten scripts from last year, already marked and moderated. Include at least one fail, one borderline, and one first. Jonsson and Svingby’s review found that exemplars support reliable use of rubrics, and these scripts are your exemplars.³

Step 2: run them. With Eduface, set up the assignment with the brief, the rubric and your feedback instructions, and choose the discipline model closest to your module. Then run the scripts through the Paper Grader.

Step 3: compare criterion by criterion. Do not compare only the total. A total that matches can hide one criterion scored too high and another too low.

Step 4: rewrite the descriptors where they disagree. If the AI scores “use of evidence” higher than your markers, read the descriptor again. Usually it rewards citing sources rather than applying them. Make the difference explicit.

Step 5: run again. Two or three rounds is typical. In a pilot at one UK university, markers who spent time tuning the model to their own standards saw the AI’s suggestion come within 98% of their grades on average, against 94% overall. Accuracy there means how small the gap is between suggestion and marker. The difference is almost entirely in the rubric and instructions.

The new-marker test

Give your rubric to a colleague who has never taught the module. Ask them to mark two of your calibration scripts using only the rubric. Every question they ask you (“what counts as a comparison?”, “is this evidence or example?”) is a descriptor that needs rewriting. An AI grader will stumble in exactly the same places.

What about feedback instructions?

A rubric tells the AI what to judge. Feedback instructions tell it how to talk to the student. They are separate, and both matter.

Good feedback instructions say:

What to prioritise: “Comment first on the criterion with the lowest score.”

What to avoid: “Do not comment on referencing format; the library covers it.”

How to phrase it: tone, length, and whether to ask questions or give directions.

What the student should do next: “End each comment with one concrete action for the next draft.”

Eduface offers four feedback styles on top of your instructions: Reflective and Socratic, Constructive and Direct, Went Well and Needs Improvement, and Supportive and Encouraging.

What if you do not have a rubric at all?

Many modules have marking criteria in the handbook but no rubric with levels. You have two options.

Write one from your criteria, using the template above and the before-and-after examples.

Let Eduface generate a first draft from your assignment brief, then edit it. A generated rubric is a starting point, not a finished product. Read every descriptor, apply the new-marker test and calibrate on marked scripts before using it on a live cohort.

What are the most common rubric mistakes?

Mistake

What goes wrong

Fix

Adjectives instead of descriptions

Markers and AI read “good” differently

Describe what the work does at each level

Two things in one criterion

Scores average what you wanted separated

Split the criterion

Language accuracy everywhere

Second-language writers penalised in every criterion

One small criterion for language

Too many criteria (10+)

Judgement fragments; weights become meaningless

Four to six criteria

Levels that overlap

“Some evidence” vs “limited evidence”

Each level adds one observable thing

No gateway where one is needed

Strong writing masks a critical failure

Add a pass/fail criterion for the thing that must be right

Never tested

Problems discovered by students and appeals

Calibrate on 5-10 marked scripts

Table 4: Common rubric mistakes and how to fix them.

For a wider view of fairness to students writing in a second language, see does AI grading disadvantage non-native English speakers?

Frequently asked questions

Does AI grading need a different kind of rubric from human marking?

No. It needs the rubric human marking always should have had: separate criteria, observable descriptors at each level, explicit weights and a test on marked work. A rubric a new marker can apply without asking questions is a rubric an AI can apply too.

How many criteria should a rubric have?

Four to six for most essays and reports. Fewer and you combine things that should be separate; more and the judgement fragments and weights lose meaning.

How many performance levels should a rubric have?

Enough to match your grading scheme, usually four or five. Each extra level adds a boundary that has to be judged, so make every boundary explicit. A 2026 study found a general-purpose model’s agreement with human graders fell as the number of levels rose.

Can AI write the rubric for me?

It can draft one. Eduface generates a rubric from the assignment brief if there is none. Treat the draft as a starting point: rewrite vague descriptors, test it with a colleague and calibrate it on marked scripts before use.

Why does my AI grader score higher than my markers on one criterion?

Usually because the descriptor rewards something easier to spot than what you meant, such as citing sources rather than applying them. Compare criterion by criterion on marked scripts, find the descriptor, and make the difference explicit.

Should grammar and spelling be part of every criterion?

No. Give language accuracy its own small criterion. Otherwise a strong argument in imperfect English is penalised several times, which is unfair to students writing in a second language.

References

1. Lancaster University. We need to talk about marking. Centre for Engagement and Development in Academia (CEDA) briefing paper. [Quotes Higher Education Academy (2018) on inconsistency in academics’ judgements; cites McConlogue (2020).]

2. Bloxham, S., Boyd, P., & Orr, S. (2011). Mark my words: the role of assessment criteria in UK higher education grading practices. Studies in Higher Education, 36(6), 655–670. [Twelve lecturers thinking aloud; holistic judgement first, criteria used to refine it.]

3. Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational Research Review, 2(2), 130–144. [Review of 75 studies; rubrics improve reliability, especially when analytic, topic-specific and supported by exemplars or rater training.]

4. Brookhart, S. M. (2018). Appropriate criteria: key to effective rubrics. Frontiers in Education, 3, 22. [A rubric requires criteria and descriptions of performance levels; rating-scale language is less useful for learning.]

5. Deng, Farber, Lee, & Tang. (2026). Rubric-conditioned LLM grading: alignment, uncertainty, and robustness. arXiv:2601.08843. [Accuracy fell from 76% to 57%, and Cohen’s kappa from 0.51 to 0.34, moving from two-level to five-level grading of short science answers.]

6. Panadero, E., & Jonsson, A. (2013). The use of scoring rubrics for formative assessment purposes revisited: a review. Educational Research Review, 9, 129–144.

Test your rubric on real scripts

Upload a rubric and a handful of marked scripts, and see where the AI and your markers disagree. Book a demo or start free.