AI GRADING · INSTRUCTOR WORKFLOW
Grading Papers With AI: An 8-Step Workflow for Instructors
How to grade a stack of papers with AI without handing over the grade: tighten the rubric, choose the right tool, calibrate on papers you graded blind, review every grade, and keep a record that holds up when a student asks why.
By Eduface · October 2026 · 16 min read
It is Sunday evening, and 120 papers are waiting in your LMS. Grades are due Friday, and you teach three other sections the same week. You have read that AI can grade an essay in seconds, and part of you wants to paste the first paper into a chatbot to see what comes back. The harder question is what you would do with the number it gives you.
How do you start grading papers with AI without losing control of the grade?
Treat the AI as the first reader and yourself as the grader of record. Tighten your rubric, choose a tool built for grading that handles student data under your institution’s rules, and calibrate it: grade eight to twelve papers yourself first, blind, then compare. Run the batch, review and approve every grade and comment, return grades through your LMS gradebook, check consistency across sections, and keep the audit trail. In the worked example below, that takes a 120-paper stack from about 40 hours to under 15.
What does grading papers with AI look like from start to finish?
Grading papers with AI is a sequence of eight steps, and the AI does its work in only one of them. The time you save comes from the AI’s first pass; the control you keep comes from everything around it.
1
Rubric (criteria a new grader could apply)
2
Tool (built for grading, cleared for student data)
3
Calibrate (grade 8 to 12 papers blind, then compare)
4
Run the batch (suggestions held in draft)
5
Review and approve (every grade, every comment)
6
Return (grades flow to the LMS gradebook)
7
Check consistency (across sections and instructors)
8
Keep the record (audit trail plus your calibration note)
Figure 1: The AI drafts in step 4. Every other step belongs to the instructor.
The same workflow applies to grading essays with AI, case studies or reports. For what happens inside the software, read how automated assessment works, step by step. If you are still choosing a tool, start with our guide to AI grading tools in higher education. The examples use a 100-point scale; if you grade with letters, convert your boundaries to points first, because calibration needs numbers you can subtract.
Step 1: Is your rubric ready for AI grading?
Your rubric is ready when a colleague who has never taught the course could grade with it and land close to you. An AI grader has two things to go on: the student’s paper and your rubric. Whatever you carry in your head about what “strong analysis” means has to be written down, or the tool cannot use it.
One thing per criterion. “Argument and evidence” as one criterion forces any grader, human or AI, to average two judgments you wanted to see separately.
Observable descriptors. “Excellent analysis” tells a grader nothing. “Compares at least two explanations and says why they differ” can be checked.
Explicit weights and boundaries. State what each criterion is worth and where the grade boundaries sit.
Our guide to writing a rubric AI can mark against has templates and before-and-after examples. If you have no rubric, Eduface can generate one from your assignment brief. Treat it as a first draft and change whatever does not match how you grade. Plan about an hour the first time; you reuse the result every term.
Step 2: Should you grade papers with ChatGPT or a dedicated grading tool?
For grades you record, use a tool built for grading that your institution has cleared for student data. A general chatbot can produce a grade and a paragraph of feedback, but it struggles on three questions: what happens to the student work, whether the same paper gets the same grade twice, and whether anything records how the grade was reached.
Student data. For a US college, the first question about any tool that receives student papers is FERPA, the Family Educational Rights and Privacy Act. One route that lets a college share personally identifiable information from education records with an outside vendor, without student consent, is the school official exception. Under the federal regulation, an outside party can count as a school official if it performs a service the institution would otherwise use employees for, is under the institution’s direct control with respect to the use and maintenance of education records, and is bound by the rules on use and redisclosure of that information.¹
It is hard to see how a personal chatbot account meets the direct-control condition, because your college has no agreement with the provider about that work. Removing the name does not settle it: a personal reflection or a workplace case study can identify a student just as well. Ask your registrar or privacy office before you paste a single paper. The Department of Education’s Privacy Technical Assistance Center publishes guidance for vendors on their FERPA responsibilities, a useful reference for that conversation.²
Consistency. In a small test published in our guide to AI grading tools, two students in the Netherlands ran six papers with known lecturer grades and the official rubrics through eight tools, on the Dutch 10-point scale.³ ChatGPT averaged 0.9 points from the lecturer on five psychology papers, and on a case study graded 6.5 it returned 8.1, praising the section the lecturer had flagged as weakest. On a law essay graded 4.4, ChatGPT, Claude and Gemini returned 7.1, 7.2 and 6.8, one dedicated tool returned a perfect 10, and Eduface returned 5.5. Six papers point in a direction; they do not make a ranking. Whatever you choose, run one paper through it twice before you trust it.
Record. A chat history is not an audit trail. When a student appeals, you need the rubric version, the suggestion, what you changed and when you approved it, attached to the grade.
Where Eduface stands: Eduface signs a Data Processing Agreement with every institution, processes student data in the EU, does not use third-party AI APIs such as OpenAI, and never uses student work to train AI models. That is our position, not a FERPA ruling. Whether a vendor qualifies as a school official is your institution’s call, so ask every vendor, us included, how its contract covers direct control, use and redisclosure.
Step 3: How do you calibrate AI grading on your own papers?
Calibrate by grading a small sample yourself first, without seeing the AI’s suggestion, and then comparing per criterion. This step decides whether you can trust the rest of the batch. It also catches the most likely failure of any grading tool: rewarding how well a student writes rather than how well they reason, as with the 6.5 case study above. Testing your rubric on last term’s papers is a good start; calibration checks whether the tool lands where you would on this cohort’s work.
A calibration protocol for one assignment:
1
Set your thresholds before you look. An example for a 100-point scale: an average gap of 5 points or less, and no paper more than 10 points off. For reference, in a June 2026 pilot at Bath Spa University (435 submissions, 6 subjects, 13 graders), the suggested grade was on average about 6 points from the lecturer’s grade, and about 2 points for lecturers who had tuned the model.⁴ That is one institution, and the pilot data does not say whether lecturers graded before or after seeing the suggestion. One more reason to run your own check blind.
2
Pick eight to twelve papers across the range. A few at random from each section, plus any you suspect will fail or excel.
3
Grade them blind, per criterion. Score every criterion and note the main weakness you would name, before you open anything the AI produced. In Eduface, blind mode does this for you: you grade first, and the suggestion appears afterwards. AI-visible mode shows it from the start, and an institution can make either mode mandatory.
4
Compare per criterion, not only the total. A matching total can hide one criterion graded too high and another too low.
5
Fix and rerun, twice at most. Rewrite the descriptor or feedback instruction behind the gap and rerun the same sample. Still outside your thresholds? Grade this assignment by hand and fix the rubric before next term.
6
Turn what you learned into review rules. A criterion the AI ran generous on gets a closer look on every paper in step 5.
What you see in the sample
What it usually means
What to do
Average gap within your threshold, no outliers
The tool reads your rubric the way you do
Run the batch with standard review
One criterion off in the same direction on most papers
That descriptor reads differently than you intended
Rewrite it, rerun the same sample
Large gaps with no pattern
The brief or rubric is missing context
Add context; if gaps persist, grade by hand
The AI misses the weakness you flagged on several papers
Feedback instructions are too general
Say what to look for, rerun
You and the AI on opposite sides of a grade boundary
A borderline paper, which is expected
Make “near a boundary” a full-review rule
Table 1: Reading a calibration sample. Set the thresholds before you compare, not after.
Ten calibration papers: your blind grade against the AI’s suggestion (example data)
Paper 1
74 / 76
Paper 2
66 / 63
Paper 3
68 / 77
Paper 4
88 / 85
Paper 5
59 / 62
Paper 6
81 / 74
Paper 7
77 / 79
Paper 8
62 / 60
Paper 9
90 / 92
Paper 10
71 / 69
Green: inside a 5-point band. Navy: outside it. Paper 3 was fluent with thin evidence; paper 6 had an unusual structure. Example data.
Figure 2: Plot your blind grades against the AI’s suggestions. Dots outside the band show which papers, and usually which criterion, to look at next.
Calibrating on ten papers takes about four hours, and those ten papers are then graded. It is the most expensive step of the first run and the cheapest insurance in the workflow.
Step 4: How do you run the batch?
Run the batch with exactly the brief, rubric and settings you calibrated on. If you change the rubric halfway, papers graded before and after the change no longer share a standard, so rerun the whole set.
In Eduface, an assignment needs three things: the assignment brief, the rubric, and your feedback instructions, which set how comments are phrased. You choose one of the Paper Grader’s six subject models (Law, Economics, Social Sciences, STEM, Humanities, Health Sciences) and a feedback style, such as Constructive and Direct. Students submit through the LMS as usual, without a separate login, and papers arrive in Eduface through LTI 1.3.
While the batch runs, nothing reaches students. Each suggested grade is held in draft, with a score and reasoning per criterion and annotations in the text, until you approve it.
Step 5: How much should you review each AI-suggested grade?
Review every grade and every comment, and decide in advance which papers get a full read. Approving without reading is the failure this workflow exists to prevent, and the hardest one to spot from outside.
Full read. Every paper the tool flags as uncertain, every paper within a few points of a grade boundary, every likely fail, and a random one in ten that you grade blind first. Eduface flags uncertainty instead of smoothing it over, so the first group arrives already marked.
Standard review. For the rest, read the paper against the per-criterion breakdown and annotations, check that the feedback names the weakness you would name, edit what is off, and approve. In the Bath Spa pilot, review took 2 to 3 minutes per submission.⁴ Expect longer for your first batch and for long papers.
Watch for anchoring
When the suggestion is on screen before you form your own view, you tend to adjust from it instead of judging from scratch. That is anchoring, and it is why Eduface has a blind mode. Use it for the full-read group, and if you have approved dozens of papers in a row without changing anything, grade the next few blind.
A “Lecturer + AI” label in Eduface tells students that the instructor reviewed each feedback point before they saw it. AI assists. Educators decide.
Step 6: How do AI grades get back to the LMS gradebook?
Approved grades and feedback should flow back to your LMS gradebook without an export, so the grade students see is the grade you approved and nobody retypes numbers between windows.
Eduface connects to Canvas, Moodle, Blackboard and Brightspace through LTI 1.3. When you approve a grade, it passes back to the LMS gradebook through LTI grade services, and the feedback syncs with it. Before your first batch, check whether your course posts grades as soon as they arrive or holds them until you release them, and set it to hold if everyone should get their grade on the same day.
For IT and learning technology teams
The LTI 1.3 connection is set up once per institution: in Canvas through Developer Keys, in Moodle as an External Tool, and through the equivalent settings in Blackboard and Brightspace. After that, instructors add Eduface to an assignment without IT involvement, and students do not need a separate login.
Step 7: How do you keep grades consistent across sections and part-time instructors?
Use one rubric, one calibration set, and a comparison across graders before grades go out. In a multi-section course taught by full-time and adjunct instructors, or across campuses, a letter-grade gap between sections is the kind of thing students compare and program reviews pick up.
1
Share the calibrated setup. Every section uses the same brief, rubric and feedback instructions, including the fixes from step 3.
2
Grade three anchor papers. Each instructor grades the same three papers blind before starting. Differences show up before forty students are affected.
3
Compare before release. Look at section averages and at each instructor’s edits against the AI suggestion.
Eduface’s grader comparison dashboard shows at a glance when an individual grader departs from the rest of the team or from the AI suggestion, without opening individual submissions. It does not tell you who is right: an instructor who edits upward may know the course best. It tells you where to have the conversation.
For program directors
Consistency across part-time and full-time instructors is easier to show when everyone grades against the same calibrated rubric and deviations are visible per grader. Check the comparison at the end of every grading window, not after the first complaint.
Step 8: What should the audit trail for an AI-assisted grade hold?
The audit trail should let someone who was not there reconstruct each grade: the rubric and brief that applied, what the AI suggested per criterion and why, what the instructor changed, and the final grade. You need it when a student appeals, a colleague takes over mid-term, or a program review samples graded work.
Eduface keeps an audit trail for summative assessments that meets EU AI Act requirements; for the Exam Grader, it covers each AI agent’s reasoning and the instructor’s decision per submission. What no software records is why you calibrated the way you did. Keep a one-page note per assignment with your thresholds, sample results, changes and review rules. It answers the question an appeal may raise: how did you know the tool was grading to your standard?
How much faster is grading papers with AI? A worked example
In this example, a 120-paper stack goes from about 40 hours to under 15, and most of the remaining time is calibration and review. All numbers are EXAMPLE assumptions except the review time, which comes from the Bath Spa pilot. Replace them with your own.
The setup (EXAMPLE). Three sections of 40 students, one 2,000-word argumentative paper, a five-criterion rubric on a 100-point scale, one instructor. Grading by hand with written feedback takes 20 minutes a paper.
Step
By hand
With the AI workflow
Basis
Rubric and setup
0 min
90 min
60 min rubric (reused next term), 30 min setup
Calibration, 10 papers
In the row below
250 min
10 × 20 min blind, plus 10 × 5 min comparing
Routine papers
120 × 20 = 2,400 min
95 × 3 = 285 min
Upper end of the Bath Spa review time
Full reads, 15 papers
Included above
15 × 12 = 180 min
Flagged, borderline and spot-check papers
Return, consistency check, audit note
Included above
60 min
EXAMPLE
Total
2,400 min (40 h)
865 min (about 14.5 h)
Table 2: A worked time example for 120 papers. EXAMPLE numbers, except the review time per routine paper.
Review is most of what remains, and it should be. Calibration and full reads take more time than the routine review. That is the instructor’s judgment at work, not overhead.
The total is sensitive to your review time. At 6 minutes per routine paper instead of 3, it rises to about 19 hours: still under half the manual time, but a different week.
Small stacks do not pay off. Calibration costs about four hours whatever the stack size. With 20 papers, the AI workflow saves little or nothing.
Hours to grade 120 papers (example data)
By hand
40 h
AI workflow, 3 min review
14.5 h
AI workflow, 6 min review
19.2 h
The AI workflow splits into rubric and setup 1.5 h, calibration 4.2 h, routine review 4.75 h, full reads 3 h, return and checks 1 h. Doubling the review time per paper moves routine review to 9.5 h.
Figure 3: In the example, calibration and review make up most of the remaining time. That is where the instructor’s judgment goes.
When should you not grade papers with AI?
Do not grade with AI when calibration costs more than it saves, when you have not decided what good looks like, or when nobody will have time to review. AI grading fits a particular shape of work: many papers, stable criteria, and an instructor with time to check.
Very small classes. With fifteen papers, the calibration sample is most of the stack. Grade by hand, or use the Paper Grader’s formative feedback on drafts instead.
Assignments without stable criteria. A new assignment, or creative work where you discover what good looks like while grading, gives the tool nothing reliable to grade against. Grade the first round by hand, then write the rubric from it.
High-stakes decisions without a full review. Capstones, papers that decide progression, and anything that may feed an academic integrity case deserve a full read. An AI grade on its own is never evidence of misconduct.
Batches you will not have time to review. If the realistic plan is to approve in bulk late on Thursday night, do not start. An unreviewed AI grade is worse than a late human one.
The quieter risk is anchoring: a grader who always sees the suggestion first tends to drift toward it over a long stack, even with a well-calibrated tool. That is a reason to keep part of every batch blind.
What Eduface does not do: it does not decide grades, it does not make a vague rubric precise on its own, and a rubric it generates from your brief is a draft, not a finished standard.
Grading by hand, with a chatbot or with a dedicated tool: how do they compare?
The three approaches differ less in speed than in what you can show afterwards about how a grade was reached.
Approach
120-paper example
Consistency and record
Student data
By hand
About 40 hours
Depends on you; your notes are the record
Stays in your LMS
General chatbot, personal account
Not estimated: rubric pasted every session, grades copied back by hand
Output can change between runs; chat history not tied to the grade
No institutional agreement; ask your registrar first
Dedicated tool with instructor approval (Eduface here)
About 14.5 hours, including calibration
Calibrated to your rubric, audit trail, grader comparison dashboard
Data Processing Agreement, EU processing; ask how it maps to FERPA
Table 3: Three ways to grade a 120-paper stack, using the EXAMPLE numbers above.
Frequently asked questions
Can teachers use AI to grade papers?
Yes, if your institution’s policy allows it, the instructor remains the person who decides the grade, and the tool handles student work under your institution’s privacy rules. Check that the tool is approved for student data before using it on a live class.
Can I use ChatGPT to help with grading at all?
Yes, for work that does not involve student papers. A general chatbot is useful for drafting rubric descriptors, writing a model answer to test your rubric against, or rephrasing your own feedback comments. Keep student work out of it unless your institution has an agreement with the provider that covers it.
Do students need to know that AI helped grade their paper?
Yes, tell them. A syllabus line that says what the AI does, that you review and approve every grade and comment, and how to ask about a grade is a good start. It also heads off the appeal that begins with “a computer graded this”.
Does this workflow work for exams as well as papers?
Yes. Eduface’s Exam Grader has three independent AI agents grade each open or closed answer without seeing each other’s results, and a fourth compares them. Agreement means high confidence; disagreement is flagged for the instructor. Eduface reports results “48% more consistent than unaided human marking” and under 4 minutes from upload to suggested grade. The grade returns to the LMS gradebook after the instructor approves it.
How much does it cost for an individual instructor to start?
Individual instructors can start with Eduface for free, with about 20 assignments a month, or use the Lecturer plan at $25 a month for about 200. Program-wide use runs on an institutional license, priced separately.
References
1. 34 CFR § 99.31(a)(1)(i)(B). Family Educational Rights and Privacy Act regulations. Read via Cornell Law School, Legal Information Institute, October 2026. [Key finding: an outside party may count as a school official if it performs a service the institution would otherwise use employees for, is under its direct control for the use and maintenance of education records, and is subject to the use and redisclosure rules of § 99.33(a).]
2. Privacy Technical Assistance Center, U.S. Department of Education. (2015). Responsibilities of Third-Party Service Providers under FERPA. studentprivacy.ed.gov. [Key finding: guidance developed to help online educational service providers, vendors and contractors understand FERPA.]
3. Eduface. (2026). AI Grading Tools in Higher Education. eduface.me/resources/blog/ai-grading-tools-higher-education-guide. [Key finding: on six papers with known lecturer grades, ChatGPT averaged 0.9 points off on five psychology papers and returned 8.1 for a 6.5 paper; on a 4.4 law essay, general tools returned 6.8 to 7.2, one dedicated tool 10, Eduface 5.5. Directional, not a large-scale study.]
4. Eduface. (2026). Pilot at Bath Spa University, June 2026. Internal pilot data. [Key finding: 435 submissions, 6 subjects, 13 graders; suggested grade on average about 6 points from the lecturer’s grade, about 2 points for lecturers who had tuned the model; review took 2 to 3 minutes per submission.]
See it on your own assignments
Eduface drafts criterion-level grades and feedback for your instructors to approve. Book a demo or start free.