VOCATIONAL EDUCATION · AI AND ASSESSMENT

AI in Vocational Education: What It Can Mark and What Stays With the Assessor

AI cannot watch a learner work, and it should not sign off competence. It can take on the written evidence around practical work, where marking hours pile up and feedback gets thin.

By Eduface · October 2026 · 15 min read

Your assessors spent the morning in the workshop, watching learners strip and refit a brake caliper. That part went fine, because they know competent work when they see it. The backlog is everything around it: sixty job sheets with short written answers, a pile of reflective logs, and a knowledge test the LMS marked overnight, wrongly in places nobody has spotted yet. The learners wrote those answers in twenty minutes, and they will wait two weeks to hear back.

AI in vocational education: where does it fit in assessment?

AI in vocational education and training (VET) is used for tutoring, simulations and admin, and at some providers for marking. For assessment the split is clear: AI cannot observe a learner wiring a circuit, and it should not sign off competence. It can draft marks and feedback on the written evidence around practical work: job sheets, lab worksheets, written knowledge tests, reflective logs and assignments. Used well, it marks against your criteria, flags answers it is unsure about, and holds every mark until an assessor approves it.

How is AI being used in vocational education and training right now?

Most AI use in vocational education and training is still on the teaching side, and the research specific to VET is thinner than the noise around it suggests.

A 2025 UNESCO-UNEVOC research brief on AI in European vocational education describes chatbots and adaptive tutoring platforms that “enable real-time feedback and skill tracking”, plus simulations and virtual labs for practising safely.¹ It sets a clear condition for assessment: tools such as “automated assessments, or exam proctoring must therefore meet strict requirements for transparency, human oversight, accuracy, and data protection.”¹ It also notes that many vocational teachers lack the digital literacy to use AI effectively.

Cedefop, the EU agency for vocational training, is blunter. Its 2026 call for papers on AI in VET states that “evidence specific to VET remains uneven”, with studies leaning on small-scale pilots or general education settings, and that “little is known” about how AI affects teachers’ practices and workload.² The research it cites points both ways: faster AI-supported feedback loops on one side; over-reliance on algorithmic recommendations, deskilling and limited transparency on the other.²

This article uses the European term, VET. In the US the same territory is called career and technical education (CTE), and the split below works the same way: hands-on competence on one side, written evidence on the other.

The practical reading: nobody has a large, VET-specific body of evidence on AI marking yet, including us. Start narrow and test on your own learners’ work before you rely on it.

What does vocational assessment actually consist of?

Vocational assessment is mostly practical and observed, but almost every observed task comes wrapped in written evidence.

Take a published end-point assessment plan for an apprenticeship in England, the Level 3 Prosthetic and Orthotic Technician.³ It has two methods. The first is a 90-minute observation of practice in the apprentice’s own workplace, followed by questions from the assessor. The second is a professional discussion underpinned by a portfolio, which can include products the apprentice made, observation records, work documentation, witness statements, and recorded questions, answers and workbooks. An independent assessor grades each method against set grading criteria, and at least 20% of each assessor’s decisions are moderated.³

Note the details: this plan excludes reflective accounts as portfolio evidence, which other programmes may accept, and its written material feeds an assessor’s judgement rather than being graded alone. Rules differ per standard, so any AI project starts with your own specification.

Outside end-point assessment, written evidence often carries more of the load. A college or private provider running programmes in motor vehicle, electrical installation, engineering, health and social care or hospitality typically collects:

Job sheets and lab worksheets: a practical task, then short written answers on what the learner did and why.

Written knowledge tests: short and open answers on principles, regulations and calculations.

Reflective logs: the learner’s account of a placement or practical session.

Assignments: reports and case studies in HNC and HND-style programmes and certificate units, marked against stated criteria.

That written part is where the hours go: many short pieces per learner per unit, across every intake, marked by the people also needed on the workshop floor. When time runs out, feedback is what gets cut: a tick, a “good”, a mark with no reason attached.

Observed: assessor only

Practical task

Workplace observation

Professional discussion

Witness testimony

Written: AI can draft, assessor approves

Job sheets and lab worksheets

Written knowledge tests

Reflective logs

Assignments

Both halves feed one decision: the assessor’s judgement on competence.

Figure 1: AI belongs on the written side of vocational assessment. The competence decision stays with the assessor.

Where can AI help with vocational assessment?

AI helps most on written evidence that is marked against stated criteria, arrives in volume, and currently gets slow or thin feedback.

Short-answer worksheets, job sheets and knowledge tests. A correct answer can be phrased in many ways, and LMS auto-marking handles that badly, as the worked example below shows. AI that marks against criteria rather than exact wording can draft the mark and a one-line reason.

HNC and HND-style or certificate assignments. Reports and case studies marked against criteria are the closest thing in VET to a higher education essay, and the same criterion-based approach applies, including feedback on drafts.

Reflective logs, formatively. A short comment on whether a log reflects or only describes. Low stakes, so a reasonable place to build staff confidence.

Consistency across sites and sessional assessors. For a multi-campus college, or a provider with rolling intakes and part-time assessors, the same criteria applied the same way to every script is a quality argument as much as a time one.

Which vocational evidence can AI mark, and who signs off?

AI can draft marks on written evidence, it has no role in judging observed performance, and a person signs off in every case.

Evidence type

AI can help

AI should not

Who signs off

Observed practical task

Nothing in the observation itself

Watch, judge or grade the performance

Assessor

Professional discussion in an end-point assessment

Little inside the assessment itself

Conduct or grade a discussion the specification assigns to an assessor

Independent assessor

Witness testimony or employer statement

Check it covers the criteria it claims

Judge whether the witness is credible

Assessor, sampled by IQA

Job sheet or lab worksheet, short answers

Draft a mark per criterion, credit correct answers in any wording, flag unclear ones

Pass a task whose practical part nobody observed

Instructor or assessor

Written knowledge test, open questions

Draft marks against the mark scheme, flag uncertain answers

Release marks before an instructor approves them

Instructor

Multiple-choice test

Nothing: the LMS quiz engine already marks it

Replace a quiz engine that works

Instructor, from the quiz results

Reflective log

Formative feedback on depth of reflection

Treat a learner’s account of safe practice as proof of it

Assessor

Written assignment (HNC/HND-style, certificate unit)

Draft criterion-level marks and feedback

Decide pass, merit or distinction alone

Assessor, sampled by IQA

Table 1: AI drafts on written evidence, a person decides, and observed performance stays out of scope.

Why does exact-match marking fail on a lab worksheet?

Exact-match marking compares wording, not meaning: a correct answer phrased differently scores zero, and a loose pattern written to catch it lets wrong answers through.

Moodle’s documentation for its Short-Answer question type is open about this. A response “must match one of your acceptable answers exactly”. You can add an asterisk as a wildcard, and Moodle advises keeping the required answer short “to avoid missing a correct answer that’s phrased differently.”⁴ That works for one-word and numeric answers. It breaks on the “what happened and why” questions most job sheets and lab worksheets ask. On another LMS, check how its short-answer question compares answers before trusting it with explanations.

Worked example (illustrative question and answers)

A Level 3 electrical engineering lab worksheet, 2 marks: “You added a second identical resistor in series with the first. The supply voltage did not change. What happened to the current, and why?” The instructor accepts the pattern halve for full marks, to catch “halves”, “halved” and “halve”. The mark scheme has two criteria. C1 (1 mark): the current halves, or an equivalent (50%, 0.5 times, I/2). C2 (1 mark): series resistances add, so total resistance doubles at the same voltage (I = V/R).

Learner answer

Exact match on halve

Criterion-based draft

What the assessor does

A: “Current halves because resistance doubles”

2/2

2/2, C1 and C2 met

Approves

B: “Dropped to 0.5 of what it was, R total = 2R and I = V/R”

0/2

2/2, both met in other words

Approves

C: “Down to 50%. Series resistors add, so double the resistance at the same voltage”

0/2

2/2, both met

Approves

D: “The current halved.”

2/2

1/2, C1 met, no reason

Approves

E: “Resistance halved so current doubled”

2/2

0/2, effect and reason wrong

Approves, adds a note on series resistance

F: “Less current, because there is more resistance now”

0/2

Flagged: effect not quantified, reason incomplete

Decides the mark, then sharpens the C2 wording

Table 2: Exact matching got four of the six wrong, in both directions, and gave no signal on the borderline answer.

B and C are false negatives: learners who understood the physics and lost marks for their wording. D and E are false positives, and worse, because nobody looks for them: E has the concept backwards and passed because “halved” appeared in the sentence. F is the honest grey zone where two experienced assessors could differ, which is exactly the case a good system hands back to a person.

Example numbers, for illustration only: 60 learners answering 10 short-answer questions is 600 answers per worksheet. If exact matching misfires on one answer in ten (an assumption, not a measurement), that is 60 wrong marks in both directions, and nothing in the gradebook tells the instructor which 60. Finding them means re-marking everything.

Criterion-based marking reads each answer against what each mark is for, which is why the mark scheme matters more than the tool. If C2 says “gives a reason”, F gets the mark; if it says “explains that total resistance doubles at constant voltage”, it does not. Your assessors and any AI both need that sentence written down.

How does Eduface handle short-answer marking with instructor approval?

Eduface’s Exam Grader marks each answer per criterion against your mark scheme, flags where it is unsure, and holds every mark until an instructor approves it.

Three independent AI agents mark each answer without seeing each other’s result, and a fourth compares them. Where they agree, the suggested mark carries high confidence; where they disagree, the answer is flagged for the instructor. Once a mark is approved, Eduface returns it to the gradebook in your LMS. Eduface connects to Moodle, Canvas, Blackboard and Brightspace through LTI 1.3, so learners submit where they already do, with no separate login. Each submission keeps an audit trail that meets AI Act requirements for summative assessments.

Eduface reports that the Exam Grader is 48% more consistent than unaided human marking and takes under 4 minutes from upload to suggested grade; the setup behind those figures is not published, so check them in your own pilot.

For longer written work, such as HNC and HND-style assignments or reflective logs, the Paper Grader marks against your rubric with an explainable grade breakdown per criterion and annotations in the text. Its formative feedback works across drafts, and the institution chooses whether that feedback goes straight to the learner or waits for the assessor’s approval. Its six subject models (Law, Economics, Social Sciences, STEM, Humanities, Health Sciences) include no trade-specific model, so for a plumbing or motor vehicle assignment the rubric you write carries the trade knowledge. A grader comparison dashboard shows when an individual marker deviates from the team or from the AI suggestion, useful when assessors work across sites.

Student data is processed in the EU and never used to train AI models, Eduface uses no third-party APIs such as OpenAI, and it signs a Data Processing Agreement with each institution.

1

Learner submits the worksheet in the LMS

2

Three AI agents mark each answer per criterion, independently

3

A fourth agent compares the results

4

Agreement: suggested mark with high confidence · Disagreement: flagged for the instructor

5

Instructor reviews, edits and approves

6

Approved mark returns to the LMS gradebook, audit trail kept

Nothing reaches the learner’s record until an instructor approves it. Disagreement between the agents becomes a flag, not a guess.

What it doesn’t do: Eduface does not observe practical work, conduct an end-point assessment discussion, or decide whether a learner is competent. AI assists. Educators decide.

What should AI not do in vocational education?

AI should not sign off competence, observe practical work, or replace the assessor’s judgement on safety-critical tasks, and those limits come from what vocational assessment is for, not from today’s technology.

Sign off competence. Competence means a person can do the job safely and to standard. In the plan above, that decision belongs to an independent assessor working from observation and discussion.³ A model reading written evidence cannot make that claim, however good its marks on the paperwork.

Observe practical work. A text model reads what a learner wrote about the task, not the task. A flawless job sheet can sit next to a poor weld.

Replace judgement on safety-critical tasks. Where a mistake can hurt someone, as in gas work, electrical isolation, patient handling or working at height, an AI-drafted mark should never be what lets a learner progress. Keep those questions with the assessor from the start.

Stand in for internal quality assurance. An AI-drafted, assessor-approved mark is still the assessor’s mark, and IQA sampling, standardisation and EQA visits apply to it as they do now.

Two more limits. If job sheets are filled in by hand, ask any vendor, including us, how scanned paper is handled; typed evidence in the LMS is the straightforward case. And check where accuracy claims come from. Eduface’s own figure comes from one pilot at one UK university: Bath Spa University, June 2026, 435 submissions across 6 subjects and 13 markers. The suggested grade was on average about 6 points from the lecturer’s grade, and about 2 points for lecturers who had tuned the model, with review taking 2 to 3 minutes per submission. Those were higher education submissions, not vocational job sheets. We do not have VET-specific accuracy data yet, which is why the pilot below starts with your own marked scripts.

What do regulators expect when AI marks vocational qualifications?

Regulators expect human oversight of the assessment decision, and the EU AI Act names vocational training explicitly.

If you deliver regulated qualifications in England, Ofqual and your awarding organisation set the rules on AI in marking, and awarding organisations can add their own guidance on top of Ofqual’s. We cover both separately: what Ofqual allows when AI marks assignments, and a JCQ and Ofqual checklist for private training providers.

If you deliver in the EU, point 3 of Annex III of the AI Act is headed “Education and vocational training”. Point 3(b) covers AI systems “intended to be used to evaluate learning outcomes, including when those outcomes are used to steer the learning process of natural persons in educational and vocational training institutions at all levels.”⁵ The text makes no exception for private providers or short courses. After the Digital Omnibus, obligations for these stand-alone Annex III systems apply from 2 December 2027.⁶ High-risk does not mean banned: it means obligations on human oversight, transparency and documentation, for the tool’s provider and for the institution using it.

In the US there is no single equivalent for CTE programmes: ask your accreditor and, for programmes leading to a licence, the relevant state board.

For heads of quality and internal quality assurers

Sample AI-drafted, assessor-approved marks in IQA like any other. Then ask one question of any tool: can you see, per submission, what the AI suggested, what the assessor changed and who approved it? Without that record, a decision is hard to defend at an EQA visit or an appeal.

How do you pilot AI marking on vocational evidence?

Pilot on one written evidence type in one unit, using scripts your assessors have already marked, and decide before you start what result would make you continue.

1

Pick the evidence type. A short-answer worksheet or a knowledge test with open questions. Not reflective logs first, because good reflection is harder to pin down in criteria, and nothing safety-critical.

2

Fix the criteria first. Write down what earns each mark, including accepted equivalents. Weak criteria produce weak AI marks and inconsistent human ones. Our guide to AI grading tools for higher education explains why the rubric is the biggest lever on accuracy.

3

Use already-marked work. Take 30 to 50 scripts from a past intake, from every site and assessor, with a data processing agreement in place before any learner work leaves your systems.

4

Compare and time it. Per answer, compare the draft with the assessor’s mark; where they differ, re-read and decide who was right. Record review time against marking from scratch.

5

Read the flags. Were the flagged answers the ones an experienced assessor would also hesitate over?

6

Decide on thresholds you set in advance. For example (illustrative): within one mark on 9 of 10 answers, and review clearly faster than marking from scratch.

How do the options compare for marking written vocational evidence?

For short and open written answers at volume, criterion-based AI with assessor approval handles varied wording and keeps a person in charge; for observed or safety-critical work, the assessor marks alone.

Approach

Handles varied wording?

Who decides the mark

Best for

LMS exact-match short-answer question

Only variants you anticipated with wildcards

The pattern the instructor wrote

One-word and numeric answers

General-purpose chatbot

Reads meaning, but no built-in criteria, approval step or audit trail

Whoever pasted the work in, with no record

Staff trying out their own example answers, not real learner work

Criterion-based AI with assessor approval

Yes, and flags uncertain answers

The assessor, who approves each mark

Short and open answers, assignments, reflective logs at volume

Assessor marks everything by hand

Yes

The assessor

Observed work, safety-critical questions, very small cohorts

Table 3: Each approach has a place. The mistake is using the first two for work that needs the third or fourth.

Frequently asked questions

Can AI be used in apprenticeship end-point assessment?

Not as the decision-maker. In the plan we read, an independent assessor grades both the observation and the professional discussion, and moderation checks those decisions. Any AI use around it has to fit the assessment plan and your end-point assessment organisation’s rules.

What do we need to set up AI marking for one vocational unit?

Three things from the instructor: the assignment brief, the rubric or mark scheme, and instructions on how feedback should be written. In Eduface, a rubric can be generated if the unit has none, and learners keep submitting through the LMS. For a pilot, add 30 to 50 already-marked scripts and a signed data processing agreement.

Does AI marking replace internal quality assurance?

No. A mark an assessor approved is the assessor’s mark, and IQA sampling and EQA visits apply to it as before. What changes is the record: Eduface keeps an audit trail per submission, which gives IQA something concrete to sample.

Can one AI model mark worksheets from different trades?

The mark scheme matters more than the model. Eduface’s Exam Grader marks each answer against the criteria you set, so a plumbing worksheet and an electrical worksheet differ in their mark schemes, not in the tool. That is another reason to write criteria that state exactly what earns each mark.

Should learners be told when AI helps mark their work?

Yes. Say it in the assessment brief: which evidence gets an AI-drafted mark or feedback, that an assessor reviews and approves every mark, and how to query a mark. Your appeals process needs the same wording, so write it once and use it in both places.

References

1. Tegelbeckers, H., Dietrich, A., Bünning, F., & Zug, S. (2025). European insights: AI integration in TVET: policies, practices and pathways for inclusive innovation. UNESCO-UNEVOC research brief. [Key finding: AI tools “enable real-time feedback and skill tracking”; automated assessments “must therefore meet strict requirements for transparency, human oversight, accuracy, and data protection”; many vocational teachers lack the digital literacy to use AI effectively.]

2. Cedefop. (2026). Cedefop Call for Papers 2026: Harnessing Artificial Intelligence and Digital Technologies for Inclusive and Resilient Vocational Education and Training (VET). [Key finding: “evidence specific to VET remains uneven”; “little is known” about AI’s effect on teachers’ practices and workload; risks include over-reliance on algorithmic recommendations, deskilling and limited transparency.]

3. Institute for Apprenticeships. (2018). End-Point Assessment Plan: Prosthetic and Orthotic Technician, Level 3 (ST0632/AP01). Published via Skills England. [Key finding: a 90-minute workplace observation of practice with questions, and a professional discussion underpinned by a portfolio; independent assessors grade; at least 20% of each assessor’s assessments moderated; simulated activities and reflective accounts not allowed as portfolio evidence.]

4. Moodle. (2026). Short-Answer question type. Moodle Docs, version 5.3. Read 5 October 2026. [Key finding: a response “must match one of your acceptable answers exactly”; the asterisk works as a wildcard; keep the required answer short “to avoid missing a correct answer that’s phrased differently”.]

5. European Parliament and Council of the EU. (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act), Annex III, point 3. [Key finding: point 3 is headed “Education and vocational training”; 3(b) covers AI systems intended to evaluate learning outcomes in educational and vocational training institutions at all levels.]

6. Gibson Dunn. (2026, 27 May). EU AI Act Omnibus Agreement: Postponed High-Risk Deadlines and Other Key Changes. [Key finding: stand-alone Annex III systems, including education, must comply by 2 December 2027.]

See it on your own assignments

Eduface drafts criterion-level grades and feedback for your instructors to approve. Book a demo or start free.