How UK Teachers Can Use AI for Designing Assessments
UK teachers can use AI to draft question banks, generate mark schemes aligned to National Curriculum objectives, and build differentiated versions of the same test in a fraction of the time manual writing takes. The teacher still sets the standard, checks every question against the specification, and decides what actually gets used — AI drafts, the professional judges.
Quick Answer: AI tools can generate first-draft assessment questions, distractors, and mark schemes mapped to National Curriculum or exam board objectives, then let a teacher edit, differentiate, and finalise in far less time than writing from a blank page. The output still requires teacher review against the specification before it reaches students.
Ofqual's own guidance on generative AI in assessment (2024) is blunt about where the line sits: AI can support preparation, but the qualification standard and final sign-off remain entirely human. This piece walks through what that looks like in practice — where AI genuinely cuts prep time, a worked example building a Year 8 science test, the tools worth comparing, and the mistakes that turn a time-saver into a liability.
Why Assessment Design Eats So Much Teacher Time
Writing a genuinely good test question is slower than it looks — a well-targeted item needs a clear stem, plausible distractors, and a mark scheme that anticipates how students will actually answer, not just how you hope they will.
- A National Foundation for Educational Research (NFER, 2023) survey found assessment design and marking together account for a substantial share of teachers' non-contact hours, second only to lesson planning
- The Department for Education's 2023 workload survey put primary and secondary teachers' average working week well above 50 hours, with planning and assessment cited as the top pressure points
- Ofsted's 2023 curriculum reviews repeatedly flag inconsistent assessment quality between classes taught by the same department as a red flag for moderation
What "Good" Actually Requires
A single well-written multiple-choice question typically needs three or four plausible wrong answers, each targeting a specific misconception rather than an obviously silly guess.
- A clear, unambiguous stem that tests one thing, not two things bundled into one question
- Distractors built from real misconceptions, not random wrong numbers
- A mark scheme that anticipates partial credit, especially for extended-response questions
- Alignment to a specific assessment objective, not just "the topic we covered this term"
That fourth point is where a lot of homemade tests quietly drift — a question feels relevant to the unit but doesn't actually map to the AO the exam board is testing. It's also the easiest thing to lose track of when a test is being written at 9pm after a full teaching day, which is exactly when most secondary assessment design actually happens.
Time pressure compounds the problem in a specific way: a rushed test tends to over-represent whatever content is freshest in the teacher's mind, rather than giving proportional coverage to the full unit. A structured drafting process — starting from a topic list rather than memory — helps regardless of whether AI is involved, but it's the step AI can make fast enough to actually stick to under time pressure.
Where AI Genuinely Speeds Up Assessment Design
The honest use case is drafting, not deciding. AI is strongest at producing volume and variety fast, weakest at judging whether a question is actually good.
- Generating a first-draft question bank from a topic list or specification reference, which a teacher then edits down
- Producing multiple distractors per question, saving the slowest part of MCQ writing — inventing plausible wrong answers
- Drafting mark schemes with model answers, giving a starting point a teacher refines against real exemplar work
- Creating differentiated versions of the same assessment — a foundation-tier and higher-tier variant, or a scaffolded version for SEND students
- Converting a written test into other formats for revision, such as flashcards from the same content bank
EduGenius can generate an MCQ quiz, worksheet, or long-format exam with an answer key from a set of topics or a class profile, which is one way to produce that first draft quickly before a teacher edits it against the specification.
Where It Genuinely Falls Short
AI does not know your students, your exam board's specific phrasing conventions, or which misconceptions your class actually holds — that judgment stays entirely with the teacher.
- It can generate a plausible-looking mark scheme that doesn't match the exact AO wording an exam board expects
- It has no visibility into which students already struggle with a topic, so differentiation suggestions are generic until you adjust them
- It cannot verify that a maths or science calculation is actually correct without a human checking the working
Which Question Types AI Handles Well — and Which It Doesn't
Not every question format benefits equally from AI drafting. Some formats are close to ready after one pass; others need heavier rework before they're usable.
| Question type | AI draft quality | Typical fix needed |
|---|---|---|
| Multiple choice with factual recall | Strong | Light distractor tuning |
| Short-answer definitions | Strong | Check exact spec wording |
| Extended-response essay prompts | Moderate | Rework mark scheme weighting |
| Multi-step maths problems | Weaker | Verify every calculation by hand |
| Data-response / graph interpretation | Weaker | Check the data set is realistic and internally consistent |
Factual recall and short-answer formats are where AI earns its keep fastest, because there's less room for a subtly wrong setup to hide. Multi-step maths and data-interpretation questions need the most scrutiny — a generated word problem can look plausible while containing numbers that don't actually work out cleanly, which is exactly the kind of error a rushed final check misses.
The pattern holds across subjects: the more a question depends on precise, verifiable facts (a definition, a labelled diagram, a single calculation step), the more reliable the first draft tends to be. The more a question depends on judgement — how much partial credit a half-correct essay paragraph deserves, or whether a data set is internally consistent — the more a teacher's review actually changes the output.
Building Distractors That Actually Target Misconceptions
Weak distractors are the single most common flaw in homemade multiple-choice tests, and it's also the area where a good AI prompt makes the biggest visible difference.
- Ask for distractors based on a specific, named misconception ("a student who thinks force and speed are the same thing") rather than a generic "plausible wrong answer"
- Request one distractor that reflects a common calculation error, such as forgetting to square a value or misreading units
- Avoid distractors that are obviously silly — if three of four options are clearly wrong, the question isn't really testing understanding
- Cross-reference generated distractors against errors you've actually seen in past marking, which sharpens them considerably
Building a Term-Long Assessment Strategy Around AI Drafting
A single test is one thing; planning assessment across a full term is where the time savings compound, provided the workflow stays organised rather than ad hoc.
- Map your assessment points at the start of term — end-of-topic tests, a mid-term formative check, and a summative assessment — before generating anything
- Generate question banks topic by topic as you teach, rather than waiting until the week before each test, so the content stays fresh from actual lessons
- Build a shared department bank if colleagues are willing, pooling reviewed AI-drafted questions rather than everyone regenerating the same topics independently
- Review generated content against real outcomes after each assessment, noting which AI-drafted questions performed well and which need retiring
This kind of forward planning also makes moderation easier. When every teacher in a department is drafting from a shared, reviewed bank rather than improvising independently, consistency between classes — the exact gap Ofsted (2023) flags as a concern — becomes much easier to maintain.
A Worked Example: Building a Year 8 Science Test
Say you teach Year 8 science and need a 30-minute end-of-topic test on forces, covering the AQA Key Stage 3 framework.
- List the specific sub-topics to cover — resultant force, friction, gravity, and Newton's first law — rather than "forces" as one vague heading
- Generate a first-draft question bank covering each sub-topic, requesting a mix of MCQ and short-answer items
- Review every question against the framework, cutting anything that tests recall of a fact you didn't actually teach this way
- Request a scaffolded variant for students who need extra support, and a stretch variant with one extended-response question for higher attainers
- Generate the mark scheme separately, then check it against two or three real student answers from a previous class before finalising
That review-against-real-work step in stage five is the one teachers skip when rushed — and it's the step that catches a mark scheme that looks fine on paper but doesn't actually match how students respond.
Comparing Approaches to Assessment Design
| Approach | Speed | Curriculum alignment | Differentiation effort |
|---|---|---|---|
| Writing every question manually | Slowest | Strong, if the teacher is experienced | High — each variant written separately |
| AI-drafted, teacher-reviewed | Fast | Requires manual check against spec | Low — AI can draft variants quickly |
| Reused past-paper questions | Fast | Strong, exam-board verified | Limited — hard to adapt without breaking security |
| Off-the-shelf test bank | Moderate | Variable, depends on publisher | Moderate — some banks include tiers |
Reused past-paper questions score highest on alignment for a reason: they were written and vetted by the exam board itself. AI-drafted content trades some of that guaranteed alignment for speed, which is exactly why the review step matters.
Adapting the Workflow Across Subjects and Key Stages
The core process — generate, review against spec, verify, differentiate — holds across subjects, but what "review" actually means shifts depending on what's being assessed.
Humanities and English
Extended-response questions dominate here, so the mark scheme matters more than the question itself. AI-drafted mark schemes for essay-style responses tend to read as generic level descriptors unless you feed in your exam board's actual assessment objectives and a couple of real exemplar scripts for calibration. Without that grounding, a generated mark scheme can reward well-structured writing over accurate content, which isn't what most humanities specifications actually want.
Maths and Sciences
Numeric accuracy is non-negotiable, so every generated calculation question needs to be worked through by hand before it reaches a student. A well-known failure mode is a generated word problem where the numbers don't divide cleanly, forcing an awkward decimal answer nobody intended. Running the maths yourself, even for a "simple" generated question, takes under a minute and catches almost all of these.
Primary Key Stage 1 and 2
At primary level, reading age matters as much as content difficulty — a question testing a Key Stage 2 science concept can accidentally use vocabulary better suited to secondary school if the prompt doesn't specify reading level explicitly. Always state the year group and, where useful, a rough reading age when generating primary assessment content, and read every question aloud to check it sounds like something a Year 4 or Year 5 pupil would actually encounter.
What to Avoid
A handful of habits turn a genuinely useful AI-assisted workflow into a marking headache or, worse, an assessment that doesn't actually measure what it claims to.
- Publishing an AI-drafted mark scheme without checking it against real student answers. Model answers written in isolation often miss how students actually phrase correct responses.
- Skipping the specification-alignment check. A question can read well and still test the wrong assessment objective.
- Reusing the exact same AI-generated question across multiple classes without variation, which risks answers circulating between groups before the later class sits the test.
- Treating AI-suggested "difficulty levels" as reliable without verifying them against your own students' actual performance data.
How This Fits Into Ofsted's Assessment Expectations
Ofsted doesn't inspect for AI use directly, but its curriculum reviews do inspect for the consistency and rigour that a well-run AI-assisted workflow, done properly, actually supports rather than undermines.
- Consistency between classes is a named concern in Ofsted's 2023 curriculum research review — a shared, reviewed AI-drafted question bank can reduce the drift that happens when every teacher in a department writes tests independently from scratch
- Assessment validity — whether a test actually measures what it claims to — depends entirely on the human review step; an AI-drafted test that skips specification alignment checking would fail an Ofsted deep dive exactly the same way a poorly written manual test would
- Evidence of impact matters too: Ofsted inspectors look for how assessment data feeds back into teaching, which is a step entirely separate from how the test itself was drafted
The practical takeaway is that AI-assisted drafting is invisible to inspection as long as the review, moderation, and follow-up steps that were always expected of good assessment practice stay firmly in place. It's a tool for the drafting stage, not a shortcut around professional judgement.
Pro Tips for UK Teachers
- Draft in batches by topic, not by lesson. Generating a full unit's question bank in one sitting is faster than doing it lesson by lesson and keeps the difficulty curve consistent.
- Keep a personal bank of misconceptions you've seen in marking, and feed those into your prompts — AI-generated distractors improve noticeably when you specify the actual wrong answer students give.
- Cross-check any AI-drafted mark scheme against your department's moderation standards before it goes anywhere near a shared assessment.
- Save strong AI-drafted questions into your own bank for reuse next year, rather than regenerating from scratch each cycle.
- Generate slightly more questions than you need, then select the strongest ones — it's faster to discard three weak AI-drafted items than to fix three flawed ones by hand.
- Time-box the review step. Give yourself a fixed 15 minutes per generated test to check alignment and calculations, rather than an open-ended editing session that eats back the time you just saved.
Key Takeaways
- AI can draft assessment questions, distractors, and mark schemes quickly, but the teacher remains responsible for checking alignment to the specification.
- Ofqual (2024) guidance treats AI as a preparation aid, not a substitute for human sign-off on qualification-standard assessments.
- The slowest part of manual test writing — inventing plausible distractors — is where AI saves the most time.
- Differentiated variants (foundation/higher, scaffolded/stretch) can be generated much faster than writing each version by hand.
- Always check an AI-drafted mark scheme against real student work before finalising it.
- Tools like EduGenius can generate a first-draft quiz, worksheet, or exam with an answer key from a topic list or class profile.
- Reused, exam-board-verified past-paper questions still score highest for guaranteed specification alignment.
FAQs
Can AI write exam-board-aligned mark schemes for UK teachers?
AI can draft a mark scheme structured around the assessment objectives you specify, but it won't match an exam board's exact wording conventions automatically — a teacher needs to check it against real specification language and student exemplar work before using it for grading.
Is it acceptable under Ofqual guidance to use AI-generated questions in formal assessments?
Ofqual's 2024 guidance treats AI as acceptable for preparation and drafting, provided a qualified teacher reviews and takes responsibility for the final content — it does not endorse publishing AI output directly without human review for any formally graded assessment.
How can AI help differentiate the same test for mixed-ability classes?
AI can generate a scaffolded version with simplified language or extra prompts, and a stretch version with an additional extended-response question, from the same underlying question bank — cutting the time it takes to build multiple tiers by hand.
Does using AI to draft assessments save time for UK teachers?
It can reduce the time spent on the initial drafting stage — generating question banks, distractors, and mark scheme starting points — though the review, alignment check, and moderation steps still take teacher time and should not be skipped.
Should I tell students their test was partly AI-drafted?
There's no regulatory requirement to disclose this to students, since the assessment itself is a teacher-reviewed, teacher-owned product regardless of how the first draft was produced — what matters for validity is that a qualified teacher checked and approved the final content before it was used.
Related Reading
- AI for Teachers and Parents: A 2026 Guide for the US, UK & UAE (pillar)
- AI Lesson Plans Aligned to Key Stage 2 (UK) (hub)
- How US Parents Can Use AI to Support an Anxious Learner (sibling)
- How UAE Teachers Can Use AI for Creating Presentations (sibling)
- A UAE Teacher's Guide to AI for Computer Science (sibling)
- Best AI Tools for US Teachers in 2026 (cross-pillar)
References
- Ofqual. (2024). Generative AI in Assessment: Guidance for Awarding Organisations.
- National Foundation for Educational Research (NFER). (2023). Teacher Workload and Wellbeing Survey.
- Department for Education. (2023). Working Lives of Teachers and Leaders Survey.
- Ofsted. (2023). Curriculum Research Review: Assessment.