Course AIAssessmentAssignment FeedbackGDPRGuide

AI Assignment Feedback: How It Works, and How Far to Trust It

August 28, 2026
11 min read

An AI can read a learner's assignment and hand you a written critique and a proposed score in about ten seconds. It is right often enough to save you hours per cohort, and wrong often enough that signing it unread is a mistake you will eventually have to explain to a student. This article is about where that line sits, exercise type by exercise type, and about the part almost nobody covers: what data protection law expects of you once a learner's submission goes through a model.

The short answer

AI is reliable at drafting assignment feedback and unreliable at deciding a grade. It reads the submission against your rubric, writes a critique that is usually specific and usually fair, and proposes a score that is roughly right and occasionally badly wrong, in ways that are hard to predict from the outside. The working model is therefore not "AI grades the assignment" but "AI writes the first pass, a human validates and signs it". That is not a hedge for legal comfort. It is the only configuration where the output is defensible to the learner who disagrees with it, and, as section 6 explains, it is also what keeps automated grading out of the strictest part of the GDPR.

How AI actually produces feedback on a submission

The mechanism matters because most of the failure modes further down are consequences of it. In a platform where this is built into the AI course creation software rather than bolted on afterwards, a submission goes through five steps:

  1. Your criteria are loaded. The rubric you wrote when you created the assignment, plus the maximum score, plus the course material the assignment belongs to.
  2. The submission is analysed against those criteria, one criterion at a time rather than as a single global impression. A model asked for "a grade out of 20" produces a vibe; a model asked to judge five named criteria produces something you can argue with.
  3. A score is proposed, with the per-criterion breakdown that produced it.
  4. Feedback is drafted in the register you set: what worked, what did not, what to do differently next time.
  5. You review, adjust and sign. Until you do, the learner sees nothing. The score and the feedback only become visible to the student once a human has validated them.

Step 5 is the whole design. A pipeline that publishes the model's output straight to the learner is a different product with a different risk profile, and the rest of this article does not apply to it.

What AI grades well, and what it grades badly

The single most useful thing to know is that reliability tracks how verifiable the exercise is, not how hard it is. A model is better at a difficult calculation than at an easy essay.

Exercise typeFirst passWhere it breaks
Closed questions (multiple choice, true/false)Not neededScored deterministically against the answer key. Involving a model here adds a failure mode and buys nothing.
Short factual answersReliableRarely. Watch for a correct answer phrased in vocabulary your course did not use.
Verifiable technical work (calculation, SQL, configuration)ReliableA valid alternative method the rubric did not anticipate gets marked down.
CodeGoodStrong on correctness, style and obvious bugs. Weak on design intent: it cannot tell a deliberate simplification from ignorance.
Argumentative essayFeedback yes, score noThe written critique is genuinely useful. The number attached to it moves with fluency more than with the quality of the argument.
Creative work (copy, pitch, design rationale)WeakConverges on the conventional answer and penalises the interesting one. Taste does not fit in a rubric.
Long project or portfolioWeakJudging it means holding weeks of context and the learner's starting point. The model has the file in front of it and nothing else.

Read down the table and a design rule falls out: put the AI first pass on the assignments where the answer is checkable, and keep the human first on the ones where judgment is the point. If you are still designing your assessments, the activity mix is worth deciding before you write them, not after; there is a walkthrough in our guide to creating an online course with AI.

Is it reliable? The four failure modes worth knowing

These are the ones that show up in practice rather than in papers. None of them are fixed by a better prompt.

1. Fluency bias

A well-written wrong answer scores higher than a badly written right one. This is the most consistent bias in AI-assisted assessment and the most consequential, because it penalises exactly the learners who are still building confidence in the language of your field, non-native speakers among them. If you use AI feedback and never look at the bottom of the distribution, this is the bias you will be reproducing.

2. Severity drift

The same submission does not always get the same score. Run a batch twice and a handful of grades move a point or two; run it in a different order and different ones move. Any single grade is inside a band of roughly plus or minus one point on a twenty-point scale, and you cannot tell from the output which grades sat on a boundary. For formative feedback this is irrelevant. For anything that gates a certificate it is not.

3. Confident wrongness on out-of-syllabus content

Ask a model to judge an answer that depends on the specific method your course teaches, and if that method differs from the internet consensus, it will confidently mark a correct answer wrong. This is the failure mode most likely to produce a complaint, because the learner did exactly what you taught. A model grading against your course material rather than against general knowledge reduces it substantially. It does not eliminate it.

4. Rubric literalism

The model applies what your rubric says, not what you meant. A criterion reading "cites at least three sources" is satisfied by three bad sources. Vague criteria are worse than strict ones here: the vaguer the wording, the more the model substitutes its own standard, and the less your feedback sounds like you.

The guardrail: you sign it

Everything above points at the same conclusion. The value of AI on assignments is the draft: the specific, structured, immediately editable critique that would have taken you twelve minutes to write from a blank page. The judgment stays yours, and it should stay visibly yours: the learner is told a human validated the grade, and it is true.

Three habits make that guardrail real rather than nominal:

  • Read the bottom of the distribution first. The lowest grades are where fluency bias and confident wrongness concentrate, and where a wrong grade does the most damage.
  • Spot-check the middle. Five submissions per cohort, read cold before you look at the proposed score. You will calibrate fast, and you will notice drift.
  • Rewrite one sentence per feedback. A learner can tell the difference between feedback addressed to them and feedback generated about them, and it is usually one sentence.

GDPR and learner submissions: the part nobody writes about

A submission is not course material. It is a document a named individual wrote about themselves, and putting it through a model is a processing operation you are responsible for. Six things follow, and the first one surprises most course creators.

You are the controller, not the platform

Inside your own workspace, you decide what learner data is collected and why. That makes you the data controller; the platform acts as your processor, on your instructions. The obligations below are yours, and a vendor's compliance page does not transfer them. What the vendor owes you is a data processing agreement that lets you meet them.

Submissions carry more than you asked for

Ask learners to "apply the framework to your own situation" and you will receive salaries, client names, health circumstances, family arrangements and, in a management course, opinions about identifiable colleagues. Some of that is special category data under Article 9 and you never requested it. Two practical consequences: word your assignment briefs so they do not invite it, and do not assume the submission store is low-sensitivity because the course is.

The model provider is a sub-processor

If feedback is generated by a third-party model, that provider processes your learners' submissions. You need to know which provider, under what agreement, and whether the processing happens outside the EU. If it does, the transfer has to rest on Standard Contractual Clauses or an adequacy mechanism. Two questions to ask any vendor, in writing: are submissions used to train models? and where is the inference performed? "We take privacy seriously" is not an answer to either.

Article 22, and why human validation is not decoration

This is the argument worth understanding properly. Article 22 of the GDPR gives people the right not to be subject to a decision based solely on automated processing that produces legal effects or otherwise significantly affects them. Whether a course grade meets that threshold depends on what rides on it: feedback on a practice exercise almost certainly does not, while a grade that decides whether someone receives a professional certification plausibly does. If you are in the second case and the grade is published automatically, you owe the learner meaningful information about the logic involved, the right to obtain human intervention, and the right to contest the outcome.

A genuine human review before publication moves the decision out of "solely automated" altogether. The word doing the work is genuine: a reviewer who clicks approve on forty grades without reading them has produced a rubber stamp, which is exactly what regulators have said does not count. This is the strongest practical reason to keep the human in the loop, and it happens to be free once your workflow already puts them there. None of this is legal advice; if grades gate a certification you sell, it is worth twenty minutes with someone who does that for a living.

Tell learners, in the privacy notice and in the course

Transparency is not satisfied by a line buried in a policy nobody opens. Say plainly, where the assignment is set, that feedback is drafted with AI assistance and validated by a human before they see it. It costs one sentence, it pre-empts the discovery that would otherwise feel like a betrayal, and in our experience it lowers the number of contested grades rather than raising it.

Storage and retention are part of the same duty

Submissions should not sit in a public bucket behind an unguessable URL. A link leaked through a shared screen or a referrer header then grants permanent access to a learner's work. Private storage plus short-lived signed links is the baseline to expect from a platform. On your side, decide how long you keep submissions and act on it: once a cohort is closed and graded, an archive you never open is a liability, not an asset. Learners can also ask for their data, including their submissions and the feedback attached to them, so know how to export and delete it before someone asks.

What it actually changes for you

Run the arithmetic on your own numbers rather than on a vendor's percentage. Take a cohort of 40 learners and one assignment that takes twelve minutes to grade properly from scratch:

Per cohort of 40From scratchReview and sign a draft
Minutes per submission124
Total8 hours2 hours 40
Realistic turnaroundA week, in practiceTwo evenings

The hours matter less than the turnaround. Feedback that arrives eight days after a submission has stopped being feedback and become an archive entry; the same feedback at 48 hours is the reason someone finishes your course. That, not the time saving, is what a first-pass draft actually buys. Completion rate is what your reviews, your referrals and your next launch rest on.

The honest limit: this only holds if you keep reading the drafts. The moment reviewing becomes clicking approve, you have not automated grading, you have stopped grading, and the first learner who reads their feedback carefully will be able to tell.

Keep reading

Criterium is built on the arrangement this article argues for: the AI drafts feedback against the criteria you wrote, and a score and feedback become visible to the learner only once a human has validated them. Nothing is ever published automatically. Submissions are held in private storage and handed over through short-lived signed links rather than public URLs. Plans, including the free one, are on the pricing page.