Every conversation I have about AI in marking begins in the same place: how much time will it save? It is a fair question. Marking is the part of assessment that swallows evenings and weekends, the part where institutional promises about turnaround quietly go to die, and the part academics are least willing to defend as a good use of their expertise. If something gives hours back to people who have none, that matters. I am not going to pretend otherwise.

What worries me is what happens when hours saved becomes the only number on the slide. Speed is the easiest thing to measure and the easiest thing to sell. It is also the one benefit that can quietly be bought at the expense of everything else marking is supposed to produce.
Earlier this month the e-Assessment Association made the case that trust, not speed, decides whether AI marking succeeds: “Automate the routine. Explain the decision. Keep a person accountable.” (e-Assessment Association, 17 September 2026) I would sharpen it slightly. The useful question is not how much time the marking process saves. It is what the marking process leaves behind.
A MARK IS A CLAIM, AND CLAIMS GET CHALLENGED
A mark is not a number. It is a claim an institution makes about a student — and like any claim it can be questioned, appealed, moderated, audited and occasionally taken to a lawyer. When that happens, speed is irrelevant. Three other things decide the outcome:
-
the evidence the marker actually had in front of them
-
the decision, and who made it
-
whether either can be shown to someone who was not in the room
We have made neighbouring arguments before — that rubrics are quality assurance rather than paperwork, and that fair, consistent and human assessment design is what digital exam platforms now have to deliver. This follows the same thread one step further, into the marking itself. Get evidence, decision and transparency right and the time savings arrive as a by-product. Chase the time savings first and there is no guarantee the other three survive the journey.
MARKING IS A JOURNEY, NOT A MOMENT
Most tooling treats marking as a single event: a script arrives, a mark comes out. In practice the marker journey has at least seven stages, and the weak links are rarely where people look. Allocation. Reading. Checking provenance. Judging against criteria. Justifying the judgement. Moderating it with a colleague. Releasing it — and, sometimes, defending it months later.
Automating stage four and leaving the rest in email, spreadsheets and memory does not make marking better. It makes one stage faster and the joins between stages weaker. My view is that the interesting problem in marking is not the judgement itself. It is the continuity between the stages, because that is where evidence quietly falls out of the process.
EVIDENCE BELONGS IN THE ROOM WHERE THE DECISION IS MADE
The clearest example is originality. For years integrity checking has lived in a separate tool, run at a separate time, often by someone who is not the marker. The report is a document that arrives before or after the judgement rather than during it.
That separation has a cost, and it is not mainly a cost in minutes. A marker who sees a similarity report in context — this passage, this source, this overlap, with correctly cited quotation excluded and visibly excluded — is making an informed academic judgement. A marker who sees a percentage in a different system a week earlier is either guessing or ignoring it.
We argued in We are not post-plagiarism that similarity checking still matters now students have AI, and in There is no quick fix for AI in assessment that AI-writing detection is not the answer. The positive version of both arguments is this: integrity evidence is only worth having where the decision is being made. If you are weighing up tools, our comparison of plagiarism checkers sets out the criteria we think matter. That is also why WISEflow Originality reports inside the marking view rather than beside it — not because it is faster, though it is, but because evidence that arrives at the wrong moment is barely evidence at all.
Reasonable people differ here. Some institutions deliberately keep integrity screening separate from marking so that suspicion does not colour the mark, and in strict anonymous-marking regimes that is a defensible position. My own view is that the larger risk today is markers forming authorship judgements with no evidence in front of them at all.
A DRAFT IS NOT A DECISION
Where does AI belong in this? We set out our line in Assist, don’t dictate: AI assists, the human decides. Three months on I would keep the line and be more specific about where it bites.
The most useful thing AI does in marking today is not scoring. It is drafting the explanation. Given a rubric and a set of marking guidelines, a model can produce a first draft of a grade justification and written feedback faster than a tired marker at eleven at night, and more consistently structured. The marker then reads it, edits it, disagrees with it, or throws it away. Nothing reaches a student that a person has not approved.
That is not a legal technicality. A drafted justification an assessor has corrected is a human decision with a better artefact attached. A generated score an assessor has waved through is a machine decision with a human alibi. The two can look almost identical in a workflow diagram. They are entirely different things in an appeal.
Ofqual’s July position on AI in regulated qualifications — that AI “may not be used as a sole marker”, and that marking decisions must be “accountable, capable of being explained and, where necessary, challenged” (Ofqual, 16 July 2026) — does not bind universities. It is still a reasonable reference point, and it points the same way.
I would also be honest about the limits. The published evidence on AI-assisted feedback quality in higher education is thin: mostly pilots, small cohorts and self-reported marker satisfaction rather than controlled comparison. That is a reason to be deliberate about where you use it, not a reason to avoid it. Low-stakes formative feedback and a final-year dissertation are not the same problem, and two thoughtful universities can read the same evidence and draw the line in different places. We expect to revisit our own line as the evidence and the regulation mature.
TRANSPARENCY IS NOT A REPORT YOU RUN AFTERWARDS
The third thing marking has to produce is the ability to show your working — to a student, an external examiner, an exam board or a regulator.
Most institutions can produce marks. Rather fewer can produce, on demand: who marked what, against which criteria, what evidence they had, where two markers diverged, how the divergence was resolved, what was released to the student and when, and which parts of the feedback were drafted by a model before a person approved them.
That last item is new, and it is the one I would not leave until someone asks. If AI contributed to feedback, the institution should be able to say so plainly, without an archaeology project. We have written about the governance side of this in Training AI on students’ work, The EU AI Act and assessment: December 2027 is not a snooze button and Getting your assessment back out. The short version: transparency built into the workflow costs very little, and transparency reconstructed after a complaint costs a great deal.
TWO ROUTES IN, ONE STANDARD
Institutions are not shaped the same way, so we have built this in two shapes rather than one. What should not change between them is the standard.
WISEflow is the route for institutions running assessment centrally. It manages the complete assessment and feedback lifecycle, with a unified marking interface, moderation tools, bulk operations, external examiner access and full audit trails inside one role-based permission model. When an appeal lands eighteen months later, the exam office produces the record from one place rather than reassembling it from three systems and an inbox.
Marking Studio is the route for institutions where the LMS is staying, by choice or by shape. It delivers the same professional marking capability — rubric-aligned marking with inline annotation and audio or text feedback, double-marking and moderation workflows, AI-assisted feedback that drafts grade justifications aligned to your own rubrics and marking guidelines, calibration and sharing so colleagues can review before release, and audit-friendly trails of decisions, progress and rationale — inside Canvas, Moodle, Blackboard, Brightspace and others through LTI 1.3. A department can adopt it without an institution-wide platform decision first.
WISEflow Originality attaches to either route, or runs on its own: AI-driven semantic similarity that identifies paraphrasing and rewritten content across more than 50 languages, source pools spanning institutional archives, the open internet and sharing groups, and an embedded report overlay inside the marking tool.
That standard has to hold for handwritten work too, and this is where evidence most often falls out of the process. Paper Submission keeps scanned scripts inside the same marking workflow on either route — unique per-candidate identification codes, scan once and auto-sort, mixed paper and digital submissions handled together — so a paper exam does not leave the audit trail behind the moment it leaves the hall. We set out the operational case for that in our business case for the Paper Submission module.
The point of offering two routes is not flexibility for its own sake. It is that a student’s mark should not be more or less defensible because their faculty walked through a different door. Whichever route an institution takes, the evidence, the human decision and the record should look the same.
We have been equally deliberate about what we have not built. We do not offer AI-writing detection, for the reasons set out in There is no quick fix for AI in assessment. And on neither route does AI submit a mark. A person does.
THE BIGGER POINT
The sector is about to spend a good deal of money on marking automation, and most of the business cases will be written in hours saved. I understand why: hours are the easiest thing to put in a spreadsheet. But no institution issues degrees on the strength of how quickly it marked. It issues them on the strength of judgements it can defend.
The institutions that get the most out of AI in marking over the next few years will, I suspect, be the ones that asked a slightly awkward question first. If we had to justify this mark in eighteen months, to someone sceptical, what would we actually be able to put on the table?
A SIMPLE TAKEAWAY
Time saved is a benefit. It is not a strategy.
Judge AI in marking by the quality of the evidence, the decision and the record it leaves behind.
If a faster mark is also a thinner one, you have not made marking better. You have made it quicker to regret.
WE ARE HERE TO HELP
If you are working out where AI belongs in your marking process — and where it does not — we would be glad to talk it through, including the parts we are still cautious about. Please reach out or request a demo.