Every week, another institution asks us the same question: “Can WISEflow Originality just tell us who used AI?“
It’s an understandable ask. Generative AI has made two very different problems collide. There is similarity - students copying each other, recycling old submissions, lifting from a source without attribution - a problem higher education has measured, documented and adjudicated for decades. And there is AI-assisted writing, a genuinely new problem that arrived almost overnight. It’s tempting to treat them as the same challenge with the same fix: run a check, get a score, act on it. They are not the same problem, and only one of them can be solved that way.
Originality detection works because it has something to point to. A similarity match against an identifiable source - a previous submission, a published article, a classmate’s essay - is falsifiable: the assessor can open the source, place it next to the submission, and see the overlap for themselves. That is evidence in the ordinary sense of the word, and it is why similarity checking has always been able to support a formal misconduct case.
AI detection has no equivalent to point to. It does not detect AI; it estimates, from statistical patterns like sentence predictability and rhythm, how much a text resembles machine-generated writing. Once that estimate is produced, there is no source document to hold it against, no way to independently confirm or deny it. It is an inference about a process that has already finished and left no trace. A recent peer-reviewed analysis of AI detectors in education sets out exactly why that gap can’t be engineered away.
That difference matters more than it sounds. Every commercial AI detector, however carefully built, is a probabilistic tool being asked a yes/no question it structurally cannot answer with certainty - because nobody, not the institution and not the vendor, can ever know the true rate of AI use in a given cohort. Change that unknown number and the same detector’s real-world reliability swings wildly, even though its advertised accuracy hasn’t moved at all. Independent researchers who tested fourteen detection tools found accuracy that looked strong in a lab and fell apart in a live cohort, and a separate study of GPT detectors found non-native English writers disproportionately caught in the gap. One well-known vendor’s own accuracy figures were even challenged by independent testing that found it got a meaningful share of a small sample wrong in the other direction, wrongly flagging genuinely human writing. And OpenAI itself withdrew its own AI text classifier, citing its low rate of accuracy - a company building the underlying models saying, in public, that reliable AI detection is not currently achievable.
Most institutions apply a balance-of-probabilities standard to misconduct: it must be more likely than not that something happened, on evidence that is credible and checkable. A percentage with no source behind it doesn’t meet that bar on its own - but in practice, it still changes who carries the burden. A flagged student is asked to explain themselves, produce drafts, prove where their ideas came from. No student can conclusively prove a negative: that a piece of writing was not produced by AI. Quietly, the presumption of innocence inverts, and it inverts for everyone, not just the students who used AI improperly. Blanket AI screening turns a targeted, proportionate response into an indiscriminate one - every submission placed under low-grade suspicion the tool cannot actually justify. Real cases of international students wrongly flagged show what that suspicion costs the people who did nothing wrong, and students already tell us it creates a climate of unease that has nothing to do with whether they did anything wrong.
There’s a second layer to this that gets far less attention than it deserves: what happens to the work itself. Submitting student writing to an AI detection service means sending personal, copyright-protected data to a third-party system - often one whose terms reserve the right to use submitted content to improve its own models, and often hosted outside the EU. The student was never asked. The institution, as the data controller, is the one left explaining under GDPR why that transfer happened, on what legal basis, and where the data ended up. That is a real exposure, not a hypothetical one, and it sits alongside a growing regulatory reality: under the EU AI Act, systems used to evaluate or monitor student performance sit in the high-risk category, with real obligations around transparency, human oversight and the right to a meaningful explanation - as we’ve set out in more detail in our take on what the Act actually requires of assessment. An unverifiable statistical score is a difficult thing to explain to anyone - least of all to the student on the other end of it.
We understand the appeal of a quick fix. A tool that returns a number feels objective, fast, and scalable in a way that redesigning an exam does not. But reaching for AI detection to solve the problem of AI misuse is like reaching for a cannon, or a spread of buckshot, to deal with a flock of sparrows: it looks decisive, it hits a lot of things, and very little of what it hits is actually the target. It doesn’t just miss - it puts the wrong people in the blast radius.
This is why WISEflow Originality doesn’t offer generative AI detection, and why we’ve made that a deliberate choice rather than a gap to fill later. What it does offer is exactly what a misconduct case needs: AI-driven semantic similarity across 50+ languages, matched against identifiable sources, with sentence-level evidence an assessor and a student can both look at and check. It also happens to catch a great deal of what institutions actually contend with day to day - our own research consistently finds that most copying is peer-to-peer, between students at the same institution. And it’s built the way European institutions need it built: EU hosting, institution-controlled sharing, full audit trails, and no data quietly leaving the region to train someone else’s model.
The contest over AI misuse in assessment won’t be won by a better detector - because there isn’t one, and by the admission of the people building the underlying AI, there may never be. It will be won the way academic integrity has always ultimately been won: through assessment design that doesn’t rely on catching people out after the fact, through diversifying how students demonstrate their learning, through clear and explicit rules about where AI help is welcome and where it isn’t, and through embracing AI as a legitimate part of the toolkit where it genuinely belongs. Our research into why students turn to plagiarism in the first place, and the underlying reasons behind rising academic misconduct, points the same way: at pressure, unclear expectations and grey zones, not at a detection gap.
That’s a longer road than running a script and getting a percentage back. It’s also the only one that actually gets somewhere — and it’s the one we’ve built WISEflow Originality to support.
Underneath all of this is a single question that outlasts any one exam or any one submission: does everyone in the room still trust the process? Students need to trust that they’ll be judged fairly, on evidence, not on a probability score they can never disprove. Assessors need to trust that the tools in front of them give them something solid to stand on when they make a judgement call. And institutions need to trust that the systems they’ve put in place will hold up - to a student, to a tribunal, to a regulator - long after the case is closed.
That trust is built slowly and spent quickly. A quick fix that looks decisive today but turns out to be unfounded, unexplainable or unfair erodes exactly the confidence it was meant to protect - and it doesn’t come back easily. Our job isn’t to hand institutions a shortcut that trades away that trust for a percentage. It’s to help build the kind of academic integrity that earns and keeps it, on the long run, one honest, evidenced decision at a time.