The question is arriving on more and more desks - Here is what is at stake
Over the past two years, several suppliers of assessment and academic integrity software have sought permission to use the work students submit as training material for artificial intelligence. Some have raised it openly with their customers. Others have changed the terms that students accept at the moment they hand in an assignment.
Either way, a growing number of European universities are being asked to take a position on it, often at short notice and often without a clear picture of what the decision involves.
This post sets that out in plain terms: what training on student work means, the legal problems it runs into in Europe, and the questions of trust that sit alongside the law.
What “training on student data” actually means
It is worth being precise here, because the phrase gets used loosely and the distinction matters later.
When a supplier stores your students’ submissions, the files sit in a database. They can be searched, exported and deleted. If you end the agreement, they can be returned or destroyed, and you can ask for confirmation that it has happened.
Training works differently. A model learns patterns from the material it is shown, and those patterns end up spread across millions of numerical values inside the model. The original essay is not sitting in there as a document you could open. But it has not disappeared either: it has shaped what the model produces.
That difference has one very practical consequence. You can delete a file. Removing one student’s work from a model that has already been trained on it is an unsolved problem. There is a whole research field devoted to it, known as machine unlearning, and it exists precisely because no one yet has a reliable method. A promise to delete data on request therefore means something quite different once training has taken place.
Why student work is such attractive material
The interest from suppliers is easy to understand.
Training a model to evaluate written work requires examples of written work that has already been evaluated. Suppliers of assessment tools hold enormous quantities of exactly that: original writing by identifiable people, paired with an expert judgement of how good it is, why, and against which criteria. Marker comments, rubric scores and moderation notes make the material more valuable still.
Very little material of that quality exists anywhere else. The question is whether suppliers are allowed to use it.
Students cannot realistically say no
European data protection law requires anyone processing personal data to have a valid legal reason for doing so. Where a supplier wants to use submissions for its own purposes rather than yours, the most obvious basis would be the student’s own consent.
Consent under the GDPR has to be freely given. The regulation itself notes that consent is presumed not to be freely given where there is a clear imbalance of power between the person giving it and the organisation asking.
Consider the situation in which the permission is usually sought. A student is submitting an assignment in a system your university requires them to use. There is a deadline. They are presented with updated terms and asked to accept them in order to hand in. There is no alternative platform, no option to submit some other way, and no realistic possibility of declining and still completing the course.
That is not a free choice in any meaningful sense. It is close to the clearest example of the imbalance the regulation has in mind.
There is a second thing worth noticing about that moment, which is easy to overlook. If the supplier is obtaining permission directly from students through its own end-user agreement, it is dealing with them without you in the room — inside a system you selected and made compulsory. A supplier processing data on your behalf has no ordinary reason to negotiate separate rights with your students. That is itself worth asking about. It is also the kind of question that makes compliance look less like red tape and more like the thing protecting the people in your care.
The work is not the university’s to give
This is the aspect most often missed, and it belongs to a different area of law altogether.
Students own the copyright in their own work. They are the authors, and authors own their work from the moment it is created. There is no employment relationship between a student and a university, so none of the rules that transfer ownership to an employer apply. Most universities state this explicitly in their own intellectual property policies: by default, the rights in work produced as a student belong to the student, including dissertations and theses.
This leads somewhere that catches many institutions by surprise. If the copyright belongs to the student, you cannot license it to anyone. It is not possible to give away something you do not own. So when a supplier’s terms take a licence over student submissions, no agreement you sign can grant it. The permission has to come from the student — which is exactly why it turns up as a click-through at the point of submission.
The two problems are therefore closely connected. The only person who can lawfully grant the licence is the one person who is not in a position to refuse it.
It is also worth reading what the licence says. Terms of this kind commonly take a licence that is perpetual and irrevocable, and state expressly that it continues even if the customer stops using the service. In practice that means a university which reviews the terms, decides it is not comfortable with them and moves to another supplier has protected future students but not past ones. The licence over work already submitted continues.
In Denmark and across much of continental Europe there is a further point. An author’s moral rights — the right to be credited and to object to uses that damage their reputation — cannot be waived except in relation to a clearly defined use. A perpetual, worldwide licence covering unspecified future product development is unlikely to meet that description.
The underlying question — who owns what, and what you can take with you — is the same one that runs through getting your assessment data back out of a platform. It is a governance question long before it is a technical one.
Anonymisation helps less than it sounds
The most common reassurance is that only anonymised submissions would be used. It sounds like a complete answer. It is not, for two reasons.
The first is that anonymity has to be demonstrated rather than stated. European data protection regulators have been clear that a model trained on personal data cannot simply be assumed to be anonymous. The test is whether personal data could realistically be recovered — either from the model itself, or by prompting it in ways that make it reproduce what it was trained on. Meeting that test requires documented evidence, including testing to show the model does not reproduce its training material. It is a reasonable thing to ask a supplier to produce, and it is the same standard of evidence we have argued for when suppliers make confident claims about AI detection.
The second reason is simpler. Copyright does not depend on the author’s name being attached. An anonymous essay is still a protected work. Removing identifying details addresses part of the data protection question and none of the ownership question.
It changes who carries the responsibility
European data protection law distinguishes between two roles. The data controller decides why and how personal data is processed, and carries the legal responsibility for it. The data processor acts on the controller’s instructions and nothing more. In this relationship your university is the controller and the supplier is the processor.
That distinction has real consequences. If a processor starts deciding its own purposes for the data, it stops being a processor and becomes a controller in its own right, with the liability that follows.
Some of the consequences land on your side of the table. Your record of processing activities — the register of what personal data your university processes and why — describes that supplier as a processor. The privacy information you give students says their work is processed for assessment and academic integrity. Your data protection impact assessment, the risk analysis sitting behind both, rests on the same description. If the supplier has taken on a purpose of its own, those documents no longer describe reality, and they are your documents.
There is also a plain question of balance. The supplier gains a lasting asset it can sell, potentially to other universities. You carry the risk, answer to your students, and answer to your data protection authority.
Where European rules are heading
Two developments are worth knowing about, without going into detail.
The first concerns copyright. European law does allow text and data mining in certain circumstances, but the exception for research organisations covers scientific research rather than commercial product development, and the broader exception can be overridden where rights have been reserved. It was written with publicly available material in mind, not unpublished coursework that students are required to hand in. Alongside it, providers of general-purpose AI models must now maintain a copyright policy and publish a summary of what their models were trained on.
The second concerns AI regulation. The EU AI Act classifies systems used to evaluate learning outcomes and determine access to education as high-risk. The obligations attached to that classification were originally due in August 2026 and have since been postponed to December 2027 — which, as we have written elsewhere, is time to prepare rather than permission to wait.
The direction in both areas is towards more documentation and more demonstrable diligence. A supplier that cannot account for how its training material was obtained is likely to find this harder over time, not easier.
The part that is not about law
The legal position matters, but it is not the whole of the objection.
What students submit is often personal. Nursing and social work students write up placements involving real patients and real families. Dissertations are frequently drawn from bereavements, diagnoses and family histories. Students write these things for an examiner, within a relationship whose limits they believe they understand. That understanding is what changes when the same material is used to build something else. It is part of the same broader question we examined in Are we assessing learning – or just grading production?
There is a wider cost too. Universities are working hard to introduce AI into assessment in ways that staff and students can trust, and that work depends on a limited supply of goodwill. Every story about student work quietly becoming training material makes the next conversation about AI harder for everyone, including the institutions doing it carefully.
Where we stand
It would be unreasonable to write all of the above without stating our own position.
Your students’ submissions are not our training data. Not now, not in anonymised form, and not later under revised terms. We will not ask your students for a licence over their work, because we do not think that permission can be meaningfully given at the moment we would be asking for it.
You are the controller of your data. We are the processor. We act on documented instructions, and if we wanted to do something outside them, the right thing is to ask, openly, with you able to say no.
We are also building AI. WISEflow’s AI-assisted feedback is moving into production, and we think it will genuinely help assessors and students. The difference is what the AI works from: our approach uses your own assessment criteria, rubric and marking guidance at the point of use, rather than a body of past student work. When we want to understand how a feature performs in practice, we ask — named institutions, an agreed scope, an evaluated pilot, and the option to stop. We then publish what we find, including the parts that did not work: the findings from our pilots at UiT and BI Norwegian Business School are on this site, and the reasoning behind the module is set out in why feedback and grade justification matter.
It is a slower way to build. It produces something an institution can explain to its students, which seems to us the only sensible basis for working in this field.
Questions worth asking any supplier - Including us
- Do your terms take a licence over student submissions, and who grants it — the institution or the student?
- Does that licence continue after our agreement ends, and what happens to work already submitted?
- Are submissions, feedback, grades or marker comments used to develop, train or evaluate any model, whether yours or a third party’s?
- If the answer is that data is anonymised, what evidence can you show us, and may we see the assessment behind it?
- If we ask you to stop and delete, what exactly can be deleted, and what would remain inside a trained model?
- Will these commitments sit in our contract rather than in terms you can change without our agreement?
And one that is easy to forget: what happens to these commitments if the company is acquired?
None of this is really about contracts. It is about whether the supplier is a partner you can rely on over years, which is why the vendor matters as much as the platform.
In summary
Students hand in their work because they are required to. They own what they write. You do not, which means you cannot pass those rights on, and neither can anyone acting for you.
Everything else is detail. If a supplier finds that proposition difficult to answer plainly, that in itself is useful information.
Rasmus Blok is Chief Evangelist and co-founder of UNIwise, the company behind WISEflow.
WE ARE HERE TO HELP
Wondering what your AI training terms actually say - or should say? We're happy to talk it through with you, no obligation. Reach out or request a demo.