Somewhere in your institution there is a slide with a number on it. Level two, perhaps, or “Intentional”, or an amber dot two-thirds of the way along a five-stage arrow. It was produced by a serious exercise, by serious people, and it is almost certainly wrong - not because the underlying work was bad, but because the question it answers is one no institution actually faces.
This is not a bad development. The sector needs a shared language for a genuinely hard problem, and these instruments give it one. But the ladder metaphor that almost all of them share carries an assumption that does not survive contact with assessment, and it is worth saying plainly what that assumption is: that an institution has a level. That readiness is a single property, held by the organisation as a whole, that moves in one direction.
It isn't, and it doesn't.
Take a pattern of the kind that turns up constantly. An institution runs coursework submission and feedback through its VLE with genuinely high adoption - nearly every module, nearly every academic, a decade of habit behind it. Integration to the student record system is solid. Accessibility has had proper investment because someone senior made it a priority in 2021. On any general digital maturity instrument, this institution scores well.
The same institution resolves second-marking disagreements by email. Its integrity policy is applied differently in three faculties because nobody has ever reconciled them. It has no consistent way of showing, after the fact, how a contested mark was arrived at. In the domains where assessment risk actually concentrates, it is nowhere near where its overall score suggests.
Now give that institution a level. Whatever number you choose, it is an average of things that should never have been averaged - and the harm is not that the number is imprecise. The harm is that averages hide exactly the cells that matter, because a strong cell and a weak cell cancel. The institution above will read “we're doing well” and invest in the areas it already leads, because those are the ones with internal champions and existing momentum. The gap that will end up in front of an appeals panel stays invisible, protected by the good score sitting next to it.
There is a second failure mode, quieter and more damaging. A ladder implies an end state, and end states become targets. Once “level three” exists as an institutional goal, work reorganises around producing the evidence of level three - the documented process, the policy on the intranet, the dashboard - rather than around the outcome the level was supposed to indicate. This is not cynicism about universities; it is the oldest finding in performance measurement, and assessment is unusually exposed to it, because so much of what makes an assessment defensible is invisible in a process document. You cannot tell, from a policy, whether two markers who disagreed actually reconciled their disagreement or simply split the difference.
Here is the sharper version of the problem. Readiness is not a property of an institution at all. It is a property of a context.
A university might be entirely ready for formative, low-stakes assessment - fast feedback loops, generous accommodations, relaxed controls, exactly as it should be - and not ready at all for a high-stakes regulated exam where a contested outcome ends up with a professional body. Those are different capability requirements, different proportionality judgements and different evidence obligations. A single institutional grade cannot distinguish them, which means it cannot help with the only question leadership actually has: given what we assess, and what happens when we get it wrong, what should we fix first?
That question has four contexts hiding inside it, and they behave differently: formative and low-stakes work; coursework-based summative assessment; exams in controlled conditions; and high-stakes, regulated or accreditation-critical assessment. The control set that is proportionate in the fourth is oppressive in the first. An institution that applies its high-stakes posture everywhere has not become more mature; it has become worse at teaching. An institution that applies its formative posture everywhere has a problem it has not met yet.
Assessment readiness, then, is not a position on a line. It is a pattern.
The honest objection to everything above is that a pattern is harder to use than a number, and that using things is the point.
A level gets an item onto a board agenda in a way that seven domain scores do not. It permits comparison with peers, which is sometimes the only argument that releases budget. It gives a change programme a destination, and change programmes without destinations tend to become permanent. And there is a real risk that a nuanced diagnostic simply produces seven amber cells and a shrug - a more accurate picture of a situation nobody now knows how to act on. Anyone who has watched a maturity exercise land in a university knows that the single number is often doing more political work than analytical work, and that this is not entirely illegitimate.
So the answer is not to refuse to score. It is to score at the level where the score is meaningful, and stop there.
Score each capability domain, separately, on a scale short enough to be honest: 0 for ad hoc, 1 for baseline stable, 2 for scaled and documented, 3 for adaptive and continuously improved. Publish those scores as a heatmap, and do not roll them up. A heatmap is not a refusal to be assessed - it is a more demanding assessment, because it removes the place where weak areas hide. What it gives leadership is not a grade but a priority, and a priority is a more useful thing to take to a board than a grade is, because it comes with an obvious next sentence.
That is the design principle behind the UNIwise Assessment Trust & Readiness Model, and it is the reason the model deliberately produces no overall label.
The model scores readiness across seven capability domains: Integrity and Authenticity; Feedback and Learning Impact; Operational Excellence; Security and Assurance; Accessibility and Equity; Interoperability and Data; and Governance and Change Capability.
They are chosen so that the ones institutions tend to be good at and the ones they tend to be weak at cannot cancel each other out. Interoperability and Data is usually the strongest domain in the set, because integration work has had budget and attention for fifteen years. Governance and Change Capability is usually the weakest, because it has no natural owner and no obvious artefact - although the EU AI Act's 2027 obligations are, for once, giving it both. In a rolled-up score, the first hides the second. Side by side, the pattern is the finding: an institution with excellent plumbing and no decision rights is not immature, it is specifically and fixably stuck, and the fix is not technical.
Two of the domains are worth a word on framing. Security and Assurance is scored on whether controls match the stakes - over-control counts against you, as it should, because blanket surveillance applied to a formative quiz is a governance failure and not a safety margin. Proportionate control is layered control, and a readiness model that rewarded maximum control would be rewarding the wrong instinct. And Accessibility and Equity is a domain rather than a compliance footnote, because an assessment some students cannot sit properly is not a defensible assessment, whatever its audit trail says.
There is one more assumption buried in most maturity frameworks, and it is the one with the clearest commercial fingerprints on it: that the route upward runs through consolidation. Standardise the platform, unify the ecosystem, retire the local tools, and maturity follows.
Sometimes that is right. Often it is not, and the exercise that concludes it should always be treated with suspicion - including when the exercise is ours. If an institution's VLE is already the operational hub for coursework, with a decade of academic habit behind it, then the highest-value, lowest-risk improvement is almost always to extend capability where the work already happens: rubrics used as quality assurance rather than paperwork, a real moderation workflow, a marking record that survives a challenge - delivered into the VLE through LTI rather than through a procurement round. That is what Marking Studio is for. Where the risk profile genuinely demands more - controlled conditions, proportionate monitoring, defensible integrity checks - those capabilities layer on afterwards, by context, in the order the heatmap says.
If an institution's VLE adoption is thin or fragmented, the honest answer is different, and a readiness model that cannot say so is not a diagnostic. It is a brochure.
Writing in HEPI this week about the cohort arriving in 2027, Roney Lima do Nascimento made an argument about students that applies just as well to institutions: diagnose on entry rather than assume. His point was that a department which can articulate what competence looks like in its discipline can teach it, while one that only knows what it wants to forbid cannot. The institutional version is the same. A university that can say precisely which of its assessment capabilities is weakest, and in which context, can fix it. One that knows only that it would like to be more mature cannot.
That is what the readiness check is for. It is live and free, at uniwise.eu/digital-assessment-readiness. It takes ten to fifteen minutes and runs in three steps - institutional context, then VLE and operations, then the seven capability domains - and it returns your heatmap on screen immediately. There is no single number at the end of it, and that is not a limitation of the tool.
What you get is a picture of where your readiness is uneven, which is the only shape institutional readiness ever actually has. What you do with it is a conversation, and a better one than “how do we get to level three” - because the answer to that question was never a level. It was a list, in order, of the specific things that would not hold up if someone asked.
External sources cited