Assessment in the Age of AI: Detection Is Not a Pedagogy

Many people in higher education remain understandably confused about how to manage assessment in the time of AI. The problem is not going away. Students are already using generative tools in varied, everyday ways: to interpret rubrics, plan work, structure arguments, obtain feedback before submission, make notes, revise, rehearse presentations and support accessibility. Jisc’s 2025 report also records students using AI for drafting and research support, while some avoid it altogether because they are unsure whether any use will be regarded as cheating (Attwell, S. 2025).
This is an important starting point. We are not dealing with a simple division between students who use AI to deceive and students who do not use it at all. We are dealing with changing study practices, uneven access to guidance, legitimate concerns about over-reliance, and a growing anxiety about what counts as authorship. Yet the institutional reflex is often to try to stop students using AI, or at least to detect it after the event.
Turnitin represents one strand of that response. Its services were built around plagiarism detection and have become woven into the assessment infrastructure of many universities. But the recent report that the University of Southampton is winding down use of Turnitin, and will not renew its contract after 2026–27, shows that the issue is now wider than detection accuracy (Resultsense. 2026). According to Resultsense’s report of a Times Higher Education story, Southampton’s concern was proposed licence terms that could permit student submissions to be used for training AI services. Turnitin says it does not train its Clarity composition tool on customer or student work, while also saying that anonymised submissions may be used to improve detection and assessment (Resultsense. 2026).
The details may yet change, and this account is necessarily based on secondary reporting. Nevertheless, the episode raises an important question. Student work is not simply a convenient dataset. It is the product of learners’ effort, often created in conditions of unequal confidence and power, and submitted because an institution requires it. Universities cannot talk about consent, privacy and ethical AI in the abstract while treating their students’ writing as an asset that may be repurposed through opaque contractual terms.
The second strand of the response is now emerging from the AI companies themselves: watermarking. Anthropic has announced that future Claude models will generate text containing a watermark, saying that the change will help it comply with the EU AI Act (Anthropic, 2026) How Claude’s text watermark works. The watermark is not a visible label or a hidden character. It is a statistical pattern created by using a secret key to steer the model’s random choices between plausible words. A detector holding the key can then estimate the likelihood that Claude was involved in generating a passage (Anthropic, 2026).
Anthropic is careful to describe the limits. The watermark does not establish that a text was written by Claude, or by any AI system; it only offers a probability of Claude’s involvement. It is less effective for short passages, factual prose, light editing and code. It also cannot identify a particular user, organisation or conversation. Anthropic says that its testing, as well as the earlier SynthID research on which its approach is based, found no measurable effect on readability, creativity or quality (Anthropic, 2026).
So far, so technically interesting. But the response to the announcement has demonstrated why assessment cannot be solved by technical detection alone. John Gruber’s objection was deliberately forceful: he argued that it is “patently offensive” for a tool to allow anything other than the user’s needs to affect its word choices (Daring Fireball, 2026). The point is not simply that a reader can see a difference in the prose - Anthropic says they cannot. It is that the system is making choices partly to enable later tracing. For critics, this turns writing generated for a user into writing that is also, silently, reporting on that user.
Tim Moon takes the argument further in his essay, A Digital Scarlet Letter (Moon, T. (2026). His concern is not that Anthropic is necessarily misrepresenting the mechanics. It is that a probabilistic signal of model involvement can become a cheap institutional sorting rule. A watermark may show that Claude contributed something to a text; it cannot tell us whether a student asked for a single sentence to be edited, used it to brainstorm, accepted a full draft unchanged, or understood the work they submitted 5. As Moon puts it, a technical flag may be accurate and still be nearly useless as evidence of misconduct.
There is a further irony here. Students determined to hide improper use may be the most willing to rewrite, translate, paraphrase or switch systems in order to remove a trace. Students who use a tool openly, perhaps for accessibility, feedback or language support, may leave more evidence behind. Detection can therefore become less a means of identifying dishonesty than a mechanism for identifying those who have not concealed their use carefully enough. That would be a very poor foundation for fair assessment.
None of this means that universities should give up on academic integrity. On the contrary, integrity matters more when the tools available to learners are powerful, ordinary and unevenly understood. But integrity is not the same thing as surveillance. A detector cannot replace academic judgement, and a watermark cannot tell the story of a student’s learning.
The alternative starts with greater openness. Universities should be explicit about what types of AI use are permitted, expected, discouraged or prohibited in each assessment. Students should be supported to explain how they have used AI, what they retained, rejected or changed, and why. Assessment design should create opportunities to demonstrate understanding through dialogue, reflection, iterative work, practical application and process evidence.
This is also where AI literacy belongs. Students need to understand how generative systems work, including their limitations, biases, incentives and tendency to produce convincing but unreliable output. They should understand why a watermark is neither a proof of cheating nor a guarantee of authorship. Staff need the same literacy, together with the confidence to discuss AI use honestly instead of relying on automated scores. And providers need to be more transparent about how data is used, how detection systems are trained, what their error rates are, and what a given signal can and cannot show.
The central question for assessment should be “How do we enable students to show what they understand, what they can do and how they have developed their work?” That is a harder question than simply banning AI use. It cannot be outsourced to Turnitin, Anthropic or any other company. But it is also the question that higher education ought to have been asking all along.
References
Anthropic. (2026). How Claude’s text watermark works. https://www.anthropic.com/news/claude-text-watermark
Attwell, S. (2025). Student perceptions of AI 2025. Jisc. https://www.jisc.ac.uk/reports/student-perceptions-of-ai-2025
Daring Fireball. (2026). Anthropic’s “watermark” text adulteration in Claude [Social media post]. Mastodon. https://mastodon.social/@daringfireball/117106850045916140
Moon, T. (2026). A digital scarlet letter. Critical-AI-Solutions. https://timmoon.substack.com/p/a-digital-scarlet-letter
Resultsense. (2026. Southampton drops Turnitin over AI training data fears. https://www.resultsense.com/news/2026-08-19-southampton-drops-turnitin-training-data/
About the Image
In order to consolidate their power and funding, elites, tech companies and organisations promote technological hype and the discourse of inevitability. Although this process is often presented as a natural phase of innovation, it is actually an intentional instrument of techno-authoritarianism which exaggerates the positive effects of technology while downplaying the negative ones. Through socio-financial speculation, accelerationism, and socio-technical fictions such as AGI, they evade democratic oversight and social regulations, transposing the socio-Darwinian mechanism of economic survival into the social, political, and cultural spheres.
