AI meets academia: a robot faces a gatekeeping professor holding keys.

Who Gets Accused? AI Detection Bias, Academic Integrity, and the Uneven Field in Higher Education

When universities reach for AI detection tools to police academic integrity, they reach for instruments that are neither accurate nor neutral. Research shows that non-native English speakers, neurodivergent students, and those whose natural writing style resembles AI output face disproportionate false accusations. This article applies a moral panic framework to ask what the response to AI in higher education actually reveals — and who it disadvantages.

Scroll through any university policy page published in the last two years and you will find the same vocabulary: academic integrity, authentic assessment, the threat of artificial intelligence. Turnitin has updated its detection algorithms. Examination boards have issued guidance. Senior academics have written op-eds warning of the death of independent thought. The tone is urgent; the stakes apparently civilisational.

Universities are not wrong to take this seriously. Generative AI represents a genuinely significant shift in what is possible when a student sits down to write an essay, and questions about how that changes the relationship between assessment and learning are legitimate and pressing. The challenge is real.

But sociologists are trained to ask a particular kind of question when intense collective anxiety breaks out around a perceived threat: not only “is the threat real?” but “what does the response itself reveal?” Who it targets, what it protects, and how it draws its lines often tells us something more illuminating than the ostensible subject of concern (Cohen, 1972; McRobbie & Thornton, 1995). Applying that lens to higher education’s encounter with AI does not dissolve the genuine pedagogical questions. It sharpens them.

This article argues that the current institutional response to AI in higher education has the character of what sociologists call a moral panic: a reaction that moralises one category of assistance whilst leaving structurally comparable advantages unexamined. It also argues that this diagnosis does not resolve the underlying challenge. No one yet has the answers, and it is unlikely any single answer exists. What HE needs is a more reflexive, evidence-led approach to adaptation — one that begins by asking better questions than current policy tends to pose.

AI and Academic Integrity as Moral Panic

The concept of moral panic was introduced by Cohen in his study of societal responses to the Mods and Rockers conflicts of the 1960s (Cohen, 1972). A moral panic occurs when a condition, episode, or group of persons comes to be defined as a threat to societal values and interests. Those identified as the threat — Cohen’s “folk devils” — face intense moral scrutiny, their behaviour exaggerated and their significance inflated relative to the actual scale of harm.

What matters here is not whether AI in education constitutes a full moral panic in Cohen’s classical sense, but the analytical value the framework provides. McRobbie and Thornton (1995), in an influential revision of Cohen’s model, argued that in contemporary multi-mediated societies, moral panics rarely take the form of unified crusades. Instead they are fragmented, contradictory, and institutionally managed — overlapping governance responses that nonetheless share the characteristic features: disproportionality, folk devil construction, and the drawing of moral boundaries that serve existing interests.

AI policy in higher education often resembles this pattern. There is no unified moral crusade; there are contradictory institutional responses, inconsistent enforcement, software procurement decisions, revised assessment regulations, and a proliferation of guidance that raises more questions than it answers. What is consistent is the moral vocabulary — integrity, authenticity, the threat to independent thought — and the implicit boundary it draws between acceptable and unacceptable forms of assistance.

That boundary deserves scrutiny. Not because integrity is a fiction, but because the history of how such boundaries are drawn in academia suggests they are rarely as neutral as they present themselves.

The Integrity Regime Was Never Neutral

The current institutional response to AI rests, often implicitly, on an assumption: that before AI, academic assessment was essentially fair. Students who worked hard and thought independently were rewarded; those who sought shortcuts were penalised. AI threatens to destroy a level playing field.

This assumption does not survive examination.

Before AI, access to the means of academic production was profoundly uneven. Students at well-resourced universities had access to comprehensive library databases; those at less prestigious institutions worked with what was available. Some students had parents who could read and comment on draft essays; others had no one. Students whose first language was not English faced barriers that had nothing to do with intellectual capacity. Wealthy students could and did purchase the services of legitimate proofreaders and editors — a practice long widespread and largely tolerated by institutions that simultaneously treat AI-assisted drafting as academic misconduct (Harwood et al., 2009; Salter-Dvorak, 2019).

The proofreading case is instructive. Research on UK university proofreading practices has shown substantial variation in what proofreaders do — from light surface corrections to substantive structural intervention — with little institutional regulation governing the boundary (Harwood et al., 2009). Salter-Dvorak (2019) demonstrated that informal language policies around proofreading in UK master’s programmes generate unequal outcomes: students with greater social and financial capital access more extensive editing support, whilst the practice remains largely invisible to formal integrity frameworks. AI has not created this asymmetry. It has made it harder to ignore.

Supervisory relationships represent perhaps the most significant and least discussed form of advantage in academic knowledge production. A doctoral student supervised by a well-connected scholar at a research-intensive university inhabits a fundamentally different world from one supervised by an overstretched colleague at a teaching-focused institution. This is not a marginal phenomenon. It shapes how academic knowledge is made and how careers are built (Bourdieu, 1990; Davies et al., 2020).

The sociology of knowledge has long established this point. Merton’s account of the norms of science described an ideal that was always aspirational rather than actual (Merton, 1973). Feminist standpoint epistemologists demonstrated that knowledge claims are always situated — produced from particular social locations that shape what questions get asked and whose experience counts as evidence (Haraway, 1988). A standpoint that presents itself as neutral tends to favour those already closest to the prevailing norm.

The integrity regime has always drawn its lines somewhere — and those lines have not historically fallen in neutral places. AI sharpens this problem. It does not create it.

When Enforcement Becomes the Problem

The problem deepens when institutional responses move from policy to enforcement. AI detection tools — now widely deployed across UK higher education — do not assess whether a student has understood their material, engaged with sources, or constructed an argument. They produce a probabilistic score based on linguistic patterns, principally a measure known as perplexity: how predictable and stylistically uniform the text appears to a language model. Writing that is direct, clear, structurally consistent, and lexically restrained scores low on perplexity and risks being flagged as AI-generated. Writing that is ornate, varied, and idiomatically flexible does not.

This has a distributional consequence that is neither accidental nor evenly spread. Research by Liang et al. (2023), testing seven widely-used detection tools, found that over half of essays by non-native English speakers were misclassified as AI-generated — a false positive rate exceeding 50% — compared to near-perfect accuracy for native English-speaking students. The mechanism is structural: non-native writers tend to produce writing with lower lexical variety and syntactic complexity, not because their thinking is less rigorous, but because producing elaborate academic prose in a second language is harder. The detector has effectively enshrined a particular kind of native-speaker fluency as the baseline for genuine intellectual work.

Neurodivergent students face a related risk through a different route. Writing characterised by direct phrasing, structural regularity, or consistent rather than varied sentence construction — features associated with some autistic and ADHD-related processing styles — can in some cases produce the same low-perplexity profile that detectors flag (Gegg-Harrison & Quarterman, 2024). Eaton (2025) situates this within a broader critique of fairness in HE assessment, arguing that detection tools misread cognitive difference as misconduct, adding another layer of disadvantage for students already navigating a system not designed for them. When a false accusation leads to a disciplinary process, the consequences — academic penalties, psychological distress, damaged trust in institutions — fall hardest on those already most marginalised. The wider implications for assessment design are explored in a companion piece on this site.

The student whose writing most closely resembles AI output is not, in many cases, the student who used AI. They are the student whose natural voice — shaped by language background, educational history, or cognitive profile — happens to fall on the wrong side of a probabilistic threshold. A policy that treats a detection score as sufficient evidence of misconduct is not defending academic integrity. It is reproducing the inequalities the system has always contained, now with the added authority of an algorithm.

Does AI Actually Harm Learning? Taking the Critique Seriously

The most thoughtful critics of AI in education are not simply defending institutional privilege. They are raising a genuinely difficult question about cognition, development, and what learning is actually for — and it deserves an honest response.

The argument runs as follows. Academic writing is not merely a means of demonstrating knowledge — it is a process through which knowledge is constructed. The struggle to articulate a half-formed idea in prose, to discover through the act of writing that one’s argument has a gap, to revise under the pressure of a reader’s imagined scepticism: these are not incidental features of academic work. They are the mechanism by which thinking becomes more rigorous. If AI short-circuits that process — producing fluent prose before the intellectual work has been done — it does not merely cheat an assessment. It forecloses the development the assessment was designed to cultivate.

This is a serious argument. Three responses are available to it.

The first is empirical. We do not yet know, with any confidence, that AI use systematically forecloses intellectual development rather than redirecting it. A recent PRISMA systematic review of empirical studies on generative AI in higher education found mixed evidence: some studies reported improved engagement and performance, others identified over-reliance and variable effectiveness across disciplines, and significant gaps remain in longitudinal evidence (Hon, 2025). Treating all AI engagement as harmful is not an empirical position; it is a moral assumption dressed as one.

The second is structural. Even if some uses of AI are genuinely harmful to intellectual development, the policy response has been consistently blunt. Blanket prohibitions treat the student who uses AI to avoid thinking and the student who uses it as a scaffold for their own cognition as equivalent. Assessment design, not surveillance, is the more appropriate institutional response.

The third is the most important. The pedagogical concern, whilst genuine, does not in itself validate the particular forms of assessment currently under defence. The unseen examination and the solo-authored essay are historically specific practices that became dominant partly for reasons of institutional scalability. They have always served some students better than others. Whether AI is harmful to learning is a real question. Whether existing assessment practices are the right vehicle for protecting learning is a separate question — and conflating the two is how structural interest operates without appearing to.

What the Calculator Debate Teaches Us About AI

The most instructive historical parallel is the introduction of the electronic calculator into mathematics education in the 1970s. Mathematicians warned that students would lose the capacity for independent numerical reasoning. Examination boards debated prohibition. Some argued that unaided arithmetic was not merely a practical skill but a marker of genuine mathematical understanding — a claim whose similarity to current arguments about AI and “authentic” writing is close enough to be instructive (Cockcroft, 1982).

With hindsight, both sides were partly right. Calculators changed what mathematical competence looked like in practice. But they did not dissolve stratification in mathematics education; they reorganised it. A student arriving at a GCSE examination with a basic four-function calculator was not equivalently equipped to one whose parents had purchased a scientific calculator with greater capability. Access was universal in principle and uneven in practice, tracking the familiar lines of class, school funding, and parental income.

The AI parallel holds. “Having access to AI” encompasses an enormous range of actual capability. A student using the free tier of a general-purpose chatbot is not equivalently equipped to one using a more sophisticated tool with refined prompting strategies and the critical literacy to evaluate and integrate AI outputs into genuinely sophisticated argument. The technology does not remove the advantage gradient; it redraws it.

The more important lesson is pedagogical. Mathematics education adapted. Calculators became tools to be understood and used appropriately, not simply banned or ignored. The curriculum shifted to foreground the kinds of mathematical reasoning that calculators could not substitute for — conceptual understanding, problem-solving, proof. That adaptation was deliberate, uneven, and imperfect. It took time. But it happened, and mathematics education did not collapse.

AI will require equivalent adaptation. The question is not whether HE will change in response, but whether that change will be deliberate and evidence-led, or reactive and driven primarily by institutional anxiety.

Undergraduate vs Doctoral Education: Why the Problem Is Not Uniform

One of the most significant weaknesses in current policy discourse is its tendency to treat “higher education” as a single entity facing a uniform challenge. It is not. Nowhere is this more clearly illustrated than in the difference between undergraduate and doctoral education.

Doctoral study has structural features that substantially reduce AI’s capacity to substitute for genuine intellectual work. The viva voce examination tests a candidate’s ownership of their research in real time, in dialogue, under sustained expert scrutiny. A thesis assembled with AI assistance rather than intellectually owned by the candidate is likely to be exposed — not incidentally, but by design. As Park (2003) establishes in the foundational account of doctoral viva best practice, the purpose of the oral is to verify independent intellectual engagement with the work. Stephenson and Jackson (2025) reinforce this, identifying confirmation of intellectual ownership as a core examiner function that has gained renewed significance in the AI era.

Sustained supervision provides a further safeguard. A supervisor who has read successive drafts and tracked a candidate’s development over years is well-placed to identify work that does not reflect the student they know. The crisis generating the most urgent policy responses is predominantly an undergraduate assessment crisis, and treating it as a problem for the whole sector obscures more than it illuminates.

The obvious question is whether oral examination and close supervision could be extended to undergraduate level. The answer is almost certainly not. The resource requirements are prohibitive, and the staff-to-student ratios that make doctoral supervision meaningful simply do not exist at undergraduate scale. Recognising this is not defeatism. It is the precondition for realistic adaptation.

Three Questions Every AI Integrity Policy Should Answer

What sociology can offer here is not a curriculum reform or an assessment framework. Those require expertise in pedagogy, institutional context, and disciplinary practice that this analysis does not claim. What it can offer is a set of diagnostic questions — ones that flow from the argument above, and that any institutional response ought to be able to answer honestly before it hardens into policy.

The first question is about consistency: does this policy apply the same standard to all forms of assistance that affect the final product? If an institution prohibits AI-assisted drafting whilst tolerating paid proofreading, ghostwritten personal statements, or the substantial editorial support available to students with well-educated parents, it is not defending academic integrity in any principled sense. It is drawing a line that falls on the newest and most visible form of assistance — and policies built on inconsistent foundations are difficult to enforce credibly.

The second question is about what is actually being assessed: does the assessment task measure what the institution claims it measures, or does it measure the ability to produce a particular kind of text under particular conditions? These are not the same thing, and they have never been. AI has made the gap between them more visible, not larger. An institution that cannot answer this question clearly for a given assessment task has a problem that predates AI and will not be solved by banning it.

The third question is about distribution: who benefits from this policy, and who is disadvantaged by it? As the calculator parallel suggests, access to AI is not uniform. As the detection bias evidence shows, enforcement tools are not neutral. Responses that treat the problem as one of individual student dishonesty, rather than one of uneven capability and uneven institutional support, are likely to reproduce existing inequalities whilst appearing to address a new one.

These are not questions with easy answers. They are, however, the questions that serious engagement with the problem requires — and the fact that current policy discourse largely avoids them is itself part of what this article has tried to explain.

Better Questions, Not Premature Answers

Moral panics pass, or they calcify into policy. The institutional response to AI in higher education will do one or the other — possibly both, at different speeds and in different contexts. Either way, it will leave a trace in assessment regulations, software procurement decisions, and a generation of academic norms.

What this article has tried to show is that the panic-like quality of some of that response is not simply about AI. It is an understandable but poorly calibrated reaction to a genuine challenge — one that, in its urgency, risks defending unequal arrangements as if they were timeless educational values, and reaching for blunt instruments where more carefully designed responses are needed.

The genuine pedagogical concerns are real. So is the unevenness of the problem across different levels of HE. So is the absence of settled answers. A response adequate to this moment needs to hold all of those things in view simultaneously — which is harder than a moral panic, and more useful.

The question to carry into every AI policy debate remains the one this article opened with: what does the response itself reveal? Which forms of help are treated as normal, and which as misconduct — and who benefits from that boundary? That is not an invitation to cynicism or to the conclusion that integrity does not matter. It is an invitation to the kind of reflexive, evidence-led thinking that the complexity of this moment genuinely requires.

The institutions that navigate this best will probably be those honest enough to treat it as an ongoing pedagogical challenge rather than a crisis to be managed into silence.

References

  • Bourdieu, P. (1990). The logic of practice (R. Nice, Trans.). Polity Press. (Original work published 1980)
  • Cockcroft, W. H. (1982). Mathematics counts: Report of the Committee of Inquiry into the Teaching of Mathematics in Schools. HMSO.
  • Cohen, S. (1972). Folk devils and moral panics: The creation of the Mods and Rockers. MacGibbon and Kee.
  • Davies, H. C., Eynon, R., & Salveson, C. (2020). The mobilisation of AI in education: A Bourdieusean field analysis. Sociology, 55(3), 539–560. https://doi.org/10.1177/0038038520967888
  • Eaton, S. E. (2025). Neurodiversity and academic integrity: Toward epistemic plurality in a postplagiarism era. Teaching in Higher Education. Advance online publication. https://doi.org/10.1080/13562517.2025.2583456
  • Gegg-Harrison, W., & Quarterman, C. (2024). AI detection's high false positive rates and the psychological and material impacts on students. In S. Mahmud (Ed.), Academic integrity in the age of artificial intelligence (pp. 199–219). IGI Global. https://doi.org/10.4018/979-8-3693-0240-8.ch011
  • Haraway, D. (1988). Situated knowledges: The science question in feminism and the privilege of partial perspective. Feminist Studies, 14(3), 575–599. https://doi.org/10.2307/3178066
  • Harwood, N., Austin, L., & Macaulay, R. (2009). Proofreading in a UK university: Proofreaders' beliefs, practices, and experiences. Journal of Second Language Writing, 18(3), 166–190. https://doi.org/10.1016/j.jslw.2009.05.002
  • Hon, K. L. (2025). Generative AI in higher education: A systematic review of its effects on learning outcomes and academic performance. Journal of Educational Computing Research. Advance online publication. https://doi.org/10.1177/00472395251400089
  • Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), Article 100779. https://doi.org/10.1016/j.patter.2023.100779
  • McRobbie, A., & Thornton, S. L. (1995). Rethinking 'moral panic' for multi-mediated social worlds. British Journal of Sociology, 46(4), 559–574. https://doi.org/10.2307/591571
  • Merton, R. K. (1973). The normative structure of science. In N. W. Storer (Ed.), The sociology of science: Theoretical and empirical investigations (pp. 267–278). University of Chicago Press. (Original work published 1942)
  • Park, C. (2003). Levelling the playing field: Towards best practice in the doctoral viva. Higher Education Review, 36(1), 47–67.
  • Salter-Dvorak, H. (2019). Proofreading: How de facto language policies create social inequality for L2 master's students in UK universities. Journal of English for Academic Purposes, 37, 30–43. https://doi.org/10.1016/j.jeap.2018.11.004
  • Stephenson, Z., & Jackson, A. (2025). Addressing the shortcomings of the closed-door viva for PhD and doctoral candidates in the UK: A voice for academic examiners. Assessment & Evaluation in Higher Education. Advance online publication. https://doi.org/10.1080/02602938.2025.2475062

Cite this article

Select a style and copy the citation into your document.

Andrew Wright
Andrew Wright

Andrew Wright is a higher education (HE) professional and PhD researcher specialising in the sociology of education. His doctoral work examines the reproduction of inequality in post-18 transitions, while his broader interests centre on how structural contexts shape life chances. He is committed to bringing sociological perspectives and research-led insight to public audiences.

Leave a Reply

Your email address will not be published. Required fields are marked *

Content continues after this advertisement