AI and Bernstein

Talking to the Machine: AI, Language Codes, and the Reproduction of Inequality

This article explores how generative AI systems reinforce class-based linguistic hierarchies by privileging elaborated codes aligned with middle-class norms, drawing on Bernstein’s theory of language and sociological critiques of educational inequality.

As artificial intelligence (AI) becomes increasingly embedded within the infrastructures of education, employment, governance, and communication, the capacity to interface effectively with such systems is emerging as a critical axis of stratification. No longer confined to speculative discourse, generative AI tools such as ChatGPT, Copilot, and Gemini are actively reconfiguring how individuals seek information, formulate arguments, and present themselves in academic and professional contexts. The social consequences of these tools extend far beyond questions of access and digital fluency; they demand a rigorous examination of the epistemic and linguistic regimes that underpin AI’s functionality. At the core of this transformation lies a question that is as linguistic as it is political: whose modes of speaking are legible to the machine, and whose are rendered opaque, deviant, or invisible?

Language Codes and Class Systems in AI Interaction

This article situates AI-human interaction within a sociological framework informed by Basil Bernstein’s theory of language codes. Bernstein’s work provides a foundational analytic for theorising the intersection of language, class, and education, particularly through his distinction between elaborated and restricted codes. These are not merely stylistic differences but socially situated communicative repertoires shaped by differential access to institutional contexts. Elaborated codes, characterised by syntactic complexity, decontextualised references, and an orientation towards abstraction, are cultivated through middle-class socialisation and institutional reinforcement. Restricted codes, by contrast, emerge from working-class social contexts and are marked by implicit references, contextual specificity, and a more relational syntax (Bernstein, 1971, 2000; Sadovnik, 2001).

These distinctions do not imply deficits on the part of restricted code users; rather, they highlight the ways in which institutional environments are structured to reward certain linguistic behaviours. When AI systems are trained primarily on elaborated code, these same institutional preferences are transferred into the digital sphere, shaping which forms of expression are deemed legitimate. This continuity between educational stratification and AI response architecture reveals the deep structural alignment between technology and existing classed communication norms (Roberts, 2011; Jenks, 2024).

AI’s Linguistic Bias Towards Elaborated Code

While AI tools hold considerable promise and can be deployed meaningfully to enhance access to knowledge, support learning, and enable more equitable participation in digital spaces, it is essential to critically examine the socio-linguistic foundations upon which their effectiveness rests. Generative AI systems, whose outputs are shaped by training datasets composed largely of formalised textual corpora—academic literature, legal documents, corporate communications, and journalistic media—display a pronounced affinity for elaborated code. The architecture of these models, particularly in their reliance on pattern recognition and probabilistic inference, privileges syntactic precision and semantic clarity (Birhane, 2021).

Consequently, user prompts articulated in elaborated code tend to elicit responses that are more contextually appropriate, informationally dense, and structurally coherent. These prompts align with the expectations of the algorithm, which has been designed to parse complex sentence structures and respond to abstract queries. Conversely, inputs grounded in restricted codes—those more conversational, elliptical, or embedded in particularistic social cues—are often misinterpreted, prompting vague, irrelevant, or overly generic responses. This disparity reproduces a latent communicative hierarchy in which certain speech forms are legitimised and others pathologised (boyd & Crawford, 2012; Shieh et al., 2024).

Reproducing Inequality: Linguistic Capital and AI

Such outcomes are not merely technical artefacts but reflect a deeply rooted sociological logic. Bernstein’s findings consistently demonstrated that the linguistic styles most valued by educational and institutional systems were those most readily accessible to middle-class students. AI communication replicates this same dynamic. As with classroom interactions that favour elaborated codes, AI systems reward inputs framed in institutionally sanctioned forms of speech. The ability to communicate effectively with AI technologies thus becomes a form of linguistic capital, reinforcing existing hierarchies by privileging those already proficient in dominant linguistic repertoires (Bourdieu, 1991; Reay, 2004).

Furthermore, this form of capital operates invisibly. Users who do not receive high-quality responses may blame themselves for lacking clarity or knowledge, rather than recognising the structural biases embedded in the system. In this sense, AI reinforces meritocratic myths, where communicative success is mistaken for individual aptitude rather than alignment with privileged linguistic norms (Littler, 2017; Warr & Heath, 2025).

Educational Consequences and Social Reproduction

In educational contexts, the stakes are particularly acute. AI tools are being rapidly adopted as pedagogical aids, homework assistants, and scaffolding devices, ostensibly designed to enhance equity and access. Yet their effectiveness is contingent upon a student’s ability to formulate questions in ways that the machine recognises as meaningful. This places students from working-class backgrounds—those more likely to utilise restricted codes—at a systemic disadvantage.

Educational inequality is therefore no longer solely mediated through teacher-student interaction or school resources, but also through algorithmic interpretation. AI systems do not simply reflect the biases of their training data; they operationalise and amplify them. Students whose linguistic style aligns with elaborated code benefit from AI as a supplementary tutor, while others face opaque or unhelpful feedback loops. This reproduces existing educational disparities while concealing them under a veneer of innovation (Williamson, 2017; Rockey & Broadfoot, 2025).

Linguistic Erasure and Epistemic Injustice

Moreover, the assumption that AI is accent-neutral or dialect-agnostic masks a deeper form of linguistic erasure. While contemporary voice-recognition systems may no longer overtly penalise regional accents, the underlying textual processing mechanisms remain attuned to standardised grammar, syntax, and lexical conventions. Dialectal variations, idiomatic expressions, and culturally specific idioms are routinely flagged as errors or paraphrased into normative formulations.

This constitutes a subtle yet pervasive form of epistemic injustice: the linguistic forms through which marginalised communities articulate meaning are rendered illegible, and thus excluded from legitimate knowledge production (Fricker, 2007; Noble, 2018; Godwin-Jones, 2023). This form of erasure has implications not only for user experience but also for cultural preservation. When AI standardises communication, it also erodes the linguistic diversity that sustains alternative worldviews and social practices.

Towards an Inclusive Digital Future

This critique does not imply a wholesale rejection of AI technologies. On the contrary, when designed and deployed reflectively, AI can offer significant opportunities for inclusion, creativity, and personalised support. It is precisely because of this potential that we must be attuned to the risks of embedding unexamined linguistic assumptions into systems that increasingly mediate access to knowledge and resources.

Efforts to democratise AI must go beyond improving access or reducing costs; they must confront the linguistic and epistemological foundations of these technologies. Developers, educators, and policymakers must work collaboratively to ensure that AI systems reflect a broad spectrum of linguistic practices. This includes diversifying training data, incorporating regional and class-based vernaculars, and embedding sociolinguistic expertise in AI design (Rockey & Broadfoot, 2025)..

More ambitiously, it may require a rethinking of what constitutes effective communication. Rather than training all users to adopt elaborated code, systems should be designed to interpret diverse linguistic inputs without imposing hierarchical valuations. This shift would re-centre communication as a relational practice rather than a technical transaction (Gross, 2023).

The Role of Sociological Research

Sociological research has a vital role to play in documenting and theorising how different social groups engage with AI, how linguistic capital is accrued or denied in these interactions, and how AI-mediated communication is reshaping the politics of voice and authority. There is an urgent need for empirical studies that centre the experiences of those on the periphery of digital and linguistic power—rural youth, racialised communities, neurodivergent individuals—whose interactions with AI may be characterised by friction, misrecognition, or exclusion.

Ethnographic and participatory research can illuminate the lived dimensions of these dynamics, offering granular insights into how language operates as a gatekeeping mechanism. By foregrounding user narratives, sociologists can challenge dominant paradigms in AI ethics and contribute to the development of more inclusive technologies (Couldry & Mejias, 2019).

So… Who Gets Heard by the Machine?

In conclusion, speaking to the machine is never a neutral act. It is shaped by the historical sedimentation of classed language practices, by the cultural hierarchies embedded in data infrastructures, and by the political economy of educational access. The findings of Bernstein remain highly relevant to this domain, as the elaborated codes that confer advantage within educational systems are similarly valorised within AI-mediated communication. Without intentional, justice-oriented design and policy, AI is likely to serve not as a democratising force, but as a new vector of symbolic domination—one that speaks in the voice of authority, yet listens only to those who have already learned its language.

References

  • Bernstein, B. (1971). Class, codes and control: Volume 1. Theoretical studies towards a sociology of language. Routledge.
  • Bernstein, B. (2000). Pedagogy, symbolic control and identity: Theory, research, critique. Rowman & Littlefield.
  • Birhane, A. (2021). Algorithmic injustice: A relational ethics approach. Patterns, 2(2), 100205. https://doi.org/10.1016/j.patter.2021.100205
  • Bourdieu, P. (1991). Language and symbolic power (J. B. Thompson, Ed.; G. Raymond & M. Adamson, Trans.). Harvard University Press.
  • boyd, d., & Crawford, K. (2012). Critical questions for big data. Information, Communication & Society, 15(5), 662–679. https://doi.org/10.1080/1369118X.2012.678878
  • Couldry, N., & Mejias, U. A. (2019). The costs of connection: How data is colonizing human life and appropriating it for capitalism. Stanford University Press.
  • Fricker, M. (2007). Epistemic injustice: Power and the ethics of knowing. Oxford University Press.
  • Godwin-Jones, R. (2023). Generative AI and the social functions of educational assessment. Language Learning & Technology. https://godwinjones.com/godwin-jones_ai_llt.pdf
  • Gross, R. (2023). Performative gender and AI communication scripts. Gender and Language, 17(2), 245–268.
  • Jenks, C. (2024). Bias and intercultural trust in generative AI. Journal of Intercultural Communication Research, 53(1), 1–18.
  • Littler, J. (2017). Against meritocracy: Culture, power and myths of mobility. Routledge.
  • Noble, S. U. (2018). Algorithms of oppression: How search engines reinforce racism. NYU Press.
  • Reay, D. (2004). Education and cultural capital: The implications of changing trends in education policies. Cultural Trends, 13(2), 73–86. https://doi.org/10.1080/0954896042000267161
  • Roberts, P. (2011). Bernstein, Bourdieu, and the sociology of pedagogy. British Journal of Sociology of Education, 32(4), 541–561. https://doi.org/10.1080/01425692.2011.578436
  • Rockey, M., & Broadfoot, P. (2025). Generative AI and the social functions of educational assessment. Assessment in Education: Principles, Policy & Practice. https://doi.org/10.1080/03054985.2025.2455549
  • Sadovnik, A. R. (2001). Basil Bernstein (1924–2000). Prospects: The Quarterly Review of Comparative Education, 31(4), 687–703. https://doi.org/10.1007/BF03220012
  • Shieh, J. H., Le, H., & Xu, T. (2024). Laissez-faire harms and subordination in language models. arXiv. https://arxiv.org/abs/2404.07475
  • Warr, K., & Heath, M. (2025). Epistemic inequality and the hidden curriculum in AI. Sociology of Education, 98(1), 41–58.
  • Williamson, B. (2017). Big data in education: The digital future of learning, policy and practice. Sage.

Cite this article

Select a style and copy the citation into your document.

Andrew Wright
Andrew Wright

Andrew Wright is a higher education (HE) professional and PhD researcher specialising in the sociology of education. His doctoral work examines the reproduction of inequality in post-18 transitions, while his broader interests centre on how structural contexts shape life chances. He is committed to bringing sociological perspectives and research-led insight to public audiences.

Leave a Reply

Your email address will not be published. Required fields are marked *

Content continues after this advertisement