AI is starting to dominate our world. According to OpenAI 200 million people receive weekly advice via ChatGPT on anything from productivity to parenting. Universities actively incorporate artificial intelligence (AI) use into assessment guidelines instead of condemning them. Households are managed by voice assistants, job applications are screened by automated algorithms, and audiobooks are narrated by artificial voices. As AI shifts from novelty to necessity, I believe it is crucial that we explore what linguistic biases we are embedding into these systems, and what sociolinguistics can tell us about developing fairer technology.
AI language models overwhelmingly privilege Standard English varieties such as Received Pronunciation (RP) and Standard Southern British English (SSBE) in their output. Training data scraped from the internet disproportionately represents formal, written registers and prestige accents, underrepresenting regional, ethnic, and working-class speech. This is not accidental. It reflects centuries-old hierarchies that sociolinguistics has long documented. One 2007 linguistic survey of 5,010 UK participants found industrial city accents like Birmingham, Liverpool, and Glasgow consistently rated lowest for prestige, whilst RP dominated positive evaluations despite representing fewer than 10% of speakers [1]. When this study was replicated in 2022, the same hierarchy persisted, with working-class regional accents still systematically downgraded [2]. Fifty years of social change, yet our linguistic prejudices remain remarkably stable.
The consequences extend far beyond aesthetic preferences, as automatic speech recognition systems show higher error rates for non-standard varieties, women, elderly speakers, and people who code-switch [4].
The UK government’s Biometrics and Forensic Ethics Group has warned that legitimate speakers with regional accents might be flagged as “suspicious” because their speech patterns are less represented in training data [5]. When AI becomes the gatekeeper to essential services such as healthcare appointments, applications, and customer support, these technical failures exclude the groups already facing a systemic disadvantage. This is not an inconvenience, but structured inequality.
The audiobook industry offers a case study. Linguistic research consistently shows that human narration outperforms synthetic voices, whether rated on enjoyment, mental imagery, narrative engagement, and information recall [6]. Participants experience lower heart rates, better attention, and predominantly positive emotions with human voices, whilst synthetic voices provoked sadness, anger, and disgust [6]. Yet when Audible announced plans to expand into AI-generated audiobooks in 2025, it seemed that the pressure to utilise synthetic voices had won. The Writers’ Guild of Great Britain responded: “AI narration means a worse product for listeners. It still sounds odd and jarring, especially at any length, which means we cannot lose ourselves in the story” [7].
Perhaps more troubling is how these systems already encode linguistic hierarchies. When researching the voice synthesis software ElevenLabs’ library, I found there were 32 RP speakers described with the terms “sophisticated”, “authoritative”, and “credible”. Yorkshire-accented voices meanwhile, are listed as “down to earth” and “perfect for a call centre agent”. The single West Country voice has “only a hint of the accent”. Scouse speakers possess “working-class grit”. These are not neutral descriptions. They are active reproductions of centuries-old prejudices, training both AI systems and users to associate particular voices with particular social roles.
Equally important is recognising that linguistic diversity is systematic, not random. Features like TH-fronting (pronouncing “think” as “fink” or “brother” as “bruvver”) and glottal stops (replacing the /t/ sound in “butter” with a catch in the throat) follow predictable geographical and social patterns [8]. Rather than treating non-standard features as errors requiring correction, AI systems could be designed to recognise legitimate linguistic variation. This requires training data representing actual speech communities, not idealised standard forms.
New research tools offer us hope to investigate these biases, the software Salient Language in Context (SLIC) can be applied in linguistic studies to capture real-time listener attention to specific speech features. Listeners are able to click whilst audio plays in order to note features or parts of speech they find relevant [9]. This and similar tools could identify which acoustic characteristics trigger bias in both human and machine processing, enabling researchers to create targeted interventions. The technology exists, what we lack is the will to prioritise linguistic diversity over convenience.
As AI becomes daily infrastructure, the linguistic ideologies embedded within these systems shape who gets access, opportunity, and inclusion. Sociolinguistics has spent decades documenting how language variation intersects with power: how accent intersects with class, how “correctness” enforces social boundaries, and how linguistic prejudice operates as a socially acceptable form of discrimination.
Allowing AI development to proceed without linguistic expertise means perpetuating accent hierarchies and social inequality. When ChatGPT advises parents, when voice assistants control smart homes, and when AI tutors educate children, the voices they speak with and the varieties they recognise matter profoundly. We can build systems that reflect linguistic reality in all its diversity, or we can dress eighteenth-century prescriptivism in twenty-first-century code.
The choice, ultimately, is in the hands of the companies that create this software, but every day, these biases become more deeply embedded in the infrastructure of modern life.
References
[1] Coupland, N., & Bishop, H. (2007). Ideologised values for British accents. Journal of Sociolinguistics, 11(1), 74-93.
[2] Sharma, D., Levon, E., & Ye, Y. (2022). 50 years of British accent bias: Stability and lifespan change in attitudes to accents. English World-Wide, 43(2), 135-166.
[3] Levon, E., Sharma, D., & Ilbury, C. (2022). Speaking up: accents and social mobility. Sutton Trust and Citi Foundation.
[4] Coto-Solano, R. (2022). Computational sociophonetics using automatic speech recognition. Language & Linguistics Compass, e12474.
[5] Biometrics and Forensic Ethics Group. (2025). Briefing note on the ethical issues arising from the public sector use of biometric voice recognition technology. Retrieved 26/01/2025, from GOV.UK: https://www.gov.uk/government/publications/public-sector-use-of-biometric-voice-recognition-technology-ethical-issues.
[6] Rodero, E., & Lucas, I. (2023). Synthetic versus human voices in audiobooks: The human emotional intimacy effect. New Media & Society, 25(7), 1746-1764.
[7] WGGB. (2025). WGGB responds to Audible AI plans. Retrieved 26/01/2025, from The Writers’ Guild of Great Britain: https://writersguild.org.uk/wggb-responds-to-audible-ai-plans/.
[8] Docherty, G., Foulkes, P., & Kerswill, P. (2024). Phonetic and Phonological Variation in England. In S. Fox (Ed.), Language in Britain and Ireland (pp. 70-97). Cambridge: Cambridge University Press.
[9] Montgomery, C., Walker, G., & Woods, H. (2025). Salient Language in Context (SLIC): a web app for collecting real-time attention data in response to audio samples. Linguistics Vanguard.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.