I have the impression that most people don’t know how AI voice mode works. This is not surprising. But in order to understand the inherent shortcomings of voice AI, it is essential to understand the following: the audio output generated by AI in voice mode is not speech and it is not generated in response to spoken input. It is text-based.
What does that mean, and why does it matter?
Let’s ask Chat.
Here’s a transcription of me asking ChatGPT about voice mode.
If you’re not paying attention, you might overlook that point about “speech recognition” and “text-to-speech.” What happens when you use an AI application in voice mode is that AI uses automated speech recognition technology to transcribe a user’s spoken input into text. Then, AI’s output is generated in textual form in response to the transcript (again, not in response to the spoken input itself). And finally, the textual response is relayed to the user as audio, using text-to-speech technology. Considering all the steps involved, and how quickly this multi-step process is completed, it is certainly impressive from the standpoint of user experience.
But as a linguist, the transcript underscores the artificiality of it. For one thing, it looks very neat… Too neat.
Human speech is messy. When people participate in spoken interactions with other people, lots of things are happening at once. The person who’s speaking may stop mid-sentence, or even mid-word, and change what they’re saying. Interruptions are common. Sometimes, an interruption results in the first person stopping, ceding the floor to the other. Other times, the first speaker doesn’t stop talking, and the two people talk over each other for a few words, or sentences. Still other times, two people talking at once isn’t treated as an interruption at all, but may instead be interpreted as agreement. While one person has the floor as the main speaker, another person may be continually back-channeling, “uh-huh, ummm, I know. Yeah, I think so too….” In some languages, like Japanese, back-channeling is expected and considered a sign of engaged listening. Rather than being seen as interrupting, back-channeling encourages the speaker to continue, and both speaker and listener treat the back-channeling (which, let’s be honest, is acoustically messy) as normal.
Sometimes the messiness of human speech looks almost like magic. One person may start a sentence, and a second person finishes it, with all the right grammar and correct words in place, almost giving the impression that the second person is a mind-reader. Or, even in the absence of an explicit interruption, a speaker will stop mid-sentence, letting another speaker take over. How did the sentence-finisher know what to say? How did the first speaker know that the second speaker wanted to talk?
These speakers are relying on a range of cues, both linguistic (or verbal) and otherwise, that help them create a communicative interaction out of the messiness of human speech. The linguistic cues are the easiest to describe. If someone says “I went to…” we can guess that a place name or location will follow. The non-linguistic cues, sometimes called paralinguistic features, are harder to describe, but no less important. Examples of paralinguistic cues include a slight raise of the eyebrows to show confusion or surprise. Or, when several people are speaking together, one person may shift their gaze to another person, conveying that a question is intended for that person specifically. Even when speakers cannot see each other, paralinguistic cues are used, such as a gentle intake of breath that conveys that the next person is getting ready to speak. The timing and placement of pauses is another example. Most people expertly interpret pauses, correctly understanding when a pause represents thinking, uncertainty, or displeasure…
These are just a few examples of the resources people use to turn the messiness of human speech into successful communication. Sociolinguists analyze speech using transcription conventions intended to capture all the elements, all the messiness, not just the words themselves. We notate these paralinguistic elements in the transcript because we recognize that in order to understand what’s happening (to analyze a conversation), we need access not only to words, but also to all of those messy elements. Transcription conventions generally include symbols to show overlap (two speakers speaking at the same time), pauses, in-breaths, laughter, rising or falling intonation… and many, many more. Not to mention transcription choices that are designed to represented phonological characteristics (e.g., an individual speaker’s pronunciation).
Consider this example:
Compared to the transcript from ChatGPT, this one is much, much messier. And harder to read, especially without a glossary of the transcription conventions being used. But unlike ChatGPT’s transcript, the sociolinguist’s transcript contains a great deal of additional information and gives us an in-depth look into how two people are able to make sense out of messiness.
What AI users need to keep in mind is the fact that voice AI is not spoken communication. Voice AI works through the step-wise process I explained above. To recap: 1) Human speech is transcribed (using automated speech recognition) into a simplified textual representation. 2) AI then uses the written transcription to generate its output, which is a response to the transcript, in other words, a response to text, not to speech. 3) AI’s textual response is then relayed to the user in audio form, using text-to-speech technology that simulates human speech.
Because voice AI is responding to a textual transcript, not to actual human speech, AI misses all the cues that humans naturally produce in the course of speaking. Once you understand this, it helps explain a lot of the shortcomings of voice AI. For example, AI can’t correctly interpret in-breaths or pauses because those are missing from the transcript. And AI can’t handle back-channeling, because back-channeling gets represented in the transcript in the same way as any other words, so AI treats back-channeling as an interruption that means it should stop talking.
If developers involved in AI research really want to create a voice AI application that accurately simulates human spoken interaction, the first thing they should do is include some sociolinguists on their teams. Otherwise, something will always be lost in transcription.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.