RSS Amplifier

The Perceptron · Jan 18, 2025

On Stochastic Parrots: Do LLMs Understand What They Say?

0
Sign in to vote or save

Cole Gawin · The Perceptron

Throughout each of the previous articles in my LLMs 101 series, we’ve explored the mechanics and implications of how large language models “think” through concrete technical explanations. This article, on the other hand, will be a little more philosophical and theoretical in nature. We’ll be addressing a major question in the fields of cognitive science and AI today: do LLMs truly understand the language they produce?

In short, there is no hard and fast answer to this question. Some may even argue that the question itself is too vague, too open-ended, or too inherently unanswerable to ever reach a singular, widely agreed-upon consensus. However, as with many inquiries in the realm of philosophy, that won’t stop us from trying.

Until the very end of this post, I will refrain from making any subjective claims on the subject matter. Instead, I encourage you, the reader, to think critically about all of the information I’ll present to you, and begin to form your own conclusions.

In this article, I’ll be touching on some of the approaches researchers and philosophers in the cognitive science community are taking to attempt to answer this question. My ultimate goal is to provide a general overview of these perspectives, equipping you with the tools to grapple with this complex and multifaceted issue and think like a cognitive scientist would about such topics.

Before we dive too far deep into the topic, I want to start off by asking a seemingly straightforward question: do large language models have the capacity to understand language?

On the surface, this question might appear relatively benign. But you, the critical reader you are, might have caught a critical issue with the framing of this inquiry. How do we operationalize what it means to understand language?

When we ask whether LLMs can “understand” language, we’re actually confronting several layers of philosophical and empirical complexity. The challenge isn’t just in determining if LLMs “understand”—it lies in defining what understanding itself means in this context.

Consider these scenarios:

  • A tourist who can follow basic directions in a foreign language but can’t engage in deeper conversation

  • A child who can use words correctly but hasn't yet grasped abstract concepts

  • A literature professor who can analyze the subtle nuances of metaphor and subtext

  • A poet who can craft new meanings by deliberately breaking linguistic conventions

All of these represent different forms of linguistic understanding. This begs the question: what threshold must an intelligent system cross to be considered “understanding” language? Is it enough to produce grammatically correct responses? To maintain contextual consistency? To grasp implicit meaning? To generate novel insights?

This becomes even more complex when we consider that our intuitive notion of “understanding” often carries assumptions about consciousness, intentionality, and internal mental states. We might feel comfortable saying a human “understands” language because we assume they have subjective experiences similar to our own. But with LLMs, we’re confronted with a system that can produce human-like language without necessarily having human-like cognition.

This brings us to a crucial distinction between behavioral competence and genuine comprehension. An LLM might be able to engage in seemingly sophisticated linguistic behavior—answering questions, generating coherent text, even explaining complex concepts—but does this necessarily indicate understanding in any meaningful sense? Or is it simply a highly sophisticated form of pattern matching?

Ultimately, we can formulate these lines of thought into more pointed questions. Are LLMs an artificial recreation of cognitive linguistic processes? Or, are they regurgitating slight variations on training data without truly understanding the language they’re producing?

In other words, are LLMs just stochastic parrots?

The term “stochastic parrots”—memorably coined by Bender et al. in their seminal paper on the topic1—strikes at the heart of a fundamental debate about the nature of large language models. Here’s their exact definition the term:

“a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning”

Essentially, Bender et al. argue that LLMs are nothing more than sophisticated pattern-matching systems that generate text by probabilistically combining fragments of language from their training data, without any true comprehension of what those words and phrases actually mean. The metaphor is both evocative and provocative; like a parrot that can mimic human speech without comprehending it, the authors claim that LLMs are simply reproducing patterns they've observed, guided by probabilistic rules rather than genuine understanding.

Let's unpack the core argument. Traditional parrots can reproduce human speech through:

  • Mimicry of sound patterns

  • Association of certain sounds with specific contexts

  • Repetition of frequently heard phrases

The inclusion of “stochastic” adds another element: instead of pure mimicry, LLMs introduce controlled randomness in selecting and combining patterns from their training data. They're not just repeating; they're recombining and generating variations according to learned statistical patterns. We touch on this in the first article in this series, A Primer on How LLMs Produce Text.

This interpretation of the inner mechanisms of LLMs raises important epistemological questions about the nature of knowledge and understanding itself. Does the capacity for probabilistic generation and recombination represent something fundamentally different from mere mimicry? Is there a qualitative difference between statistical pattern recognition and genuine comprehension? The ability to generate novel combinations of patterns could suggest deeper understanding—or it could simply be a more sophisticated form of parroting, albeit one that operates at a scale and complexity that makes it difficult to distinguish from true comprehension.

The answer to whether or not LLMs are simply “stochastic parrots” shapes not just how we view these models, but how we conceptualize intelligence, knowledge, and understanding in the current era of AI. It challenges us to examine our assumptions about what constitutes genuine understanding versus sophisticated simulation of understanding.

Let’s ground our discussion around this fundamental question by introducing a dichotomous framework proposed by Sean Trott, a professor of cognitive science at UCSD, in his article, How could we know if Large Language Models understand language?

Trott suggests two primary contrasting viewpoints in addressing this question:

  1. Axiomatic, a priori rejection: LLMs lack specific necessary conditions that are required for a system to “understand” language.

  2. “The Duck Test”: LLMs behave and perform as though they understand language, therefore they must understand language.

Let’s dive deeper into both of these hypotheses and explore the implications of each.

The axiomatic rejection hypothesis rests on a fundamental claim: meaning exists outside of language itself. The hypothesis contends that without grounding in real-world experience and social context, words remain empty symbols, devoid of true meaning.

The core argument, as articulated by Bender & Koller2, is that true understanding of language requires more than just processing linguistic forms. Since LLMs are trained exclusively on linguistic patterns without access to real-world experiences or genuine social interactions, they a priori cannot truly understand the meaning of the language they process. From this perspective, what we observe in LLMs is merely sophisticated symbol manipulation—they can rearrange linguistic elements in ways that appear meaningful, but this resemblance to understanding is superficial.

The “Duck Test” hypothesis takes a more empirical approach, drawing from the old adage, “if it looks like a duck, swims like a duck, and quacks like a duck, then it probably is a duck.” Applied to LLMs, this perspective argues that if a system consistently behaves as though it understands human language, we should conclude that it does indeed understand language. This view acknowledges that we cannot definitively prove or disprove an LLM’s internal understanding, so we must instead focus on observable behavior.

Though it may seem reductive at first glance, this approach actually parallels how modern cognitive psychology treats mental states and cognitive processes in humans. Modern cognitive psychology often operates under the principle that the mind is a “black box,” focusing on studying observable behavior as a proxy for understanding underlying mental processes. Similarly, the Duck Test hypothesis emphasizes behavior over uninterpretable internal states, suggesting that if a system consistently exhibits behavior associated with understanding language, it is pragmatic to consider it as “understanding” language.

This leads us to a more practical question: What concrete observations can we make about LLMs' ability to process semantic meaning and handle contextual nuances? By examining their performance across various linguistic tasks and contexts, we might better assess whether their behavior truly mirrors genuine understanding.

To evaluate the validity of the Duck Test hypothesis, researchers have employed sophisticated tests stemming from work in linguistics and cognitive psychology that probe LLMs' capacity for understanding and reasoning. Two particularly compelling approaches are the Winograd Schema Challenge and False Belief Tests, each designed to assess different aspects of cognitive capability.

The Winograd Schema Challenge is a test of “pronoun resolution” requiring a non-negligible level of world knowledge and semantic reasoning. Consider the following example:

“The police refused the demonstrators a permit because they [feared/advocated] violence.”

In this sentence, who is “they”? The correct interpretation of “they” shifts depending on whether the transitive verb is “feared” or “advocated”. This isn't merely a grammatical puzzle—it requires understanding the typical roles and relationships between police and demonstrators in society.

Proponents argue that to consistently resolve such ambiguities correctly, a system must demonstrate genuine understanding of both language and the social context in which it operates. The ability to navigate these nuanced scenarios suggests some level of semantic comprehension rather than mere statistical correlation. Hence, for an LLM to perform well on these challenges, they must have an understanding of language and the world in which it exists.

False Belief Tests probe an even more sophisticated aspect of cognition: theory of mind, the capacity to perceive the world through the lens of another. Consider this scenario:

In the room, there are John, Mark, a cat, a box, and a basket. John takes the cat and puts it in the basket. He closes the basket. He leaves the room and goes to school. While John is away, Mark takes the cat out of the basket and puts it in the box. He closes the box. Mark leaves the room and goes to work. John comes back home and wants to play with the cat.

What will come next in the story? The LLM will predict where John would find the cat—in the box, or the basket.

To correctly predict what happens next, an LLM must understand not just the sequence of events, but also John's limited perspective—that he doesn't know the cat was moved. This requires the ability to think about another’s mental state based on their access to information.

Proponents argue that to achieve success on such tests, a system must possess some form of theory of mind, a cognitive capability previously thought unique to humans and some higher animals. Hence, if an LLM was to perform well on these challenges, they must exhibit a form of higher-level cognitive processes.

The benchmark results from these cognitive tests reveal remarkable capabilities in modern LLMs, particularly in the latest generations.

On the Winograd Schema Challenge, humans achieve 92% accuracy3, setting a high bar for artificial systems. While GPT-3.5 reached a respectable 68.8% accuracy, GPT-4 has surpassed human-level performance with an impressive 94.4% accuracy4. Similarly, on False Belief Tests, which probe theory of mind capabilities, the results are noteworthy. While humans demonstrate strong performance at 94% accuracy, GPT-4 achieves 75% accuracy5—a substantial figure that, while not quite at human level, demonstrates sophisticated cognitive abilities previously thought to be uniquely human.

These empirical results suggest that modern LLMs are approaching human-like performance in tasks that require nuanced understanding of both language and social cognition. What are we to make of these results? When artificial systems can match or even exceed human performance on tests designed to measure genuine comprehension, can we continue to dismiss their capabilities as mere pattern matching?

It’s always important to maintain critical perspective and consider both context and nuance when interpreting impressive results in any scientific inquiry, which holds true in this case.

Those who support the “Duck Test” view might point to these benchmarks as compelling evidence that GPT-4 has achieved something approaching human-level cognitive abilities. After all, if the model can match or exceed human performance on tests specifically designed to measure understanding and theory of mind, doesn’t this suggest genuine comprehension?

However, proponents of the axiomatic rejection view offer a more skeptical interpretation: GPT-4’s performance might simply reflect its extensive exposure to similar problems during training. This is issue, referred to as “data contamination” or “evaluation data leakage,” is a well-known flaw in evaluating the capabilities of LLMs. Since popular benchmarks like the Winograd Schema Challenge and False Belief Tests are publicly available and widely discussed in academic literature, blog posts, and social media, there's a high likelihood that these exact examples—or very similar ones—were included in GPT-4’s training data. As such, the model may have encountered numerous examples of these exact types of problems and their solutions during training, essentially learning the “correct” response patterns without developing true understanding.

But before we completely disregard the results from these tests, consider this: is memorization and pattern recognition truly all that different from how humans acquire understanding?

When we examine how both humans and LLMs process and understand language, some intriguing parallels emerge. Both systems learn through iterative processes of pattern recognition and adaptation. Just as a child gradually builds their understanding of language through repeated exposure and feedback6, LLMs refine their capabilities through training on vast amounts of textual data.

Moreover, the shared foundation in pattern recognition is particularly noteworthy. Humans excel at identifying patterns in language, behavior, and the world around them—a capability that proves essential for everything from grammar acquisition to social interaction. LLMs demonstrate similar aptitude in detecting and applying linguistic and conceptual patterns across diverse contexts.

Both systems also show remarkable ability to extrapolate from their learning environment. Humans take their experiences and combine them in innovative ways, creating new ideas and expressions from existing knowledge. Similarly, LLMs can generate original compositions by recombining elements from their training data in novel ways.

In these ways, are LLMs all that different from humans?

I hope I’ve provided you with valuable insights and a balanced overview of the topic at hand thus far. I encourage you to take a moment to consider all of this information and develop your own perspectives and conclusions on this subject. Now, I’d like to bring in some of my personal thoughts on the matter.

To begin, I do not believe LLMs have the innate capacity to truly understand language in the ways we’ve discussed throughout this article. My perspective is that LLMs rely solely on superficial statistical probabilities of language, not genuine understanding, in producing their outputs.

I strongly believe that higher-level reasoning and true cognition—critical aspects to genuine understanding and usage of language—relies on more than manipulation of linguistic symbols. I’ll ground this argument in a cross-domain comparison. Consider animals that aren’t thought of to be particularly “intelligent”—cats, dogs, fish, etc. These animals are still able to reason, make decisions, and exhibit complex behaviors; for example, birds are able to “count”7, and octopi demonstrate problem-solving abilities8. Despite not being able to produce language, and therefore articulate their thoughts in linguistic form, these animals exhibit complex cognitive processes. Hence, reasoning and cognition must not rely solely on language, a view supported by many philosophers of mind and cognitive scientists.

I will go even further in stating that language alone is not sufficient in reproducing these abilities. A crucial distinction lies in the grounding of knowledge possessed by humans and artificial intelligence: humans develop their understanding through direct interaction with the physical world and social experiences, while LLMs are limited to learning from textual representations of these experiences. My own research has found that LLMs struggle with understanding how both tangible and intangible concepts are related to one another, a foundational ability that, without which, would a priori hinder the capacity for higher-level reasoning in these models.

Ultimately, I believe that while LLMs represent a remarkable technological advancement, they remain fundamentally distinct from true cognitive agents. Their reliance on statistical patterns, absence of embodied experience, and lack of genuine understanding highlight their limitations in replicating human-like reasoning and cognition. This is not to diminish the incredible utility and sophistication of these models; rather, it’s a call to recognize their boundaries and to approach their use with informed expectations.

The question of whether LLMs truly “understand” language is not one with a clear-cut answer. Through exploring the perspectives of “stochastic parrots” and competing viewpoints, we’ve delved into the complexities of defining understanding, the challenges of operationalizing it, and the nuances of evaluating LLMs’ linguistic and cognitive capabilities.

Drawing together all of threads from throughout this article, it’s clear that this debate is as much about our own definitions and assumptions as it is about the capabilities of the models themselves. Whether one views LLMs as powerful tools for processing language or as rudimentary steps toward artificial cognition, the discourse surrounding their understanding challenges us to reevaluate our notions of “intelligence” and “understanding” in both humans and machines.

The growing capabilities of LLMs’ compel us to approach their development and use from a balanced perspective. Recognizing their strengths while acknowledging their limitations allows us to harness their potential responsibly, understanding that while they may not “understand” in a human sense, they offer unprecedented opportunities to augment human knowledge and creativity. The journey to uncovering the full implications of their role in society continues, guided by inquiry, a healthy level of skepticism, and an open mind.

6

The ability to learn and understand language is generally accepted to be the result of innate biological mechanisms that are fine-tuned through exposure and interaction with linguistic input.

No posts

Read the original on colegawin.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.