Every six months, Twitter and Bluesky circle back to whether LLMs are “stochastic parrots,” though the debate has become increasingly confusing as proponents of the metaphor have started using it in very different ways. This blog post argues that the metaphor is counterproductive to modern AI discourse, even in its mildest forms, and gives an overview of the debate around it.
For the uninitiated, “Stochastic Parrots 🦜” is a term from a highly influential paper by Bender et al., published at FAccT 2021, a flagship AI Fairness conference. The paper raised various concerns about the then-emerging trend toward ever-larger language models, e.g., their environmental costs, the biases embedded in training data, and the risks of deploying systems that produce remarkably fluent text. (Famously, this paper led to the firing of two of the paper’s authors, Timnit Gebru and Margaret Mitchell, from Google.)
But what really stuck from this paper was the Stochastic Parrot metaphor: that a useful mental model for thinking about LLMs is that they are next-token predictors. In their words, an LLM would be “a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning.” Hence, a stochastic parrot.
The argument here draws from another paper, Bender and Koller 2020, which distinguishes between form, the observable “stuff” of language, and meaning, which connects those forms to something in the real world, like people’s communicative intentions. Using this distinction, understanding means recovering meaning from form. The authors then argue that, because LLMs only ever observe form, they can never recover meaning from linguistic form alone.
But it is worth remembering what “LLM” meant when these papers were written. GPT-3 had just been released, and it had been trained only to predict the next token. These systems, however, would undergo several revolutions in the years that followed. Most relevant to the debate at hand, models would: 1) receive “post-training” using reinforcement learning methods; 2) become embedded in the real (virtual) world, being given access to tools that interact and shape online environments and multimodal, learning from and operating over images and audio in addition to text; and 3) become components of increasingly elaborate agentic systems with scaffoldings that allow them to keep episodic and long-term memory, and to span and coordinate with other sub-agents.
Thus, the big question is: in this ladder of increasingly sophisticated models, where does the Stochastic Parrots metaphor stop being useful? These enumerated developments do not form a strict hierarchy, but they give us several places to ask that question (which, to be fair, many people have been doing, even before any sort of development). Let me walk you through some of what I consider to be the key evidence contra this argument at each rung:
Rung #0: Pretraining. Even for models trained solely on next-token prediction, the training objective doesn’t tell us what algorithms or internal representations the model will learn to accomplish it. Models trained only to predict sequences have been shown to learn surprisingly rich internal representations of the processes that generate those sequences. For example, a model trained to predict moves in the game “Othello” learned an internal representation of the board state. My Princeton colleague Sanjeev Arora developed, alongside Anirudh Goyal, a theoretical framework suggesting that, as models scale and improve at predicting text, they can combine skills in ways they never encountered during training.
Rung #1: Post-training. During post-training, methods like RLHF (or DPO) and RLVR optimize models against human judgments or externally verifiable outcomes. Thus, model behavior is not merely shaped by “what word is most likely to come next.” For example, in the DeepSeek-R1-Zero paper, applying reinforcement learning increased performance on math problems (AIME 2024) from 15.6% to 71%. More broadly, rewarding correct answers with RL seems to improve the intermediate reasoning used to reach correct answers.
Rung #2: Tool use and multimodal capacities. The original argument for Stochastic Parrots rested on the premise that LLMs had access only to linguistic form and could not connect language to anything outside language. Yet, modern models can use tools to search the web, execute code, query databases, and interact with external environments. In systems like ReAct, models alternate between reasoning and taking actions. Multimodal models such as Flamingo learn from images and text. Thus, models arguably “ground” language expressions in something other than language.
Rung #3: Agentic systems. Models are now embedded in multi-agent systems where a lead agent can spawn sub-agents that investigate different parts of a problem in parallel, then synthesize their findings. They have memory scaffolding that allows these systems to retain plans and intermediate results. These systems have achieved striking results that defy the parrots metaphor. For example, an advanced version of Gemini Deep Think achieved gold-medal-level performance at the 2025 International Mathematical Olympiad, solving five of six newly posed problems, while Anthropic reported that Claude Opus 4.6 discovered 22 Firefox vulnerabilities during a collaboration with Mozilla.
The original authors of the papers have acknowledged limits to the metaphor’s scope in various places. Bender notes in a FAQ about the term that “image/text models [...] can be argued to meet the definition of understanding in Bender and Koller, 2020,” albeit in an “extremely thin” sense. In a recent interview, Gebru draws the boundary elsewhere, maintaining that LLMs remain stochastic parrots even when post-trained or embedded within agentic systems, but distinguishes the LLM itself from the larger system around it. Mitchell, in a blog post, argues that applying the metaphor to larger stateful agentic systems is a “category error,” although “Large language models (LLMs) are.”
Yet, authors have largely held that this metaphor remains useful to understand the behavior of LLMs because it still accurately describes the output mechanism of language models, e.g., in a podcast for Teaching in Higher Ed, Bender says: “the point of the phrase stochastic parrots is to try to make vivid what’s happening when you use a large language model to synthesize text to repeatedly answer the question, what’s a likely next word?”
But here’s the catch. In practice, use of the “Stochastic Parrots” metaphor repeatedly slides between describing how models generate text and making a stronger claim about what they can understand or do. A motte-and-bailey, if you will. The motte is the trivially true observation that LLMs generate text by repeatedly predicting probability distributions over the next token. The bailey is the claim that originally gave the metaphor its “bite,” that LLMs would be “haphazardly stitching together” linguistic forms, without meaning or understanding.
Importantly, the bailey argument is central to a central thesis of some of the movement against the political and economic project of AI: that recent AI progress is mostly bullshit, mere “parroting” dressed up as intelligence. We can find evidence of this back-and-forth very easily in public discourse around AI. E.g., Bender, in her FAQ, e.g., presents the metaphor as a description of what LLMs do, but in her book with Alex Hannah, includes the term among the deliberately ridiculous phrases readers are invited to substitute for “AI”:
“Every time we write “AI”, imagine we have a set of scare quotes around it. Or if you prefer, replace it with a ridiculous phrase. Some of our favorites include “mathy maths”, “a racist pile of linear algebra”, “stochastic parrots” (referring to large language models specifically), or Systemic Approaches to Learning Algorithms and Machine Inferences (aka SALAMI).”
In her recent interview, Gebru also alternates between the two frames. When challenged on the metaphor, she says: “these are definitions of what large language models are.” Yet she also invokes the framing to explain the factual unreliability of Google’s AI overviews, attributing mistakes to the fact that models generating the overviews “are stochastic parrots trained to give you the most likely sequences of text based on their training data.”
To me, personally, calling out this slippage matters because we need an informed public more than ever. I believe that the tremendous energy that the general public has around AI right now must be channeled to drive development and deployment to be safer and saner. In this context, the careless use of the “Stochastic Parrots” metaphor will likely do the opposite: lead the public to misunderstand the capabilities and the nature of frontier AI systems and make people more likely to support policies that will not be effective in minimizing their potential harms (and maximizing the potential benefits, too!).
As I often parrot myself, it is productive to treat AI as normal technology; transformative, but deeply shaped by people and institutions. We should care deeply about and oppose the AI industry’s political and economic project (see here for more). But this opposition does not hinge on dismissing the underlying technology. If anything, taking those capabilities seriously makes accountability all the more necessary.


They going to have to admit, and bite the bullet that llms have utility/value or go back again.