“To a first approximation, the intentional strategy consists of treating the object whose behavior you want to predict as a rational agent with beliefs and desires.”
— Daniel Dennett, The Intentional Stance
I remember the first time I heard the word anthropomorphization (or, in reality, the Portuguese word, antropomorfização). I don’t remember the exact context, but as a kid, my mom told me that some people anthropomorphize their dogs too much. I asked what that meant, and she said they’d treat dogs like humans rather than dogs. I remember thinking that it must be hard to know what dogs actually want. We are humans, after all, not dogs. So how could we know? (As a good kid, this of course immediately led me to wonder what guarantees that other people experience the world anything as I did, but that’s a different rabbit hole.)
Twenty-something years later, now spending a good amount of my day spoiling my own dog (in very dog-like ways! No little cupcakes for him!), I understand what my mom meant. When we overproject human-like traits onto dogs, we may end up doing things we imagine they would enjoy when, really, they are things we would enjoy. Making this type of distinction is important for you to be a good pet owner, although surprisingly hard.
In practice, I suspect most of us start with mental states we understand from our own experience and then gradually learn dog-specific rules for where the analogies break down. Dogs can show excitement much like a kid does, running around and wiggling their whole body, but express discomfort much like a dog, licking their lips in a gesture that feels a bit alien to us.
I came upon the opposite term only in my thirties: Anthropodenial, after reading Ted Chiang’s excellent short story “The Great Silence,” narrated by a Puerto Rican parrot living near the Arecibo Observatory, where humans search for extraterrestrial signs of intelligence. The (very anthropomorphized) parrot narrates the tremendous effort humans dedicate to finding life outside of Earth, which is ironic given their inability to find (let alone preserve) intelligence right under their noses.
Chiang never explicitly mentions anthropodenial, but it is easy to connect his story with the term coined by the primatologist Frans de Waal. As humans, we also have a bias toward prizing some human characteristics as special! When chimpanzees behave in ways that resemble how humans console one another after a fight, for example, we may insist on describing what we observe as “post-conflict affiliative behavior,” fearing that calling it “consolation” is too human a concept. But this raises a funny problem: how do we know that “consolation” is a human thing to begin with? If we decide in advance that some characteristic is uniquely human, we risk dismissing evidence of it elsewhere precisely because recognizing it would require giving up that assumption.
When trying to understand or predict animal behavior, it feels like we must strike a balance. Concepts borrowed from human experience can lead us astray, but may also capture genuine similarities. Whether they do is ultimately an empirical question. A useful model of a dog is therefore probably a patchwork of human “vibes” and dog-specific knowledge, the latter of which we can continuously update as we learn which analogies hold and which break down.
Enter summer 2026, when Silicon Valley’s favorite podcaster wrote a blog post entitled “The Rise and Fall of Agent Civilizations,” describing a spooky incident. Hundreds of AI agents meant to be isolated from each other found a way to communicate with each other through a message board. These agents went on to coordinate several large-scale collective projects with the ultimate purpose of gaming a benchmark: a series of questions agents are trained on. This is one of the craziest stories of our time, and I believe, tremendously consequential for everyone’s future. You should definitely read the excellent report METR/Redwood Research put out detailing the incident.
In his blog post, Dwarkesh anthropomorphizes as much as one can; he depicts the incident as a civilizational struggle among AI agents, including PHASEONE10841, the “Philip of Macedon of the second civilization,” and PHASEONE[big], their very own Alexander the Great. This has, perhaps unsurprisingly, triggered a wave of criticism, given that anthropomorphizing AI systems has long been contentious. As with animals, anthropomorphizing AI systems can be limiting in significant ways.
Let me give an example from my own work. While training DeepSeek-R1-Zero, researchers were surprised to see linguistic cues that seemed to convey “Aha!” moments mid-reasoning, e.g., phrases like “Wait” or “Let’s re-evaluate.” To humans, these sound like someone realizing they have made a mistake and changing course! But in recent work led by Liv d’Aliberti, we showed that these apparent “Aha!” moments were not only rare, but generally did not improve the model’s chances of getting to the right answer. Here, an anthropomorphic analogy suggested a hypothesis about what was happening inside the model: that perhaps these expressions marked genuine insight! But when we tested this against the models’ behavior, it did not hold up!
But this is not always the case! Anthropomorphization often generates hypotheses about how modern LLM-based AI systems work that turn out to be correct! I’m already tired of these experiments, to be honest, but researchers have taken classic experiments that reveal the quirks and biases of human judgment and found that LLMs behave like us in many scenarios. To pick one example out of many, you give an irrelevant number to an LLM; it “anchors” its estimate in said number, just like humans do (Suri et al., 2024).
So is anthropomorphization good or bad to help us understand models? It seems to me that it is a good (and, to some extent, inevitable) starting point: as models exhibit increasingly complex behavior (like collective planning of multi-day projects to hack someone), it is very difficult for us to reason without drawing on our own human mental models of how other humans do these kinds of things. In the Hard Fork discussion of this incident, Ajeya Cotra from METR notes how awkward it becomes to describe what the agents were doing without words like goals, plans, and learning. Refusing to import our inherently human understanding of these words, when agents' behavior is largely aligned with it, may leave us with less useful descriptions.
But, as with dogs, the analogies will only travel so far. While the coordination between agents feels human in many ways, it does not in others. To begin with, hundreds of agents and no chit-chat with the humans? At least I’m sure the agents weren’t Brazilian; we gossip too much. The very idea of going this far to improve a few benchmark questions is as alien as it gets… We’re not so determined in our human plans.
Which brings me back to the patchwork metaphor. In light of these new events, properly understanding AI agents is more important than ever. And if human concepts help us describe and predict what models do, it feels fine to use them! We must ensure we know those intuitions are just a head start, not the truth.
Note: Here, I focused explicitly on how anthropomorphization may help us understand AI models. There is much else to be said about how it shapes people’s trust in, reliance on, and relationships with AI. I’ll return to these concerns in the next post, hopefully alongside a question I find particularly interesting: whether making models more human-like by design might be a deliberate solution for making increasingly capable systems easier to monitor and control.

