
There's an assumption baked into most conversational AI roadmaps that nobody really questions anymore: more human-like is better. Friendlier persona, more natural turn-taking, more emotional mirroring, all treated as obvious improvements, the same way "more accurate" or "faster" would be. But is that actually true? Or have we been chasing a proxy for quality and calling it the goal itself?
Push a chatbot far enough toward human-like, and something strange happens: it can start working against you. This is often described through the idea of the uncanny valley, the discomfort people feel toward something that looks or behaves almost, but not quite, human. A study combining psychophysiology and questionnaires found that a more human-like chatbot produced a stronger uncanny valley effect than a more clearly artificial one, even though people still said they'd want to interact with both again. An e-commerce study found something similar with real business stakes attached: making a chatbot avatar more visually realistic and animated increased feelings of eeriness, which then lowered trust and purchase intent. The humanization was meant to help. But instead, it worked against the goal.
This isn't only a visual effect, either. A study had people chat with a bot deliberately built to sit right in the uncanny valley versus one built to converse naturally, and the uncanny bot scored lowest not just on likability but on perceived intelligence too. The discomfort shows up in plain text, without any avatar involved.

So why does this happen? Gray and Wegner's mind perception research offers a useful answer: what unsettles people about humanlike machines isn't really the appearance on its own, it's the moment people start attributing experience (the capacity to feel) to something, rather than just agency (the capacity to act). A machine that clearly does things doesn't bother us. A machine that seems to feel things does. That's a sharper target than adjusting how the interface looks, because it points at something conversational AI does constantly: language that sounds emotionally attuned invites exactly this kind of attribution, with no face required.
Which raises a bigger question: maybe human-likeness was never the right goal to aim for in the first place. Bender and colleagues' widely cited "Stochastic Parrots" paper makes a related point from a different angle, cautioning against attributing real understanding or meaning to fluent, human-sounding text generated by a language model. The more human a response feels, the more likely people are to assume the system understood something it didn't, or to trust output that's really just a well-shaped guess. Human-like framing doesn't just risk discomfort, it risks misplaced confidence.
Brynjolfsson's "Turing Trap" argument makes a similar case at the level of the whole field: treating human-equivalence as AI's goal is a trap, both economically and in terms of what the technology could actually be good for. AI aimed at imitating humans tends to just replace what people already do. AI aimed at augmenting human ability tends to create genuinely new value while keeping people involved. Applied to conversational AI specifically, this is a real choice to make: Are we building something meant to convincingly pass as a person, or something that's simply excellent at the job it's there to do?
That points toward a more useful design target: competence, transparency, and clear signals about what the system can and can't do, instead of personality and emotional mimicry. A system that's upfront about what it is sets expectations it can actually meet, which connects to a point worth remembering from research on user expectations more broadly: the gap between what people expect and what a system delivers is often where trust breaks down, and an honest, competent tool has an easier time closing that gap than a persona pretending to be more than it is.
None of this means human-like design is always the wrong call. Certain contexts, such as companionship apps or specific accessibility use cases, genuinely seem to benefit from warmth and humanlike framing, and researchers are still debating exactly what drives the uncanny valley effect in the first place. The point isn't that human-likeness is bad. It's that it should be a deliberate choice suited to a specific context, not a default setting applied everywhere because it seemed like the obviously better option.
The interesting design question was never "how human can we make this." It's "what does this system actually need to be good at, and does human-likeness serve that, or just perform it." Chasing the wrong target can feel like progress for a while, even when it isn't one.