Google Duplex Was Unsettling. That's the Point.

At Google’s developer conference in May, they played a recording of their assistant telephoning a hair salon to book an appointment. It handled an awkward exchange, adjusted when the salon offered a different time, and used the small verbal noises people make while thinking.

The room applauded. Then, over the following days, a substantial number of people said they found it disturbing, and Google clarified that the system would identify itself in deployment.

I think both reactions were correct, and the gap between them is worth sitting with.

What Was Actually Impressive

Set the ethics aside for a paragraph, because the engineering deserves acknowledgement.

Telephone conversations with strangers are genuinely hard. The audio is poor. People interrupt. They answer a different question than the one asked. They use fragments rather than sentences. Restaurant and salon staff are busy and speak quickly.

Handling that in a narrow domain, well enough that the person on the other end didn’t notice, is a serious result. The disfluencies were the detail everyone latched onto, and they’re not decoration. Silence on a phone call reads as a dropped connection. Filling it is what makes the interaction work.

Why People Recoiled

The discomfort wasn’t about capability. It was about consent.

The person at the salon was in a conversation they did not know the nature of. They were being polite to something that does not experience politeness. They had no opportunity to decide whether they wanted to speak to a machine.

“The problem was never that it sounded human. It’s that the person on the other end had no way to know they were the only human in the conversation.” — Sameer Gupta

There’s something else underneath, which I think explains the strength of the reaction. If machines can hold convincing telephone conversations, then the phone call stops being evidence of anything. A voice you don’t recognise is no longer a person. Every business process that treats a phone conversation as verification, and there are many, quietly weakens.

I flagged something similar a few years ago about generative models and imagery. The pattern repeats: a capability arrives, and its second-order effect is that a form of evidence people relied on stops working.

The Principle Is Simple

I don’t think this is a hard ethical question, which is what makes the initial omission notable.

A system that could be mistaken for a person should say that it isn’t, at the start, unprompted.

That’s it. Not on request. Not in the terms of service. At the beginning of the interaction, in the channel the interaction is happening in.

The objections are predictable and weak:

  • “It ruins the experience.” It slightly changes the experience. California is already legislating on bot disclosure and the sky has not fallen.
  • “People will hang up.” Some will. That’s them exercising a preference you were previously overriding by concealment.
  • “It’s obvious it’s a machine.” Today, in most cases. The entire point of the demo was that it wasn’t.

What This Means If You’re Building One

The chatbot wave of the last two years produced a lot of systems that pretend, in small ways, to be people. Human names. Typing indicators. Deliberate delays before responding. Deflection when asked directly.

Every one of those choices is now worth re-examining, because the standard is moving and it’s moving one way.

  • Disclose at the start of every session, not once during onboarding.
  • Answer honestly when asked. A bot that dodges “am I talking to a person” is making a decision on behalf of a company, and it’s not a defensible one.
  • Give a route to a human, clearly, without requiring the customer to guess a magic word.
  • Don’t use a human name and a human photograph unless you’re prepared to explain that choice publicly.
  • Log what was said. If a system speaks on your behalf, you need a record of what it committed you to.

“If your product needs the customer not to know what it is, you don’t have a product problem. You have a disclosure problem you’ve decided to solve by hiding.” — Sameer Gupta


The Broader Governance Point

What strikes me most is that this failure was entirely foreseeable, and the fix was trivial, and it still shipped as a demo without it.

That’s the same shape as Tay two years ago. Not a technical failure. A review process that assessed whether the system worked and didn’t ask what it would feel like to be on the other end of it.

I’d add that question to any launch review for a system that interacts with people who haven’t chosen to interact with it:

Who is affected by this who is not our customer, and would they consent if asked?

The salon receptionist wasn’t Google’s user. Nobody in the design process was representing them.

Final Thoughts

Google corrected course quickly and I’d give them credit for that. The demo will be remembered as a milestone, and rightly.

But the useful lesson isn’t about speech synthesis. It’s that we’re entering a period where a series of things that used to serve as evidence, a voice, a photograph, a document, a face, are going to stop serving as evidence, one at a time, faster than institutions adapt.

Disclosure is the cheap fix available now, while it’s still a choice. It won’t stay a choice.