Two robots with identical capability and identical information can leave completely different impressions. The difference is voice and phrasing, and both are design choices that are usually left at their defaults.
Voice choice matters more than expected
It is the first thing noticed. Before content, before capability, a visitor forms an impression from how the robot sounds.
Default voices are usually poor fits. They are chosen for broad acceptability, which means they are specific to nothing.
Options, in ascending cost. Selecting a better stock voice. Recording a person for fixed lines and using synthesis for the rest. Generating a custom voice from a recorded sample.
What to listen for when choosing. How it reads numbers and proper nouns — the common failure points. How it sounds over the robot's actual speakers rather than through headphones. And how it sounds after five minutes rather than after one sentence.
If using a real person's voice. Secure a clear usage agreement, and consider that the person may not remain available or associated with the organisation.
Consistency beats quality. One adequate voice used everywhere is better than a good voice mixed with a different one for some sections — a common outcome when several people build content.
Register and phrasing
Decisions that should be made deliberately and written down.
Formality. Match the setting. A clinic, a café and a bank call for different levels, and the difference is noticed immediately by visitors.
Length. Shorter than written content. A listening person has considerably less patience than a reading one — three sentences with detail on request beats a complete answer delivered at once.
Pace. Default settings are often slightly fast. Where older visitors are common, slower is better.
How it handles not knowing. A brief acknowledgement and a redirection is far better than an evasive non-answer. This phrasing should be written deliberately rather than left to a system default.
Confirmation. Briefly restating what was asked reassures the visitor that it understood, and it surfaces misrecognitions early.
Interruptibility. Whether a visitor can speak over the robot. Not being able to is one of the most consistently reported frustrations.
The opening and closing
The only two lines every visitor certainly hears.
The opening does three jobs. Signals that the robot is addressing this person, states briefly what it can help with, and invites a response. Failing any of the three loses the interaction.
Timing matters. Greeting someone still approaching leaves them unsure who is speaking; greeting too late means they have walked past.
Naming the capability is essential. Not knowing what to ask is the most common reason visitors standing beside a robot do not engage with it.
The closing should be definite. Ask whether anything else is needed, then close. A conversation that trails off leaves the visitor unsure whether to stay.
Both should be short. The temptation is to include everything in the greeting. Doing so loses people before the useful part.
Test both with real visitors. Ten minutes of watching people encounter the greeting identifies problems that no amount of drafting will.
Should a robot sound human
A question worth deciding deliberately rather than by default.
The general direction is no. A robot attempting to pass as a person creates unease, and the effect when the pretence is noticed is worse than the benefit of the pretence.
People are not bothered by the robot being a robot. They are bothered by being misled about it.
Acknowledging it helps. It sets accurate expectations, and users extend more tolerance to a machine that is honest about being one.
Personality is different from pretence. A robot with a name, a consistent manner and occasional light self-awareness about being a machine is generally well received.
Where the question becomes serious. Voice calls and messaging. An AI system contacting customers should identify itself as automated — both as a matter of honesty and because the alternative destroys trust when discovered.
The summary. Aim for a robot that is pleasant to deal with, not for a convincing imitation of a person.
Practical setup
A checklist for getting voice and manner right.
- Choose the voice by listening through the robot's own speakers, in the intended space.
- Write the register down so everyone producing content follows it.
- Draft the greeting and closing carefully, then test them with real visitors.
- Read all content aloud during drafting and cut anything that does not read naturally.
- Check numbers, names and foreign words by listening to the robot read them.
- Write the not-knowing response deliberately.
- Verify interruptibility and enable it if the product allows.
- Have several people unfamiliar with the project try it and describe how it felt.
The whole exercise takes a day and it affects visitor perception more than most technical decisions in the deployment.
Frequently asked questions
Why does voice choice matter so much?
Because it is the first thing a visitor notices, before content or capability. Default voices are chosen for broad acceptability, which means they suit nothing in particular.
What is the most consistently reported frustration?
Not being able to interrupt the robot — having to wait through a complete wrong answer before asking again. Where a product allows interruptibility it should be enabled.
Why do the greeting and closing matter most?
They are the only two lines every visitor certainly hears. The greeting must signal it is addressing this person, say what it can help with, and invite a response — failing any of the three loses the interaction.
Should a robot pretend to be human?
No. People are not bothered by a robot being a robot; they are bothered by being misled. Acknowledging it sets accurate expectations and earns more tolerance for imperfect answers.
More in Local ecosystem and Robots and jobs.