VietRobots.comVietRobots
Voice and language

Speech in a real room is a different problem

Recognition accuracy quoted by vendors is measured in quiet conditions. What happens in an actual lobby is another matter.

Person speaking to a robot in a busy indoor space

Speech recognition accuracy figures are measured with a single speaker in a quiet room at close range. A hotel lobby at check-in time is a different acoustic problem, and the gap between the two determines whether a voice interface works in practice.

What actually degrades recognition

In rough order of impact in real deployments.

Background noise. The dominant factor, and it affects every system similarly. Above a certain level, accuracy falls sharply regardless of which product is used.

Distance. Recognition quality drops with distance from the microphone. A person standing two metres away is a substantially harder problem than one standing at half a metre.

Reverberation. Hard floors, glass walls and high ceilings — exactly the materials in modern lobbies — reflect sound and blur the signal.

Multiple speakers. Groups arrive together and talk over one another. Separating who is addressing the robot is difficult.

Quiet speakers. Many people are reluctant to speak loudly to a machine in public. The robot fails to hear, they become more self-conscious, and they give up.

Accents and non-native speech. Systems perform best on the accents most represented in their training data.

Notably, the first three are properties of the room rather than of the robot — which is why a site survey matters more than a specification comparison.

What can be fixed in the environment

Often cheaper and more effective than changing equipment.

Move away from noise sources. Entrance doors, air handling, food service, music speakers. Several metres can make a substantial difference.

Add soft surfaces nearby. A rug, an upholstered bench, an acoustic panel. Reduces reverberation cheaply.

Create a natural close-range position. A counter, a marked spot, a design that brings people to within comfortable speaking distance.

Reduce competing audio. Background music at a robot's position works against it.

Use a directional microphone. Focused on the expected speaker position rather than picking up the whole room.

Consider a semi-enclosed position. A recess or partial screen reduces both noise and self-consciousness, and the second effect is larger than people expect.

These measures cost very little and they typically improve recognition more than moving to a higher-specification product would.

Design so speech is not the only path

The most important principle for public deployments.

Every voice action must have a touch equivalent. This is not a fallback for failure — it is a primary path for the substantial proportion of users who will not speak aloud in public.

Show what was heard. Displaying the recognised text lets the user see a misrecognition and correct one word rather than repeating the whole utterance.

Display answers as well as speaking them. In noisy places people read rather than listen.

Offer suggested questions on screen. Three to five common ones, selectable. This solves both the noise problem and the more common problem of users not knowing what to ask.

Keep required utterances short. Short phrases are recognised more accurately. Design the interaction as short steps rather than one open question.

Make retry easy. A clear repeat option that does not require sitting through a wrong answer first.

Never require speech for numbers. Phone numbers, reference codes, amounts. One wrong digit invalidates the whole string.

Testing before you commit

A short protocol that answers the question properly.

Test in the actual position at the actual busy hour. Not in a showroom and not when the space is empty.

Use several speakers. Different accents, different ages, different volumes. Include at least one quiet speaker and one older person.

Test from realistic distances. Where people will actually stand, not where a demonstrator stands.

Count failures. Out of thirty utterances, how many were recognised correctly. Write it down as it happens.

Test the recovery path. When it mishears, how easy is it to correct.

Test the touch path fully. Can everything be done without speaking at all.

Test with background audio running. Music, announcements, whatever is normally present.

Half a day of this produces a clearer answer than any amount of specification comparison, and it is a test no supplier can prepare for in advance.

Setting expectations correctly

Both internally and with users.

Internally. Expect a meaningful proportion of interactions to happen entirely by touch. That is not a failure of the deployment; it is how public spaces work.

With users. A brief prompt indicating both options — speak or touch — removes the assumption that speaking is required.

When it mishears. The robot should respond in a way that does not imply the user was at fault. Wording matters here more than people expect, particularly with older users who readily blame themselves.

Track the split. What proportion of interactions use voice versus touch. If voice use is very low, that is diagnostic — usually of position or noise rather than of the product.

And accept the ceiling. Above a certain noise level, no product performs well. Recognising that and designing around it produces a better deployment than continuing to look for a system that solves it.

Frequently asked questions

What degrades speech recognition most in real deployments?

Background noise, followed by distance and reverberation. All three are properties of the room rather than of the robot, which is why a site survey matters more than comparing product specifications.

What environmental fixes help most?

Moving away from noise sources, adding soft surfaces to reduce reverberation, creating a natural close-range speaking position, and using a directional microphone. These usually beat upgrading to a higher-specification product.

Why must every voice action have a touch equivalent?

Because a substantial proportion of users will not speak aloud in a public space regardless of how well recognition works. Touch is a primary path rather than a fallback for failure.

How should speech quality be tested before buying?

In the actual position at the actual busy hour, with several speakers of different accents and volumes, from realistic distances, counting failures out of thirty utterances. No supplier can prepare for this in advance.

More in Myths and questions and AI robots for real work.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
Call now 0926 138 138