Voice-to-Wheels: Turning Speech into ROS 2 Motion
A robot you control by voice lives and dies on two numbers: how long from speaking to moving, and how often it does the wrong thing. Our first version was slow and occasionally wrong, and the fix was not a better model, it was a better structure. Here is how we ended up building voice-to-wheels.
Key Insights
- The voice chain is speech-to-text, then intent, then a ROS 2 motion command, and on our robot all of it runs on-device
- Routing every command through a language model was both slower and occasionally wrong, for example mapping “move forward” to a sideways strafe
- A deterministic rule layer for the known command vocabulary runs instantly and reliably, with the model kept only as a fallback for open phrasing
- Safety commands like stop need their own fast path that bypasses transcription entirely
- The right structure matters more than the model choice for a responsive, trustworthy voice interface
The naive version, and why it disappoints
The obvious design is a pipeline: transcribe the speech, hand the text to a small on-device language model to figure out what the person meant, and send the resulting action to the robot. It works, and it disappoints in two specific ways.
It is slow, because the language step adds one to three seconds on top of transcription for even a trivial command. And it is occasionally wrong in ways that feel dumb to a user: our model would take “move forward” and produce a sideways strafe, or take “come here” and produce nothing. When a person says a plain command and the robot does something odd, the whole thing feels broken, no matter how clever the model is elsewhere.
The fix: a rule layer in front of the model
The insight is that the commands that matter most for a demo or a product are a small, known vocabulary: stop, forward, back, turn, follow, play. Those do not need a language model at all. We put a deterministic rule layer in front of the model that resolves the known vocabulary directly, in effectively zero time, and only hands the sentence to the language model when nothing in the rules matches.
That single change fixed both problems. The core commands became instant and correct, because they no longer depended on a model guessing at them, and the language model was reserved for the open-ended phrasing it is actually good at. The user-facing result is a robot that responds to plain commands immediately and predictably, which is the whole point.
Stop is special
There is one more structural rule worth stating on its own: a safety command cannot wait in line. On our robot a spoken “stop” that went through the full transcription pipeline took several seconds, which is far too long when the robot is moving toward a table edge. So stop got its own path that bypasses transcription entirely and publishes the halt directly, alongside a dedicated stop control on the dashboard. The direct path responds in well under a tenth of a second versus seconds for the spoken route. When it matters, you use the fast path.
The general principle: the more important and time-critical a command is, the fewer stages it should have to pass through.
What we keep for a conversation
This post is the shape of the design, not the implementation. The reason is that the interesting decisions for your product, where the rule layer ends and the model begins, how the intent maps to your specific actions, how you keep it responsive on your hardware, are judgement calls we make with you, not a snippet to paste. Structuring a voice interface that is fast and trustworthy on-device is the work we do.
If you are building voice or command interfaces on an edge robot or device, on the MediaTek Genio or on NVIDIA Jetson, tell us what you are building and we will help you get the structure right. For the speech model behind this, see Whisper on the Genio NPU, and for the whole robot, the ROSOrin build overview.
Relevant Services
MediaTek Genio Expert Support
Building on MediaTek Genio?
BSP bring-up, GStreamer pipelines, NeuroPilot integration, we've shipped it. Get unblocked fast. One call to scope it, fixed bid to deliver it.
Frequently Asked Questions
How do you turn a spoken command into robot motion?
The chain is speech-to-text, then an intent step that maps the words to an action, then a control message to the robot. On our ROSOrin robot all of it runs on-device: Whisper transcribes on the NPU, an intent layer decides the action, and a ROS 2 message drives the mecanum wheels. Keeping the whole chain local is what makes it work with no network.
Should you use an LLM to interpret robot voice commands?
Not for the core command set. A small language model is useful for open-ended phrasing, but it is slower and can map a simple command wrongly, for example turning move forward into a sideways motion. A deterministic rule layer that handles the known vocabulary first is faster and more reliable, with the model as a fallback for anything it does not match.
How do you make a robot stop instantly by voice?
Do not route a stop through the full transcription pipeline, because that can take several seconds. Give stop its own fast path that bypasses transcription, and pair it with a physical or on-screen stop control. On our robot a spoken stop took seconds while the direct stop path responded in well under a tenth of a second.
Written by
Andrés CamposCo-Founder & CTO · ProventusNova
8 years deep in embedded systems, from underwater ROVs to edge AI. Andrés leads every technical delivery personally.
Connect on LinkedInRelated Articles
MediaTek Genio for robotics edge AI: inference, camera, BSP reality
Is MediaTek Genio viable for robotics edge AI? Honest assessment of inference latency, camera pipeline, ROS 2 support, and BSP limitations for robotics builds.
MediaTek Genio vs Jetson for a ROS 2 Robot
We built a ROS 2 robot on the MediaTek Genio. Here is how the Genio and NVIDIA Jetson actually compare for robotics, from first-hand experience on both.
Running ROS 2 on the MediaTek Genio
ROS 2 runs well on the MediaTek Genio, and containers on a custom Yocto image make it reproducible. What the stack looks like and the gotchas to expect.
Voice-Controlled Robot on the MediaTek Genio: the ROSOrin Build
We built a voice-controlled, self-driving robot on the MediaTek Genio and demoed it at Embedded World NA. The stack, the on-NPU AI, and what it took.