Voice-Controlled Robot on the MediaTek Genio: the ROSOrin Build
At Embedded World North America this year we walked the floor with a robot small enough to hold in one hand. You could tell it to follow you, drive it from a tablet, or set a ball down and watch it play. Everything it did ran on a single MediaTek Genio 520, offline, with no cloud behind it.
We call it ROSOrin. This post is the overview of what it is and what it took to build. The deep dives on each part link out from here.
Watch the demo from the show floor:
Key Insights
- ROSOrin runs voice control, person and ball detection, navigation, and driving on one MediaTek Genio 520, fully offline
- The two AI workloads that matter, speech-to-text and object detection, both run on the Genio NPU, not the CPU
- Moving object detection from the CPU to the NPU took it from unusable to real-time, and dropped the board temperature enough to run fanless
- The software is ROS 2 in a container stack on a custom Yocto image, which is what makes the whole thing reproducible instead of a one-off demo
- The hard parts were rarely the parts you would expect, and that is the theme of every deep dive below
What the robot does

You speak to it. “Follow me”, “move forward”, “turn left”, “stop”, “let’s play”. A microphone on the robot turns your voice into text, a small on-device language step turns that text into an action, and the wheels respond. No phone app, no server, no internet.
It sees. A depth camera feeds an object detector that finds people and a ball. In floor mode it can follow a person or chase the ball. In table mode it stays put and just turns to keep you in frame, so it can sit on a stand at a booth and never drive off the edge.
You can also drive it yourself. A tablet loads an installable dashboard, a first-person camera view with a joystick, plus a live radar scope from the lidar so you can see what the robot sees around it.
And it holds still until told otherwise. A trade-show robot that wanders on its own is a liability, so the default is passive: it does nothing until a voice command or the dashboard engages it.
The stack, in one picture

- Compute: a MediaTek Genio 520 on a MiTwell PicoITX board, our partner hardware for this build.
- AI on the NPU: speech-to-text and object detection both run on the Genio’s neural accelerator. This is the difference between a demo that works and one that overheats. See running YOLOv8 object detection on the Genio NPU and Whisper speech-to-text on the Genio NPU.
- Software: ROS 2 Humble, running as a set of Docker containers on a custom Yocto image. See running ROS 2 on the MediaTek Genio.
- Sensing and motion: an LD06 lidar for 360-degree awareness, a depth camera for vision, and a mecanum drive that can move in any direction.
What it actually took
The interesting part of a build like this is never the block diagram. It is the week you lose to a camera that browns out the whole board, the driver bug that makes your booth WiFi look dead when it is fine, and the AI model that runs at three frames per second until you move it onto the right silicon.
We wrote those stories up, because they are the parts nobody publishes and the parts that actually decide whether an edge AI product ships. Each one names the problem and the point in the build where it bites. The exact fix is the work we do for clients, so we keep that part for a conversation.
Embedded World North America

We showed ROSOrin at the MiTwell booth alongside MediaTek, the two partners behind the platform. It drew a crowd, which was the point: a hand-sized robot doing real on-device speech and vision is a good way to make “edge AI on the Genio” concrete for a founder who is deciding what to build their product on.
If you are building a computer-vision or robotics product on the MediaTek Genio or on NVIDIA Jetson and you want it to run real AI on-device, that is exactly the work we do. Tell us what you are building and we will tell you the fastest path to a working result.
Relevant Services
MediaTek Genio Expert Support
Building on MediaTek Genio?
BSP bring-up, GStreamer pipelines, NeuroPilot integration, we've shipped it. Get unblocked fast. One call to scope it, fixed bid to deliver it.
Frequently Asked Questions
What is ROSOrin?
ROSOrin is a voice-controlled, self-driving mecanum-wheel robot ProventusNova built on the MediaTek Genio 520 and demoed at Embedded World North America 2026. It runs ROS 2, understands spoken commands, follows people, and plays with a ball, all on-device with no cloud. Speech recognition and object detection both run on the Genio NPU.
What hardware is ROSOrin built on?
The compute is a MediaTek Genio 520 on a MiTwell PicoITX board. The chassis is a mecanum-wheel platform based on a HiWonder kit, with an LD06 lidar for obstacle avoidance and a depth camera for vision. Everything runs from a single battery pack with no external compute.
Does it need a network or cloud to work?
No. Speech-to-text, the language step that maps a command to an action, object detection, navigation, and control all run on the Genio itself. The robot works fully offline, which is the point of doing the AI on the NPU rather than in the cloud.
Written by
Andrés CamposCo-Founder & CTO · ProventusNova
8 years deep in embedded systems, from underwater ROVs to edge AI. Andrés leads every technical delivery personally.
Connect on LinkedInRelated Articles
MediaTek Genio for robotics edge AI: inference, camera, BSP reality
Is MediaTek Genio viable for robotics edge AI? Honest assessment of inference latency, camera pipeline, ROS 2 support, and BSP limitations for robotics builds.
MediaTek Genio vs Jetson for a ROS 2 Robot
We built a ROS 2 robot on the MediaTek Genio. Here is how the Genio and NVIDIA Jetson actually compare for robotics, from first-hand experience on both.
Running ROS 2 on the MediaTek Genio
ROS 2 runs well on the MediaTek Genio, and containers on a custom Yocto image make it reproducible. What the stack looks like and the gotchas to expect.
Voice-to-Wheels: Turning Speech into ROS 2 Motion
Routing every spoken command through an LLM made our robot slow and wrong. A rule layer in front of the model fixed both. How we structured voice control.