Ito: Streaming Neural Speech on a $5 Chip
A 4.9 MB text-to-speech model built for the ESP32-S3, a microcontroller with no neural accelerator. It streams, runs without a cloud, and is verified in emulation. Board timings are coming.
Text-to-speech on a microcontroller has mostly meant one of two things: a robotic voice from the 1990s, or a neural model that takes tens of seconds to say a sentence. Neither is something you’d put in a product people talk to.
Today we’re releasing Ito: a neural text-to-speech model built for the ESP32-S3, a $5-class microcontroller with a 240 MHz dual-core CPU and no neural accelerator. It has 4.4 million parameters, takes 4.9 MB in int8, speaks with two voices (female D and male G) at 24 kHz, and needs no cloud.
You can hear it next to other small systems on the demo page.
What we claim, and what we don’t
We want to be precise, because the field is crowded and several people got here first.
- Ito is not the first neural TTS on a microcontroller, and it is not the smallest. sanoTTS, Moonshine Micro and others came earlier, and sanoTTS’s voices are much smaller.
- What we believe is true: Ito is the most natural complete neural TTS built for a microcontroller without an NPU, and the first streaming neural TTS for the ESP32-S3. Both are qualified by one fact below.
- That fact: it is verified bit-exact in Espressif’s QEMU emulator, and no physical board has run it yet. Time to first audio and real-time factor are estimates from exact operation counts, not measurements.
How it compares
We scored Ito and 22 public small and on-device systems on the same 54 English prompts, none of which are in Ito’s training text. Audio passed through identical post-processing, then through UTMOS22, UTMOSv2, DNSMOS and Whisper large-v3 for word errors.
Among the systems built for microcontrollers:
| System | Runs on | UTMOS22 | UTMOSv2 | WER | Speed on its chip |
|---|---|---|---|---|---|
| Ito (int8, chip-exact) | ESP32-S3 (emulated) | 4.44 | 3.15 | 0.4% | estimated, not measured |
| Inflect Nano v2 | ESP32-P4 (third-party port) | 4.41 | 3.08 | 1.1% | about 3.5x slower than real time |
| TinyTTS | ESP32-S3 | 3.66 | 2.45 | 6.8% | 22.9x slower than real time |
| sanoTTS heart-nano | MCU-sized, no published timing for this voice | 2.17 | 1.33 | 1.7% | not published |
Ito is ahead on UTMOS22 and has the lowest word error rate. It ties Inflect Nano v2 on UTMOSv2 (the difference is inside the confidence interval) and on DNSMOS. The Inflect speeds are third-party measurements from that project’s README, and Ito’s are not, so the speed column is not like for like.
Larger models score higher on UTMOSv2: Kokoro, Piper, Supertonic and others. They are 9 million parameters and up, and they need a CPU or GPU. We’re not claiming to beat them. The claim is about what fits on this chip.
The listening test is small. One listener, blind to system, four sentences each: Ito 4.0 out of 5, sanoTTS amy 2.0, sanoTTS heart-nano 1.0, and our full-size reference model 4.75. That is evidence, not a MOS study, and the listener is the founder. Amy is also not the voice sanoTTS runs on a chip, so the like-for-like gap is a different question. We’ll run a larger test with outside listeners.
Streaming
Ito produces a 25 ms first chunk and then 100 ms chunks, so the wait before the first sound does not grow with the length of the sentence. Models that synthesize a whole utterance first take longer on a long sentence; Ito’s first chunk takes the same work whether the sentence is five words or fifty.
Our estimate, from exact instruction counts in QEMU plus a model of PSRAM traffic, is that the first chunk needs about 23 million instructions and 4.5 MB of weights read from PSRAM. That puts time to first audio at 124 to 127 ms in the optimistic case, 176 to 179 ms in the central case and 266 to 269 ms in the pessimistic case. Real-time factor comes out at 0.72 to 0.78, 1.15 to 1.23 and 1.9 to 2.0, so real-time playback is not established: in the central case the board is slightly slower than real time. We first published 130 to 210 ms and real time; that was written from operation counts at an assumed throughput and was too optimistic. These are models, not benchmarks. The real risk is the weights: 4.9 MB can’t sit in the chip’s 512 KB of fast memory, so they stream from PSRAM, which is much slower. If that costs more than we’ve assumed, the numbers will be worse, and we’ll publish them anyway.
Where it stands
- Verified: bit-exact in QEMU, matching the host build sample for sample. Quantization to int8 costs nothing measurable on any metric we ran.
- Not verified: speed on silicon.
- Limits: English only, two voices, and automatic scores such as UTMOS cannot see prosody.
- Paper: not published. It will follow the board measurements.
Why we built it
Voice is moving off the cloud. A device that streams to a server for every sentence pays for it forever, goes quiet when the connection drops, and sends its words somewhere else. Ito is the speaking half of the on-device voice stack we’re building at Lokutor. The listening half, Oído, runs on the same chip.
Try it
Code: github.com/lokutor-ai/ito. Weights: huggingface.co/lokutor-ai/ito-tts-v3, a gated download where you accept the terms. The README has the build and flashing steps.
Licensing
The code is GPLv3, with commercial licenses for products that can’t meet its terms. The weights are CC BY-NC-SA 4.0 plus additional terms. We offer free licenses for small companies, makers, education and research. For anything else, write to contact@lokutor.com. The training code is not released.
If you flash it onto a board, please send us the timings. They decide what we say next.
Run voice AI where your users are.
Free plan, no card. Or talk to us about running it inside your own environment.