Docs Book a Demo Sign in
On-device

Voice AI on a $5 chip.

Open-source speech recognition and speech synthesis for the ESP32-S3, a microcontroller with a 240 MHz dual-core CPU and no neural accelerator. No cloud, no command list, and no inference bill that grows with every sentence.

The two halves

Listening and speaking, both on the chip.

  • Oído Speech recognition for any English or Spanish sentence. English word error rate on LibriSpeech is 3.7% on test-clean, against 6.3% for Whisper tiny and 5.0% for Moonshine tiny running on a laptop, and 8.5% for Espressif's own command recognizer on the same chip. It runs NVIDIA's openly licensed 13-million-parameter Conformer-CTC Small on an inference engine we wrote for the chip's vector instructions. GitHub · Hugging Face
  • Ito Streaming neural text-to-speech: 4.4 million parameters, 4.9 MB in int8, two voices at 24 kHz. It produces a 25 ms first chunk and then 100 ms chunks, so the wait before the first sound does not grow with the length of the sentence. On 54 held-out prompts it scores 4.44 on UTMOS22 with a 0.4% word error rate. GitHub · Hugging Face · Listen
Where it stands

Verified in emulation. Board timings are coming.

Both run in Espressif's QEMU emulator: Ito is bit-exact with its host build, and Oído gives identical transcripts on most utterances. Speed on the chip is estimated from exact instruction counts, not measured: Oído at roughly 0.7 to 0.95 times real time, and Ito in a central case slightly slower than real time. We say so in both posts and we will publish the numbers from physical boards, whatever they turn out to be.

Status as of 5 October 2026. The posts carry the detail and any later updates. Ito is English only with two voices; Oído handles English and Spanish.

Why on-device

The voice stays in the product.

  • Unit costNo per-unit cloud inference line item after the sale.
  • OfflineVoice keeps working when connectivity doesn't: warehouses, vehicles, outdoors.
  • PrivacyAudio never has to leave the device.

More in Voice is moving off the cloud.

Licensing

Open source, with a commercial option.

The code is GPLv3. Model licenses differ: Oído's int8 and Spanish models are CC BY 4.0 and its int4 model is CC BY-SA 4.0; Ito's weights are CC BY-NC-SA 4.0 plus additional terms, with free licenses for small companies, makers, education and research. For products that can't meet those terms we offer commercial licenses, models for other languages and integration support. Write to contact@lokutor.com.

Building a voice agent in the cloud instead?

60 minutes a month on the free plan, no card.