We train our own speech models.
Versa, our speech foundation model, gives every agent its voice in nine languages. Turno decides whose turn it is, Psst hears the voice and not the room, and Conv 2.0 is our next recogniser. All of them are small enough to run in real time on an ordinary CPU.
Five models, one stack.
-
Versa 2.0
Speech synthesis Speaks on every callOur speech foundation model. Non-autoregressive flow matching: it generates each chunk of audio in parallel instead of token by token. Nine languages, ten voices, and any new voice from a few seconds of reference audio.
First audio in a median of 64 ms on our GPU servers. Before our cloud moved it there to serve more calls per server, it served live calls from CPU-only servers with first audio at a median of 139 ms.
-
Turno
Turn-taking and barge-in Decides every turnDecides on the audio itself whether the caller has finished, is interrupting, or is only saying “mhm”, instead of waiting on a silence timer.
Average precision 0.931 against 0.876 for Silero VAD on a neutral real-world test set (TEN-VAD leads on voice activity alone at 0.985), flat from clean audio down to 0 dB, at about half Silero’s cost per inference.
-
Psst
Noise suppression Cleans every callRemoves background noise and other voices before recognition, tuned for what matters to a voice agent: the words, not how pretty the audio sounds.
Word error after cleaning 36% lower than DeepFilterNet3, the strongest published real-time denoiser, and the only system we tested that improves on no denoising at all. One CPU core cleans 18 calls at once.
-
Conv 2.0
Streaming speech recognition In developmentOur own streaming recogniser, a quarter smaller than the Parakeet model it replaces. Production recognition today is Conv 1.0, which runs NVIDIA’s open Parakeet on our servers; Conv 2.0 is our next release.
5.59% word error on a conversational benchmark scored like Artificial Analysis’s streaming test, and a final transcript 14 ms after speech ends (median, on one NVIDIA L4 GPU).
-
Ito
Speech synthesis for microcontrollers Research, open weightsA streaming neural voice built for the ESP32-S3, a $5 chip with no neural accelerator. English, two voices, no cloud.
Scores 4.44 on UTMOS22 with 0.4% word error on 54 held-out prompts. Verified bit for bit in Espressif’s emulator; no physical board has run it yet.
Small models go further.
- PriceLess compute per call is how a voice agent costs about 5¢ a minute, language model included.
- AnywhereThe same models run in our cloud, on your servers or inside a device. Running on an ordinary CPU is the proof, not the point.
- SpeedA small model answers fast: Versa’s first audio in about 65 ms, and a voice agent’s reply about 0.6 s after the caller stops talking.
The parts we don’t build.
- Language modelFrom low-latency providers, plugged into our pipeline.
- Conv 1.0Production speech recognition runs NVIDIA’s open Parakeet TDT 0.6B on our own servers, with no audio sent to a third-party cloud, until Conv 2.0 replaces it.
- OídoOur inference engine for the ESP32-S3 runs NVIDIA’s open Conformer-CTC Small. The engine is ours; the network is NVIDIA’s, and our Spanish version is a fine-tune of it.
Common questions.
Does Lokutor build its own AI models?
Yes. We design and train the speech models that do the listening and the speaking: Versa for the voice, Turno for turn-taking, Psst for noise suppression, Conv 2.0 for speech recognition (in development) and Ito for microcontrollers. The language model comes from a provider, and production speech recognition currently runs NVIDIA’s open Parakeet model on our own servers.
Why are Lokutor’s models so small?
Because a model that is cheap to run can be run anywhere and sold for less. Versa has 91.4 million deployable parameters and Turno 18,651. They run in real time on an ordinary CPU, so the same models serve our cloud, a customer’s own servers or a device.
Do Lokutor’s models need a GPU?
No. Our hosted API serves Versa on GPUs because one GPU server carries many more simultaneous calls (first audio there takes a median of 64 ms), but the models do not need one: until September 2026 Versa served live calls from CPU-only servers, with first audio at a median of 139 ms.
Where are the models trained?
On cloud GPUs and on the Barcelona Supercomputing Center’s cluster, through its AI Factory program for European AI startups.
Use them in your product.
Through our API, on your own servers, or inside your device.