Speech to text for a fraction of a cent.
Live or one-shot transcription in Spanish, English, Portuguese, French, Italian, German, for $0.084 an hour, billed per second. The audio is transcribed on our own servers in the EU.
Spoken words, written down.
Speech to text, or automatic speech recognition, turns audio of people talking into text. A voice agent needs it to understand the caller; on its own it transcribes calls, meetings, voice notes and dictation.
Ours runs NVIDIA’s open Parakeet TDT 0.6B model on our own GPUs, about 45 ms to transcribe a four-second sentence, with no audio sent to a third-party cloud. Our own recogniser, Conv 2.0, is in development and will replace it.
- LiveWebSocket streaming with automatic utterance detection, for transcripts while people talk.
- One-shotSend a recording over REST, get the text and timed segments back. Any sample rate.
- LanguagesSpanish, English, Portuguese, French, Italian, German. Not yet Catalan, Galician, Basque.
- Price$0.0014 a minute, $0.084 an hour, billed per second from a prepaid balance.
One request.
curl -X POST "https://api.lokutor.com/stt/transcribe" \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "audio=@recording.wav" \
-F "language=es" For live transcription, use WS /ws/stt or the SDKs.
Common questions.
How much does Lokutor’s speech to text cost?
$0.0014 a minute ($0.084 an hour), billed per second from a prepaid balance with no plan. New accounts get $5 of credit with no card.
Which languages does it transcribe?
Spanish, English, Portuguese, French, Italian, German. Catalan, Galician and Basque are not recognised yet, although our voice speaks them.
Which model does Lokutor use for speech recognition?
Production runs NVIDIA’s open Parakeet TDT 0.6B on our own GPUs in the EU. Our own streaming recogniser, Conv 2.0, is in development: in our benchmark it returns a final transcript in about 8 ms, and its accuracy on conversational speech is still catching up.
Is my audio sent to another company?
No. Transcription runs on our own servers; the audio is not sent to a third-party speech service.
Transcribe your first file.
There is a speech-to-text studio in the dashboard. $5 of credit to start.