Access Ito, a non-commercial voice model
Ito's weights are free for hobby, research, education and personal use. Tell us a little about your project: it helps us decide what to build next, and we answer every request for a commercial or collaboration license.
The Ito voice model is licensed under CC BY-NC-SA 4.0 together with Lokutor's terms of use (TERMS.md in this repository). Without a written license from Lokutor you may not use the weights, weights derived from them, or the audio they produce commercially, and you may not use Ito's output to train or improve a text-to-speech or voice model that is offered or used commercially. Ito's output is synthetic speech in the voice of a real (LibriTTS-R) speaker: say that it is synthetic when you share it, and never use it to impersonate or deceive. Small companies, startups, makers, schools and research groups can ask for a no-cost license at contact@lokutor.com.
Log in or Sign Up to review the conditions and access this model content.
Ito: natural-sounding streaming TTS for the ESP32-S3
Ito is English text-to-speech built to run entirely on an ESP32-S3 microcontroller (240 MHz dual-core, 8 MB PSRAM, 16 MB flash), with no cloud and no neural accelerator. It streams: audio starts after a short first chunk (25 ms), and the work before it does not grow with the sentence length. Its 4.4 M parameters fit in 4.9 MB of int8 weights.
Watch the 50 s video: lokutor-ai.github.io/ito/demo.mp4 (with sound) Β· Listen next to sanoTTS: lokutor-ai.github.io/ito Β· Code, firmware and tools: github.com/lokutor-ai/ito (GPLv3, commercial licenses available). From Lokutor, the makers of OΓdo, speech recognition on the same chip.
Status (3 October 2026). The on-chip engine is verified on a laptop and in Espressif's QEMU emulator: the firmware's audio is bit-identical to the host build of the engine, and every sample below is that engine's exact output. Nothing has run on a physical board yet. Time to first audio and real-time factor are estimated from exact QEMU instruction counts and an assumed PSRAM bandwidth. Real-time playback is not established on silicon: the optimistic and central estimates are faster than real time (RTF 0.53β0.54 and 0.77β0.80), the pessimistic one is still slower (1.18β1.27). An earlier version of this card (and our first GitHub README) said 130β210 ms and real time at 0.5β1 GOPS; that was too optimistic and is corrected below. Board measurements are coming.
Voices
Two voices, each distilled from StyleTTS 2 conditioned on recordings of one LibriTTS-R reader (LibriTTS-R, CC BY 4.0). Each voice is a separate model with the same architecture, size and speed.
| Voice | Speaker | PyTorch | Chip (flash at 0x200000) |
|---|---|---|---|
| female (default) | LibriTTS-R speaker 4970 | ito_female.pt |
ito_female_esp32s3.bin |
| male | LibriTTS-R speaker 5105 | ito_male.pt |
ito_male_esp32s3.bin |
The male voice is new (2 October 2026). Its chip file passes the same host and QEMU checks as the female voice (quantised vs float PESQ 4.53; firmware PCM bit-identical to the host engine). It has not been in a blind listening test yet; the results below are for the female voice.
Listen
The female voice, the chip engine's exact output, for eight sentences Ito never saw in training. The male voice reads the same sentences on the demo page.
| Sentence | Ito (chip-exact) | |
|---|---|---|
| 01 | Hey, are you still coming over for dinner tonight, or should I save you a plate? | |
| 02 | Your package should arrive on Friday, October 9th, sometime before noon. | |
| 03 | When I finally got to the station, the last train had already left, so I ended up sharing a taxi with two strangers who turned out to be surprisingly good company. | |
| 04 | The pharmacist recommended an anti-inflammatory, but honestly, I'd rather try physiotherapy first. | |
| 05 | Thanks so much for calling. I'll check the schedule and get back to you first thing tomorrow morning. | |
| 06 | Could you grab some quinoa and Worcestershire sauce on your way home? | |
| 07 | It's about 23 degrees outside, so you probably won't need a jacket. | |
| 08 | I know it sounds strange, but I actually enjoy the quiet hours before everyone else wakes up. |
The players stream from the public demo page, so they work before you accept the gate. The bit-exact WAVs are in
samples/ in this repository. The demo page plays the same sentences from
sanoTTS and from Ito's teacher, with a blind mode.
Results
Blind listening test #9. One expert listener rated naturalness from 1 to 5, with system names hidden. Each system read four new conversational sentences.
| System | Parameters | Runs on | Mean (4 clips) |
|---|---|---|---|
| Teacher: StyleTTS 2 (LibriTTS model) | large | GPU / laptop | 4.75 |
| Ito | 4.4 M | ESP32-S3 (bit-exact in emulation) | 4.00 |
| sanoTTS amy | 1.46 M | browser / desktop (not run on an MCU) | 2.00 |
| sanoTTS heart-nano | 0.29 M | ESP32-S3 | 1.00 |
Automatic metrics on the eight sentences above:
| Teacher | Ito | sanoTTS amy | sanoTTS heart-nano | |
|---|---|---|---|---|
| UTMOS | 4.49 | 4.46 | 3.98 | 2.07 |
| WER, Whisper medium.en / base | 0 / 0 % | 0 / 0 % | 0 / 1.0 % | 1.0 / 1.0 % |
Please read these with their limits:
- One listener (Lokutor's founder, so not a neutral party) and four sentences per system. This is a strong direction, not a MOS study. With n = 4, differences under about half a point are noise.
- The rated Ito clips are the float model through the same streaming path as the chip. The quantised chip engine scores PESQ 4.52 against it, but the exact chip configuration has not been blind-rated yet.
- UTMOS cannot hear intonation. Treat it as a check, not a verdict.
What we think this supports, and no more: the highest automatically measured naturalness of any complete text-to-waveform neural TTS built for a microcontroller without an NPU (UTMOS 4.46 on 8 sentences; the best previously published on-chip model scores 2.80), and the first streaming neural TTS designed for the ESP32-S3. Both are verified in emulation, not yet on a board. Ito is not the first or the smallest TTS on a microcontroller.
A broader benchmark against other small and embedded TTS systems is being finalised; it will be in
bench/ on GitHub.
Size and compute
| Parameters | 4.40 M: acoustic front 1.62 M + vocoder 2.79 M |
| Chip weights | 4.89 MB (ito_female_esp32s3.bin): int8 mel head and vocoder, int16 pitch path |
| Memory | peak 6.5 of 8 MB PSRAM, 340 of 384 KB internal SRAM (QEMU; about 353 KB on the chip with the I2S buffers) |
| Work before the first audio (125 ms chunk) | 25.3β26.3 M instructions and 5.1 MB of weights read from PSRAM, for any sentence length (exact counts from QEMU) |
| Time to first audio | estimated, not measured: 137β143 ms optimistic, 200β207 ms central, 311β318 ms pessimistic. The first chunk is 125 ms of audio and the chunks behind it are sized so that playback can start with it without a gap |
| Real-time factor | estimated, not measured: 0.53β0.54 optimistic, 0.77β0.80 central, 1.18β1.27 pessimistic (below 1 is faster than real time). The pessimistic case (40 MB/s PSRAM, nothing overlapped) is still slower than real time |
| Start delay for gapless speech | estimated: 137β143 ms optimistic (the same as the time to first audio), 215β240 ms central, 880 ms or more pessimistic (the real-time factor is above 1 there, so the delay grows with the sentence) |
| On a laptop | RTF β 0.01 on an Apple M4 Max CPU |
All timings are estimated from exact instruction counts, not measured on silicon. The open question is the effective PSRAM bandwidth and how much of it overlaps compute; the pessimistic column depends on both. A weight pass now serves 24-frame (300 ms) chunks, which halves the weight traffic (38 to 17 MB per second of audio) with bit-identical output. The firmware benchmarks itself at boot
and prints BOARD_SUMMARY. If you flash a board, please
open an issue with that line.
How it works
Text becomes phonemes on the host (espeak-ng), and token ids go to the chip over serial. On the chip, a streaming acoustic front (1.62 M; a forward-only GRU predicts durations, pitch, energy and a mel spectrogram) feeds a Vocos-style vocoder (2.79 M; ConvNeXt blocks, a harmonic pitch source and an inverse STFT) that outputs 24 kHz audio in 100 ms chunks. Ito learned its voice from StyleTTS 2 speaking as a LibriTTS-R reader. The engine is new portable C99 code with the chip's integer arithmetic; the same source builds on a laptop. The training code and recipe are not public.
Use
git clone https://github.com/lokutor-ai/ito && cd ito && pip install -e .
hf download lokutor-ai/ito --include "ito_*" --local-dir models # both voices (accept the terms first)
ito-tts "Good morning! The coffee is ready." -o hello.wav # PyTorch, CPU is fine; female voice
ito-tts --voice male "Good morning! The coffee is ready." -o hello_male.wav # male voice
cd esp32/host && make && cd ../.. # the chip engine, built for your laptop
python esp32/tools/chip_wav.py "Good morning! The coffee is ready." hello_chip.wav # chip-exact audio
esp32/tools/flash.sh /dev/ttyUSB0 # ESP32-S3-DevKitC-1 N16R8 + I2S DAC/amp (female voice)
VOICE=male esp32/tools/flash.sh /dev/ttyUSB0 # ... or the male voice
python esp32/tools/say.py "Hello from a five dollar chip." --port /dev/ttyUSB0
from ito import Ito
tts = Ito.load("models/ito_female.pt") # female voice; Ito.load(voice="male") or "models/ito_male.pt" for the male voice
wav = tts.synthesize("Could you grab some quinoa on your way home?") # float32 numpy, 24 kHz
for chunk in tts.stream("A sentence of any length."): # 100 ms chunks, as on the chip
...
Files
ito_female_esp32s3.bin: the female voice's chip weights (4.89 MB), for the firmware and the host build of the engine.ito_female.pt: the female voice's PyTorch inference checkpoint (18 MB): front, vocoder, the fixed style the chip uses, and an optional text-to-style predictor.ito_male_esp32s3.bin,ito_male.pt: the same for the male voice. The male checkpoint has no text-to-style predictor; it uses the fixed style, as the chip does.samples/01.wavβ¦samples/08.wav: the chip engine's exact output for the sentences above (24 kHz, 16-bit).LICENSE-WEIGHTS,TERMS.md,NOTICE: the license, the terms of use, and third-party attributions.
Limitations
- English only, two voices (one per weights file), one fixed speaking style on the chip.
- Speed on the chip is estimated until board measurements are published.
- The pitch range is slightly narrower than the teacher's (female 0.94, male 0.97).
- Needs an ESP32-S3 with 8 MB PSRAM (N16R8 recommended). Text-to-phoneme conversion runs on the host.
License and attribution
The weights and the Ito audio are licensed CC BY-NC-SA 4.0 together with Lokutor's terms of use. Commercial use needs a written license from Lokutor. That covers products and devices, paid services and APIs, internal business use and commercial content, and also weights derived from Ito. The terms also exclude training commercial TTS or voice models on Ito's output. The engine, firmware and Python package are GPLv3, with commercial licenses available.
Ito learned its voice from StyleTTS 2 (MIT) conditioned on LibriTTS-R speakers
4970 (female) and 5105 (male) (LibriTTS-R, Koizumi et al. 2023, CC BY 4.0;
derived from LibriTTS and LibriVox public-domain recordings). Ito is not affiliated with or endorsed by those
speakers, LibriVox or Google. No third-party weights or recordings are included. Its
output is synthetic speech: please say so when you share it. Full attributions are in NOTICE.
Attribution: "Ito voice model by Lokutor (lokutor.com), CC BY-NC-SA 4.0".
We'd love to hear what you build. Lokutor offers no-cost licenses to small companies, startups, makers selling small batches, schools and research groups, and has more voices, more languages and speech recognition for the same chip. Write to contact@lokutor.com.
