Access Ito, a non-commercial voice model

Ito's weights are free for hobby, research, education and personal use. Tell us a little about your project: it helps us decide what to build next, and we answer every request for a commercial or collaboration license.

The Ito voice model is licensed under CC BY-NC-SA 4.0 together with Lokutor's terms of use (TERMS.md in this repository). Without a written license from Lokutor you may not use the weights, weights derived from them, or the audio they produce commercially, and you may not use Ito's output to train or improve a text-to-speech or voice model that is offered or used commercially. Ito's output is synthetic speech in the voice of a real (LibriTTS-R) speaker: say that it is synthetic when you share it, and never use it to impersonate or deceive. Small companies, startups, makers, schools and research groups can ask for a no-cost license at contact@lokutor.com.

Log in or Sign Up to review the conditions and access this model content.

Ito: natural-sounding streaming TTS for the ESP32-S3

Ito: natural speech from a $5 chip

Ito is English text-to-speech built to run entirely on an ESP32-S3 microcontroller (240 MHz dual-core, 8 MB PSRAM, 16 MB flash), with no cloud and no neural accelerator. It streams: audio starts after a short first chunk (25 ms), and the work before it does not grow with the sentence length. Its 4.4 M parameters fit in 4.9 MB of int8 weights.

Watch the 50 s video: lokutor-ai.github.io/ito/demo.mp4 (with sound) Β· Listen next to sanoTTS: lokutor-ai.github.io/ito Β· Code, firmware and tools: github.com/lokutor-ai/ito (GPLv3, commercial licenses available). From Lokutor, the makers of OΓ­do, speech recognition on the same chip.

Status (3 October 2026). The on-chip engine is verified on a laptop and in Espressif's QEMU emulator: the firmware's audio is bit-identical to the host build of the engine, and every sample below is that engine's exact output. Nothing has run on a physical board yet. Time to first audio and real-time factor are estimated from exact QEMU instruction counts and an assumed PSRAM bandwidth. Real-time playback is not established on silicon: the optimistic and central estimates are faster than real time (RTF 0.53–0.54 and 0.77–0.80), the pessimistic one is still slower (1.18–1.27). An earlier version of this card (and our first GitHub README) said 130–210 ms and real time at 0.5–1 GOPS; that was too optimistic and is corrected below. Board measurements are coming.

Voices

Two voices, each distilled from StyleTTS 2 conditioned on recordings of one LibriTTS-R reader (LibriTTS-R, CC BY 4.0). Each voice is a separate model with the same architecture, size and speed.

Voice Speaker PyTorch Chip (flash at 0x200000)
female (default) LibriTTS-R speaker 4970 ito_female.pt ito_female_esp32s3.bin
male LibriTTS-R speaker 5105 ito_male.pt ito_male_esp32s3.bin

The male voice is new (2 October 2026). Its chip file passes the same host and QEMU checks as the female voice (quantised vs float PESQ 4.53; firmware PCM bit-identical to the host engine). It has not been in a blind listening test yet; the results below are for the female voice.

Listen

The female voice, the chip engine's exact output, for eight sentences Ito never saw in training. The male voice reads the same sentences on the demo page.

Sentence Ito (chip-exact)
01 Hey, are you still coming over for dinner tonight, or should I save you a plate?
02 Your package should arrive on Friday, October 9th, sometime before noon.
03 When I finally got to the station, the last train had already left, so I ended up sharing a taxi with two strangers who turned out to be surprisingly good company.
04 The pharmacist recommended an anti-inflammatory, but honestly, I'd rather try physiotherapy first.
05 Thanks so much for calling. I'll check the schedule and get back to you first thing tomorrow morning.
06 Could you grab some quinoa and Worcestershire sauce on your way home?
07 It's about 23 degrees outside, so you probably won't need a jacket.
08 I know it sounds strange, but I actually enjoy the quiet hours before everyone else wakes up.

The players stream from the public demo page, so they work before you accept the gate. The bit-exact WAVs are in samples/ in this repository. The demo page plays the same sentences from sanoTTS and from Ito's teacher, with a blind mode.

Results

Blind listening test #9. One expert listener rated naturalness from 1 to 5, with system names hidden. Each system read four new conversational sentences.

System Parameters Runs on Mean (4 clips)
Teacher: StyleTTS 2 (LibriTTS model) large GPU / laptop 4.75
Ito 4.4 M ESP32-S3 (bit-exact in emulation) 4.00
sanoTTS amy 1.46 M browser / desktop (not run on an MCU) 2.00
sanoTTS heart-nano 0.29 M ESP32-S3 1.00

Automatic metrics on the eight sentences above:

Teacher Ito sanoTTS amy sanoTTS heart-nano
UTMOS 4.49 4.46 3.98 2.07
WER, Whisper medium.en / base 0 / 0 % 0 / 0 % 0 / 1.0 % 1.0 / 1.0 %

Please read these with their limits:

  • One listener (Lokutor's founder, so not a neutral party) and four sentences per system. This is a strong direction, not a MOS study. With n = 4, differences under about half a point are noise.
  • The rated Ito clips are the float model through the same streaming path as the chip. The quantised chip engine scores PESQ 4.52 against it, but the exact chip configuration has not been blind-rated yet.
  • UTMOS cannot hear intonation. Treat it as a check, not a verdict.

What we think this supports, and no more: the highest automatically measured naturalness of any complete text-to-waveform neural TTS built for a microcontroller without an NPU (UTMOS 4.46 on 8 sentences; the best previously published on-chip model scores 2.80), and the first streaming neural TTS designed for the ESP32-S3. Both are verified in emulation, not yet on a board. Ito is not the first or the smallest TTS on a microcontroller.

A broader benchmark against other small and embedded TTS systems is being finalised; it will be in bench/ on GitHub.

Size and compute

Parameters 4.40 M: acoustic front 1.62 M + vocoder 2.79 M
Chip weights 4.89 MB (ito_female_esp32s3.bin): int8 mel head and vocoder, int16 pitch path
Memory peak 6.5 of 8 MB PSRAM, 340 of 384 KB internal SRAM (QEMU; about 353 KB on the chip with the I2S buffers)
Work before the first audio (125 ms chunk) 25.3–26.3 M instructions and 5.1 MB of weights read from PSRAM, for any sentence length (exact counts from QEMU)
Time to first audio estimated, not measured: 137–143 ms optimistic, 200–207 ms central, 311–318 ms pessimistic. The first chunk is 125 ms of audio and the chunks behind it are sized so that playback can start with it without a gap
Real-time factor estimated, not measured: 0.53–0.54 optimistic, 0.77–0.80 central, 1.18–1.27 pessimistic (below 1 is faster than real time). The pessimistic case (40 MB/s PSRAM, nothing overlapped) is still slower than real time
Start delay for gapless speech estimated: 137–143 ms optimistic (the same as the time to first audio), 215–240 ms central, 880 ms or more pessimistic (the real-time factor is above 1 there, so the delay grows with the sentence)
On a laptop RTF β‰ˆ 0.01 on an Apple M4 Max CPU

All timings are estimated from exact instruction counts, not measured on silicon. The open question is the effective PSRAM bandwidth and how much of it overlaps compute; the pessimistic column depends on both. A weight pass now serves 24-frame (300 ms) chunks, which halves the weight traffic (38 to 17 MB per second of audio) with bit-identical output. The firmware benchmarks itself at boot and prints BOARD_SUMMARY. If you flash a board, please open an issue with that line.

How it works

Text becomes phonemes on the host (espeak-ng), and token ids go to the chip over serial. On the chip, a streaming acoustic front (1.62 M; a forward-only GRU predicts durations, pitch, energy and a mel spectrogram) feeds a Vocos-style vocoder (2.79 M; ConvNeXt blocks, a harmonic pitch source and an inverse STFT) that outputs 24 kHz audio in 100 ms chunks. Ito learned its voice from StyleTTS 2 speaking as a LibriTTS-R reader. The engine is new portable C99 code with the chip's integer arithmetic; the same source builds on a laptop. The training code and recipe are not public.

Use

git clone https://github.com/lokutor-ai/ito && cd ito && pip install -e .
hf download lokutor-ai/ito --include "ito_*" --local-dir models   # both voices (accept the terms first)

ito-tts "Good morning! The coffee is ready." -o hello.wav     # PyTorch, CPU is fine; female voice
ito-tts --voice male "Good morning! The coffee is ready." -o hello_male.wav   # male voice
cd esp32/host && make && cd ../..                              # the chip engine, built for your laptop
python esp32/tools/chip_wav.py "Good morning! The coffee is ready." hello_chip.wav   # chip-exact audio
esp32/tools/flash.sh /dev/ttyUSB0                              # ESP32-S3-DevKitC-1 N16R8 + I2S DAC/amp (female voice)
VOICE=male esp32/tools/flash.sh /dev/ttyUSB0                    # ... or the male voice
python esp32/tools/say.py "Hello from a five dollar chip." --port /dev/ttyUSB0
from ito import Ito
tts = Ito.load("models/ito_female.pt")                 # female voice;  Ito.load(voice="male") or "models/ito_male.pt" for the male voice
wav = tts.synthesize("Could you grab some quinoa on your way home?")   # float32 numpy, 24 kHz
for chunk in tts.stream("A sentence of any length."):                  # 100 ms chunks, as on the chip
    ...

Files

  • ito_female_esp32s3.bin: the female voice's chip weights (4.89 MB), for the firmware and the host build of the engine.
  • ito_female.pt: the female voice's PyTorch inference checkpoint (18 MB): front, vocoder, the fixed style the chip uses, and an optional text-to-style predictor.
  • ito_male_esp32s3.bin, ito_male.pt: the same for the male voice. The male checkpoint has no text-to-style predictor; it uses the fixed style, as the chip does.
  • samples/01.wav … samples/08.wav: the chip engine's exact output for the sentences above (24 kHz, 16-bit).
  • LICENSE-WEIGHTS, TERMS.md, NOTICE: the license, the terms of use, and third-party attributions.

Limitations

  • English only, two voices (one per weights file), one fixed speaking style on the chip.
  • Speed on the chip is estimated until board measurements are published.
  • The pitch range is slightly narrower than the teacher's (female 0.94, male 0.97).
  • Needs an ESP32-S3 with 8 MB PSRAM (N16R8 recommended). Text-to-phoneme conversion runs on the host.

License and attribution

The weights and the Ito audio are licensed CC BY-NC-SA 4.0 together with Lokutor's terms of use. Commercial use needs a written license from Lokutor. That covers products and devices, paid services and APIs, internal business use and commercial content, and also weights derived from Ito. The terms also exclude training commercial TTS or voice models on Ito's output. The engine, firmware and Python package are GPLv3, with commercial licenses available.

Ito learned its voice from StyleTTS 2 (MIT) conditioned on LibriTTS-R speakers 4970 (female) and 5105 (male) (LibriTTS-R, Koizumi et al. 2023, CC BY 4.0; derived from LibriTTS and LibriVox public-domain recordings). Ito is not affiliated with or endorsed by those speakers, LibriVox or Google. No third-party weights or recordings are included. Its output is synthetic speech: please say so when you share it. Full attributions are in NOTICE.

Attribution: "Ito voice model by Lokutor (lokutor.com), CC BY-NC-SA 4.0".

We'd love to hear what you build. Lokutor offers no-cost licenses to small companies, startups, makers selling small batches, schools and research groups, and has more voices, more languages and speech recognition for the same chip. Write to contact@lokutor.com.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support