WakeHuBERT tiny

A 0.64M-parameter streaming speech feature extractor for wake-word detection, distilled from HuBERT-base. It turns 16 kHz audio into 128-dimensional features at 50 frames per second, so a small classifier trained on those features, from synthetic speech alone, can detect a wake word typed as text. It is the extractor behind wakeforge's wakehubert featurizer.

Using it

The ONNX graph takes waveform [batch, samples] (float32, 16 kHz, -1..1) and returns features [batch, samples // 320, 128]. The extractor is strictly causal: frame t depends only on audio before sample 320·(t+1), within a receptive field of 2.5 s. It was trained to reproduce each HuBERT frame five frames (100 ms) later, so its features describe the speech with a delay of about 100 ms and a detector built on it reacts that much after the word ends. For streaming, keep 40,000 samples (2.5 s) of context so streamed frames equal offline ones. The log-mel front end is part of the graph. wakehubert_int8.onnx is a static int8 version, with the front end kept in float, whose features agree with float at a mean cosine of 0.998.

import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
path = hf_hub_download("TigreGotico/wakehubert-tiny", "wakehubert.onnx")
sess = ort.InferenceSession(path)
feats = sess.run(None, {"waveform": np.zeros((1, 24000), np.float32)})[0]  # (1, 75, 128)

In wakeforge: ww_trainer-train --featurizer wakehubert ... downloads and uses it.

How it was made

The student is a fixed causal log-mel front end (64 bins), a strided convolution to 50 frames per second, eight dilated depthwise-separable convolution blocks with 256 channels, and a 1×1 projection to 128 features: convolution, batch norm and ReLU only, which quantises well. It was trained for 30,000 steps to predict standardised HuBERT-base layers 4, 8 and 12 (L1 plus log-sigmoid cosine, as in DistilHuBERT), with the teacher hearing clean speech while the student heard it with AudioSet, MUSAN and non-speech noise, room reverberation and one to three background talkers never louder than the voice. A quarter of the training items were non-speech sounds heard identically by both. Spans of the student's input log-mel were masked during training (probability 0.065 per frame, 10-frame spans), which was the largest single gain in robustness found in the experiments.

Training speech: LibriSpeech (960 h), Multilingual LibriSpeech (seven languages) and a language-balanced sample of Multilingual Spoken Words (41 languages), cut into 600,000 two-second crops.

Results

Single runs. A GRU classifier trained only on 900 synthetic "alexa" clips (TTS voices cloned with voice conversion, with noise, babble, reverberation, speed and gain augmentation) was scored on the real speakers of the Picovoice wake-word benchmark (315 recordings), with the threshold chosen on separate calibration audio (LibriSpeech dev-clean and babble made from it) for 0.5 false activations per hour, and false activations then measured on 6.5 h of held-out streams (LibriSpeech test-clean, three-talker test-other babble, held-out non-speech):

Extractor Recall, quiet (95% interval) Recall in babble at 10 / 5 / 0 dB False activations per hour measured
WakeHuBERT tiny (0.64M) 95% (92–97) 94 / 85 / 56% 0.31
Same architecture without masking 92% (89–95) 90 / 80 / 39% 0.46

The recall intervals come from resampling the 315 recordings; differences of a few points between extractors are within noise, and two classifier types on the same extractor can differ by several points.

Limitations

The evaluation covers one English wake word from one public benchmark; results for other words, languages and devices are not measured here. Agreement with the teacher is a poor predictor of detection quality, so the extractor should be judged by a detector trained on it.

Licence and attribution

Released under the Apache License 2.0, the licence of its teacher (HuBERT-base, Apache 2.0). Trained with LibriSpeech, Multilingual LibriSpeech and Multilingual Spoken Words (all CC BY 4.0), MUSAN, AudioSet-derived noise and room impulse responses, and distilled from facebook/hubert-base-ls960.

Downloads last month
136
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including TigreGotico/wakehubert-tiny