WakeHuBERT wake words
Ready English wake-word models for OpenVoiceOS, trained with wakeforge. Each model is a small GRU classifier that reads features from WakeHuBERT tiny, a 0.64M-parameter speech feature extractor distilled from HuBERT-base. A model scores the last 1.5 s of audio (75 feature frames) and outputs a logit; its sigmoid is the probability that the window holds the wake word.
Try them in your browser: WakeHuBERT wake-words Space runs every model here on your microphone or an uploaded file, with a live score and the calibrated threshold. Everything runs locally in the browser.
The repository holds one ONNX file per word under models/, and models.json, which lists each model's word,
featurizer, calibration, default threshold, SHA-256 and measured results. Every model carries the same facts in
its own ONNX metadata (wake_word, pretrained_featurizer, default_threshold, window_frames, license,
training_data, and on calibrated models calibrated, calib_a and calib_b).
| model | word | featurizer | calibrated | default threshold |
|---|---|---|---|---|
wakehubert_jarvis |
jarvis | wakehubert-int8 | yes | 0.57 |
wakehubert_alexa |
alexa | wakehubert-int8 | yes | 0.40 |
wakehubert_hey_jarvis |
hey jarvis | wakehubert-int8 | yes | 0.16 |
wakehubert_hey_marvin |
hey marvin | wakehubert-int8 | yes | 0.34 |
wakehubert_home_assistant |
home assistant | wakehubert-int8 | yes | 0.19 |
wakehubert_okay_nabu |
okay nabu | wakehubert-int8 | yes | 0.45 |
wakehubert_hello_nabu |
hello nabu | wakehubert-int8 | yes | 0.49 |
wakehubert_hey_chatterbox |
hey chatterbox | wakehubert-int8 | yes | 0.13 |
wakehubert_hey_floyd |
hey floyd | wakehubert-int8 | yes | 0.43 |
wakehubert_hey_rhasspy |
hey rhasspy | wakehubert-int8 | yes | 0.40 |
wakehubert_hey_robin |
hey robin | wakehubert-int8 | yes | 0.05 |
wakehubert_marvin |
marvin | wakehubert-int8 | yes | 0.06 |
wakehubert_sheila |
sheila | wakehubert-int8 | yes | 0.47 |
wakehubert_stop |
stop | wakehubert-int8 | yes | 0.14 |
wakehubert_android |
android | wakehubert-int8 | yes | 0.42 |
wakehubert_hey_computer |
hey computer | wakehubert-int8 | yes | 0.31 |
wakehubert_hey_k9 |
hey k9 | wakehubert-int8 | yes | 0.06 |
wakehubert_hey_scout |
hey scout | wakehubert-int8 | yes | 0.36 |
wakehubert_wake_up |
wake up | wakehubert-int8 | yes | 0.36 |
wakehubert_hey_ziggy |
hey ziggy | wakehubert-int8 | yes | 0.27 |
wakehubert_hey_stemcom |
hey stemcom | wakehubert-int8 | yes | 0.28 |
wakehubert_computer |
computer | wakehubert (float32) | no | 0.99 |
wakehubert_hey_mycroft |
hey mycroft | wakehubert (float32) | no | 0.965 |
Use with OpenVoiceOS
The models run in ovos-ww-plugin-wakeforge, which
downloads them from this repository on first use, at a pinned revision, checks each file against its sha256 in
models.json, and keeps them in the Hugging Face cache; the featurizer comes from
TigreGotico/wakehubert-tiny the same way. Install the plugin
and name the model in mycroft.conf:
pip install --pre ovos-ww-plugin-wakeforge
{
"listener": {
"wake_word": "hey_jarvis"
},
"hotwords": {
"hey_jarvis": {
"module": "ovos-ww-plugin-wakeforge",
"model": "wakehubert_hey_jarvis",
"listen": true
}
}
}
A model file from this repository also loads by path: set model to the local .onnx file. The plugin reads
the featurizer and the default threshold from the model's metadata.
Use in your own code
The models need only numpy, onnxruntime and huggingface_hub; nothing from OpenVoiceOS. Each model is a small classifier that reads WakeHuBERT-tiny features, so you run two ONNX files: the featurizer, then the wake-word model.
import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download
word = "jarvis"
head = ort.InferenceSession(hf_hub_download("OpenVoiceOS/wakehubert-wakewords", f"models/wakehubert_{word}.onnx"))
meta = head.get_modelmeta().custom_metadata_map
featurizer_file = "wakehubert_int8.onnx" if meta["pretrained_featurizer"] == "wakehubert-int8" else "wakehubert.onnx"
feat = ort.InferenceSession(hf_hub_download("TigreGotico/wakehubert-tiny", featurizer_file))
threshold = float(meta["default_threshold"])
WINDOW = 24000 # 1.5 s of 16 kHz audio: the window the models were trained on (75 feature frames)
BLOCK = 1280 # score every 80 ms
DEBOUNCE_BLOCKS = 25 # ignore 2 s after a detection
class WakeWordDetector:
def __init__(self):
self.buf = np.zeros(WINDOW, np.float32)
self.cooldown = 0
def push(self, chunk):
"""chunk: 1280 float32 samples at 16 kHz in -1..1. Returns (score, detected)."""
self.buf = np.concatenate([self.buf, chunk.astype(np.float32)])[-WINDOW:]
features = feat.run(None, {"waveform": self.buf[None]})[0] # [1, 75, 128]
logit = head.run(None, {"features": features})[0][0]
score = float(1.0 / (1.0 + np.exp(-logit)))
self.cooldown = max(0, self.cooldown - 1)
detected = score >= threshold and self.cooldown == 0
if detected:
self.cooldown = DEBOUNCE_BLOCKS
return score, detected
From a microphone, for example with sounddevice:
import sounddevice as sd
detector = WakeWordDetector()
with sd.InputStream(samplerate=16000, channels=1, dtype="float32", blocksize=BLOCK) as stream:
while True:
chunk, _ = stream.read(BLOCK)
score, detected = detector.push(chunk[:, 0])
if detected:
print(f"{word} detected (score {score:.2f})")
Each block featurizes the last 1.5 s on its own, exactly as the models were trained and scored. The featurizer is causal, the per-block cost is about 1–2 ms on one CPU core for the int8 featurizer, and one featurizer run can feed any number of wake-word models: run feat once per block and pass the same features to each model. Replace threshold to change sensitivity (see below). Checked against the plugin on real recordings: the snippet fires on the same clips.
Choosing a threshold
Set "threshold" in the hotword config to trade missed wake words against false activations. A lower value
fires more readily and falsely more often; a higher value misses more wake words and fires falsely less often.
On a calibrated model the number means the same thing for every word. Its score is mapped so that a threshold of 0.5 gives about one false activation per hour on held-out speech and noise. The default threshold is the point that maximises F2, which weighs recall above precision. Raising the threshold toward 0.8 or 0.9 trades recall for fewer false activations, and lowering it does the opposite.
The uncalibrated models (wakehubert_computer and wakehubert_hey_mycroft) output a probability too, but their
score is not mapped to a false-activation rate. Their threshold is not a calibrated knob: the same number gives a
different trade-off on each of them, and useful values sit close to 1.
Results
Measured through the plugin at the default threshold and at 0.8. Recall is the share of test clips detected; false activations are counted per hour of negative audio.
| model | recall at default | false activations/h at default | recall at 0.8 | false activations/h at 0.8 | recall test set |
|---|---|---|---|---|---|
wakehubert_jarvis |
98.4% | 0.69 | 93.8% | 0.15 | 384 Picovoice recordings of real speakers |
wakehubert_alexa |
86.7% | 0.69 | 73.3% | 0.13 | 315 Picovoice recordings of real speakers |
wakehubert_hey_jarvis |
94.5% | 0.09 | 86.2% | 0.02 | 384 clips in held-out synthetic voices |
wakehubert_hey_marvin |
97.4% | 0.95 | 90.7% | 0.24 | 386 clips in held-out synthetic voices |
wakehubert_home_assistant |
91.1% | 0.30 | 84.2% | 0.15 | 380 clips in held-out synthetic voices |
wakehubert_okay_nabu |
92.0% | 0.26 | 75.4% | 0.02 | 386 clips in held-out synthetic voices |
wakehubert_hello_nabu |
80.9% | 0.39 | 65.2% | 0.02 | 382 clips in held-out synthetic voices |
wakehubert_hey_chatterbox |
82.8% | 0.09 | 61.2% | 0.00 | 116 OVOS community recordings of real speakers |
wakehubert_hey_floyd |
91.7% | 0.47 | 82.3% | 0.02 | 96 OVOS community recordings of real speakers |
wakehubert_hey_rhasspy |
100.0% | 0.60 | 97.6% | 0.11 | 374 clips in held-out synthetic voices |
wakehubert_hey_robin |
99.5% | 0.39 | 97.9% | 0.06 | 380 clips in held-out synthetic voices |
wakehubert_marvin |
73.3% | 0.77 | 71.8% | 0.67 | 195 Speech Commands test recordings of real speakers |
wakehubert_sheila |
88.7% | 2.08 | 84.4% | 0.54 | 212 Speech Commands test recordings of real speakers |
wakehubert_stop |
86.6% | 2.06 | 76.9% | 0.45 | 411 Speech Commands test recordings of real speakers |
wakehubert_android |
98.7% | 0.95 | 96.4% | 0.11 | 390 clips in held-out synthetic voices |
wakehubert_hey_computer |
96.4% | 0.47 | 93.3% | 0.04 | 390 clips in held-out synthetic voices |
wakehubert_hey_k9 |
99.4% | 0.19 | 93.5% | 0.04 | 338 clips in held-out synthetic voices |
wakehubert_hey_scout |
95.9% | 0.09 | 93.8% | 0.04 | 390 clips in held-out synthetic voices |
wakehubert_wake_up |
98.1% | 1.10 | 96.8% | 0.19 | 378 clips in held-out synthetic voices |
wakehubert_hey_ziggy |
93.6% | 0.64 | 89.3% | 0.24 | 374 clips of OmniVoice and held-out edge-tts voices |
wakehubert_hey_stemcom |
89.0% | 0.21 | 83.1% | 0.06 | 337 clips of OmniVoice and held-out edge-tts voices |
The false activations are counted over 46.5 h of negative audio (speech, non-speech and household audio). The held-out synthetic voices are text-to-speech voices that no training clip uses, converted to the voices of speakers who appear in no training clip. No figures are published for the uncalibrated models.
Training data
The calibrated models were trained on synthetic speech only, with six edge-tts voices held out of training for
testing. The positives are an edge-tts voice grid, further edge-tts and Google Translate TTS voices and OmniVoice
clips, and for wakehubert_jarvis, wakehubert_hey_chatterbox, wakehubert_hey_floyd, wakehubert_marvin,
wakehubert_sheila and wakehubert_stop also voice-converted copies of edge-tts clips. Each model's training_data
metadata names its own sources. Most of these clips are published in the TigreGotico/synthetic-wakeword-<word>
datasets listed above. Speech from LibriSpeech train-clean-100 is mixed into training clips as background babble.
The negatives are the wakeforge negative list and half of an AudioSet noise sample; the other half is held out.
wakehubert_computer was trained on synthetic speech only, from TigreGotico/synthetic-wakeword-computer, with
negatives from TigreGotico/not-wake-words-speech-en and AudioSet-derived clips. wakehubert_hey_mycroft was trained on human
recordings.
Calibration
A calibrated model has an affine map folded into its graph, applied to the GRU's logit before the sigmoid. The map
is fitted on a stream of held-out noise and speech that the model did not train on, so that a probability of 0.5
falls at about one false activation per hour of that stream. The default threshold is then the F2-optimal point on
the calibrated scale. The fitted slope and offset are in each model's calib_a and calib_b metadata.
Limitations
- Twelve of the nineteen calibrated models are scored on held-out synthetic voices, so recall on real speech can
be lower for those words.
wakehubert_jarvisandwakehubert_alexaare scored on the Picovoice recordings,wakehubert_hey_chatterboxandwakehubert_hey_floydon OVOS community recordings, andwakehubert_marvin,wakehubert_sheilaandwakehubert_stopon Speech Commands, all real speakers. wakehubert_sheilaandwakehubert_stopgive about two false activations per hour at their default threshold; raise it toward 0.8 for about one every two hours.wakehubert_marvinhas a steep calibration, so its recall and false-activation rate change little between its 0.06 default and 0.8.- The calibration is fitted on about 10 h of audio with few false activations in it (2 to 14 per model), so the map is extrapolated, and "0.5 is about one false activation per hour" is approximate.
- On real jarvis recordings,
wakehubert_jarvispeaks just above its 0.57 default, so a quiet or distant speaker has little margin. - Similar-sounding words trigger each other's model.
wakehubert_jarvisfires on "hey jarvis",wakehubert_marvinon "hey marvin", andwakehubert_hey_jarvison "hey chatterbox".wakehubert_hello_nabu,wakehubert_hey_marvinandwakehubert_okay_nabucan fire on each other's words,wakehubert_hey_marvinalso on "hey rhasspy", "hey robin" and "marvin",wakehubert_okay_nabuon "hey rhasspy",wakehubert_hey_robinon "hey marvin", "hey rhasspy" and "okay nabu",wakehubert_okay_nabusometimes on "hey k9",wakehubert_androidsometimes on "hey floyd",wakehubert_computerandwakehubert_hey_computeron each other's words, andwakehubert_sheilaon "computer", all at their default thresholds. Raise the threshold when two of these models run side by side. - The models are English only.
License
Apache-2.0. The featurizer, TigreGotico/wakehubert-tiny, is
Apache-2.0. The synthetic-wakeword-* and not-wake-words-speech-en datasets are CC BY 4.0; LibriSpeech is
CC BY 4.0; AudioSet labels are CC BY 4.0 and its audio comes from YouTube videos under their uploaders' terms. The
voice-conversion and voice-cloning folders of the synthetic-wakeword-* datasets take their voices from Mozilla
Common Voice contributors, through the MLCommons Multilingual Spoken Words Corpus (CC BY 4.0).
Model tree for OpenVoiceOS/wakehubert-wakewords
Base model
facebook/hubert-base-ls960