WakeHuBERT wake words

Ready English wake-word models for OpenVoiceOS, trained with wakeforge. Each model is a small GRU classifier that reads features from WakeHuBERT tiny, a 0.64M-parameter speech feature extractor distilled from HuBERT-base. A model scores the last 1.5 s of audio (75 feature frames) and outputs a logit; its sigmoid is the probability that the window holds the wake word.

Try them in your browser: WakeHuBERT wake-words Space runs every model here on your microphone or an uploaded file, with a live score and the calibrated threshold. Everything runs locally in the browser.

The repository holds one ONNX file per word under models/, and models.json, which lists each model's word, featurizer, calibration, default threshold, SHA-256 and measured results. Every model carries the same facts in its own ONNX metadata (wake_word, pretrained_featurizer, default_threshold, window_frames, license, training_data, and on calibrated models calibrated, calib_a and calib_b).

model word featurizer calibrated default threshold
wakehubert_jarvis jarvis wakehubert-int8 yes 0.57
wakehubert_alexa alexa wakehubert-int8 yes 0.40
wakehubert_hey_jarvis hey jarvis wakehubert-int8 yes 0.16
wakehubert_hey_marvin hey marvin wakehubert-int8 yes 0.34
wakehubert_home_assistant home assistant wakehubert-int8 yes 0.19
wakehubert_okay_nabu okay nabu wakehubert-int8 yes 0.45
wakehubert_hello_nabu hello nabu wakehubert-int8 yes 0.49
wakehubert_hey_chatterbox hey chatterbox wakehubert-int8 yes 0.13
wakehubert_hey_floyd hey floyd wakehubert-int8 yes 0.43
wakehubert_hey_rhasspy hey rhasspy wakehubert-int8 yes 0.40
wakehubert_hey_robin hey robin wakehubert-int8 yes 0.05
wakehubert_marvin marvin wakehubert-int8 yes 0.06
wakehubert_sheila sheila wakehubert-int8 yes 0.47
wakehubert_stop stop wakehubert-int8 yes 0.14
wakehubert_android android wakehubert-int8 yes 0.42
wakehubert_hey_computer hey computer wakehubert-int8 yes 0.31
wakehubert_hey_k9 hey k9 wakehubert-int8 yes 0.06
wakehubert_hey_scout hey scout wakehubert-int8 yes 0.36
wakehubert_wake_up wake up wakehubert-int8 yes 0.36
wakehubert_hey_ziggy hey ziggy wakehubert-int8 yes 0.27
wakehubert_hey_stemcom hey stemcom wakehubert-int8 yes 0.28
wakehubert_computer computer wakehubert (float32) no 0.99
wakehubert_hey_mycroft hey mycroft wakehubert (float32) no 0.965

Use with OpenVoiceOS

The models run in ovos-ww-plugin-wakeforge, which downloads them from this repository on first use, at a pinned revision, checks each file against its sha256 in models.json, and keeps them in the Hugging Face cache; the featurizer comes from TigreGotico/wakehubert-tiny the same way. Install the plugin and name the model in mycroft.conf:

pip install --pre ovos-ww-plugin-wakeforge
{
  "listener": {
    "wake_word": "hey_jarvis"
  },
  "hotwords": {
    "hey_jarvis": {
      "module": "ovos-ww-plugin-wakeforge",
      "model": "wakehubert_hey_jarvis",
      "listen": true
    }
  }
}

A model file from this repository also loads by path: set model to the local .onnx file. The plugin reads the featurizer and the default threshold from the model's metadata.

Use in your own code

The models need only numpy, onnxruntime and huggingface_hub; nothing from OpenVoiceOS. Each model is a small classifier that reads WakeHuBERT-tiny features, so you run two ONNX files: the featurizer, then the wake-word model.

import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download

word = "jarvis"
head = ort.InferenceSession(hf_hub_download("OpenVoiceOS/wakehubert-wakewords", f"models/wakehubert_{word}.onnx"))
meta = head.get_modelmeta().custom_metadata_map
featurizer_file = "wakehubert_int8.onnx" if meta["pretrained_featurizer"] == "wakehubert-int8" else "wakehubert.onnx"
feat = ort.InferenceSession(hf_hub_download("TigreGotico/wakehubert-tiny", featurizer_file))
threshold = float(meta["default_threshold"])

WINDOW = 24000          # 1.5 s of 16 kHz audio: the window the models were trained on (75 feature frames)
BLOCK = 1280            # score every 80 ms
DEBOUNCE_BLOCKS = 25    # ignore 2 s after a detection


class WakeWordDetector:
    def __init__(self):
        self.buf = np.zeros(WINDOW, np.float32)
        self.cooldown = 0

    def push(self, chunk):
        """chunk: 1280 float32 samples at 16 kHz in -1..1. Returns (score, detected)."""
        self.buf = np.concatenate([self.buf, chunk.astype(np.float32)])[-WINDOW:]
        features = feat.run(None, {"waveform": self.buf[None]})[0]       # [1, 75, 128]
        logit = head.run(None, {"features": features})[0][0]
        score = float(1.0 / (1.0 + np.exp(-logit)))
        self.cooldown = max(0, self.cooldown - 1)
        detected = score >= threshold and self.cooldown == 0
        if detected:
            self.cooldown = DEBOUNCE_BLOCKS
        return score, detected

From a microphone, for example with sounddevice:

import sounddevice as sd

detector = WakeWordDetector()
with sd.InputStream(samplerate=16000, channels=1, dtype="float32", blocksize=BLOCK) as stream:
    while True:
        chunk, _ = stream.read(BLOCK)
        score, detected = detector.push(chunk[:, 0])
        if detected:
            print(f"{word} detected (score {score:.2f})")

Each block featurizes the last 1.5 s on its own, exactly as the models were trained and scored. The featurizer is causal, the per-block cost is about 1–2 ms on one CPU core for the int8 featurizer, and one featurizer run can feed any number of wake-word models: run feat once per block and pass the same features to each model. Replace threshold to change sensitivity (see below). Checked against the plugin on real recordings: the snippet fires on the same clips.

Choosing a threshold

Set "threshold" in the hotword config to trade missed wake words against false activations. A lower value fires more readily and falsely more often; a higher value misses more wake words and fires falsely less often.

On a calibrated model the number means the same thing for every word. Its score is mapped so that a threshold of 0.5 gives about one false activation per hour on held-out speech and noise. The default threshold is the point that maximises F2, which weighs recall above precision. Raising the threshold toward 0.8 or 0.9 trades recall for fewer false activations, and lowering it does the opposite.

The uncalibrated models (wakehubert_computer and wakehubert_hey_mycroft) output a probability too, but their score is not mapped to a false-activation rate. Their threshold is not a calibrated knob: the same number gives a different trade-off on each of them, and useful values sit close to 1.

Results

Measured through the plugin at the default threshold and at 0.8. Recall is the share of test clips detected; false activations are counted per hour of negative audio.

model recall at default false activations/h at default recall at 0.8 false activations/h at 0.8 recall test set
wakehubert_jarvis 98.4% 0.69 93.8% 0.15 384 Picovoice recordings of real speakers
wakehubert_alexa 86.7% 0.69 73.3% 0.13 315 Picovoice recordings of real speakers
wakehubert_hey_jarvis 94.5% 0.09 86.2% 0.02 384 clips in held-out synthetic voices
wakehubert_hey_marvin 97.4% 0.95 90.7% 0.24 386 clips in held-out synthetic voices
wakehubert_home_assistant 91.1% 0.30 84.2% 0.15 380 clips in held-out synthetic voices
wakehubert_okay_nabu 92.0% 0.26 75.4% 0.02 386 clips in held-out synthetic voices
wakehubert_hello_nabu 80.9% 0.39 65.2% 0.02 382 clips in held-out synthetic voices
wakehubert_hey_chatterbox 82.8% 0.09 61.2% 0.00 116 OVOS community recordings of real speakers
wakehubert_hey_floyd 91.7% 0.47 82.3% 0.02 96 OVOS community recordings of real speakers
wakehubert_hey_rhasspy 100.0% 0.60 97.6% 0.11 374 clips in held-out synthetic voices
wakehubert_hey_robin 99.5% 0.39 97.9% 0.06 380 clips in held-out synthetic voices
wakehubert_marvin 73.3% 0.77 71.8% 0.67 195 Speech Commands test recordings of real speakers
wakehubert_sheila 88.7% 2.08 84.4% 0.54 212 Speech Commands test recordings of real speakers
wakehubert_stop 86.6% 2.06 76.9% 0.45 411 Speech Commands test recordings of real speakers
wakehubert_android 98.7% 0.95 96.4% 0.11 390 clips in held-out synthetic voices
wakehubert_hey_computer 96.4% 0.47 93.3% 0.04 390 clips in held-out synthetic voices
wakehubert_hey_k9 99.4% 0.19 93.5% 0.04 338 clips in held-out synthetic voices
wakehubert_hey_scout 95.9% 0.09 93.8% 0.04 390 clips in held-out synthetic voices
wakehubert_wake_up 98.1% 1.10 96.8% 0.19 378 clips in held-out synthetic voices
wakehubert_hey_ziggy 93.6% 0.64 89.3% 0.24 374 clips of OmniVoice and held-out edge-tts voices
wakehubert_hey_stemcom 89.0% 0.21 83.1% 0.06 337 clips of OmniVoice and held-out edge-tts voices

The false activations are counted over 46.5 h of negative audio (speech, non-speech and household audio). The held-out synthetic voices are text-to-speech voices that no training clip uses, converted to the voices of speakers who appear in no training clip. No figures are published for the uncalibrated models.

Training data

The calibrated models were trained on synthetic speech only, with six edge-tts voices held out of training for testing. The positives are an edge-tts voice grid, further edge-tts and Google Translate TTS voices and OmniVoice clips, and for wakehubert_jarvis, wakehubert_hey_chatterbox, wakehubert_hey_floyd, wakehubert_marvin, wakehubert_sheila and wakehubert_stop also voice-converted copies of edge-tts clips. Each model's training_data metadata names its own sources. Most of these clips are published in the TigreGotico/synthetic-wakeword-<word> datasets listed above. Speech from LibriSpeech train-clean-100 is mixed into training clips as background babble. The negatives are the wakeforge negative list and half of an AudioSet noise sample; the other half is held out.

wakehubert_computer was trained on synthetic speech only, from TigreGotico/synthetic-wakeword-computer, with negatives from TigreGotico/not-wake-words-speech-en and AudioSet-derived clips. wakehubert_hey_mycroft was trained on human recordings.

Calibration

A calibrated model has an affine map folded into its graph, applied to the GRU's logit before the sigmoid. The map is fitted on a stream of held-out noise and speech that the model did not train on, so that a probability of 0.5 falls at about one false activation per hour of that stream. The default threshold is then the F2-optimal point on the calibrated scale. The fitted slope and offset are in each model's calib_a and calib_b metadata.

Limitations

  • Twelve of the nineteen calibrated models are scored on held-out synthetic voices, so recall on real speech can be lower for those words. wakehubert_jarvis and wakehubert_alexa are scored on the Picovoice recordings, wakehubert_hey_chatterbox and wakehubert_hey_floyd on OVOS community recordings, and wakehubert_marvin, wakehubert_sheila and wakehubert_stop on Speech Commands, all real speakers.
  • wakehubert_sheila and wakehubert_stop give about two false activations per hour at their default threshold; raise it toward 0.8 for about one every two hours. wakehubert_marvin has a steep calibration, so its recall and false-activation rate change little between its 0.06 default and 0.8.
  • The calibration is fitted on about 10 h of audio with few false activations in it (2 to 14 per model), so the map is extrapolated, and "0.5 is about one false activation per hour" is approximate.
  • On real jarvis recordings, wakehubert_jarvis peaks just above its 0.57 default, so a quiet or distant speaker has little margin.
  • Similar-sounding words trigger each other's model. wakehubert_jarvis fires on "hey jarvis", wakehubert_marvin on "hey marvin", and wakehubert_hey_jarvis on "hey chatterbox". wakehubert_hello_nabu, wakehubert_hey_marvin and wakehubert_okay_nabu can fire on each other's words, wakehubert_hey_marvin also on "hey rhasspy", "hey robin" and "marvin", wakehubert_okay_nabu on "hey rhasspy", wakehubert_hey_robin on "hey marvin", "hey rhasspy" and "okay nabu", wakehubert_okay_nabu sometimes on "hey k9", wakehubert_android sometimes on "hey floyd", wakehubert_computer and wakehubert_hey_computer on each other's words, and wakehubert_sheila on "computer", all at their default thresholds. Raise the threshold when two of these models run side by side.
  • The models are English only.

License

Apache-2.0. The featurizer, TigreGotico/wakehubert-tiny, is Apache-2.0. The synthetic-wakeword-* and not-wake-words-speech-en datasets are CC BY 4.0; LibriSpeech is CC BY 4.0; AudioSet labels are CC BY 4.0 and its audio comes from YouTube videos under their uploaders' terms. The voice-conversion and voice-cloning folders of the synthetic-wakeword-* datasets take their voices from Mozilla Common Voice contributors, through the MLCommons Multilingual Spoken Words Corpus (CC BY 4.0).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenVoiceOS/wakehubert-wakewords

Quantized
(3)
this model

Datasets used to train OpenVoiceOS/wakehubert-wakewords

Space using OpenVoiceOS/wakehubert-wakewords 1