hv-intent-router-word-2048
Parameters: 106 Γ 2048 Γ 2 = 434,176 trainable bits (434.2K)
Codebook size: 53 KB
Total model size: 71,863 bytes (~70 KB)
5-fold CV accuracy: ~73% (varies 71β76% across random seeds)
Random baseline: 20.0%
Training time: 0.6 seconds on a single CPU core
Dependencies: NumPy only
Model Card for hv-intent-router-word-2048
Model Description
A 70 KB hypervector intent classifier for 5-class routing (greeting, math, reminder, time, weather). The model uses word-level hyperdimensional computing with supervised word weighting and weighted k-nearest-neighbour voting. No neural networks, no gradient descent, no GPU.
The full model fits in L2 cache. Training completes in under one second. Inference is one similarity comparison per bank item.
Architecture
- Tokens: Whitespace-split words. Vocabulary is 106 unique words drawn from the training phrases.
- Encoding: Each word is assigned a 2048-bit bipolar hypervector. A phrase is encoded by summing weighted word hypervectors, then normalizing to unit length.
- Supervised weights: Each word gets a weight in [0, 1] based on how discriminative it is across the 5 intents. Function words like "the" and "what" get weight near 0; discriminating words like "weather" and "calculate" get weight near 1. Words below 0.15 are dropped entirely.
- Augmented bank: Each training phrase is expanded into 13 variants (1 original + 12 word-drop augments), giving a bank of 650 items.
- Two codebooks: Two independent random codebooks, each encoding the full augmented bank.
- Classifier: For each codebook, cosine similarity between query and every bank item. The top-7 nearest neighbours vote for their intent, weighted by similarity. Votes from both codebooks are summed. Argmax over 5 intents is the prediction.
- Query augmentation: At inference, the query is dropped-augmented twice. All three variants vote.
Accuracy and Scaling
The model was tested at three codebook dimensions under identical conditions (5-fold CV, same augmentation, same weights, same classifier). Only D changed.
| D | Parameters | Codebook size | 5-fold CV | Wall time |
|---|---|---|---|---|
| 2,048 | 434,176 | 53 KB | 71.0% | 0.3s |
| 20,480 | 4,341,760 | 530 KB | 72.0% | 1.2s |
| 204,800 | 43,417,600 | 5.3 MB | 72.0% | 18.7s |
The model's accuracy is flat across 100Γ variation in parameter count. Going from 434 K to 43 M parameters buys one additional correct classification out of 100. Runtime scales linearly with D, as expected, but accuracy does not.
The same model run twice with different random seeds varies by 2β3 examples on the 50-phrase test set. The true accuracy of the D=2048 model is approximately 73% Β± 3%. The single-run number of 76% reported in an earlier version of this card was one sample from that distribution.
What this scaling result means
For this dataset, dimension is not the bottleneck. The model's errors come from phrases that share nearly all their vocabulary with a competing intent:
- "when does the store close" (time) vs "what is the weather today" (weather) β both start with a wh-question word
- "set an alarm for six" (reminder) vs "when is the meeting" (time) β both describe a scheduled event
- "how hot is it" (weather) vs "how late is it" (time) β identical structure, different function word
Adding more dimensions cannot separate these phrases, because their word-level overlap is the same regardless of embedding size. Only additional training data or a different tokenisation scheme can help.
Where larger D does help
At D=2,048, the augmented bank of 650 items is near the measured capacity of a single hypervector (~64 items at 90% retrieval per hypervector, though here the bank uses one hypervector per bank item rather than bundling). Increasing D would be useful if the training set were 5β10Γ larger, because the number of distinct discriminative patterns would exceed what D=2,048 can represent.
The recommendation: use D=2,048 for datasets up to ~100 training phrases. Above that, increase D proportionally to the number of distinct phrases per intent.
Training Data
| Intent | Phrases |
|---|---|
| greeting | 10 |
| math | 10 |
| reminder | 10 |
| time | 10 |
| weather | 10 |
| Total | 50 |
Training phrases are short English commands, e.g. "what is the weather today", "remind me to call mom", "what is two plus two".
Evaluation
| Model | Accuracy | Size | Training Time |
|---|---|---|---|
| Random codebook | 20.0% | 53 KB | 0s |
| This model (5-fold CV) | ~73% | 70 KB | 0.6s |
| This model at 100Γ D | 72.0% | 5.3 MB | 18.7s |
| DistilBERT (reference, 66M params) | ~95% | 250 MB | hours |
This model reaches ~73% with roughly 4000Γ fewer parameters than a distilled transformer. It fits comfortably in an always-on memory region of any modern CPU.
History of this model
| Version | Method | 5-fold CV |
|---|---|---|
| v1 | Character-level, single memory, GA-evolved codebook | 44.0% |
| v2 | Word-level, single memory | 52.0% |
| v3 | Word + word-drop augmentation | 68.0% |
| v4 | Word + augmentation + supervised word weights | 73β76% |
| v4 at D=20,480 | Same, 10Γ parameters | 72.0% |
| v4 at D=204,800 | Same, 100Γ parameters | 72.0% |
Each row is measured under identical 5-fold CV. The jump from 44% to 73% came from changing the representation (character to word, then adding weights and augmentation). The jump from 73% to 72% is noise.
Intended Use
- Edge intent routing β classify short commands on microcontrollers, DSPs, or always-on co-processors where a transformer cannot fit.
- Pre-filtering β route requests to the correct downstream model before invoking it, saving compute on out-of-scope queries.
- Few-shot classification β train in under a second on a new label set without GPU or autograd.
- On-device personalisation β retrain per user with a handful of example phrases.
Limitations
- Closed vocabulary. The 106-word codebook covers only words that appear in the training data. Out-of-vocabulary words are silently dropped.
- Small training set. 50 phrases across 5 intents. The 5-fold CV estimate has a 95% confidence interval of roughly 62β84%. The true generalisation accuracy is likely 71β76%.
- No semantics. The model learns word-level co-occurrence, not meaning. Paraphrases that share no words with the training set will not be classified correctly.
- Bag of words. Word order is discarded. "what time is it" and "is it time" produce similar encodings.
- No handling of negation. "Is it not sunny" and "is it sunny" produce nearly identical encodings.
- Parameter scaling does not improve accuracy on this dataset. See the scaling table above.
How to Use
from hv_intent import load_v2, predict
model = load_v2("zeechimp/zee")
predict(model, "what is the weather today") # -> "weather"
predict(model, "remind me to call mom") # -> "reminder"
predict(model, "calculate ten times three") # -> "math"
- Downloads last month
- 53
Evaluation results
- 5-fold CV Accuracy (mean over seeds) on 5-class intent routingself-reported73.000
- Random Baseline on 5-class intent routingself-reported20.000