ruvector-typesafe β€” CLINC150 bank

A trained example bank for @ruvector/typesafe: typed decisions over text, locally, with no API bill and no network in the decision path. 150 intents across 10 domains, plus 1,000 deliberately out-of-scope utterances.

npm install @ruvector/typesafe
curl -LO https://huggingface.co/ruvnet/ruvector-typesafe-clinc150/resolve/main/bank.json
curl -LO https://huggingface.co/ruvnet/ruvector-typesafe-clinc150/resolve/main/questions.json

echo "what expression would i use to say i love you in italian" \
  | npx typesafe decide --questions questions.json --bank bank.json --embedder onnx \
      --engine-options '{"probeIterations":4000,"probeClassBalanced":true,"head":"probe"}'

The --engine-options are not optional. A bank stores examples and their frozen splits β€” never hyperparameters β€” and the head is refit from it on load. Omit them and you refit with the library defaults (400 iterations, head: auto), which is a different and materially worse model than the one measured below. These are the exact options this bank was trained under.

What this artifact is

Labelled examples with frozen split assignments β€” not weights. The heads (nearest-prototype, or a multinomial probe once a class has enough examples) and the temperature calibration are refit from the bank when the engine loads it.

Two consequences worth knowing:

  • The bank is encoder-independent β€” it holds text and content-hashed split tags, nothing encoder-derived. Verified directly on the banking77 bank, where both bundled encoders exported byte-identical banks; here only one encoder was run. Only the accuracy below is encoder-specific.
  • The first decision after loading is slow. importBankJson just admits the examples; the head is fitted lazily on the first decide, and at probeIterations: 4000 over 15,000 examples and 150 classes that fit is minutes, not milliseconds. Every later call is the steady-state latency in the table. Import once, warm it with one throwaway decision, and keep the engine alive β€” do not load a bank per request.

A reloaded bank reproduces the trained engine's answers exactly β€” that round-trip is asserted by test/bank-roundtrip.test.mjs, not assumed.

Accuracy

Held-out test split, 4,500 utterances, 150 classes.

encoder accuracy p50 latency p95 latency
bge-small-en-v1.5 91.0% 4 ms 6 ms

Only bge-small-en-v1.5 was measured for this bank. The bank itself is encoder-independent β€” the other bundled encoder will load it, it just has no number here.

Trained on 15,000 labelled examples (splitsHash: cfe29e0a95ec973c…).

Out-of-scope detection

CLINC150 ships 1000 utterances that belong to no intent. @ruvector/typesafe answers with an abstain mass that the Jev API it replaces has no field for at all. Ranking out-of-scope against in-scope text by that one number:

metric value
Abstain AUROC (out-of-scope vs in-scope) 0.9011
Mean abstain, out-of-scope 5.07e-7
Mean abstain, in-scope 2.55e-8

An AUROC of 0.5 would mean abstain carries no signal about scope.

Rank it, do not threshold it. Look at those two means: the separation is about 20x, which is what earns the AUROC β€” but both numbers are minuscule in absolute terms. A fixed cut-off like the engine's default abstainTau of 0.35 will never fire here. Sort by abstain, or calibrate a threshold against your own traffic; do not copy a constant from another deployment.

Training

node scripts/typesafe-banks/build-bank.mjs --dataset clinc150 --encoder bge-small-en-v1.5

The probe head is full-batch gradient descent with a fixed iteration count, and that count is the thing to tune when you add data: the default 400 iterations fits ~1k examples well and underfits ~10k badly. Raise it through EngineOptions:

createTypesafe({ embedder: …, engine: { probeIterations: 4000 } })

Limitations

  • English only; both bundled encoders are English sentence encoders.
  • The label set is closed. New intents need new examples and a refit.
  • Accuracy is reported on this dataset's own test split β€” it is not a claim about your traffic.

Credit

Larson et al., An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction (EMNLP 2019). Dataset licence: CC-BY-3.0; this bank redistributes the utterance text under that licence. The @ruvector/typesafe code is MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train ruvnet/ruvector-typesafe-clinc150