BANKING77 MiniLM — 77 intents, a frozen encoder, and a small linear head

Route English banking-support messages locally. 92.76% accuracy on the 3,080-example BANKING77 test split, versus 88.47% for our TF-IDF baseline. The encoder has 22.7M parameters; only the 29,645-parameter classifier head was trained. This is a simple reproducible baseline, not a new architecture or state-of-the-art claim.

Weights use safetensors and the head uses plain JSON. Inference needs no scikit-learn, pickle, hosted API, or custom remote model code. The complete encoder is included, so the downloaded artifact works offline.

Try it in your browser · Compact ONNX export

Lightweight CPU option (no PyTorch)

Use the 45.4 MB ONNX artifact with ONNX Runtime and the native tokenizer: installation, Python API, and offline deployment. The helper is tested in an isolated environment without PyTorch or Transformers, and reproduces all 3,080 released test predictions. It supports local files, pinned Hub downloads, and cached offline loading. No model weights changed.

from onnx_predict import ONNXIntentClassifier
router = ONNXIntentClassifier(threads=2)
print(router.predict("How can I change my PIN?", top_k=3))

See the linked quickstart to install dependencies and download the helper first.

PyTorch option (CPU)

pip install "torch==2.14.0" "transformers==5.17.0" "huggingface_hub==1.32.0"
hf download Future-Labs/banking77-minilm frozen_predict.py --local-dir .
import torch
from frozen_predict import IntentClassifier

torch.set_num_threads(2)
model = IntentClassifier("Future-Labs/banking77-minilm")
print(model.predict("Where is my card?", top_k=3))
print(model.predict(["How can I change my PIN?", "I want to close my account."]))

predict always returns a list of per-message top-k results, each containing label and score. Read the short frozen_predict.py helper in this repository; it uses standard Transformers loading and explicit mean pooling/linear algebra. This artifact does not use transformers.pipeline(): the JSON head must be applied to the normalized mean-pooled embedding, as the helper demonstrates.

To download for offline use:

hf download Future-Labs/banking77-minilm --local-dir banking77-minilm
python frozen_predict.py ./banking77-minilm "Where is my card?"

All 77 label names and the exact head coefficients are in classifier.json. Softmax scores are uncalibrated. There is no out-of-domain rejection: unrelated inputs still receive banking labels. Input is truncated to 256 tokens.

Measured results and development history

Candidate Validation accuracy Test accuracy Test macro F1
Word TF-IDF + logistic regression 88.53% 88.47% Not recorded
BERT Tiny, full fine-tuning 86.00% 85.16% 84.71%
MiniLM-L12, initial full fine-tuning schedule 85.27% 83.12% 81.19%
MiniLM-L12, higher classifier LR 86.87% Not evaluated Not evaluated
Frozen all-MiniLM-L6-v2 + linear head (this release) 92.53% 92.76% 92.74%

The first three neural candidates failed our predeclared quality criteria and were not released. The latter two candidates required at least 90.53% validation accuracy before test evaluation. No checkpoint or regularization parameter was selected on test accuracy. Several earlier test results were seen during project development, so the benchmark is not an untouched holdout for the entire project. We report the failed runs to make this limitation clear. Only one seed was used.

Data protocol: normalize case/whitespace to remove duplicate training texts, exclude conflicting normalized labels (zero found), and stratify 9,999 unique training examples into 8,499 training / 1,500 validation with random seed 42. Official test data are unchanged (40 examples per class). Exact indices and file checksums are included. The same split was used for every candidate and baseline.

Seven test texts overlap the gradient/head-training split after normalization. Accuracy on the 3,073 nonoverlapping examples is 92.74%. We report only exact normalized text overlap, not a semantic near-duplicate audit or upstream encoder pretraining-contamination audit.

Lowest class F1 scores include pending_transfer (78.26%), balance_not_updated_after_bank_transfer (79.45%), and transfer_not_received_by_recipient (79.52%). Review the full class report; confusing closely related transfer intents remains a meaningful limitation.

A short ambiguous message such as "Where is my card?" scored 38.71% for lost_or_stolen_card and 37.26% for card_arrival in the manual check. Applications should request context or route to a human rather than treating such close scores as a reliable decision. These scores are not calibrated.

CPU artifact verification reloaded the saved encoder/JSON head and reproduced all 3,080/3,080 original test predictions. The verification also exercises empty/long inputs and ordinary/unrelated messages. Details and an illustrative shared-node latency measurement are in artifact_verification.json; latency is machine-specific and is not a deployment promise.

Reproduce

The encoder is unchanged from sentence-transformers/all-MiniLM-L6-v2 at revision 1110a243fdf4706b3f48f1d95db1a4f5529b4d41. Pool token states using the attention-mask-weighted mean and L2 normalize. Train scikit-learn multinomial logistic regression on those embeddings with C in [0.1, 1, 10, 100], max_iter=2000; validation selects C=10. The plain-JSON coefficients exactly reproduced sklearn validation predictions. No encoder gradients or contrastive fine-tuning were used for this release.

# In an isolated CUDA environment matching environment.json:
pip install "torch==2.14.0" "transformers==5.17.0" "scikit-learn==1.7.2" "numpy==2.2.6"
export OUTPUT_DIR=/your/output/directory
python source/intent77/train_frozen.py
python verify_frozen.py "$OUTPUT_DIR"

Full source, pinned configuration, package freeze, split indices, checksums, validation scores, class reports, and test predictions accompany the artifact. Prior fine-tuning code/configurations/results are in prior_experiments/. All training and full-dataset evaluation ran on budget-gated cluster compute. Training job: b278095f-742a-4777-83d3-ba66f3755cdc.

Intended use and limitations

English, single-intent support routing and reproducible intent-classification experiments. This model does not answer questions or perform banking actions. Evaluate on representative, consented local data and provide a human fallback.

  • No out-of-domain class or calibrated confidence threshold.
  • Not evaluated for multilingual, adversarial, multi-intent, or current bank traffic.
  • The benchmark is older and narrow; scores do not establish production readiness.
  • Not a fraud detector, authorization mechanism, financial adviser, or autonomous transaction system. Do not use a label as permission to execute an action.

Attribution and licenses

Encoder: sentence-transformers/all-MiniLM-L6-v2, Apache 2.0. The Apache 2.0 license text (LICENSE) and original model card (ENCODER_README.md) are included; encoder weights are unchanged. Our head and helper code are MIT (HEAD_CODE_LICENSE.txt). Bundle metadata reflects the Apache-licensed encoder.

BANKING77 data: PolyAI, revision 57ec275d8078af65b7731c2a98be812d844a6d6b, CC BY 4.0 (DATA_LICENSE.txt). Training texts were deduplicated; test examples unchanged.

Casanueva, I., Temčinas, T., Gerz, D., Henderson, M., and Vulić, I. (2020). Efficient Intent Detection with Dual Sentence Encoders. Sentence-BERT: Reimers and Gurevych (2019), Sentence Embeddings using Siamese BERT-Networks.

Published by Future-Labs. See SHA256SUMS.json for file integrity checks.

Browser demo and compact ONNX export

Try the live browser demo. It runs in a web worker on your device, with no inference server or API key. Your message is not submitted for server-side classification. The first use downloads a 45.4 MB model plus runtime libraries; your browser may cache files.

The ONNX export stores large matrices in FP16 and casts them back to FP32 for computation. It retains all 3,080 original test labels in CPU ONNX evaluation, with unchanged validation/test accuracy. Runtime memory is not necessarily reduced by half. A smaller int8 candidate failed our fidelity gate and was rejected; its measurements are disclosed in the export evaluation.

Real Chromium verification covered normal, empty, ambiguous, long, and unrelated messages; mobile layout; eight tokenizer/numerical parity fixtures; and network requests with no message payload. See browser verification. This is a browser smoke check; full benchmark parity was measured in the CPU ONNX runtime. Browser speed and compatibility can vary by device.

The Space source is plain HTML/CSS/JavaScript, pinned to the evaluated ONNX artifact and library versions. The earlier Gradio source remains as an optional local example; Gradio startup was not validated and the live demo uses the static app.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Future-Labs/banking77-minilm

Dataset used to train Future-Labs/banking77-minilm

Space using Future-Labs/banking77-minilm 1

Papers for Future-Labs/banking77-minilm

Evaluation results