Instructions to use willykeenan/banking-intent-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Scikit-learn
How to use willykeenan/banking-intent-classifier with Scikit-learn:
# ⚠️ Model filename not specified in config.json
- Notebooks
- Google Colab
- Kaggle
Banking Intent Classifier
A downloadable, CPU-friendly classifier for 77 English banking intents, trained by William Keenan as part of Waggle and Kea. It combines word unigram/bigram and character 3–5-gram TF-IDF features with a seeded, log-loss SGD classifier.
The package contains the fitted feature vocabulary, IDF values, classifier weights, class labels, inference example and reproduction code. This is a classical supervised model, not an LLM or a transformer fine-tune.
Results
Evaluation uses 3,050 unique, unambiguous test queries after excluding normalized train/test overlap. Training uses 9,971 filtered BANKING77 examples.
| Metric | Word + character model | Published word-only baseline |
|---|---|---|
| Macro-F1 | 0.9119 | 0.8915 |
| Accuracy | 0.9115 | 0.8915 |
| Top-3 accuracy | 0.9744 | 0.9708 |
| Log loss | 0.4434 | 0.5848 |
| Expected calibration error, 15 bins | 0.1467 | 0.2106 |
The model's fresh export check reproduces every reference top-1 label and top-3 probability rounded to parts per million, then checks identical probabilities after saving and reloading all 3,050 cases. Exact results and the artifact hash are in verification.json. The baseline is a historical reference; this package does not retrain or ship that baseline.
These results are exploratory. The original design was informed by aggregate test performance. This is not independent replication, a confirmatory experiment, or a claim of state-of-the-art performance. The official test split has 3,080 rows; these scores apply to the documented filtered 3,050-row population and should not be compared directly with unfiltered leaderboard scores.
Use
Use Python 3.12 and the supplied dependency versions:
python -m pip install -r requirements.txt
from huggingface_hub import hf_hub_download
import skops.io as sio
path = hf_hub_download("willykeenan/banking-intent-classifier", "model.skops")
unknown_types = sio.get_untrusted_types(file=path)
if unknown_types:
raise RuntimeError(f"Unexpected model types: {unknown_types}")
model = sio.load(path, trusted=[])
queries = ["My card has not arrived yet", "How can I change my PIN?"]
labels = model.predict(queries)
probabilities = model.predict_proba(queries)
print(labels)
The probability columns follow labels.json, which is identical to model.classes_. predict.py also provides a local command-line example. A hosted inference endpoint or Transformers pipeline() integration is not included.
Reproduce
Download this repository, install requirements.txt, and run:
python reproduce.py --cache ./source-cache --output ./reproduced
The command downloads only the three public source files named in SOURCE.json, verifies their sizes and SHA-256 hashes, applies the original filtering, fits the frozen canonical seed 20260805, checks every reference prediction, and saves model.skops. Use --offline after populating the source cache. Use a new output folder for each run.
The training recipe is preserved from public commit 54041045c82953cbbabac155b150a0f7b7ed5603. PROTOCOL.md describes the filtering, matched comparison, controls and original experiment. The export script refits only this model and verifies the export; it does not rerun the original multi-seed, bootstrap or handoff experiments.
Agent evaluation companion
The companion Agent Handoff Benchmark packages the public frozen predictions, full probability vectors and 36,600 decisions under 12 synthetic cost policies. It supports research into whether an agent receives sufficient information to preserve a routing decision. The model's classification quality and the handoff method's decision agreement are separate questions.
Intended use and limitations
- Suitable for reproducible intent-classification experiments, teaching, baselines and agent handoff research.
- English only; no demonstrated generalization to other languages, current bank products, natural customer noise or out-of-distribution inputs.
- Every input receives one of 77 labels. There is no trained unknown-intent detector or built-in abstention mechanism. Probabilities are not reliably calibrated risk estimates.
- Not validated for customer-facing financial decisions or autonomous banking actions. No private bank, employer or customer data were used.
- TF-IDF vocabularies contain features learned from the public source text. The model is not a privacy-preserving transformation of that source.
Data and licensing
William Keenan's model package and code are released under the MIT license. BANKING77 remains PolyAI's dataset, released under CC BY 4.0; William's contribution is the fitted classifier, filtering/evaluation and packaging. Preserve THIRD_PARTY_NOTICES.md and upstream attribution when using or redistributing the package. No raw training or test query text is included in this repository.
Upstream citation: Inigo Casanueva, Tadas Temcinas, Daniela Gerz, Matthew Henderson and Ivan Vulic. 2020. Efficient Intent Detection with Dual Sentence Encoders. Proceedings of the 2nd Workshop on NLP for ConvAI. Original source.
- Downloads last month
- -