airline-intent-fasttext

A 19-class intent classifier for airline customer-support messages written in code-mixed Bengali / Banglish (romanised Bengali) / English.

This is the recommended model for this domain. It outperforms the DistilBERT student trained on the same data by +17 points on the held-out human-written test set, while being 20x smaller and ~230x faster.

artifact size use
intent.ftz 6.0 MB ship this — quantized, negligible accuracy cost
intent.bin 81.5 MB full precision, ~1.7 pts better

Why fastText beats a transformer here

The corpus is transliterated Banglish, where the same word arrives spelled many ways — koto / kotoo / kt. fastText's character n-grams (minn=3, maxn=5) share weight across those spellings by construction. A subword-tokenised transformer has to learn the equivalences from data it does not have enough of.

Intents (19)

greeting, goodbye, thanks, flight_status, booking_new, booking_manage, checkin_boarding, baggage_policy, baggage_issue, refund_compensation, fare_payment, airport_info, special_assistance, loyalty_program, travel_documents, disruption, complaint, agent_request, out_of_scope

out_of_scope is the reject class — adjacent-service questions (train tickets, hotels, visa agencies) that the airline does not handle.

Usage

import fasttext

from huggingface_hub import hf_hub_download

path = hf_hub_download("Badhon/airline-intent-fasttext", "intent.ftz")
model = fasttext.load_model(path)

labels, probs = model.predict("amar flight ta koto tay chare janaben?", k=3)
# ('__label__flight_status', ...)  0.94

intent = labels[0].removeprefix("__label__")
confidence = probs[0]

Inputs should be lowercased and stripped of trailing newlines; fastText treats a newline as a document boundary.

Confidence gating. Route to a human below ~0.6. Two known failure shapes sit above a 0.5 threshold and will pass it silently:

  • adjacent-service questions leak into booking_new with high confidence (apnara ki train er ticket o katenbooking_new 0.88, gold out_of_scope)
  • bare social tokens are low-confidence (assalamu alaikumgreeting 0.56)

out_of_scope recall on the synthetic split is 0.358. Do not rely on the reject class alone to catch off-domain traffic.

Evaluation

Two evaluation sets. Rank on h-test — the hand-written adversarial holdout, unseen during any tuning. The synthetic test split is generated from the same templates as training and overstates every model.

model test h-dev h-test p50 size
fasttext (quantized) 0.772 0.869 0.729 0.03 ms 6.0 MB
fasttext (full) 0.774 0.746 0.02 ms 81.5 MB
distilbert-int8 (student) 0.846 0.656 0.559 6.95 ms 121.6 MB

Full 120-item holdout accuracy 0.800, macro-F1 0.798.

Note DistilBERT scores higher on the synthetic split (0.846) and far lower on real phrasing (0.559). That divergence is the whole argument for ranking on h-test.

Intended use

Front-line intent routing for an airline support bot or IVR, with a confidence threshold and a human fallback. Not a decision system — it selects a reply template or a queue, nothing more.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support