Alben

An open, self-hosted decision model. Give it a piece of text and a set of typed questions, and it answers them — a label, a rating on an ordered scale, or a probability. No API key, no per-token billing, no data leaving your machine.

Alben answers questions you define at request time. You are not restricted to a fixed taxonomy: pass any set of options, with a description of what each one means, and the model scores the text against them. Up to 255 options, or an ordered scale of 2–10 levels, or a yes/no.

It speaks the same HTTP protocol as TypeSafe's System One API, so existing clients — including the official Python and JavaScript SDKs — work against it unchanged. That is a compatibility property, deliberately maintained; it is not the point of the project. The point is a small, inspectable, self-hostable model that stays useful when the question is one nobody anticipated.

pip install alben
from alben import Alben

a = Alben()                    # fetches this checkpoint on first use

answer = a.predict(
    "We were billed twice for March. Refund today or we cancel.",
    {
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this?",
            "criteria": {"billing": "invoices and refunds",
                         "tech": "bugs and outages"},
        },
        "urgency": {
            "type": "score",
            "instructions": "How urgent is this?",
            "criteria": ["can wait", "this week", "today", "immediately"],
        },
        "churn_risk": {
            "type": "noul",
            "instructions": "Did they threaten to cancel?",
        },
    },
)

print(answer["answers"]["team"]["choice"])        # billing
print(answer["answers"]["urgency"]["score"])       # 1.35 — between two levels
print(answer["answers"]["churn_risk"]["noul"])     # 0.38

The response is the same shape the System One API returns (model / answers / usage, probabilities rounded to two decimals) — deliberately, so the official SDKs parse it without changes.


What is in this checkpoint

bge-small-zh-v1.5-based sentence encoder with a trained projection head, fine-tuned end-to-end on questions the model had to answer from the option descriptions alone.

Base encoder bge-small-zh-v1.5 — BAAI, MIT. Fetched at runtime, not redistributed here
Training examples 56160
Distinct question sets 1882 — each with its own label names and wording
Answer types choice (≤255 options) · ordered score (2–10 levels) · yes/no
Calibration temperature scaling, fitted on a text-disjoint held-out set
Footprint ~92 MB checkpoint + ~92 MB encoder, runs on CPU or CUDA
License Apache-2.0 (this checkpoint)

Training data

Assembled from the train splits of public research datasets — MASSIVE (10 languages), BANKING77, AG News, DAIR Emotion, SST-5 — plus synthetic question sets generated for this project. Evaluation splits were held out and decontaminated; 3 upstream duplicates between SST-5 train and test were removed from the training side.


Measured accuracy

Public benchmarks, test splits only:

Benchmark Alben Random Release gate
MASSIVE intent, English (60 intents) 0.6975 0.0303 ≥ 0.60
MASSIVE intent, 9 other languages 0.3883 0.0303 ≥ 0.30
BANKING77 (77 classes) 0.8817 0.0500 ≥ 0.70
SST-5 (5-level ordered) 0.3133 0.2000 ≥ 0.25

Per language, MASSIVE intent — the spread is worth seeing before you deploy:

Language MASSIVE intent
en-US 0.6975
zh-CN 0.8025
de-DE 0.4600
fr-FR 0.4550
es-ES 0.5050
ja-JP 0.6350
ko-KR 0.1475
ar-SA 0.2125
hi-IN 0.1125
th-TH 0.1650

On held-out question sets from the same task family (0% text overlap with training, label-balanced, real question wording): 0.6465 average, 0.6178 worst.

Every gate above is checked by a script before release — python -u release.py verify, which exits non-zero on failure. The gates were fixed before the model was trained.


Calibrated confidence

Confidence is temperature-scaled on a held-out set that shares no text with training. Accuracy is unchanged by calibration; only the confidence moves.

Accuracy Mean confidence ECE ↓
Raw 0.7019 0.8167 0.1273
Calibrated (T = 1.65) 0.7019 0.7089 0.0508

Expected calibration error falls by 60%. A confidence threshold is therefore meaningful for routing: act above it, escalate below it.


Honest limitations

We would rather you know these before you build on it.

Scenario Works? Measured
Questions whose labels resemble what it trained on Yes 0.62–0.88 on unseen text
Same task, different language Yes, unevenly 0.70 (en) down to 0.11 (hi) across 10 languages
A genuinely new question type No — near chance 0.3150 vs 0.3211 random
New question type + a few dozen labelled examples Yes 0.4917 → 0.9917

In one line: it generalises across text and language, not across question types.

We checked the second row rather than assuming it. Splitting BANKING77's 77 labels in half — train on 38, evaluate on the other 39 with zero training examples — gives 0.5208, against 0.4725 for the frozen encoder and 0.9158 for a model that saw all 77. So most of the headline accuracy comes from having seen that label set, not from generalising to a new one. That is why the next section exists.

Teaching a new question type

This is a first-class operation, not a workaround:

python -m alben.teach --examples my_labelled_rows.jsonl --out my-dim.pt
a.teach(examples=my_labelled_rows, out="my-dim.pt", epochs=12)
# -> {"before": 0.513, "after": 0.663, "forgetting_delta": -0.020}

A few dozen examples are enough. Two properties matter, and both are measured:

  • Reported honestly — teach returns before, after, and a forgetting probe computed on old data that was not replayed into training. Reporting only after would make any model look good.
  • Local gains — untrained control questions move by +0.0000. Teaching one question type does not teach another. Teach the type you need.

Reproduce every number

git clone https://github.com/XiaoBinGan/Alben && cd Alben
pip install -e ".[dev]"

python -u release.py verify     # the six acceptance gates, exits non-zero on failure
python -u release.py status     # what is deployed right now
python -u corpus/am_i_general.py  # the three senses of "general", with evidence

Attribution

  • This checkpoint: Apache-2.0.
  • Base encoder: BAAI/bge-small-zh-v1.5 (Beijing Academy of Artificial Intelligence), MIT. Fetched at runtime; not redistributed in this repository.
  • Training text: MASSIVE (CC BY 4.0), BANKING77 (CC BY 4.0), AG News, DAIR Emotion, SST-5 — train splits only, under their respective licenses.
  • Protocol compatibility: Alben implements the publicly documented request and response shapes of the TypeSafe System One API so existing clients interoperate. It is an independent implementation: no TypeSafe weights, no TypeSafe training data, and no affiliation or endorsement. Outputs are not expected to match the hosted service.

中文说明

Alben 是一个可自托管的决策模型。给它一段文本和一组你自己定义的问题,它返回 经过校准的答案 —— 分类、有序打分、或是/否概率。不用 API key,不按 token 计费, 数据不出你的机器。

重点是「问题由你在运行时定义」:传任意一组选项、每个选项的描述,模型据此给 文本打分。最多 255 个选项,或 2–10 级有序打分,或是/否。

它说同一套 HTTP 协议,所以现成客户端(含官方 SDK)不用改就能用。兼容是刻意维护的 特性,但不是这个项目的重点。

项 值
训练样本 56160
不同问题集 1882
底模 bge-small-zh-v1.5(BAAI,MIT,不随本仓库分发)
许可 Apache-2.0

能力边界(必须一起读):它跨文本和语言泛化,不跨问题类型泛化。 遇到全新类型的问题时接近随机(0.3150 vs 随机 0.3211);但给几十条标注做热启动微调 就能学会(0.4917 → 0.9917),而且收益只限教过的那个类型(未教过的对照问题 +0.0000)。

完整数据、复现命令、以及我们明确排除掉的方向,见 GitHub 仓库。

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support