Alben
An open, self-hosted decision model. Give it a piece of text and a set of typed questions, and it answers them — a label, a rating on an ordered scale, or a probability. No API key, no per-token billing, no data leaving your machine.
Alben answers questions you define at request time. You are not restricted to a fixed taxonomy: pass any set of options, with a description of what each one means, and the model scores the text against them. Up to 255 options, or an ordered scale of 2–10 levels, or a yes/no.
It speaks the same HTTP protocol as TypeSafe's System One API, so existing clients — including the official Python and JavaScript SDKs — work against it unchanged. That is a compatibility property, deliberately maintained; it is not the point of the project. The point is a small, inspectable, self-hostable model that stays useful when the question is one nobody anticipated.
pip install alben
from alben import Alben
a = Alben() # fetches this checkpoint on first use
answer = a.predict(
"We were billed twice for March. Refund today or we cancel.",
{
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "invoices and refunds",
"tech": "bugs and outages"},
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["can wait", "this week", "today", "immediately"],
},
"churn_risk": {
"type": "noul",
"instructions": "Did they threaten to cancel?",
},
},
)
print(answer["answers"]["team"]["choice"]) # billing
print(answer["answers"]["urgency"]["score"]) # 1.35 — between two levels
print(answer["answers"]["churn_risk"]["noul"]) # 0.38
The response is the same shape the System One API returns (model / answers /
usage, probabilities rounded to two decimals) — deliberately, so the official SDKs
parse it without changes.
What is in this checkpoint
bge-small-zh-v1.5-based sentence encoder with a trained projection head, fine-tuned
end-to-end on questions the model had to answer from the option descriptions alone.
| Base encoder | bge-small-zh-v1.5 — BAAI, MIT. Fetched at runtime, not redistributed here |
| Training examples | 56160 |
| Distinct question sets | 1882 — each with its own label names and wording |
| Answer types | choice (≤255 options) · ordered score (2–10 levels) · yes/no |
| Calibration | temperature scaling, fitted on a text-disjoint held-out set |
| Footprint | ~92 MB checkpoint + ~92 MB encoder, runs on CPU or CUDA |
| License | Apache-2.0 (this checkpoint) |
Training data
Assembled from the train splits of public research datasets — MASSIVE (10 languages), BANKING77, AG News, DAIR Emotion, SST-5 — plus synthetic question sets generated for this project. Evaluation splits were held out and decontaminated; 3 upstream duplicates between SST-5 train and test were removed from the training side.
Measured accuracy
Public benchmarks, test splits only:
| Benchmark | Alben | Random | Release gate |
|---|---|---|---|
| MASSIVE intent, English (60 intents) | 0.6975 | 0.0303 | ≥ 0.60 |
| MASSIVE intent, 9 other languages | 0.3883 | 0.0303 | ≥ 0.30 |
| BANKING77 (77 classes) | 0.8817 | 0.0500 | ≥ 0.70 |
| SST-5 (5-level ordered) | 0.3133 | 0.2000 | ≥ 0.25 |
Per language, MASSIVE intent — the spread is worth seeing before you deploy:
| Language | MASSIVE intent |
|---|---|
en-US |
0.6975 |
zh-CN |
0.8025 |
de-DE |
0.4600 |
fr-FR |
0.4550 |
es-ES |
0.5050 |
ja-JP |
0.6350 |
ko-KR |
0.1475 |
ar-SA |
0.2125 |
hi-IN |
0.1125 |
th-TH |
0.1650 |
On held-out question sets from the same task family (0% text overlap with training, label-balanced, real question wording): 0.6465 average, 0.6178 worst.
Every gate above is checked by a script before release — python -u release.py verify,
which exits non-zero on failure. The gates were fixed before the model was trained.
Calibrated confidence
Confidence is temperature-scaled on a held-out set that shares no text with training. Accuracy is unchanged by calibration; only the confidence moves.
| Accuracy | Mean confidence | ECE ↓ | |
|---|---|---|---|
| Raw | 0.7019 | 0.8167 | 0.1273 |
| Calibrated (T = 1.65) | 0.7019 | 0.7089 | 0.0508 |
Expected calibration error falls by 60%. A confidence threshold is therefore meaningful for routing: act above it, escalate below it.
Honest limitations
We would rather you know these before you build on it.
| Scenario | Works? | Measured |
|---|---|---|
| Questions whose labels resemble what it trained on | Yes | 0.62–0.88 on unseen text |
| Same task, different language | Yes, unevenly | 0.70 (en) down to 0.11 (hi) across 10 languages |
| A genuinely new question type | No — near chance | 0.3150 vs 0.3211 random |
| New question type + a few dozen labelled examples | Yes | 0.4917 → 0.9917 |
In one line: it generalises across text and language, not across question types.
We checked the second row rather than assuming it. Splitting BANKING77's 77 labels in half — train on 38, evaluate on the other 39 with zero training examples — gives 0.5208, against 0.4725 for the frozen encoder and 0.9158 for a model that saw all 77. So most of the headline accuracy comes from having seen that label set, not from generalising to a new one. That is why the next section exists.
Teaching a new question type
This is a first-class operation, not a workaround:
python -m alben.teach --examples my_labelled_rows.jsonl --out my-dim.pt
a.teach(examples=my_labelled_rows, out="my-dim.pt", epochs=12)
# -> {"before": 0.513, "after": 0.663, "forgetting_delta": -0.020}
A few dozen examples are enough. Two properties matter, and both are measured:
- Reported honestly —
teachreturns before, after, and a forgetting probe computed on old data that was not replayed into training. Reporting only after would make any model look good. - Local gains — untrained control questions move by +0.0000. Teaching one question type does not teach another. Teach the type you need.
Reproduce every number
git clone https://github.com/XiaoBinGan/Alben && cd Alben
pip install -e ".[dev]"
python -u release.py verify # the six acceptance gates, exits non-zero on failure
python -u release.py status # what is deployed right now
python -u corpus/am_i_general.py # the three senses of "general", with evidence
Attribution
- This checkpoint: Apache-2.0.
- Base encoder:
BAAI/bge-small-zh-v1.5(Beijing Academy of Artificial Intelligence), MIT. Fetched at runtime; not redistributed in this repository. - Training text: MASSIVE (CC BY 4.0), BANKING77 (CC BY 4.0), AG News, DAIR Emotion, SST-5 — train splits only, under their respective licenses.
- Protocol compatibility: Alben implements the publicly documented request and response shapes of the TypeSafe System One API so existing clients interoperate. It is an independent implementation: no TypeSafe weights, no TypeSafe training data, and no affiliation or endorsement. Outputs are not expected to match the hosted service.
中文说明
Alben 是一个可自托管的决策模型。给它一段文本和一组你自己定义的问题,它返回 经过校准的答案 —— 分类、有序打分、或是/否概率。不用 API key,不按 token 计费, 数据不出你的机器。
重点是「问题由你在运行时定义」:传任意一组选项、每个选项的描述,模型据此给 文本打分。最多 255 个选项,或 2–10 级有序打分,或是/否。
它说同一套 HTTP 协议,所以现成客户端(含官方 SDK)不用改就能用。兼容是刻意维护的 特性,但不是这个项目的重点。
| 项 | 值 |
|---|---|
| 训练样本 | 56160 |
| 不同问题集 | 1882 |
| 底模 | bge-small-zh-v1.5(BAAI,MIT,不随本仓库分发) |
| 许可 | Apache-2.0 |
能力边界(必须一起读):它跨文本和语言泛化,不跨问题类型泛化。 遇到全新类型的问题时接近随机(0.3150 vs 随机 0.3211);但给几十条标注做热启动微调 就能学会(0.4917 → 0.9917),而且收益只限教过的那个类型(未教过的对照问题 +0.0000)。
完整数据、复现命令、以及我们明确排除掉的方向,见 GitHub 仓库。