BeeNara 🐝
Put a document in the right folder, or admit that none fits.
BeeNara reads a document together with your own list of folder names and returns one of them, or "none fits". Its confidence is calibrated. When it is not sure enough, it tells you to ask the user instead of guessing. It is the category decider of PollySort, a local, bilingual document archivist.
Key points
- Any folder names, no retraining. Folder names are plain text, like "Rechnungen", "Tax 2025" or "Bank statements". You can add a short description to each.
- Knows when nothing fits. It recognizes 96.8 % of documents whose folder is missing from the list. Small local LLMs rarely do this: Qwen3.5-4B managed it in 1 % of cases and Qwen3.5-9B in 44 %.
- Sorts only when it is sure. Split-conformal prediction decides when to sort automatically. On the benchmark it sorts 64 % of documents on its own, and 99.6 % of those are correct. The rest go to the user.
- Small and fast on a laptop CPU. It is a 332 MB ONNX file, about 0.2–0.3 s per document, with no GPU and no PyTorch.
- German and English. Documents and folder names can be in either language, or mixed.
- Private by design. It was trained only on synthetic documents, with no user data. It runs fully offline.
What it is good for, and what not
| Good for | Not for |
|---|---|
| Sorting letters, invoices, contracts and statements into a user's folders | Decisions with legal, medical or financial consequences |
| Deciding "choose one of these, or none" over short, changing label lists | Fixed large taxonomies with hundreds of classes |
| Laptops and offline tools where a local LLM is too slow | Languages other than German and English |
| Pipelines that need an "ask the human" signal | Extracting fields such as amounts or dates |
Quickstart
pip install onnxruntime tokenizers numpy huggingface_hub
hf download Kwokou/BeeNara beenara.py --local-dir .
from beenara import BeeNara
bn = BeeNara.from_pretrained("Kwokou/BeeNara") # downloads onnx/ and tokenizer/, 366 MB
d = bn.decide(
"Rechnung Nr. 4711\nBetrag: 119,00 EUR, zahlbar bis 30.10.2026.",
["Rechnungen", "Verträge", "Arztbriefe"],
descriptions={"Rechnungen": "Rechnungen von Lieferanten"}, # optional, helps with unusual names
)
d.category # 'Rechnungen', or None when the user should decide
d.ask # False: the prediction set holds exactly this one folder
d.distribution # e.g. {'Rechnungen': 0.97, 'Verträge': 0.001, 'Arztbriefe': 0.001, '__none__': 0.026}
bn.decide("Mietvertrag über die Wohnung …", ["Rechnungen", "Arztbriefe"]).ask # True, none fits
language="en" switches the internal question to English. The document language does not
matter for the categories: German documents may go into English folders, and the other way
round.
Results
Benchmark: 1,200 fixed decision tasks on 600 held-out documents, measured on the shipped ONNX file. Every document appears twice: once with training folder names, and once with folder names the model never saw, like "Rechnungskram" or "Payables". 25.8 % of the tasks have no fitting folder. laya-multilingual has no trained abstention output, so "none fits" was offered to it as an extra option.
| Metric | BeeNara v0.2 | BeeNara v0.1 | laya-multilingual |
|---|---|---|---|
| Accuracy, folders and "none fits" | 94.0 % | 87.5 % | 55.8 % |
| Accuracy on unseen folder names | 88.0 % | 75.0 % | – |
| "None fits" recall | 96.8 % | 98.4 % | 21.7 % |
| "None fits" precision | 84.2 % | 69.1 % | – |
| Wrongly said "none fits" | 6.3 % | 15.3 % | 5.8 % |
| Calibration error (ECE) | 0.030 | 0.065 | 0.262, uncalibrated |
| Sorted automatically, of which correct | 64.1 %, 99.6 % | 60.9 %, 98.8 % | – |
| CPU latency per document, 4 threads, median | 0.2–0.28 s | 0.19 s | 0.34 s |
The 95 % bootstrap interval for the v0.2 accuracy is 92.7–95.3 %. The ONNX file matches the PyTorch weights within 0.3 points on every metric and makes the same choice on all 64 parity-checked tasks. It was calibrated on itself after quantization. Latency was 205 ms on mains power, measured on v0.1 with the same architecture and file type, and 279 ms for v0.2 on battery.
| Slice, v0.2 | Accuracy | Wrongly said "none fits" |
|---|---|---|
| Training folder names | 100 % | 0 % |
| Unseen folder names | 88.0 % | 12.8 % |
| With a short description | 100 % | 0 % |
| Name only | 88.6 % | 12.2 % |
| Folder names in the other language | 91.8 % | 8.5 % |
Independent test, 70 documents from a different generator (PollySort's test projects), with 8 folders, measured on the PyTorch weights. In C2 the right folder is missing. C3 uses English names in shuffled order.
| Model | C1 German | C2 "none fits" | C3 English | Time per document |
|---|---|---|---|---|
| BeeNara v0.2 | 100 % | 94.3 % | 100 % | 0.26 s, CPU |
| Qwen3.5-9B via Ollama | 99 % | 44 % | 100 % | 0.51 s on GPU, 9.7 s on CPU |
| Qwen3.5-4B via Ollama | 91 % | 1 % | 100 % | 0.43 s on GPU, 5.4 s on CPU |
| laya-multilingual, with descriptions | 92.9 % | 0 % | 64.3 % | 0.48 s, CPU |
How it works
BeeNara is a cross-encoder with option markers. The question, every folder name and the document go into one sequence and through one forward pass:
[CLS] question [SEP] [MASK] folder 1 [MASK] folder 2 … [SEP] document [SEP]
- Each folder is scored at its own
[MASK]. - A separate head reads
[CLS]plus the shape of the folder scores (entropy, top-2 gap, count) and gives p("none fits"). - Both combine into one distribution:
p(none) = σ(e) p(folder i) = (1 − p(none)) · softmax(s)ᵢ
Decision rule. Temperature scaling calibrates the probabilities. Split-conformal prediction at α = 0.1 then builds a prediction set. If the set holds exactly one folder, the document is sorted. If it holds several folders, none, or "none fits", the user is asked.
| Part | Choice |
|---|---|
| Encoder | mmBERT-base, a multilingual ModernBERT: 22 layers, 768 dims, 256k vocabulary |
| Head | one Transformer layer, a folder scorer and a "none" scorer |
| Input | up to 512 tokens; 192 for question and folders, the rest for the document start |
| Shipped file | ONNX, weights in 8 bit (MatMulNBits, block 32), embeddings in INT8, 332 MB |
Training
- Supervised, not reinforcement learning. laya trains with RLCD, a REINFORCE method against a proper scoring rule. The log score already is such a rule, and our labels are known exactly. So we use plain cross-entropy for the folders and binary cross-entropy for "none", with label smoothing 0.05. Calibration follows in a separate step.
- Episodes instead of fixed classes. Every step draws a new task for each document:
- 3 to 9 folders in random order, with hard distractors such as "reminder" next to "invoice";
- umbrella terms such as "Finance";
- names with or without a description;
- in 25 % of tasks the right folder is removed, and the target becomes "none fits".
- Many ways to name a folder. This is what v0.2 added, and it is where the accuracy on
unseen names came from, 75 % → 88 %:
- spelling variants such as
rechnungen_2025,KONTOAUSZUEGEand03 Spesen; - 319 extra names from a local teacher model;
- description-only options;
- 20 % English mail-triage tasks from Open-Jev.
- spelling variants such as
- Kept apart by tests. Three disjoint sets of folder names are used: one for training, one for checkpoint selection and calibration, and one only for the benchmark.
- Setup:
- 28,524 synthetic documents in 20 business types, German, English and mixed, with simulated OCR noise;
- 3 epochs, 5,346 steps at batch 16;
- AdamW with encoder at 2.5e-5 and head at 1e-4, 5 % warmup, then cosine;
- bf16, frozen embeddings, 118M trainable parameters.
- Hardware: one laptop GPU, NVIDIA RTX PRO 1000 Blackwell with 8 GB. 34 minutes, peak 5.8 GB.
Environmental impact
We did not have a power meter. The numbers below are an upper-bound estimate following Lacoste et al. 2019: logged GPU time × power × grid intensity.
| Compute | Energy, at most | CO₂eq, at most | |
|---|---|---|---|
| Shipped v0.2 run | 34 min GPU + 6 min CPU calibration | 0.07 kWh | 25 g |
| All training runs of the project | about 2.1 h GPU: v0.1, a run lost to a crash, v0.2 and one comparison run | 0.21 kWh | 80 g |
| Evaluation, teacher names, quantization experiments | roughly 6 h laptop CPU | 0.35 kWh | 130 g |
Assumptions:
- Power: at most 100 W for the whole laptop during GPU training. The GPU is capped at 45–60 W. During CPU work, at most 60 W.
- Grid: 0.38 kg CO₂eq/kWh. The German electricity mix in 2023 and 2024 was about 0.36–0.38 kg/kWh (Umweltbundesamt).
- Not included: the pretraining of mmBERT and of the teacher model, which are far larger.
In total, all of this is at most about 0.2 kg CO₂eq, roughly one to two kilometres by car.
Credits and inspiration
BeeNara stands on published work. The implementation is our own.
| Source | What BeeNara takes from it |
|---|---|
| mmBERT (paper), JHU CLSP, MIT | the encoder weights: every weight in BeeNara starts from it |
| ModernBERT, Warner et al. 2024 | the encoder architecture behind mmBERT |
| laya, convaiinnovations, Apache-2.0 | the idea of one [MASK] marker per option, scored in one pass; the "choose or abstain" framing; the fine-tuning learning rates. We replaced its RLCD training with supervised learning |
| GLiClass | all labels in one sequence, one forward pass |
| Guo et al. 2017 | temperature scaling |
| Desai & Durrett 2020 | evidence that fine-tuned Transformers calibrate well after temperature scaling |
| Angelopoulos & Bates 2021 | split-conformal prediction and the "ask when unsure" rule |
| Szegedy et al. 2016, Müller et al. 2019 | label smoothing, and its known side effects |
| Open-Jev, ZefanCai, CC0-1.0 | about 20 % of the training tasks, from mailroom-control-v1. The customer-support configurations were excluded |
| Qwen3.5-9B, Qwen team, Apache-2.0 | proposed extra folder names, run locally and filtered against every held-out name |
| ONNX Runtime | weight-only 8-bit quantization (MatMulNBits) |
| Lacoste et al. 2019 | the method for the CO₂ estimate |
Limitations
- Unseen names without a description. Here BeeNara still says "none fits" in about 12 % of cases where a folder fits. In PollySort that means asking instead of sorting: safe, but noisy. A short description removes this on the benchmark.
- Missed "none fits". On the independent test, 4 of 70 documents whose folder was missing were sorted into a wrong one when only names were given. With descriptions, none were.
- Coverage below target. The conformal sets reach 86–87 % coverage instead of the targeted 90 %, because the calibration names are somewhat easier than truly new ones.
- Confidence ranks errors only moderately. The AUROC for telling right answers from wrong ones is 0.75. It was 0.89 in v0.1, and label smoothing is the likely cause. The sort-or-ask rule is not affected much.
- Synthetic training data. The documents come from templates with simulated noise. Real scans, letterheads and handwriting are messier, so expect a gap until the model is checked on real, consented documents.
- The benchmark shares the generator with the training data: separate seeds, but the same templates. The 70-document test is independent, but small.
- 20 document types in training. Unusual folders, such as "Garden", work only as far as the name generalization carries.
- German and English only.
- BeeNara supports a decision; it does not make one. Use it where a wrong folder can be undone.
Files
| File | Content | Size |
|---|---|---|
onnx/beenara.onnx |
the model | 332 MB |
onnx/calibration.json |
temperatures and conformal threshold, fitted on this file | < 1 KB |
tokenizer/ |
tokenizer of mmBERT-base | 34 MB |
beenara.py |
inference: input format, calibration, decision rule | 6 KB |
License
Apache-2.0, see LICENSE. The weights derive from mmBERT-base, which is MIT-licensed; see
NOTICE. Developed by Akinara as part of PollySort.
@misc{beenara2026,
title = {BeeNara: a calibrated, abstaining category decider for document sorting},
author = {Akinara},
year = {2026},
url = {https://huggingface.co/Kwokou/BeeNara}
}
Model tree for Kwokou/BeeNara
Base model
jhu-clsp/mmBERT-base