BeeNara 🐝

Put a document in the right folder, or admit that none fits.

BeeNara reads a document together with your own list of folder names and returns one of them, or "none fits". Its confidence is calibrated. When it is not sure enough, it tells you to ask the user instead of guessing. It is the category decider of PollySort, a local, bilingual document archivist.

Key points

  • Any folder names, no retraining. Folder names are plain text, like "Rechnungen", "Tax 2025" or "Bank statements". You can add a short description to each.
  • Knows when nothing fits. It recognizes 96.8 % of documents whose folder is missing from the list. Small local LLMs rarely do this: Qwen3.5-4B managed it in 1 % of cases and Qwen3.5-9B in 44 %.
  • Sorts only when it is sure. Split-conformal prediction decides when to sort automatically. On the benchmark it sorts 64 % of documents on its own, and 99.6 % of those are correct. The rest go to the user.
  • Small and fast on a laptop CPU. It is a 332 MB ONNX file, about 0.2–0.3 s per document, with no GPU and no PyTorch.
  • German and English. Documents and folder names can be in either language, or mixed.
  • Private by design. It was trained only on synthetic documents, with no user data. It runs fully offline.

What it is good for, and what not

Good for Not for
Sorting letters, invoices, contracts and statements into a user's folders Decisions with legal, medical or financial consequences
Deciding "choose one of these, or none" over short, changing label lists Fixed large taxonomies with hundreds of classes
Laptops and offline tools where a local LLM is too slow Languages other than German and English
Pipelines that need an "ask the human" signal Extracting fields such as amounts or dates

Quickstart

pip install onnxruntime tokenizers numpy huggingface_hub
hf download Kwokou/BeeNara beenara.py --local-dir .
from beenara import BeeNara

bn = BeeNara.from_pretrained("Kwokou/BeeNara")          # downloads onnx/ and tokenizer/, 366 MB

d = bn.decide(
    "Rechnung Nr. 4711\nBetrag: 119,00 EUR, zahlbar bis 30.10.2026.",
    ["Rechnungen", "Verträge", "Arztbriefe"],
    descriptions={"Rechnungen": "Rechnungen von Lieferanten"},   # optional, helps with unusual names
)
d.category        # 'Rechnungen', or None when the user should decide
d.ask             # False: the prediction set holds exactly this one folder
d.distribution    # e.g. {'Rechnungen': 0.97, 'Verträge': 0.001, 'Arztbriefe': 0.001, '__none__': 0.026}

bn.decide("Mietvertrag über die Wohnung …", ["Rechnungen", "Arztbriefe"]).ask   # True, none fits

language="en" switches the internal question to English. The document language does not matter for the categories: German documents may go into English folders, and the other way round.

Results

Benchmark: 1,200 fixed decision tasks on 600 held-out documents, measured on the shipped ONNX file. Every document appears twice: once with training folder names, and once with folder names the model never saw, like "Rechnungskram" or "Payables". 25.8 % of the tasks have no fitting folder. laya-multilingual has no trained abstention output, so "none fits" was offered to it as an extra option.

Metric BeeNara v0.2 BeeNara v0.1 laya-multilingual
Accuracy, folders and "none fits" 94.0 % 87.5 % 55.8 %
Accuracy on unseen folder names 88.0 % 75.0 % –
"None fits" recall 96.8 % 98.4 % 21.7 %
"None fits" precision 84.2 % 69.1 % –
Wrongly said "none fits" 6.3 % 15.3 % 5.8 %
Calibration error (ECE) 0.030 0.065 0.262, uncalibrated
Sorted automatically, of which correct 64.1 %, 99.6 % 60.9 %, 98.8 % –
CPU latency per document, 4 threads, median 0.2–0.28 s 0.19 s 0.34 s

The 95 % bootstrap interval for the v0.2 accuracy is 92.7–95.3 %. The ONNX file matches the PyTorch weights within 0.3 points on every metric and makes the same choice on all 64 parity-checked tasks. It was calibrated on itself after quantization. Latency was 205 ms on mains power, measured on v0.1 with the same architecture and file type, and 279 ms for v0.2 on battery.

Slice, v0.2 Accuracy Wrongly said "none fits"
Training folder names 100 % 0 %
Unseen folder names 88.0 % 12.8 %
With a short description 100 % 0 %
Name only 88.6 % 12.2 %
Folder names in the other language 91.8 % 8.5 %

Independent test, 70 documents from a different generator (PollySort's test projects), with 8 folders, measured on the PyTorch weights. In C2 the right folder is missing. C3 uses English names in shuffled order.

Model C1 German C2 "none fits" C3 English Time per document
BeeNara v0.2 100 % 94.3 % 100 % 0.26 s, CPU
Qwen3.5-9B via Ollama 99 % 44 % 100 % 0.51 s on GPU, 9.7 s on CPU
Qwen3.5-4B via Ollama 91 % 1 % 100 % 0.43 s on GPU, 5.4 s on CPU
laya-multilingual, with descriptions 92.9 % 0 % 64.3 % 0.48 s, CPU

How it works

BeeNara is a cross-encoder with option markers. The question, every folder name and the document go into one sequence and through one forward pass:

[CLS] question [SEP] [MASK] folder 1 [MASK] folder 2 … [SEP] document [SEP]
  • Each folder is scored at its own [MASK].
  • A separate head reads [CLS] plus the shape of the folder scores (entropy, top-2 gap, count) and gives p("none fits").
  • Both combine into one distribution:
p(none) = σ(e)        p(folder i) = (1 − p(none)) · softmax(s)ᵢ

Decision rule. Temperature scaling calibrates the probabilities. Split-conformal prediction at α = 0.1 then builds a prediction set. If the set holds exactly one folder, the document is sorted. If it holds several folders, none, or "none fits", the user is asked.

Part Choice
Encoder mmBERT-base, a multilingual ModernBERT: 22 layers, 768 dims, 256k vocabulary
Head one Transformer layer, a folder scorer and a "none" scorer
Input up to 512 tokens; 192 for question and folders, the rest for the document start
Shipped file ONNX, weights in 8 bit (MatMulNBits, block 32), embeddings in INT8, 332 MB

Training

  • Supervised, not reinforcement learning. laya trains with RLCD, a REINFORCE method against a proper scoring rule. The log score already is such a rule, and our labels are known exactly. So we use plain cross-entropy for the folders and binary cross-entropy for "none", with label smoothing 0.05. Calibration follows in a separate step.
  • Episodes instead of fixed classes. Every step draws a new task for each document:
    • 3 to 9 folders in random order, with hard distractors such as "reminder" next to "invoice";
    • umbrella terms such as "Finance";
    • names with or without a description;
    • in 25 % of tasks the right folder is removed, and the target becomes "none fits".
  • Many ways to name a folder. This is what v0.2 added, and it is where the accuracy on unseen names came from, 75 % → 88 %:
    • spelling variants such as rechnungen_2025, KONTOAUSZUEGE and 03 Spesen;
    • 319 extra names from a local teacher model;
    • description-only options;
    • 20 % English mail-triage tasks from Open-Jev.
  • Kept apart by tests. Three disjoint sets of folder names are used: one for training, one for checkpoint selection and calibration, and one only for the benchmark.
  • Setup:
    • 28,524 synthetic documents in 20 business types, German, English and mixed, with simulated OCR noise;
    • 3 epochs, 5,346 steps at batch 16;
    • AdamW with encoder at 2.5e-5 and head at 1e-4, 5 % warmup, then cosine;
    • bf16, frozen embeddings, 118M trainable parameters.
  • Hardware: one laptop GPU, NVIDIA RTX PRO 1000 Blackwell with 8 GB. 34 minutes, peak 5.8 GB.

Environmental impact

We did not have a power meter. The numbers below are an upper-bound estimate following Lacoste et al. 2019: logged GPU time × power × grid intensity.

Compute Energy, at most CO₂eq, at most
Shipped v0.2 run 34 min GPU + 6 min CPU calibration 0.07 kWh 25 g
All training runs of the project about 2.1 h GPU: v0.1, a run lost to a crash, v0.2 and one comparison run 0.21 kWh 80 g
Evaluation, teacher names, quantization experiments roughly 6 h laptop CPU 0.35 kWh 130 g

Assumptions:

  • Power: at most 100 W for the whole laptop during GPU training. The GPU is capped at 45–60 W. During CPU work, at most 60 W.
  • Grid: 0.38 kg CO₂eq/kWh. The German electricity mix in 2023 and 2024 was about 0.36–0.38 kg/kWh (Umweltbundesamt).
  • Not included: the pretraining of mmBERT and of the teacher model, which are far larger.

In total, all of this is at most about 0.2 kg CO₂eq, roughly one to two kilometres by car.

Credits and inspiration

BeeNara stands on published work. The implementation is our own.

Source What BeeNara takes from it
mmBERT (paper), JHU CLSP, MIT the encoder weights: every weight in BeeNara starts from it
ModernBERT, Warner et al. 2024 the encoder architecture behind mmBERT
laya, convaiinnovations, Apache-2.0 the idea of one [MASK] marker per option, scored in one pass; the "choose or abstain" framing; the fine-tuning learning rates. We replaced its RLCD training with supervised learning
GLiClass all labels in one sequence, one forward pass
Guo et al. 2017 temperature scaling
Desai & Durrett 2020 evidence that fine-tuned Transformers calibrate well after temperature scaling
Angelopoulos & Bates 2021 split-conformal prediction and the "ask when unsure" rule
Szegedy et al. 2016, Müller et al. 2019 label smoothing, and its known side effects
Open-Jev, ZefanCai, CC0-1.0 about 20 % of the training tasks, from mailroom-control-v1. The customer-support configurations were excluded
Qwen3.5-9B, Qwen team, Apache-2.0 proposed extra folder names, run locally and filtered against every held-out name
ONNX Runtime weight-only 8-bit quantization (MatMulNBits)
Lacoste et al. 2019 the method for the CO₂ estimate

Limitations

  • Unseen names without a description. Here BeeNara still says "none fits" in about 12 % of cases where a folder fits. In PollySort that means asking instead of sorting: safe, but noisy. A short description removes this on the benchmark.
  • Missed "none fits". On the independent test, 4 of 70 documents whose folder was missing were sorted into a wrong one when only names were given. With descriptions, none were.
  • Coverage below target. The conformal sets reach 86–87 % coverage instead of the targeted 90 %, because the calibration names are somewhat easier than truly new ones.
  • Confidence ranks errors only moderately. The AUROC for telling right answers from wrong ones is 0.75. It was 0.89 in v0.1, and label smoothing is the likely cause. The sort-or-ask rule is not affected much.
  • Synthetic training data. The documents come from templates with simulated noise. Real scans, letterheads and handwriting are messier, so expect a gap until the model is checked on real, consented documents.
  • The benchmark shares the generator with the training data: separate seeds, but the same templates. The 70-document test is independent, but small.
  • 20 document types in training. Unusual folders, such as "Garden", work only as far as the name generalization carries.
  • German and English only.
  • BeeNara supports a decision; it does not make one. Use it where a wrong folder can be undone.

Files

File Content Size
onnx/beenara.onnx the model 332 MB
onnx/calibration.json temperatures and conformal threshold, fitted on this file < 1 KB
tokenizer/ tokenizer of mmBERT-base 34 MB
beenara.py inference: input format, calibration, decision rule 6 KB

License

Apache-2.0, see LICENSE. The weights derive from mmBERT-base, which is MIT-licensed; see NOTICE. Developed by Akinara as part of PollySort.

@misc{beenara2026,
  title  = {BeeNara: a calibrated, abstaining category decider for document sorting},
  author = {Akinara},
  year   = {2026},
  url    = {https://huggingface.co/Kwokou/BeeNara}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kwokou/BeeNara

Finetuned
(157)
this model

Dataset used to train Kwokou/BeeNara

Papers for Kwokou/BeeNara