BeeNara Big 🐝 v0.3 (preview)
Suggest the right folder in a whole folder tree, or say that none fits.
BeeNara Big reads a document (its text and, if available, the page image) together with the complete folder tree of a user (from 3 to several hundred folders) and returns a calibrated probability for every folder and for "none fits". It is the large sibling of BeeNara, which decides between at most 9 folders on a CPU.
Status: preview, suggestion mode. v0.3 ranks folders well, but it is not certified to sort automatically: the risk calibration found no confidence threshold that keeps automatic decisions at ≤ 1 % wrong (see Limitations). Use it to suggest folders and let the user confirm.
Quickstart
pip install torch "transformers>=5.17" peft safetensors tokenizers huggingface_hub pillow
hf download Kwokou/BeeNara-Big beenara_big.py --local-dir .
from beenara_big import BeeNaraBig
bb = BeeNaraBig.from_pretrained("Kwokou/BeeNara-Big") # adapter 164 MB + Qwen3.5-2B-Base 4.6 GB, GPU if available
s = bb.suggest(
"Einkommensteuerbescheid 2025\nFinanzamt Musterstadt\nFestgesetzte Einkommensteuer: 1.234,00 EUR",
["Finanzen/Steuer/2024", "Finanzen/Steuer/2025", "Finanzen/Bank", "Wohnung/Miete", "Arbeit/Verträge"],
)
s.suggestions # [('Finanzen/Steuer/2025', 0.91), ('Finanzen/Steuer/2024', ...), ...]
s.none # p("none of these folders fits"), here 0.07
s.ask # True: in v0.3 the user always confirms
- Folder names are free text; paths like
Finanzen/Steuer/2025work best. A description can follow the name:"Wohnung/Miete — Mietvertrag, Nebenkosten". images=[...]adds page images (file paths or PIL images);language="en"switches the internal prompt to English.- On a GPU the base model loads in bf16, on a CPU in fp32. The loader gives exactly the same scores as the training code (checked on the same inputs).
Key points
- Whole folder trees. One forward pass over up to ~300 folders (trained up to 500), with
folder paths such as
Finanzen/Steuer/2025. - Good where the small model cannot see enough. On held-out trees with 25–299 folders it finds the right folder in 79–88 % of the cases; the small BeeNara with a 9-folder pre-selection reached about 53 % on comparable large trees.
- Knows when nothing fits: 82–84 % of the documents whose folder is missing are recognized.
- Reads the page. The page image (letterhead, logo, layout) is used together with the text in half of the training tasks; it also works with text only.
- Calibrated. Expected calibration error 0.045.
- Same quality as the 4B model at half the size. A paired comparison with a Qwen3.5-4B version (same data, same seeds) found differences of at most ±1 point per list size.
Results
Held-out "world": folder trees, folder names and documents (written by gpt-oss-120b instead of the training teacher Qwen3.8-27B) that were never used for training. 2,344 tasks, 400 per list size (fewer for the largest trees), including tasks where no folder fits. Calibrated, PyTorch bf16. This repository ships seed 1.
| Folders in the tree | Tasks | Accuracy, seed 1 | Accuracy, seed 2 |
|---|---|---|---|
| 3–9 | 400 | 91.7 % | 91.2 % |
| 10–24 | 400 | 90.0 % | 88.7 % |
| 25–49 | 400 | 88.2 % | 87.5 % |
| 50–99 | 400 | 82.8 % | 82.3 % |
| 100–199 | 400 | 78.5 % | 78.2 % |
| 200–299 | 293 | 79.2 % | 77.8 % |
| 300 and more | 51 | 70.6 % | 60.8 % |
| All | 2,344 | 85.0 % | 84.1 % |
| Metric | Seed 1 | Seed 2 |
|---|---|---|
| "None fits" recall | 84.0 % | 82.3 % |
| "None fits" precision | 72.4 % | 70.7 % |
| Wrongly said "none fits" | 9.5 % | 10.1 % |
| Calibration error (ECE) | 0.045 | 0.045 |
| Decision changes when the folder order is shuffled | 2.3 % | 2.3 % |
2B against 4B (paired, two seeds each, same tasks): accuracy difference 4B − 2B of +0.6, −0.9, +0.4 and −0.9 points for 3–9, 10–24, 25–49 and 50–99 folders. The 2B model was chosen by the pre-registered rule "at most 2 points behind the 4B in every list size up to 99 folders".
Speed and memory.
| GPU memory, 8-bit weights, 300 folders | 4.2 GB |
| CPU, 4 threads, 7 folders, text only (337 tokens) | 3.0 s |
| CPU, 4 threads, 89 folders, text only (2,702 tokens) | 31 s |
| CPU, 4 threads, 35 folders, with page image (4,525 tokens) | 99 s |
| CPU, 4 threads, 375 folders (16,107 tokens) | 217 s |
On a CPU it is too slow for interactive use; plan for a GPU. GPU latency was not measured yet.
How it works
- Base: Qwen3.5-2B-Base (revision
b1485b2f), a hybrid model with linear and full attention and a vision encoder. - Input (variant "L2"):
[task] [page image] [document text] [folder list] [folder list with one marker per folder] [none marker]. The folder list appears twice so that every folder marker sees the document and the complete list. Markers are existing special tokens; no new vocabulary. - Head: a small decision head reads the hidden states at the markers and gives one score per folder and one for "none fits"; together they form one probability distribution.
- Limits: context up to 16,384 tokens; document text up to 2,048 tokens (start and end are kept); up to 32 tokens per folder name; page images up to about 1024 × 1448 pixels.
- Calibration: temperature scaling on 7,115 tasks from the held-out world (1,837 for the
smallest trees, 15 for trees above 300 folders); separate calibration files for bf16 and for
8-bit weights (
kalibrierung.json,kalibrierung_8bit.json), each bound to its weights by SHA-256.
Training
- Data. Only synthetic documents, no user data. 45,322 documents placed in 3,000 synthetic
folder trees of 10–500 folders (German, English and mixed; private households, freelancers,
small firms, construction offices, associations, law offices):
- 15,813 new documents written by Qwen3.8-27B for concrete leaf folders, each checked against the facts it was built from;
- 29,509 template documents assigned to trees by their metadata;
- page images rendered for all documents; about 20 % degraded (scan, phone photo, skew) and read with Tesseract for OCR text.
- Episodes. Every step draws a task per document: the whole tree or a part of it (3–300 folders, 5 % up to 500), with folder descriptions in 30 % and spelling variants in 20 % of the tasks; in 25 % the right folder is missing, half of those with a similar neighbour present; in 15 % only the parent of the right folder is in the list, and the parent is the answer.
- Kept apart by tests. Separate names and trees for training and for the held-out world; normalized and substring overlap checks, near-duplicate and 13-gram checks found no leakage.
- Setup: LoRA (r = 32, α = 64, dropout 0.05) on attention, linear attention and MLP of the language model; vision encoder frozen. Learning rate 1e-4 for LoRA and 3e-4 for the head, 3 % warmup, 2 epochs = 5,664 steps of 16 tasks, label smoothing 0.05, gradient checkpointing, bf16. Best checkpoint by validation NLL (step 5,000).
- Hardware: one NVIDIA RTX A6000 (48 GB), about 30 hours.
Limitations
- No automatic sorting in v0.3. The risk calibration (Learn-then-Test, ≤ 1 % wrong among automatic decisions, δ = 0.1) found no threshold for any list size. The likely cause is the label smoothing used in training, which compresses the confidence of the surest answers. Use the model for suggestions; v0.4 will be trained without label smoothing.
- Large trees are harder. Accuracy falls from about 92 % (3–9 folders) to about 79 % (100–299) and further above 300 folders, where the test set is small (51 tasks).
- "None fits" is less reliable than in the small model, especially in large trees.
- Slow on a CPU (seconds to minutes per document), heavy with a page image.
- Synthetic training data. Real archives are messier. Expect a gap until the model is checked on real, consented documents.
- German and English only.
Files
| File | Content | Size |
|---|---|---|
lora.safetensors |
LoRA weights for the language model | 125 MB |
kopf.safetensors |
decision head | 19 MB |
big_config.json |
base model and revision, LoRA settings, input limits, variant | < 1 KB |
kalibrierung.json |
calibration for bf16 weights | 6 KB |
kalibrierung_8bit.json |
calibration for 8-bit weights | 6 KB |
tokenizer/ |
Qwen3.5 tokenizer | 20 MB |
bildprozessor/ |
image preprocessing settings | < 1 KB |
beenara_big.py |
loader and inference: input format, head, calibration | 14 KB |
The base model weights are not included; they come from
Qwen/Qwen3.5-2B-Base at the revision in
big_config.json.
Credits
| Source | What BeeNara Big takes from it |
|---|---|
| Qwen3.5-2B-Base, Qwen team, Apache-2.0 | the base model |
| LoRA, Hu et al. 2021 | low-rank fine-tuning |
| laya, convaiinnovations, Apache-2.0 | the idea of one marker per option, scored in one pass |
| Guo et al. 2017 | temperature scaling |
| Angelopoulos et al. 2021, Learn then Test | the risk calibration |
| Qwen3.8-27B, Qwen team, Apache-2.0 | wrote training documents and proposed training folder names, run locally |
| gpt-oss-120b, OpenAI, Apache-2.0 | wrote the held-out documents for selection, calibration and evaluation, and one family of training names, run locally |
| Tesseract OCR, Apache-2.0 | the OCR text in training |
License
Apache-2.0, see LICENSE. The adapter derives from Qwen3.5-2B-Base, which is Apache-2.0
licensed; see NOTICE. Developed by Akinara as part of PollySort.
Model tree for Kwokou/BeeNara-Big
Base model
Qwen/Qwen3.5-2B-Base