BeeNara Big 🐝 v0.3 (preview)

Suggest the right folder in a whole folder tree, or say that none fits.

BeeNara Big reads a document (its text and, if available, the page image) together with the complete folder tree of a user (from 3 to several hundred folders) and returns a calibrated probability for every folder and for "none fits". It is the large sibling of BeeNara, which decides between at most 9 folders on a CPU.

Status: preview, suggestion mode. v0.3 ranks folders well, but it is not certified to sort automatically: the risk calibration found no confidence threshold that keeps automatic decisions at ≤ 1 % wrong (see Limitations). Use it to suggest folders and let the user confirm.

Quickstart

pip install torch "transformers>=5.17" peft safetensors tokenizers huggingface_hub pillow
hf download Kwokou/BeeNara-Big beenara_big.py --local-dir .
from beenara_big import BeeNaraBig

bb = BeeNaraBig.from_pretrained("Kwokou/BeeNara-Big")   # adapter 164 MB + Qwen3.5-2B-Base 4.6 GB, GPU if available

s = bb.suggest(
    "Einkommensteuerbescheid 2025\nFinanzamt Musterstadt\nFestgesetzte Einkommensteuer: 1.234,00 EUR",
    ["Finanzen/Steuer/2024", "Finanzen/Steuer/2025", "Finanzen/Bank", "Wohnung/Miete", "Arbeit/Verträge"],
)
s.suggestions     # [('Finanzen/Steuer/2025', 0.91), ('Finanzen/Steuer/2024', ...), ...]
s.none            # p("none of these folders fits"), here 0.07
s.ask             # True: in v0.3 the user always confirms
  • Folder names are free text; paths like Finanzen/Steuer/2025 work best. A description can follow the name: "Wohnung/Miete — Mietvertrag, Nebenkosten".
  • images=[...] adds page images (file paths or PIL images); language="en" switches the internal prompt to English.
  • On a GPU the base model loads in bf16, on a CPU in fp32. The loader gives exactly the same scores as the training code (checked on the same inputs).

Key points

  • Whole folder trees. One forward pass over up to ~300 folders (trained up to 500), with folder paths such as Finanzen/Steuer/2025.
  • Good where the small model cannot see enough. On held-out trees with 25–299 folders it finds the right folder in 79–88 % of the cases; the small BeeNara with a 9-folder pre-selection reached about 53 % on comparable large trees.
  • Knows when nothing fits: 82–84 % of the documents whose folder is missing are recognized.
  • Reads the page. The page image (letterhead, logo, layout) is used together with the text in half of the training tasks; it also works with text only.
  • Calibrated. Expected calibration error 0.045.
  • Same quality as the 4B model at half the size. A paired comparison with a Qwen3.5-4B version (same data, same seeds) found differences of at most ±1 point per list size.

Results

Held-out "world": folder trees, folder names and documents (written by gpt-oss-120b instead of the training teacher Qwen3.8-27B) that were never used for training. 2,344 tasks, 400 per list size (fewer for the largest trees), including tasks where no folder fits. Calibrated, PyTorch bf16. This repository ships seed 1.

Folders in the tree Tasks Accuracy, seed 1 Accuracy, seed 2
3–9 400 91.7 % 91.2 %
10–24 400 90.0 % 88.7 %
25–49 400 88.2 % 87.5 %
50–99 400 82.8 % 82.3 %
100–199 400 78.5 % 78.2 %
200–299 293 79.2 % 77.8 %
300 and more 51 70.6 % 60.8 %
All 2,344 85.0 % 84.1 %
Metric Seed 1 Seed 2
"None fits" recall 84.0 % 82.3 %
"None fits" precision 72.4 % 70.7 %
Wrongly said "none fits" 9.5 % 10.1 %
Calibration error (ECE) 0.045 0.045
Decision changes when the folder order is shuffled 2.3 % 2.3 %

2B against 4B (paired, two seeds each, same tasks): accuracy difference 4B − 2B of +0.6, −0.9, +0.4 and −0.9 points for 3–9, 10–24, 25–49 and 50–99 folders. The 2B model was chosen by the pre-registered rule "at most 2 points behind the 4B in every list size up to 99 folders".

Speed and memory.

GPU memory, 8-bit weights, 300 folders 4.2 GB
CPU, 4 threads, 7 folders, text only (337 tokens) 3.0 s
CPU, 4 threads, 89 folders, text only (2,702 tokens) 31 s
CPU, 4 threads, 35 folders, with page image (4,525 tokens) 99 s
CPU, 4 threads, 375 folders (16,107 tokens) 217 s

On a CPU it is too slow for interactive use; plan for a GPU. GPU latency was not measured yet.

How it works

  • Base: Qwen3.5-2B-Base (revision b1485b2f), a hybrid model with linear and full attention and a vision encoder.
  • Input (variant "L2"): [task] [page image] [document text] [folder list] [folder list with one marker per folder] [none marker]. The folder list appears twice so that every folder marker sees the document and the complete list. Markers are existing special tokens; no new vocabulary.
  • Head: a small decision head reads the hidden states at the markers and gives one score per folder and one for "none fits"; together they form one probability distribution.
  • Limits: context up to 16,384 tokens; document text up to 2,048 tokens (start and end are kept); up to 32 tokens per folder name; page images up to about 1024 × 1448 pixels.
  • Calibration: temperature scaling on 7,115 tasks from the held-out world (1,837 for the smallest trees, 15 for trees above 300 folders); separate calibration files for bf16 and for 8-bit weights (kalibrierung.json, kalibrierung_8bit.json), each bound to its weights by SHA-256.

Training

  • Data. Only synthetic documents, no user data. 45,322 documents placed in 3,000 synthetic folder trees of 10–500 folders (German, English and mixed; private households, freelancers, small firms, construction offices, associations, law offices):
    • 15,813 new documents written by Qwen3.8-27B for concrete leaf folders, each checked against the facts it was built from;
    • 29,509 template documents assigned to trees by their metadata;
    • page images rendered for all documents; about 20 % degraded (scan, phone photo, skew) and read with Tesseract for OCR text.
  • Episodes. Every step draws a task per document: the whole tree or a part of it (3–300 folders, 5 % up to 500), with folder descriptions in 30 % and spelling variants in 20 % of the tasks; in 25 % the right folder is missing, half of those with a similar neighbour present; in 15 % only the parent of the right folder is in the list, and the parent is the answer.
  • Kept apart by tests. Separate names and trees for training and for the held-out world; normalized and substring overlap checks, near-duplicate and 13-gram checks found no leakage.
  • Setup: LoRA (r = 32, α = 64, dropout 0.05) on attention, linear attention and MLP of the language model; vision encoder frozen. Learning rate 1e-4 for LoRA and 3e-4 for the head, 3 % warmup, 2 epochs = 5,664 steps of 16 tasks, label smoothing 0.05, gradient checkpointing, bf16. Best checkpoint by validation NLL (step 5,000).
  • Hardware: one NVIDIA RTX A6000 (48 GB), about 30 hours.

Limitations

  • No automatic sorting in v0.3. The risk calibration (Learn-then-Test, ≤ 1 % wrong among automatic decisions, δ = 0.1) found no threshold for any list size. The likely cause is the label smoothing used in training, which compresses the confidence of the surest answers. Use the model for suggestions; v0.4 will be trained without label smoothing.
  • Large trees are harder. Accuracy falls from about 92 % (3–9 folders) to about 79 % (100–299) and further above 300 folders, where the test set is small (51 tasks).
  • "None fits" is less reliable than in the small model, especially in large trees.
  • Slow on a CPU (seconds to minutes per document), heavy with a page image.
  • Synthetic training data. Real archives are messier. Expect a gap until the model is checked on real, consented documents.
  • German and English only.

Files

File Content Size
lora.safetensors LoRA weights for the language model 125 MB
kopf.safetensors decision head 19 MB
big_config.json base model and revision, LoRA settings, input limits, variant < 1 KB
kalibrierung.json calibration for bf16 weights 6 KB
kalibrierung_8bit.json calibration for 8-bit weights 6 KB
tokenizer/ Qwen3.5 tokenizer 20 MB
bildprozessor/ image preprocessing settings < 1 KB
beenara_big.py loader and inference: input format, head, calibration 14 KB

The base model weights are not included; they come from Qwen/Qwen3.5-2B-Base at the revision in big_config.json.

Credits

Source What BeeNara Big takes from it
Qwen3.5-2B-Base, Qwen team, Apache-2.0 the base model
LoRA, Hu et al. 2021 low-rank fine-tuning
laya, convaiinnovations, Apache-2.0 the idea of one marker per option, scored in one pass
Guo et al. 2017 temperature scaling
Angelopoulos et al. 2021, Learn then Test the risk calibration
Qwen3.8-27B, Qwen team, Apache-2.0 wrote training documents and proposed training folder names, run locally
gpt-oss-120b, OpenAI, Apache-2.0 wrote the held-out documents for selection, calibration and evaluation, and one family of training names, run locally
Tesseract OCR, Apache-2.0 the OCR text in training

License

Apache-2.0, see LICENSE. The adapter derives from Qwen3.5-2B-Base, which is Apache-2.0 licensed; see NOTICE. Developed by Akinara as part of PollySort.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kwokou/BeeNara-Big

Adapter
(36)
this model

Collection including Kwokou/BeeNara-Big

Papers for Kwokou/BeeNara-Big