Jev topic classifier — local replacement

📄 Review & open questions live here: https://huggingface.co/spaces/Kodep/jev-review — our read of the spike report, the numbers that matter, and the questions to confirm in the code.

🔒 Data is kept private & separate: https://huggingface.co/datasets/Kodep/jev-topic-data — public code/weights, but the raw item text (school-email-derived) is gated. The data/ folder in this repo is only a pointer.

Replaces the hosted Jev step of the school-email digest pipeline: classify each item into one of 56 topics so the fixed rules table can decide show/hide.

Chosen approach (per the baseline report): Qwen/Qwen3-Embedding-0.6B + LoRA (rank 16) + a prototype head (56 learned vectors; score = cosine similarity × learned scale), running on CPU — ~0.2 s/item, ~2.2 GiB RAM, 0 GPU.

STATUS: SCAFFOLD. This repo is ready to receive the code, data, and weights from the owner. Nothing owner-side has been pushed yet. The card metadata above is provisional until the real LoRA adapter + head weights land. The numbers in docs/report_claude_v1.md are Claude-written and unverified — treat them as claims to confirm against the code and a re-run of evaluate.py, not as established facts.

Intended layout

jev-topic-head/
├── README.md
├── docs/report_claude_v1.md     # baseline claims (Claude-written, unverified)
├── config.yaml                  # all hyperparameters          ← to push
├── data/                        # → pointer to private repo Kodep/jev-topic-data (not stored here)
├── src/                         # train_head.py, evaluate.py, rules_table.json ← to push
├── checkpoints/                 # LoRA adapter + head prototypes (seed 0)   ← to push
└── eval/                        # per-item test dump + confusion matrix    ← to push

Data, the topic taxonomy, and the human ("gold") labels live in the private dataset repo Kodep/jev-topic-data. To pull them for a run (collaborators only): hf download Kodep/jev-topic-data --repo-type dataset --local-dir data.

What unblocks verification (priority order)

  1. src/train_head.py + config.yaml — confirm the actual training matches the report's claims (checklist below).
  2. eval/per_item_test.jsonl — recompute top-1 accuracy, macro-F1, and show/hide without re-running the model.
  3. checkpoints/ + data (private repo) — reproduce end to end and re-confirm the CPU latency/RAM.

Claims to verify against the code

  • LoRA on attention + feed-forward layers, rank 16 / alpha 32 / dropout 0.05; rest frozen
  • Head = 56 prototype vectors, score = cosine(item, proto) × learned scale (scale init 20)
  • Prototypes initialized from the embedding of each topic description (14 topics have no training item)
  • Loss = softmax cross-entropy, label smoothing 0.05, single-label (not multi-label)
  • AdamW: LoRA lr 3e-4 (wd 0.01), head lr 2e-3 (no wd); warmup 10% → linear decay to 5%
  • Batch 16, ≤ 12 epochs, early stop on validation macro-F1, bf16
  • Class-balanced sampling (weight 1/√count), hard topic pairs ×2, no augmentation / no synthetic items
  • 3 seeds; macro-F1 spanned 0.49–0.59 across seeds; shipped run = seed 0
  • "87.6%" = top-1 topic accuracy vs Jev's fresh answer on test (92/105); macro-F1 0.572; show/hide 95/105 (90.5%)

Headline caveat

All scores are measured against Jev, not ground truth. Jev matches itself only 97.1% of the time and matched a human label 11/20 on one 20-item set. Only 31 items have independent human labels, and the test split was used for decisions (treated as dev data). A fresh, unopened 100-item set is the intended final benchmark.

Reproduction

TBD once src/ lands — will document exact commands (env, base-model path, seed, config) to reproduce train + eval.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kodep/jev-topic-head

Adapter
(31)
this model

Space using Kodep/jev-topic-head 1