YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

UnBlur ModernBERT β€” Multi-Task Model

Fine-tuned from answerdotai/ModernBERT-base. One shared backbone, three classification heads: clickbait (2-class), political leaning (3-class), sentiment (3-class).

Files

  • config.json + model.safetensors β€” backbone weights (HuggingFace format)
  • tokenizer.json + tokenizer files β€” tokenizer
  • task_heads.pt β€” classification head weights only
  • model_full.pt β€” full checkpoint (backbone + heads)

Usage

Copy this folder to backend/models/ and start the backend. Or point MODEL_REPO_ID at a private HF Hub repo and it'll pull from there.

Performance

Numbers below are from backend/evaluate.py on a 30-example hand-labelled set covering the corner cases (ambiguous leaning, mixed sentiment, emotive- but-legit headlines). The old numbers in this file were training-time validation accuracy on the per-task datasets β€” not comparable, and misleadingly high.

Task Accuracy Macro F1
Clickbait 83% 0.75
Political leaning 70% 0.64
Sentiment 57% 0.51

Latency: ~63ms avg, ~130ms p95 (CPU).

Test set is tiny (n=30), so treat these as ballpark β€” roughly Β±15 points per task. Real eval needs a few hundred held-out examples. That's on the list.

Why it's not better

  • Sentiment (57%) β€” trained on the wrong domain. The head learned on tweet_eval, which is tweets. News headlines aren't tweets. Any headline with a strong word ("crash", "devastate", "threat") gets read as negative, even when the story is neutral or good news. Most of the errors are neutral/positive headlines called negative. Used this dataset as a placeholder.

  • Clickbait (83%) β€” conflates tone with clickbait. Every error is a false positive: real headlines that happen to be loud ("GOP Tax Cuts Devastate Working Families", "Democrats Slam Republican Plan") get flagged. The training set was Buzzfeed-style listicle bait vs. clean wire copy, so the model keys on emotional language instead of the actual curiosity-gap pattern.

  • Political leaning (70%) β€” collapses toward center, barely sees right. Only 2 of 6 right-leaning examples classified right; the rest went center. Labels come from AllSides source ratings (config.py), so the head learned publisher style, not article content, and the training mix leans left-heavy. Need to use a larger test set to understand model choices better.

Planned for v2

  • Sentiment: retrain on news-domain data. Either distill labels from cardiffnlp/twitter-roberta onto an actual news corpus (already wired up in datasets_loader.distill_sentiment_labels), or use a headline sentiment set. Add a proper "neutral" bias since most news is neutral.
  • Clickbait: rebuild the negative class from strongly-worded real headlines, not just wire copy, so the model has to learn the structure and not the volume. Lower the flag threshold's reliance on the raw probability.
  • Leaning: stop labelling by source. Hand-label or LLM-label a content-based set, balance left/center/right, and add right-leaning framing examples specifically.
  • Eval: grow the test set to 300+ and version it, so v2 vs v1 is an honest comparison instead of a vibe check.
Downloads last month
8
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support