YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
UnBlur ModernBERT β Multi-Task Model
Fine-tuned from answerdotai/ModernBERT-base. One shared backbone, three
classification heads: clickbait (2-class), political leaning (3-class),
sentiment (3-class).
Files
config.json+model.safetensorsβ backbone weights (HuggingFace format)tokenizer.json+ tokenizer files β tokenizertask_heads.ptβ classification head weights onlymodel_full.ptβ full checkpoint (backbone + heads)
Usage
Copy this folder to backend/models/ and start the backend. Or point
MODEL_REPO_ID at a private HF Hub repo and it'll pull from there.
Performance
Numbers below are from backend/evaluate.py on a 30-example hand-labelled
set covering the corner cases (ambiguous leaning, mixed sentiment, emotive-
but-legit headlines). The old numbers in this file were training-time
validation accuracy on the per-task datasets β not comparable, and
misleadingly high.
| Task | Accuracy | Macro F1 |
|---|---|---|
| Clickbait | 83% | 0.75 |
| Political leaning | 70% | 0.64 |
| Sentiment | 57% | 0.51 |
Latency: ~63ms avg, ~130ms p95 (CPU).
Test set is tiny (n=30), so treat these as ballpark β roughly Β±15 points per task. Real eval needs a few hundred held-out examples. That's on the list.
Why it's not better
Sentiment (57%) β trained on the wrong domain. The head learned on
tweet_eval, which is tweets. News headlines aren't tweets. Any headline with a strong word ("crash", "devastate", "threat") gets read as negative, even when the story is neutral or good news. Most of the errors are neutral/positive headlines called negative. Used this dataset as a placeholder.Clickbait (83%) β conflates tone with clickbait. Every error is a false positive: real headlines that happen to be loud ("GOP Tax Cuts Devastate Working Families", "Democrats Slam Republican Plan") get flagged. The training set was Buzzfeed-style listicle bait vs. clean wire copy, so the model keys on emotional language instead of the actual curiosity-gap pattern.
Political leaning (70%) β collapses toward center, barely sees right. Only 2 of 6 right-leaning examples classified right; the rest went center. Labels come from AllSides source ratings (
config.py), so the head learned publisher style, not article content, and the training mix leans left-heavy. Need to use a larger test set to understand model choices better.
Planned for v2
- Sentiment: retrain on news-domain data. Either distill labels from
cardiffnlp/twitter-robertaonto an actual news corpus (already wired up indatasets_loader.distill_sentiment_labels), or use a headline sentiment set. Add a proper "neutral" bias since most news is neutral. - Clickbait: rebuild the negative class from strongly-worded real headlines, not just wire copy, so the model has to learn the structure and not the volume. Lower the flag threshold's reliance on the raw probability.
- Leaning: stop labelling by source. Hand-label or LLM-label a content-based set, balance left/center/right, and add right-leaning framing examples specifically.
- Eval: grow the test set to 300+ and version it, so v2 vs v1 is an honest comparison instead of a vibe check.
- Downloads last month
- 8