YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

RoBERTa-base FOMC Stance Classifier (temporal split)

Classifies sentences from Federal Reserve communications as hawkish, dovish, or neutral.

Trained on the Trillion Dollar Words annotated corpus (Shah et al., ACL 2023) under a strict temporal split rather than the random split shipped with the source dataset.

Why the temporal split matters

The upstream train/test split spans 1996–2022 on both sides. FOMC minutes recycle near-identical boilerplate across meetings, so a random split lets a model memorise sentences it is later evaluated on.

This model is trained on 1996–2014 and evaluated on 2015–2022, with zero overlap. The reported score is therefore lower than commonly cited numbers on this dataset and not directly comparable to them β€” the temporal task is harder.

Results

macro F1
5-fold CV (inside train block) 0.664 Β± 0.026
Test 2015–2022, per seed 0.659 Β± 0.009
Test, 3-seed ensemble 0.673

Per-class F1: neutral 0.744, hawkish 0.659, dovish 0.617. Dovish language is more hedged and conditional, making it the hardest class.

CV and test agree within 0.005 β€” evidence that model selection did not overfit, since all tuning ran inside the training block.

Usage

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

tok = AutoTokenizer.from_pretrained("LazyPranay01/ROBERTa_FOMC")
model = AutoModelForSequenceClassification.from_pretrained("LazyPranay01/ROBERTa_FOMC")

text = "The Committee judges that inflation remains elevated and further tightening may be warranted."
enc = tok(text, return_tensors="pt", truncation=True, max_length=128)
probs = torch.softmax(model(**enc).logits, dim=1)[0]
# index 0 = dovish, 1 = hawkish, 2 = neutral

Document-level aggregation

To score a whole document, average probabilities, not predicted labels:

stance = (probs[:, 1] - probs[:, 0]).mean()   # positive = hawkish

Averaging predicted label integers is a silent bug β€” with dovish=0, hawkish=1, neutral=2, the mean places neutral above hawkish and inverts the signal.

Validation

Applied to 75 FOMC meeting minutes (2015–2024), the resulting stance index tracks known policy regimes:

Regime Mean stance
2020 COVID βˆ’0.277 most dovish
2019 cuts βˆ’0.206 dovish
2015–2018 hiking βˆ’0.060
2022 hiking +0.113 most hawkish

Training

RoBERTa-base, 3 seeds (5768 / 78516 / 944601), class-weighted cross-entropy (neutral is ~2Γ— the minority classes), dynamic padding, max length 128, LR 2e-5, batch 16, early stopping on validation macro F1. Single T4, ~35 minutes.

Training data was deduplicated before splitting β€” 50 verbatim-repeated sentences (mandate boilerplate) removed β€” and the build asserts zero cross-split sentence overlap.

Limitations

  • Trained on minutes, speeches and press-conference transcripts. FOMC statements are a different genre (heavily templated, ~12 sentences) and scores on them are not calibrated to the same scale.
  • Sentence-level only; no document-level fine-tuning.
  • Dovish recall is the weakest dimension.
  • Downstream testing found no evidence this stance signal predicts 2Y Treasury direction beyond macro state (Ξ”AUC βˆ’0.13, 95% CI [βˆ’0.29, +0.02], n=74). It is a validated measurement instrument, not a trading signal.

Citation

Source dataset:

Shah, A., Paturi, S., & Chava, S. (2023). Trillion Dollar Words: A New Financial Dataset, Task & Market Analysis. ACL 2023.

Code and full analysis: https://github.com//fomc-stance-pipeline

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support