YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
RoBERTa-base FOMC Stance Classifier (temporal split)
Classifies sentences from Federal Reserve communications as hawkish, dovish, or neutral.
Trained on the Trillion Dollar Words annotated corpus (Shah et al., ACL 2023) under a strict temporal split rather than the random split shipped with the source dataset.
Why the temporal split matters
The upstream train/test split spans 1996β2022 on both sides. FOMC minutes recycle near-identical boilerplate across meetings, so a random split lets a model memorise sentences it is later evaluated on.
This model is trained on 1996β2014 and evaluated on 2015β2022, with zero overlap. The reported score is therefore lower than commonly cited numbers on this dataset and not directly comparable to them β the temporal task is harder.
Results
| macro F1 | |
|---|---|
| 5-fold CV (inside train block) | 0.664 Β± 0.026 |
| Test 2015β2022, per seed | 0.659 Β± 0.009 |
| Test, 3-seed ensemble | 0.673 |
Per-class F1: neutral 0.744, hawkish 0.659, dovish 0.617. Dovish language is more hedged and conditional, making it the hardest class.
CV and test agree within 0.005 β evidence that model selection did not overfit, since all tuning ran inside the training block.
Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tok = AutoTokenizer.from_pretrained("LazyPranay01/ROBERTa_FOMC")
model = AutoModelForSequenceClassification.from_pretrained("LazyPranay01/ROBERTa_FOMC")
text = "The Committee judges that inflation remains elevated and further tightening may be warranted."
enc = tok(text, return_tensors="pt", truncation=True, max_length=128)
probs = torch.softmax(model(**enc).logits, dim=1)[0]
# index 0 = dovish, 1 = hawkish, 2 = neutral
Document-level aggregation
To score a whole document, average probabilities, not predicted labels:
stance = (probs[:, 1] - probs[:, 0]).mean() # positive = hawkish
Averaging predicted label integers is a silent bug β with dovish=0, hawkish=1, neutral=2, the mean places neutral above hawkish and inverts the signal.
Validation
Applied to 75 FOMC meeting minutes (2015β2024), the resulting stance index tracks known policy regimes:
| Regime | Mean stance | |
|---|---|---|
| 2020 COVID | β0.277 | most dovish |
| 2019 cuts | β0.206 | dovish |
| 2015β2018 hiking | β0.060 | |
| 2022 hiking | +0.113 | most hawkish |
Training
RoBERTa-base, 3 seeds (5768 / 78516 / 944601), class-weighted cross-entropy (neutral is ~2Γ the minority classes), dynamic padding, max length 128, LR 2e-5, batch 16, early stopping on validation macro F1. Single T4, ~35 minutes.
Training data was deduplicated before splitting β 50 verbatim-repeated sentences (mandate boilerplate) removed β and the build asserts zero cross-split sentence overlap.
Limitations
- Trained on minutes, speeches and press-conference transcripts. FOMC statements are a different genre (heavily templated, ~12 sentences) and scores on them are not calibrated to the same scale.
- Sentence-level only; no document-level fine-tuning.
- Dovish recall is the weakest dimension.
- Downstream testing found no evidence this stance signal predicts 2Y Treasury direction beyond macro state (ΞAUC β0.13, 95% CI [β0.29, +0.02], n=74). It is a validated measurement instrument, not a trading signal.
Citation
Source dataset:
Shah, A., Paturi, S., & Chava, S. (2023). Trillion Dollar Words: A New Financial Dataset, Task & Market Analysis. ACL 2023.
Code and full analysis: https://github.com//fomc-stance-pipeline