cad-reviews — Deceptive Review Classifier (RoBERTa-base on Ott corpus)
Part of the Content Authenticity Detector project.
Model description
roberta-base fine-tuned on Ott et al.'s Deceptive Opinion Spam Corpus for binary deceptive-review classification.
Dataset: 1,600 hotel reviews — balanced across deceptive/truthful × positive/negative from 20 Chicago hotels.
Label mapping:
0= truthful1= deceptive
Training uses hotel-stratified 5-fold cross-validation (4 hotels held out per fold) to prevent hotel-identity leakage.
Results
5-fold CV (hotel-stratified):
| Fold | Accuracy | Macro F1 |
|---|---|---|
| Fold 1 | 91.9% | 0.919 |
| Fold 2 | 90.6% | 0.906 |
| Fold 3 | 87.8% | 0.877 |
| Fold 4 | 90.3% | 0.903 |
| Fold 5 | 88.4% | 0.884 |
| Mean ± std | 89.8% ± 1.7% | 0.898 ± 0.017 |
Baseline (LR + TF-IDF): 5-fold CV acc=87.5%±2.2%, F1=0.875±0.022. RoBERTa: +2.3% accuracy, +2.3% F1.
Model weights are from the best single fold (Fold 1, macro-F1=0.919).
Training
- Base:
roberta-base - Epochs: 3 (early stopping, patience=2)
- Batch size: 16
- Max length: 256 tokens (Ott reviews average ~150 words; 256 covers 99%+)
- Learning rate: 2e-5
- Hardware: NVIDIA RTX 2060 (6GB)
- CV strategy: Hotel-stratified 5-fold (4 hotels per fold as held-out set)
Intended use & limitations
Intended: Research and educational demonstrations of NLP-based deceptive review detection.
Limitations:
- Trained exclusively on hotel reviews from Chicago, written by Mechanical Turk workers (Ott et al. 2011). Will not generalise to product reviews, restaurant reviews, or other domains.
- Small corpus (1,600 examples). Performance estimates have moderate variance across folds.
- Deceptive reviews in the wild are stylistically different from deliberately written Turk deception.
- Not suitable for production content moderation. Treat all verdicts as signals, not ground truth.
Attention caveat
The accompanying demo uses last-layer attention weights as token highlights. Per Jain & Wallace (2019), attention weights are not faithful explanations — high-attention tokens are not necessarily causal. Highlights are visualisation hints only.
Citation
@inproceedings{ott2011finding,
title={Finding Deceptive Opinion Spam by Any Stretch of the Imagination},
author={Ott, Myle and Choi, Yejin and Cardie, Claire and Hancock, Jeffrey T.},
booktitle={ACL 2011},
}
@inproceedings{ott2013negative,
title={Negative Deceptive Opinion Spam},
author={Ott, Myle and Cardie, Claire and Hancock, Jeffrey T.},
booktitle={NAACL-HLT 2013},
}
- Downloads last month
- 3
Model tree for omkarwaikar/cad-reviews
Base model
FacebookAI/roberta-base