LieLine

LieLine is a fine-tuned RoBERTa-base classifier that detects disinformation / lie allegations — instances where a political speaker explicitly or implicitly accuses another actor of lying, deception, or spreading disinformation — in political speech text. It was developed as part of the pipeline described in "Finding the Needle in a Haystack: Using Large Language Models to Detect Rare Speech Acts" (Mochtak & Meijers, 2026).

Model Description

  • Base model: roberta-base (English)
  • Task: Binary sentence/snippet-level text classification (1/0)
  • Fine-tuning library: simpletransformers
  • Language: English (source speeches in other EU languages were machine-translated to English prior to processing, using Meta's NLLB-200)
  • Developed by: Michal Mochtak (Radboud University) and Maurits J. Meijers (University of Antwerp)
  • Funded by: European Research Council Starting Grant "Deception in Democracy: Political Lying Accusations and Their Effects on Democratic Citizenship" (DEMO-LIES), Grant agreement ID: 101164535

Intended Uses

LieLine is intended for research use in computational social science and political communication research, specifically for:

  • Flagging sentences or short text snippets in political speech corpora that likely contain disinformation or lie allegations
  • Large-scale, exploratory measurement of the prevalence and trends of this rare speech act (e.g., across time, speakers, or party groups)
  • Serving as a starting point / fine-tuning base for related rare-speech-act detection tasks in political text

Training Data

The training data was constructed through a multi-phase pipeline rather than through direct random sampling of the underlying corpus, to address the rarity of the target speech act:

  1. Source corpus: 523,983 European Parliament speeches (1999–2024, 5th–9th terms), non-English speeches translated to English via NLLB-200.
  2. Keyword filtering: An LLM-generated dictionary of 376 deception-related lemmas was used to identify 4,344 speeches with ≥5 keyword occurrences.
  3. LLM few-shot extraction: ChatGPT-3.5-Turbo extracted 1,816 candidate sentences/snippets likely to contain lie allegations.
  4. Manual annotation: Two trained annotators labeled the 1,816 snippets (yes/no) for explicit disinformation allegations (Krippendorff's α = 0.72 on a 100-item pilot); 905 were confirmed positive. Cross-coder consistency checks flagged and reconciled 356 systematically inconsistent instances (19.6%).
  5. Negative augmentation: 1,816 additional sentences containing none of the dictionary keywords were added as (validated, near-certain) negative examples.

The final training set used for the released model contains 3,632 sentences/snippets, of which 985 (27%) are positive instances of disinformation/lie allegations. Class weighting (1:2.17) was applied to address the residual imbalance.

Training Procedure

  • Architecture: RoBERTa-base, fine-tuned for sequence classification
  • Epochs: 5
  • Learning rate: 2e-5 (final production model trained on the full dataset without a held-out split); 4e-5 was used during the cross-validation/consistency-check phase
  • Batch size: 8
  • Class weighting: 1:2.17 (minority class up-weighted)
  • All other hyperparameters left at simpletransformers defaults.

The final released model was trained on the entire 3,632-instance dataset (no evaluation split held out), after validation was completed via a separate 100-bootstrap cross-validation procedure (see below).

Evaluation Results

Performance was estimated via 100-bootstrap 80/20 train/evaluation splits:

Version MCC Accuracy F1 AUROC AUPRC
No class weights 0.82 (0.03) 0.93 (0.01) 0.91 (0.01) 0.97 (0.01) 0.92 (0.02)
Weighted (1:2.17) 0.82 (0.02) 0.93 (0.01) 0.91 (0.01) 0.98 (0.01) 0.92 (0.02)

Values are means across 100 bootstrapped models; standard deviations in parentheses.

Usage example

from transformers import AutoModelForSequenceClassification, TextClassificationPipeline, AutoTokenizer, AutoConfig

MODEL = "mmochtak/lieline"
tokenizer = AutoTokenizer.from_pretrained(MODEL)
config = AutoConfig.from_pretrained(MODEL)
model = AutoModelForSequenceClassification.from_pretrained(MODEL)

pipe = TextClassificationPipeline(model=model, tokenizer=tokenizer, task='classification', device=0)

result = pipe([
    "You are a moron.",
    "You, sir, are a liar.",
    "I do not know; I was not there.",
    "You are giving us very misleading information!"
])

print(result)

Please cite the model as follows:

@misc{parlasent-model,
    author       = {Mochtak, Michal and Meijers, Maurits},
    title        = {LieLine},
    year         = {2026},
    url          = {https://huggingface.co/mmochtak/lieline},
    publisher    = {Hugging Face}
}
Downloads last month
14
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support