DEBATE-kor-base

DEBATE-kor-base is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text.

The model is initialized from mlburnham/Political_DEBATE_DeBERTa_base_v1.1, the original DeBERTa-based Political DEBATE checkpoint, and subsequently fine-tuned on jongrock17/PolNLI-kor, a Korean translation and adaptation of PolNLI.

The adaptation pipeline is:

Political DEBATE DeBERTa-base → PolNLI-kor → DEBATE-kor-base

Unlike the PolNLI-kor-RoBERTa model family, which starts from Korean-pretrained KLUE-RoBERTa encoders, DEBATE-kor directly adapts the original Political DEBATE checkpoint to Korean political NLI.

Labels

DEBATE-kor formulates NLI as a binary classification problem.

ID Label
0 not_entailment
1 entailment

not_entailment combines non-entailment cases into a single binary class.

Intended Use

DEBATE-kor-base is intended for research on:

  • Korean natural language inference
  • Korean political text
  • cross-lingual adaptation of political language models
  • political text classification and measurement
  • entailment-based analysis of political language

The model may be useful as a component in downstream political text analysis, but its performance should be validated when applied to new domains, genres, or time periods.

Evaluation

The model was evaluated on the full PolNLI-kor test set containing 15,366 premise-hypothesis pairs.

Overall Performance

Metric Score
Accuracy 0.9009
Balanced Accuracy 0.8903
Macro F1 0.8958
Weighted F1 0.9000
Weighted F1 95% CI [0.8954, 0.9047]
MCC 0.7946
AUROC 0.9620
AUPRC 0.9523
Brier Score ↓ 0.0908
ECE ↓ 0.0857
Test N 15,366

Comparison with Alternative Models

Model Weighted F1 95% CI
PolNLI-kor-RoBERTa-base 0.9109 [0.9062, 0.9155]
DEBATE-kor-base 0.9000 [0.8954, 0.9047]
PolNLI-kor-RoBERTa-large 0.8946 [0.8897, 0.8996]
DEBATE-kor-large 0.8750 [0.8697, 0.8803]

DEBATE-kor-base achieves strong performance despite being adapted from an English Political DEBATE checkpoint.

The comparison with the PolNLI-kor-RoBERTa models provides evidence about two alternative adaptation strategies:

  1. starting from a Korean-pretrained encoder and adapting it to political NLI; and
  2. starting from an English political-NLI-specialized model and adapting it to Korean.

On the full evaluation set, the Korean-pretrained RoBERTa-base model achieves the highest overall weighted F1, while DEBATE-kor-base remains competitive.

Performance by Task

Task N Weighted F1 Macro F1 MCC
Event extraction 2,864 0.9116 0.9111 0.8315
Hate speech & toxicity 3,002 0.9173 0.8786 0.7596
Stance detection 4,993 0.9012 0.8979 0.7962
Topic classification 4,507 0.8802 0.8770 0.7596

The model performs above 0.90 weighted F1 on event extraction, hate speech & toxicity, and stance detection. Topic classification is comparatively more challenging for this adaptation.

MCC Distribution Across Tasks

Paired Comparison with PolNLI-kor-RoBERTa-base

Predictions were additionally compared using a continuity-corrected McNemar test on the same 15,366 test examples.

Paired outcome Count
RoBERTa-base correct / DEBATE-kor-base wrong 982
DEBATE-kor-base correct / RoBERTa-base wrong 819
Statistic Value
McNemar χ² 14.57
p-value 1.35 × 10⁻⁴

The paired difference is statistically significant, although the absolute difference in weighted F1 is approximately 1.1 percentage points. Statistical significance should therefore be interpreted together with the magnitude of the performance difference and the task-specific results.

Transformers Usage

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

repo_id = "jongrock17/DEBATE-kor-base"

tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSequenceClassification.from_pretrained(repo_id)

model.eval()

premise = "정부는 해당 법안을 국회에 제출했다."
hypothesis = "정부가 법안을 제출했다."

inputs = tokenizer(
    premise,
    hypothesis,
    return_tensors="pt",
    truncation=True,
    max_length=256,
)

with torch.no_grad():
    logits = model(**inputs).logits
    probs = torch.softmax(logits, dim=-1)[0]

prediction = int(probs.argmax())

print("Prediction:", model.config.id2label[prediction])
print("P(entailment):", float(probs[1]))

Training and Evaluation Context

  • Starting checkpoint: mlburnham/Political_DEBATE_DeBERTa_base_v1.1
  • Adaptation dataset: PolNLI-kor
  • Task: binary Korean natural language inference
  • Number of labels: 2
  • Entailment label: 1
  • Non-entailment label: 0
  • Maximum sequence length used in evaluation: 256

Model Lineage

Original Political DEBATE
DeBERTa-base
        │
        ▼
    PolNLI-kor
        │
        ▼
 DEBATE-kor-base

DEBATE-kor-base therefore retains the model lineage of the original Political DEBATE framework while extending it to Korean through additional supervised adaptation on PolNLI-kor.

Relationship to PolNLI-kor-RoBERTa

Two different Korean adaptation strategies are evaluated in this project:

Korean-pretrained approach
KLUE-RoBERTa
     │
     ▼
Korean general NLI
     │
     ▼
  PolNLI-kor
     │
     ▼
PolNLI-kor-RoBERTa


Political-domain transfer approach
Political DEBATE
     │
     ▼
  PolNLI-kor
     │
     ▼
  DEBATE-kor

This distinction makes it possible to compare language-specific pretraining with political-domain-specific pretraining/adaptation under the same Korean political NLI evaluation setting.

Limitations

DEBATE-kor-base is a research model and may inherit limitations from both the original Political DEBATE checkpoint and PolNLI-kor.

In particular:

  • the starting checkpoint was originally developed for English political text rather than Korean;
  • PolNLI-kor is derived from translated political NLI data and may not fully represent linguistic and contextual features specific to Korean politics;
  • cross-lingual adaptation may introduce tokenization or representation limitations;
  • performance may vary across political topics, genres, actors, and time periods;
  • model predictions should not automatically be interpreted as substantive political measurements without downstream validation;
  • probability estimates should not be assumed to remain calibrated under domain shift;
  • the binary formulation collapses different forms of non-entailment into a single class.

Researchers using the model for substantive measurement should evaluate domain shift, classification error, and uncertainty in their specific application.

Citation

A paper citation will be added when the associated manuscript or preprint becomes publicly available.

If you use this model before then, please cite the Hugging Face repository:

DEBATE-kor-base
https://huggingface.co/jongrock17/DEBATE-kor-base

Please also cite the original Political DEBATE work and model where appropriate.

License

This model is derived from an existing Political DEBATE checkpoint and is additionally fine-tuned on PolNLI-kor.

Users should review the licenses and terms associated with:

  • the original Political DEBATE model,
  • its upstream pretrained model,
  • PolNLI,
  • PolNLI-kor,
  • and any other upstream training resources

before redistribution or downstream use.

Downloads last month
26
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jongrock17/DEBATE-kor-base

Finetuned
(1)
this model

Dataset used to train jongrock17/DEBATE-kor-base

Collection including jongrock17/DEBATE-kor-base