DEBATE-kor-large

DEBATE-kor-large is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text.

The model is initialized from mlburnham/Political_DEBATE_DeBERTa_large_v1.1, the original DeBERTa-based Political DEBATE checkpoint, and subsequently fine-tuned on jongrock17/PolNLI-kor, a Korean translation and adaptation of PolNLI.

The adaptation pipeline is:

Political DEBATE DeBERTa-large → PolNLI-kor → DEBATE-kor-large

Unlike the PolNLI-kor-RoBERTa model family, which starts from Korean-pretrained KLUE-RoBERTa encoders, DEBATE-kor directly adapts the original Political DEBATE checkpoint to Korean political NLI.

Labels

DEBATE-kor formulates NLI as a binary classification problem.

ID Label
0 not_entailment
1 entailment

not_entailment combines non-entailment cases into a single binary class.

Intended Use

DEBATE-kor-large is intended for research on:

  • Korean natural language inference
  • Korean political text
  • cross-lingual adaptation of political language models
  • political text classification and measurement
  • entailment-based analysis of political language

The model may be useful as a component in downstream political text analysis, but its performance should be validated when applied to new domains, genres, or time periods.

Evaluation

The model was evaluated on the full PolNLI-kor test set containing 15,366 premise-hypothesis pairs.

Overall Performance

Metric Score
Accuracy 0.8765
Balanced Accuracy 0.8626
Macro F1 0.8694
Weighted F1 0.8750
Weighted F1 95% CI [0.8697, 0.8803]
MCC 0.7440
AUROC 0.9344
AUPRC 0.9264
Brier Score ↓ 0.1015
ECE ↓ 0.0727
Test N 15,366

Comparison with Alternative Models

Model Weighted F1 95% CI
PolNLI-kor-RoBERTa-base 0.9109 [0.9062, 0.9155]
DEBATE-kor-base 0.9000 [0.8954, 0.9047]
PolNLI-kor-RoBERTa-large 0.8946 [0.8897, 0.8996]
DEBATE-kor-large 0.8750 [0.8697, 0.8803]

The large DEBERTa adaptation performs below DEBATE-kor-base and both Korean-pretrained RoBERTa adaptations in aggregate performance.

However, its task-level results are heterogeneous, with its strongest performance observed on topic classification.

Performance by Task

Task N Weighted F1 Macro F1 MCC
Event extraction 2,864 0.8707 0.8704 0.7614
Hate speech & toxicity 3,002 0.8867 0.8284 0.6729
Stance detection 4,993 0.8477 0.8426 0.6855
Topic classification 4,507 0.8998 0.8971 0.7996

Among the four tasks, DEBATE-kor-large performs best on topic classification, reaching approximately 0.90 weighted F1, while stance detection is the most challenging task.

MCC Distribution Across Tasks

Paired Comparison with PolNLI-kor-RoBERTa-base

Predictions were additionally compared using a continuity-corrected McNemar test on the same 15,366 test examples.

Paired outcome Count
RoBERTa-base correct / DEBATE-kor-large wrong 1,276
DEBATE-kor-large correct / RoBERTa-base wrong 739
Statistic Value
McNemar χ² 142.58
p-value 7.27 × 10⁻³³

The paired comparison indicates a clear difference in prediction errors between the models, consistent with the higher aggregate performance of PolNLI-kor-RoBERTa-base.

Transformers Usage

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

repo_id = "jongrock17/DEBATE-kor-large"

tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSequenceClassification.from_pretrained(repo_id)

model.eval()

premise = "정부는 해당 법안을 국회에 제출했다."
hypothesis = "정부가 법안을 제출했다."

inputs = tokenizer(
    premise,
    hypothesis,
    return_tensors="pt",
    truncation=True,
    max_length=256,
)

with torch.no_grad():
    logits = model(**inputs).logits
    probs = torch.softmax(logits, dim=-1)[0]

prediction = int(probs.argmax())

print("Prediction:", model.config.id2label[prediction])
print("P(entailment):", float(probs[1]))

Training and Evaluation Context

  • Starting checkpoint: mlburnham/Political_DEBATE_DeBERTa_large_v1.1
  • Adaptation dataset: PolNLI-kor
  • Task: binary Korean natural language inference
  • Number of labels: 2
  • Entailment label: 1
  • Non-entailment label: 0
  • Maximum sequence length used in evaluation: 256

Model Lineage

Original Political DEBATE
DeBERTa-large
        │
        ▼
    PolNLI-kor
        │
        ▼
 DEBATE-kor-large

DEBATE-kor-large therefore retains the model lineage of the original Political DEBATE framework while extending it to Korean through additional supervised adaptation on PolNLI-kor.

Relationship to PolNLI-kor-RoBERTa

Two different Korean adaptation strategies are evaluated in this project:

Korean-pretrained approach
KLUE-RoBERTa
     │
     ▼
Korean general NLI
     │
     ▼
  PolNLI-kor
     │
     ▼
PolNLI-kor-RoBERTa


Political-domain transfer approach
Political DEBATE
     │
     ▼
  PolNLI-kor
     │
     ▼
  DEBATE-kor

This distinction makes it possible to compare language-specific pretraining with political-domain-specific pretraining/adaptation under the same Korean political NLI evaluation setting.

Limitations

DEBATE-kor-large is a research model and may inherit limitations from both the original Political DEBATE checkpoint and PolNLI-kor.

In particular:

  • the starting checkpoint was originally developed for English political text rather than Korean;
  • PolNLI-kor is derived from translated political NLI data and may not fully represent linguistic and contextual features specific to Korean politics;
  • cross-lingual adaptation may introduce tokenization or representation limitations;
  • performance may vary across political topics, genres, actors, and time periods;
  • model predictions should not automatically be interpreted as substantive political measurements without downstream validation;
  • probability estimates should not be assumed to remain calibrated under domain shift;
  • the binary formulation collapses different forms of non-entailment into a single class.

Researchers using the model for substantive measurement should evaluate domain shift, classification error, and uncertainty in their specific application.

Citation

A paper citation will be added when the associated manuscript or preprint becomes publicly available.

If you use this model before then, please cite the Hugging Face repository:

DEBATE-kor-large
https://huggingface.co/jongrock17/DEBATE-kor-large

Please also cite the original Political DEBATE work and model where appropriate.

License

This model is derived from an existing Political DEBATE checkpoint and is additionally fine-tuned on PolNLI-kor.

Users should review the licenses and terms associated with:

  • the original Political DEBATE model,
  • its upstream pretrained model,
  • PolNLI,
  • PolNLI-kor,
  • and any other upstream training resources

before redistribution or downstream use.

Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jongrock17/DEBATE-kor-large

Finetuned
(2)
this model

Dataset used to train jongrock17/DEBATE-kor-large

Collection including jongrock17/DEBATE-kor-large