BERT for Grammatical Error Detection (BLiMP + CoLA)

bert-base-uncased fine-tuned for binary grammatical error detection: given one English sentence, decide whether it contains a grammatical error.

label meaning
0 grammatical
1 ungrammatical

Note the orientation: 1 means "has an error." This is the inverse of CoLA's native convention (where 1 = acceptable), and the training labels were flipped accordingly.

Usage

from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch

model_id = "yuvstk/bert-grammar-error-detection"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

sentences = ["Katherine can't help himself.", "She went home early."]
inputs = tokenizer(sentences, padding=True, truncation=True,
                   max_length=64, return_tensors="pt")

with torch.no_grad():
    probs = torch.softmax(model(**inputs).logits, dim=-1)

for sentence, prob in zip(sentences, probs):
    label = int(prob.argmax())
    print(f"{'UNGRAMMATICAL' if label else 'GRAMMATICAL'} "
          f"({prob[label]:.1%}) — {sentence}")

Or with a pipeline (LABEL_1 = ungrammatical):

from transformers import pipeline
clf = pipeline("text-classification", model=model_id)
clf("The professor talked us.")

Training data

Two sources merged into a single 4000 + 9594 sentence corpus:

source rows (train) what it contributes
BLiMPanaphor_gender_agreement + anaphor_number_agreement 3200 synthetic minimal pairs; reflexive pronoun agreement; exactly 50/50 balanced
CoLA (GLUE) 7695 real linguistics-literature sentences; many error types; ~70/30 imbalanced

Splitting differs per source, because the sources need different treatment:

  • BLiMP is split by pair_id, never by row. The two sentences of a minimal pair differ by exactly one word, so a row-level split would put a near-duplicate of a test sentence into training.
  • CoLA is split by stratified rows. GLUE's test split is unlabelled (all -1), so GLUE validation is used as the test set and the validation set is carved out of GLUE train.

Total: 10,895 train / 1,256 validation / 1,443 test.

Results

Evaluated separately per source, because the two halves differ enormously in difficulty — a single pooled number would mostly reflect the mixture ratio.

test set n accuracy precision recall F1 MCC
BLiMP 400 1.000 1.000 1.000 1.000 1.000
CoLA 1043 0.837 0.833 0.590 0.691 0.601
pooled 1443 0.882 0.911 0.747 0.821 0.743

CoLA MCC of 0.601 is in the normal published range for BERT-base (~0.55–0.60).

Merging helped. The same model trained on CoLA alone reached MCC 0.576; adding BLiMP raised it to 0.601 while BLiMP itself stayed at 1.000 — positive transfer, not interference.

Baselines, for scale

method accuracy MCC
majority class (CoLA) 0.691 0.000
bag-of-words logistic regression (CoLA) 0.718 0.092
pronoun-only rule (BLiMP) 0.688
zero-shot bert-base-uncased, no fine-tuning (BLiMP) 0.973

That last row is worth dwelling on: masking the pronoun and asking the raw pretrained model which word it prefers already solves BLiMP at 97.3%. Fine-tuning on BLiMP mostly attaches an output head to knowledge the model already had. CoLA is where fine-tuning does real work.

Training procedure

hyperparameter value
base model bert-base-uncased (109.5M parameters)
epochs 4
learning rate 2e-5
warmup ratio 0.06
batch size 32
max sequence length 64
weight decay 0.01
optimizer AdamW
seed 42
best checkpoint by validation MCC (not accuracy — the data is imbalanced)

Per-epoch validation MCC: 0.727 → 0.735 → 0.771 → 0.775.

Limitations

Measured on 40 hand-written test sentences (33/40 correct, 82.5%), the failure modes are systematic rather than random:

  1. Blind to omissions. "Although it was raining, we decided go for a walk." is judged correct. Both training sets create errors by substituting a word, never deleting one, so the model never learned to notice something missing.
  2. Over-flags correct sentences. "He is an honest man." and "She arrived at the airport." are both flagged as errors. Recall on CoLA is 0.590 while precision is 0.833 — it misses more errors than it invents, but its false alarms land on perfectly ordinary sentences.
  3. Untrained phenomena fail. Determiner–noun agreement ("Raymond is selling this sketch.") is flagged as an error. On BLiMP's determiner_noun_agreement_1 — a phenomenon never seen in training — the model scores 0.675, far below its 1.000 on trained phenomena.
  4. Confidence is not reliability. Several wrong predictions are made at 100% confidence. Do not treat the softmax score as a calibrated probability.
  5. English only, and short sentences only — training data averaged well under 20 words.

This is a coursework model built to study what fine-tuning contributes, not a production grammar checker.

Intended use

Educational and research use: demonstrating grammatical acceptability classification, and comparing fine-tuned versus zero-shot versus from-scratch transformers. Not suitable for grading student writing, automated proofreading, or any decision affecting a person.

Citation

@misc{bert-grammar-error-detection,
  title  = {BERT for Grammatical Error Detection (BLiMP + CoLA)},
  author = {Your Real Name},
  year   = {2026},
  url    = {https://huggingface.co/yuvstk/bert-grammar-error-detection}
}

Datasets: BLiMP (Warstadt et al., TACL 2020) and CoLA (Warstadt et al., TACL 2019).

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yuvstk/bert-grammar-error-detection

Finetuned
(6868)
this model

Datasets used to train yuvstk/bert-grammar-error-detection