BERT for Grammatical Error Detection (BLiMP + CoLA)
bert-base-uncased fine-tuned for binary grammatical error detection: given
one English sentence, decide whether it contains a grammatical error.
| label | meaning |
|---|---|
0 |
grammatical |
1 |
ungrammatical |
Note the orientation: 1 means "has an error." This is the inverse of CoLA's native convention (where 1 = acceptable), and the training labels were flipped accordingly.
Usage
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
model_id = "yuvstk/bert-grammar-error-detection"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
sentences = ["Katherine can't help himself.", "She went home early."]
inputs = tokenizer(sentences, padding=True, truncation=True,
max_length=64, return_tensors="pt")
with torch.no_grad():
probs = torch.softmax(model(**inputs).logits, dim=-1)
for sentence, prob in zip(sentences, probs):
label = int(prob.argmax())
print(f"{'UNGRAMMATICAL' if label else 'GRAMMATICAL'} "
f"({prob[label]:.1%}) — {sentence}")
Or with a pipeline (LABEL_1 = ungrammatical):
from transformers import pipeline
clf = pipeline("text-classification", model=model_id)
clf("The professor talked us.")
Training data
Two sources merged into a single 4000 + 9594 sentence corpus:
| source | rows (train) | what it contributes |
|---|---|---|
BLiMP — anaphor_gender_agreement + anaphor_number_agreement |
3200 | synthetic minimal pairs; reflexive pronoun agreement; exactly 50/50 balanced |
| CoLA (GLUE) | 7695 | real linguistics-literature sentences; many error types; ~70/30 imbalanced |
Splitting differs per source, because the sources need different treatment:
- BLiMP is split by
pair_id, never by row. The two sentences of a minimal pair differ by exactly one word, so a row-level split would put a near-duplicate of a test sentence into training. - CoLA is split by stratified rows. GLUE's
testsplit is unlabelled (all-1), so GLUEvalidationis used as the test set and the validation set is carved out of GLUEtrain.
Total: 10,895 train / 1,256 validation / 1,443 test.
Results
Evaluated separately per source, because the two halves differ enormously in difficulty — a single pooled number would mostly reflect the mixture ratio.
| test set | n | accuracy | precision | recall | F1 | MCC |
|---|---|---|---|---|---|---|
| BLiMP | 400 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| CoLA | 1043 | 0.837 | 0.833 | 0.590 | 0.691 | 0.601 |
| pooled | 1443 | 0.882 | 0.911 | 0.747 | 0.821 | 0.743 |
CoLA MCC of 0.601 is in the normal published range for BERT-base (~0.55–0.60).
Merging helped. The same model trained on CoLA alone reached MCC 0.576; adding BLiMP raised it to 0.601 while BLiMP itself stayed at 1.000 — positive transfer, not interference.
Baselines, for scale
| method | accuracy | MCC |
|---|---|---|
| majority class (CoLA) | 0.691 | 0.000 |
| bag-of-words logistic regression (CoLA) | 0.718 | 0.092 |
| pronoun-only rule (BLiMP) | 0.688 | — |
zero-shot bert-base-uncased, no fine-tuning (BLiMP) |
0.973 | — |
That last row is worth dwelling on: masking the pronoun and asking the raw pretrained model which word it prefers already solves BLiMP at 97.3%. Fine-tuning on BLiMP mostly attaches an output head to knowledge the model already had. CoLA is where fine-tuning does real work.
Training procedure
| hyperparameter | value |
|---|---|
| base model | bert-base-uncased (109.5M parameters) |
| epochs | 4 |
| learning rate | 2e-5 |
| warmup ratio | 0.06 |
| batch size | 32 |
| max sequence length | 64 |
| weight decay | 0.01 |
| optimizer | AdamW |
| seed | 42 |
| best checkpoint by | validation MCC (not accuracy — the data is imbalanced) |
Per-epoch validation MCC: 0.727 → 0.735 → 0.771 → 0.775.
Limitations
Measured on 40 hand-written test sentences (33/40 correct, 82.5%), the failure modes are systematic rather than random:
- Blind to omissions. "Although it was raining, we decided go for a walk." is judged correct. Both training sets create errors by substituting a word, never deleting one, so the model never learned to notice something missing.
- Over-flags correct sentences. "He is an honest man." and "She arrived at the airport." are both flagged as errors. Recall on CoLA is 0.590 while precision is 0.833 — it misses more errors than it invents, but its false alarms land on perfectly ordinary sentences.
- Untrained phenomena fail. Determiner–noun agreement ("Raymond is
selling this sketch.") is flagged as an error. On BLiMP's
determiner_noun_agreement_1— a phenomenon never seen in training — the model scores 0.675, far below its 1.000 on trained phenomena. - Confidence is not reliability. Several wrong predictions are made at 100% confidence. Do not treat the softmax score as a calibrated probability.
- English only, and short sentences only — training data averaged well under 20 words.
This is a coursework model built to study what fine-tuning contributes, not a production grammar checker.
Intended use
Educational and research use: demonstrating grammatical acceptability classification, and comparing fine-tuned versus zero-shot versus from-scratch transformers. Not suitable for grading student writing, automated proofreading, or any decision affecting a person.
Citation
@misc{bert-grammar-error-detection,
title = {BERT for Grammatical Error Detection (BLiMP + CoLA)},
author = {Your Real Name},
year = {2026},
url = {https://huggingface.co/yuvstk/bert-grammar-error-detection}
}
Datasets: BLiMP (Warstadt et al., TACL 2020) and CoLA (Warstadt et al., TACL 2019).
- Downloads last month
- -
Model tree for yuvstk/bert-grammar-error-detection
Base model
google-bert/bert-base-uncased