AraSpellX

Arabic spelling and OCR error correction with a small character-level transformer trained from scratch.

AraSpellX reads Arabic text, typed by people or produced by OCR, and predicts for every character whether to keep it, delete it, replace it or insert something after it. The AraSpellX package turns these predictions into the corrected text and a list of corrections, each with its position in the input, a calibrated confidence and an error category, so that a pipeline can apply only confident edits, show suggestions or flag words.

Task Arabic spelling and OCR error correction, as character-level edit tagging
Architecture BERT encoder: 8 layers, hidden size 384, 6 heads, feed-forward 1536, relative positions (relative_key), 512 positions
Parameters 15.1M
Input โ†’ output characters (210-token vocabulary) โ†’ one of 178 edit labels per character
Training from scratch: masked-character pretraining, then correction training
Language Arabic: modern standard and classical text are corrected; dialect and diacritized text are left alone
Model class standard BertForTokenClassification (loads in plain transformers)
License MIT
Code github.com/mahmoudalrefaey/AraSpellX (branch v1)

Example

Output of this model for a typed sentence:

ุฐู‡ุจุช ุงู„ูŠ ุงู„ุฌุงู…ุนู‡ ุตุจุงุญุง ู„ูƒูŠ ุงุญุถุฑ ุงู„ู…ุญุงุถุฑู‡ ุงู„ุงูˆู„ูŠุŒ ุซู… ู‚ุงุจู„ุช ุตุฏูŠู‚ูŠ ููŠ ุงู„ู…ูƒุชุจู‡.
ุฐู‡ุจุช ุฅู„ู‰ ุงู„ุฌุงู…ุนุฉ ุตุจุงุญุง ู„ูƒูŠ ุฃุญุถุฑ ุงู„ู…ุญุงุถุฑุฉ ุงู„ุฃูˆู„ู‰ุŒ ุซู… ู‚ุงุจู„ุช ุตุฏูŠู‚ูŠ ููŠ ุงู„ู…ูƒุชุจุฉ.

  ุงู„ูŠ -> ุฅู„ู‰            98%   mixed
  ุงู„ุฌุงู…ุนู‡ -> ุงู„ุฌุงู…ุนุฉ     100%   ta_marbuta
  ุงุญุถุฑ -> ุฃุญุถุฑ          94%   hamza_alef
  ุงู„ู…ุญุงุถุฑู‡ -> ุงู„ู…ุญุงุถุฑุฉ   100%   ta_marbuta
  ุงู„ุงูˆู„ูŠุŒ -> ุงู„ุฃูˆู„ู‰ุŒ     99%   mixed
  ุงู„ู…ูƒุชุจู‡. -> ุงู„ู…ูƒุชุจุฉ.   100%   ta_marbuta

On OCR output it corrects what it is sure about and leaves the rest: in ุชุนุชุจุฑ ุงู„ู…ุฏูŠู†ู‡ ู…ู† ุฃู‡ู… ุงู„ู…ุฑุงูƒุฒ ุงู„ู†ุฌุงุฑูŠุฉ ููŠ ุงู„ู…ู†ุทูุฉุŒ ุญูŠุซ ูŠู‚ุตุฏู‡ุง ุงู„ุฒูˆุงุฑ ู…ู† ุญู…ูŠุน ุฃู†ุญุงุก ุงู„ุนุงู„ู…. it fixes ุงู„ู…ุฏูŠู†ู‡ โ†’ ุงู„ู…ุฏูŠู†ุฉ and ุญู…ูŠุน โ†’ ุฌู…ูŠุน and leaves the garbled ุงู„ู†ุฌุงุฑูŠุฉ and ุงู„ู…ู†ุทูุฉ as they are.

Intended use

  • Search and retrieval-augmented generation: clean Arabic text before indexing it, so that misspelled words match their correct forms.
  • OCR post-processing of printed Arabic: fix common OCR confusions without risking the text that was read correctly.
  • Text review: suggest or flag likely spelling errors using each correction's confidence and position.

Out of scope: grammar (agreement, case endings), punctuation and style, converting dialect to standard spelling, adding or fixing diacritics, handwriting and layout. The target spelling is standard modern Arabic orthography, as in edited publications and Arabic Wikipedia.

How to use

With the AraSpellX package (recommended)

The package normalizes the input, reads long text in overlapping windows, applies the confidence threshold and the editing rules, maps every correction back to the input, and calibrates the confidences.

git clone -b v1 https://github.com/mahmoudalrefaey/AraSpellX.git
cd AraSpellX
uv sync
from araspellx.correct.corrector import Corrector

corrector = Corrector("mahmoudalrefaey/AraSpellX")              # downloads this repository once
result = corrector.correct("ุฐู‡ุจุช ุงู„ูŠ ุงู„ุฌุงู…ุนู‡", source="typed")    # source="ocr" for OCR output

print(result.text)   # ุฐู‡ุจุช ุฅู„ู‰ ุงู„ุฌุงู…ุนุฉ
for c in result.corrections:
    print(c.start, c.end, c.original, c.replacement, round(c.confidence, 2), c.category)

A local web page for trying the model (right-to-left, nothing leaves the computer):

python -m araspellx.correct.demo --model mahmoudalrefaey/AraSpellX

With transformers only

The model is a standard BertForTokenClassification and predicts one label per character:

import torch
from transformers import AutoModelForTokenClassification, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("mahmoudalrefaey/AraSpellX")
model = AutoModelForTokenClassification.from_pretrained("mahmoudalrefaey/AraSpellX").eval()

text = "ุฐู‡ุจุช ุงู„ูŠ ุงู„ุฌุงู…ุนู‡"
inputs = tokenizer(text, return_tensors="pt")   # one token per character, plus [CLS] and [SEP]
with torch.no_grad():
    probs = model(**inputs).logits.softmax(-1)[0]
for char, p in zip(text, probs[1:-1]):
    label = model.config.id2label[p.argmax().item()]
    if label != "K":
        print(char, label, round(p.max().item(), 2))   # ุง R:ุฅ 0.96 ยท ูŠ R:ู‰ 0.95 ยท ู‡ R:ุฉ 0.99

Labels: K keep, D delete, R:x replace with x, and โ€ฆ+s insert s after the character; the [CLS] position carries insertions before the first character. This raw output is not yet a corrected text: use the package, or reproduce its normalization, windowing, threshold and editing rules.

Outputs of the package

Field Meaning
Result.text the corrected text (normalized input with the corrections applied)
Correction.start, Correction.end span of the original word in the input
Correction.original, Correction.replacement the word before and after
Correction.confidence calibrated probability that the correction is right
Correction.category hamza_alef, hamza_seat, ta_marbuta, alef_maqsura, alef_fariqa, dots, spacing, typo or mixed

How it works

  • Normalization: look-alike Persian and Urdu letters are folded into Arabic ones, presentation-form ligatures expanded, letters stored as a base letter plus a combining hamza or madda composed, and tatweel and invisible characters removed. Every character keeps a link to its span in the input.
  • Edit labels: one label per character, from the 178 most frequent edits of the training data (99.5% of all edits). Correct text stays correct by default, and every change is an explicit decision with a probability.
  • Editing rules: only Arabic letters, diacritics, tatweel and spaces are edited. Digits, Latin text, symbols and existing punctuation are never changed; a letter that OCR made of a punctuation mark may be turned back into it. Spacing next to punctuation follows Arabic typography.
  • Threshold and calibration: an edit is applied only above a confidence of 0.9, chosen on development data to keep damage on clean text within 0.05%. calibration.json maps raw confidences to observed accuracy, with one fit for typed and one for OCR input.

Training data

All training and test data are public. Training text is filtered against every test set (any paragraph sharing an 8-word passage with a test text is dropped), and test pages are held out by a hash of their page id.

Source Role Size License
Arabic, Egyptian and Moroccan Wikipedia, Arabic Wikisource (2026-10-01 dumps) pretraining text; clean and noisy correction text 2.09B characters CC BY-SA
Typed-error noise generated on the fly dropped hamza, ุฉ/ู‡, ู‰/ูŠ, hamza seats, ุธ/ุถ, dropped alef after waw, keyboard typos, merged and split words โ€” โ€”
Wikipedia (70%) and Wikisource (30%) paragraphs rendered in five open fonts, degraded like scans and read by Tesseract OCR errors 93,499 pairs CC BY-SA text
Yarmouk Arabic OCR Dataset: real scans, read by Tesseract and by ABBYY real OCR errors 104,553 pairs listed as CC0; the text is Wikipedia's
Arabic Wikipedia edit history real spelling fixes, learned only on the words the editor changed 253,294 pairs CC BY-SA

Training procedure

Pretraining Correction
Objective restore 15% hidden characters (whole words and 1โ€“5 character spans) one edit label per character
Steps 80,000 ร— 16,384 characters (128-character windows, then 512) 30,000 ร— 32 windows of 512 characters (best checkpoint: step 28,000)
Mixture โ€” clean 35%, typed noise 25%, OCR 30%, real edits 10%
Optimizer AdamW (ฮฒ 0.9/0.98), learning rate 5e-4, cosine decay AdamW (ฮฒ 0.9/0.98), learning rate 2e-4, cosine decay
Hardware and time one RTX 3060 Laptop GPU (6 GB), bf16, about 8 hours same GPU, about 4 hours

During both stages each attention head's largest score is capped at 50 by scaling its query; this prevents a runaway head that collapsed an earlier pretraining run, and leaves the saved model a standard BERT.

Evaluation

Measured on frozen test sets of real text at the threshold of 0.9:

Test set Result
T-1: 7,627 Wikipedia paragraphs before and after real spelling fixes precision 0.962 on the words editors fixed, recall 0.082; elsewhere 3.4 edits per 1,000 words, 76% of them fixes editors made on other pages
Benchmark: 27 hand-corrected typed sentences precision 1.000, recall 0.900, word error rate 38.4% โ†’ 3.8%
T-7: clean text, correct words changed modern 0.058%, classical 0.010%, diacritized 0.020%, dialect 0.007%
T-4: real Yarmouk scans, word error rate 24.0% โ†’ 20.8% (โˆ’13.2%); ABBYY output โˆ’17.2%, Tesseract output โˆ’7.8%
T-4: pages no worse than the raw OCR 99.7% (589 of 591)
T-4: harmful edits (a correct word changed, or a wrong word not brought closer) 5.1% of the model's edits
T-6: classical books (OpenITI), error rates CER 12.6% โ†’ 12.3%, WER 43.3% โ†’ 41.9%
Calibration (T-4 and T-6) expected calibration error 0.048 (0.085 before calibration)
Speed, laptop CPU (Intel i5-10500H, 4 threads) 410 words/s (ONNX int8), 337 (ONNX fp32); with PyTorch about 20 ms per sentence

T-1 is scored on the words the editors fixed, because a single Wikipedia edit leaves a paragraph's other errors in place; strict precision against such references counts the model's fixes of those errors as damage. References in general contain spelling errors, so some edits counted as harmful are fixes of the reference itself. Full method and per-set results: evaluation.

Limitations

  • OCR recall is low. On real scans the model rarely makes a page worse, but it removes only 13% of word errors (the project's target is 20%): most OCR errors are badly garbled words, and it fixes about 1 in 10 of them.
  • Errors that produce another valid word need meaning rather than spelling, which a 15M-parameter character model knows little about.
  • Quranic text and accepted spelling variants are not protected: verses are corrected like any other text, and either form of a word with two accepted spellings (ู…ุณุคูˆู„/ู…ุณุฆูˆู„) may be changed.
  • Dialect and diacritized text are left alone, not corrected.
  • The training text is encyclopedic (Wikipedia) and classical (Wikisource); informal registers such as social media were not evaluated.

Files

File Content
config.json, model.safetensors the model (BertForTokenClassification, 178 labels)
tokenizer.json, tokenizer_config.json, special_tokens_map.json the character tokenizer
labels.json the edit labels, in the order of the classifier
calibration.json the threshold and the confidence calibration for typed and OCR input

License and attribution

The model and the code are released under the MIT License. The training text comes from Wikimedia projects (CC BY-SA) and from the Yarmouk Arabic OCR Dataset (Abu Doush, AlKhateeb and Gharibeh, CSIT 2018, doi:10.1109/CSIT.2018.8486162); test data also include NOD (CC BY 4.0) and the OpenITI OCR gold standard (CC BY-NC-SA 4.0, evaluation only). OCR training pairs were produced with Tesseract (Apache-2.0). The project started from AraSpell (Salhab and Abu-Khzam); the edit-label approach follows GECToR and Alhafni and Habash's Arabic text editing.

Citation

@software{alrefaey2026araspellx,
  author = {Al-Refaey, Mahmoud},
  title  = {{AraSpellX}: Arabic Spelling and {OCR} Error Correction with a Character-Level Transformer},
  year   = {2026},
  url    = {https://huggingface.co/mahmoudalrefaey/AraSpellX}
}
Downloads last month
22
Safetensors
Model size
15.1M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Papers for mahmoudalrefaey/AraSpellX

Evaluation results

  • Precision on the words editors fixed on T-1: real spelling fixes from Arabic Wikipedia's edit history
    self-reported
    0.962
  • Recall on the words editors fixed on T-1: real spelling fixes from Arabic Wikipedia's edit history
    self-reported
    0.082
  • Word error rate after correction, % (24.0 before) on T-4: Yarmouk real scans (Tesseract and ABBYY output)
    self-reported
    20.800