Egyptian Document OCR

A LoRA adapter fine-tuned on top of unsloth/Qwen3-VL-4B-Instruct-unsloth-bnb-4bit for full-page OCR of scanned Egyptian government documents (Official Gazette / الجريدة الرسمية — tax laws and Ministry of Finance decisions), in Arabic.

The model reads a full page image and transcribes its text, preserving the original reading order, punctuation and spacing conventions, and rendering tables as inline HTML - matching the exact transcription conventions of the training dataset (see below).

This is a LoRA adapter, not a merged standalone model. It must be loaded on top of the base model (see Usage below). This keeps the download small (tens of MB instead of several GB) and makes clear this is a fine-tune, not a repackaged copy of the base model.

Example

Input image:

example page

Model output:

الجريدة الرسمية - العدد ٢٦ مكرر (أ) فى ٣ يونية سنة ٢٠١٤
٥
المادة (٥٩/ مكرر):
"على الجهات المنصوص عليها فى البند (١) من الفقرة الأولى من المادة (٥٩) من هذا القانون التى تتولى بيع أو توزيع أى سلع أو منتجات صناعية أو حاصلات زراعية محلية أو مستوردة إلى أشخاص القطاع الخاص للإتجار فيها أو تصنيعها أن تخطر المصلحة ببيان عن التعاملات والبالغ التى تحصل عليها من هؤلاء الأشخاص.
المادة (٥٩/ مكرر ٢٦٦):
"على الجهات المنصوص عليها فى البندين (١)، (٢) من الفقرة الأولى من المادة (٥٩) من هذا القانون، أن تخطر المصلحة ببيان عن التعاملات والبالغ والإيجارات التى تحصلها من المستأجرين للأماكن المملوكة لها والمعدة للإتجار أو التصنيع فيها أو تقديم أو إعداد أية خدمات أو مأكولات أو مشروبات."
المادة (٥٩/ مكرر ٢٦):
"تحدد بقرار من الوزير السلع والمنتجات الصناعية والحاصلات الزراعية وأوجه النشاط وأنواع الإيجارات التى تسرى عليها أحكام المادتين (٥٩ مكرر)، (٥٩ مكرر أ) من هذا القانون، وعلى الجهات والمنشآت المشار إليها فى البندين (١)، (٢) من الفقرة الأولى من المادة (٥٩) من هذا القانون إخطار المصلحة ببيان بقيمة السلع والمنتجات الصناعية والحاصلات الزراعية والتعاملات والبالغ والإيجارات التى حصلت عليها من كل ممول فى موعد أقصاه أواخر أبريل ويوليو وأكتوبر ويناير من كل عام عن الثلاثة أشهر السابقة، وذلك طبقًا لإجراءات التى تحددها اللائحة التنفيذية."
المادة (٧٢):
"تلتزم الجهات المنصوص عليها فى المواد (٦٦)، (٦٧)، (٦٨)، (٦٩)، (٧٠)، (٧١)، من هذا القانون، بتوريد قيمة ما حصلته تحت حساب الضريبة إلى المصلحة، وذلك طبقًا للإجراءات وفى المواعيد التى تحددها اللائحة التنفيذية."
وفى حالة عدم خصم أو توريد المبالغ الواجب خصمها تلتزم الجهة بأن تؤدى للمصلحة هذه المبالغ بالإضافة إلى ما يستحق عليها من مقابل تأخير"

Ground truth (for comparison):

الجريدة الرسمية - العدد ٢٦ مكرر (أ) فى ٣٠ يونية سنة ٢٠١٤
٥
المادة (٥٩ / مكرراً):
"على الجهات المنصوص عليها فى البند (١) من الفقرة الأولى من المادة (٥٩) من هذا القانون التى تتولى بيع أو توزيع أى سلع أو منتجات صناعية أو حاصلات زراعية محلية أو مستوردة إلى أشخاص القطاع الخاص للاتجار فيها أو تصنيعها أن تخطر المصلحة ببيان عن التعاملات والمبالغ التى تحصل عليها من هؤلاء الأشخاص".
المادة (٥٩ / مكرراً "١"):
"على الجهات المنصوص عليها فى البندين (١)، (٢) من الفقرة الأولى من المادة (٥٩) من هذا القانون، أن تخطر المصلحة ببيان عن التعاملات والمبالغ والإيجارات التى تحصلها من المستأجرين للأماكن المملوكة لها والمعدة للاتجار أو التصنيع فيها أو تقديم أو إعداد أية خدمات أو مأكولات أو مشروبات".
المادة (/٥٩ مكرراً "٢"):
"تحدد بقرار من الوزير السلع والمنتجات الصناعية والحاصلات الزراعية وأوجه النشاط وأنواع الإيجارات التى تسرى عليها أحكام المادتين (٥٩ مكرراً)، (٥٩ مكرراً ١) من هذا القانون، وعلى الجهات والمنشآت المشار إليها فى البندين (١)، (٢) من الفقرة الأولى من المادة (٥٩) من هذا القانون إخطار المصلحة ببيان بقيمة السلع والمنتجات الصناعية والحاصلات الزراعية والتعاملات والمبالغ والإيجارات التى حصلت عليها من كل ممول فى موعد أقصاه أواخر أبريل ويوليو وأكتوبر ويناير من كل عام عن كل الثلاثة أشهر السابقة، وذلك طبقًا للإجراءات التى تحددها اللائحة التنفيذية".
المادة (٧٢):
"تلتزم الجهات المنصوص عليها فى المواد (٦٦)، (٦٧)، (٦٨)، (٦٩)، (٧٠)، (٧١)، من هذا القانون، بتوريد قيمة ما حصلته تحت حساب الضريبة إلى المصلحة، وذلك طبقًا للإجراءات وفى المواعيد التى تحددها اللائحة التنفيذية".
وفى حالة عدم خصم أو توريد المبالغ الواجب خصمها تلتزم الجهة بأن تؤدى للمصلحة هذه المبالغ بالإضافة إلى ما يستحق عليها من مقابل تأخير"

This example scores CER 2.20% / WER 9.26% (the median result among non-table test pages - a representative example, not the single best result).

Training data

Fine-tuned on mohamedwasef/egyptian-official-documents-ocr, a page-level OCR dataset of scanned Egyptian tax legislation and ministerial decisions, transcribed to preserve exact printed spelling, spacing, and layout. 280 pages used for training, 56 held out for validation.

Training details

  • Base model: unsloth/Qwen3-VL-4B-Instruct-unsloth-bnb-4bit
  • Method: LoRA (r=16, alpha=16), language-model layers only - the vision encoder was kept frozen
  • Trainable parameters: 33,030,144 (0.74% of the 4.47B total)
  • Hardware: single NVIDIA T4 (Kaggle, free tier)
  • Epochs: 3, ~37 minutes total training time
  • Precision: fp16 (T4 has no native bf16 support)
  • Training target: the page's plain transcribed text (no added structural markers) - section/document metadata and any downstream chunking are built separately from this model's output, not predicted by the model itself

Evaluation

Evaluated on the dataset's official 30-page test set (documents never seen during training or validation), using CER/WER after normalizing diacritics and unifying ى/ي (so formatting variation isn't counted as a reading error; digits are evaluated separately and exactly, with no leniency).

Subset Pages CER WER
All test pages 30 3.82% 9.39%
Pages without a table 28 2.33% 8.09%
Pages with a table 2 24.69% 27.56%

Digit-run error rate on non-table pages: 11.69% (digits are harder for the model than surrounding text - see Limitations).

Known limitations

  • Tables. Only 18 of 336 training pages (5.4%) contained a table, and CER on table pages rises to roughly 10x the non-table rate. The model can sometimes lose the <table> HTML structure or misorder columns/rows on dense tables. Treat table output as unreliable without manual review.
  • Digits. Digit accuracy is measurably weaker than surrounding text accuracy, even on non-table pages. Numeric values (amounts, dates, legal reference numbers) should be spot-checked rather than trusted blindly.
  • Low-resolution scans. The training data includes scans below 1200px wide; performance on very low-quality scans has not been isolated in this evaluation (the official test set only contains higher-quality scans) and may be weaker than the numbers above suggest.
  • Labels in the training dataset are LLM-generated with spot checks, not fully human-verified - see the dataset card for details.

Usage

import torch
from unsloth import FastVisionModel
from PIL import Image

BASE_MODEL = "unsloth/Qwen3-VL-4B-Instruct-unsloth-bnb-4bit"
ADAPTER = "mohamedwasef/egyptian-document-ocr"

# 1. Load the base model, then attach the adapter
model, processor = FastVisionModel.from_pretrained(
    ADAPTER,  # loading the adapter repo directly also pulls in the base model
    load_in_4bit=True,
    max_seq_length=2048,
    dtype=None,
)
FastVisionModel.for_inference(model)

# 2. Load your page image
image = Image.open("your_page.png").convert("RGB")

# 3. Use the SAME instruction the model was trained with
instruction = (
    "اقرأ هذه الصفحة ونسخ نصها كاملاً بالضبط كما هو مكتوب، "
    "بنفس ترتيب القراءة وبدون أي إضافة أو تلخيص."
)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": image},
            {"type": "text", "text": instruction},
        ],
    }
]

input_text = processor.apply_chat_template(messages, add_generation_prompt=True)
inputs = processor(
    image, input_text, add_special_tokens=False, return_tensors="pt"
).to("cuda")

with torch.inference_mode():
    output_ids = model.generate(
        **inputs,
        max_new_tokens=1024,
        use_cache=True,
        temperature=0.0,
        do_sample=False,
    )

generated_ids = output_ids[:, inputs["input_ids"].shape[1]:]
prediction = processor.batch_decode(generated_ids, skip_special_tokens=True)[0].strip()
print(prediction)

Alternative loading without Unsloth (plain transformers + peft):

import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
from peft import PeftModel
from PIL import Image

base_model = AutoModelForImageTextToText.from_pretrained(
    "unsloth/Qwen3-VL-4B-Instruct-unsloth-bnb-4bit", torch_dtype=torch.float16, device_map="auto"
)
model = PeftModel.from_pretrained(base_model, "mohamedwasef/egyptian-document-ocr")
processor = AutoProcessor.from_pretrained("mohamedwasef/egyptian-document-ocr")

image = Image.open("your_page.png").convert("RGB")
instruction = (
    "اقرأ هذه الصفحة ونسخ نصها كاملاً بالضبط كما هو مكتوب، "
    "بنفس ترتيب القراءة وبدون أي إضافة أو تلخيص."
)
messages = [
    {"role": "user", "content": [
        {"type": "image", "image": image},
        {"type": "text", "text": instruction},
    ]}
]
input_text = processor.apply_chat_template(messages, add_generation_prompt=True)
inputs = processor(image, input_text, return_tensors="pt").to(model.device)
output_ids = model.generate(**inputs, max_new_tokens=1024, do_sample=False)
prediction = processor.batch_decode(
    output_ids[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True
)[0].strip()
print(prediction)

License

Apache 2.0, matching the base model. The training data annotations are CC BY 4.0 (see the dataset card); source documents are official Egyptian government publications.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mohamedwasef/Docora-lite