Egyptian Document OCR (Qwen3-VL-8B)

A LoRA adapter fine-tuned on top of unsloth/Qwen3-VL-8B-Instruct-unsloth-bnb-4bit for full-page OCR of scanned Egyptian government documents (Official Gazette / الجريدة الرسمية — tax laws and Ministry of Finance decisions), in Arabic.

This is a larger sibling to mohamedwasef/egyptian-document-ocr (Qwen3-VL-4B-Instruct) and mohamedwasef/egyptian-document-ocr-gemma4 (Gemma 4 E4B), fine-tuned on the exact same dataset and train/validation split for a direct size/architecture comparison. This 8B version is the strongest of the three.

This is a LoRA adapter, not a merged standalone model - load it on top of the base model (see Usage below).

Example

Input image:

example page

Model output:

الجريدة الرسمية - العدد ٢٦ مكرر (أ) فى ٣٠ يونية سنة ٢٠١٤
٨
كما يُضاف إلى المادة (١٩) من القانون المشار إليه فقرة ثانية، نصحها الآتى:
"كما تسرى الضريبة على الأرباح الناتجة عن الاستثمار فى الأوراق المالية فى الخارج
أو التصرف فيها".
ويضاف إلى المادة (٥٩) من ذات القانون فقرة ثالثة، نصحها الآتى:
"وفى جميع الأحوال تلتزم الجهات والمنشآت المنصوص عليها فى البندين (١، ٢) من الفقرة الأولى من هذه المادة بأن تخطر المصلحة ببيان التعاملات والمبالغ المدفوعة لأى شخص من أشخاص القطاع الخاص إذا زادت قيمة التعامل خلال كل فترة ربع سنوية على ثلاثمائة جنيه، وذلك فى موعد أقصاه أوآخر أبريل ويوليو وأكتوبر ويناير من كل عام عن المعاملات خلال الأشهر السابقة، وذلك طبقاً للإجراءات التى تحددها الثلاثة التنفيذية للقانون".
(المادة الثالثة)
يُضاف إلى قانون الضريبة على الدخل المشار إليه مادة جديدة برقم (٢٩ مكرراً)، وبابان جديدين للكتاب الثانى "الباب السادس - توزيعات الأرباح" ويتكون من ثلاث مواد
(٤٦ مكرراً)، (٤٦ مكرراً ١)، (٤٦ مكرراً ٢) "والباب السابع - أرباح بيع المحصص
أو الأوراق المالية" ويتكون من أربعة مواد (٤٦ مكرراً ٣)، (٤٦ مكرراً ٤)، (٤٦ مكرراً ٥)، (٤٦ مكرراً ٦)، كما يضاف إلى ذلك القانون أربع مواد جديدة أرقام (٤٩ مكرراً)، (٥٦ مكرراً)، (٩٢ مكرراً)، و(١٣٥ مكرراً) نصوصها كالآتى.
المادة (٢٩ مكرراً):
"استثناءً من حكم المادة (٢٩) من هذا القانون تخصم الخسائر الرأسمالية المحققة نتيجة التصرف فى الأوراق المالية فى حدود الأرباح الرأسمالية المحققة من التصرف فى أوراق مالية خلال السنة الضريبية ذاتها.
وفى حالة زيادة الخسائر الرأسمالية المحققة وفقاً لأحكام الفقرة السابقة من هذه المادة عن الأرباح الرأسمالية المحققة خلال السنة الضريبية يسمح بترحيل الزيادة فى الخسائر من الأرباح المحققة نتيجة التصرف فى الأوراق المالية فى السنوات التالية حتى السنة الثالثة".

Ground truth (for comparison):

الجريدة الرسمية - العدد ٢٦ مكرر (أ) فى ٣٠ يونية سنة ٢٠١٤
٨
كما يُضاف إلى المادة (١٩) من القانون المشار إليه فقرة ثانية، نصها الآتى:
"كما تسرى الضريبة على الأرباح الناتجة عن الاستثمار فى الأوراق المالية فى الخارج أو التصرف فيها".
ويضاف إلى المادة (٥٩) من ذات القانون فقرة ثالثة، نصها الآتى:
"وفى جميع الأحوال تلتزم الجهات والمنشآت المنصوص عليها فى البندين (١، ٢) من الفقرة الأولى من هذه المادة بأن تخطر المصلحة ببيان التعاملات والمبالغ المدفوعة لأى شخص من أشخاص القطاع الخاص إذا زادت قيمة التعامل خلال كل فترة ربع سنوية على ثلاثمائة جنيه، وذلك فى موعد أقصاه أواخر أبريل ويوليو وأكتوبر ويناير من كل عام عن المعاملات خلال الأشهر السابقة، وذلك طبقًا للإجراءات التى تحددها اللائحة التنفيذية للقانون".
(المادة الثالثة)
يُضاف إلى قانون الضريبة على الدخل المشار إليه مادة جديدة برقم (٢٩ مكرراً)، وبابان جديدان للكتاب الثانى "الباب السادس - توزيعات الأرباح" ويتكون من ثلاث مواد (٤٦ مكرراً)، (٤٦ مكرراً ١)، (٤٦ مكرراً ٢) "والباب السابع - أرباح بيع الحصص أو الأوراق المالية" ويتكون من أربعة مواد (٤٦ مكرراً ٣)، (٤٦ مكرراً ٤)، (٤٦ مكرراً ٥)، (٤٦ مكرراً ٦)، كما يضاف إلى ذلك القانون أربع مواد جديدة أرقام (٤٩ مكرراً)، (٥٦ مكرراً)، (٩٢ مكرراً)، و(١٣٥ مكرراً) نصوصها كالآتى.
المادة (٢٩ مكرراً):
"استثناءً من حكم المادة (٢٩) من هذا القانون تخصم الخسائر الرأسمالية المحققة نتيجة التصرف فى الأوراق المالية فى حدود الأرباح الرأسمالية المحققة من التصرف فى أوراق مالية خلال السنة الضريبية ذاتها.
وفى حالة زيادة الخسائر الرأسمالية المحققة وفقًا لأحكام الفقرة السابقة من هذه المادة عن الأرباح الرأسمالية المحققة خلال السنة الضريبية يسمح بترحيل الزيادة فى الخسائر من الأرباح المحققة نتيجة التصرف فى الأوراق المالية فى السنوات التالية حتى السنة الثالثة".

This example scores CER 0.60% / WER 2.19% (the median result among non-table test pages). Two test pages scored a perfect CER=0.00% / WER=0.00% (exact character-for-character match): law_201_2014_p001.png, law_53_2014_p008.png.

Training data

Fine-tuned on mohamedwasef/egyptian-official-documents-ocr, 280 pages for training, 56 held out for validation (identical split used for the 4B and Gemma 4 companion models).

Training details

  • Base model: unsloth/Qwen3-VL-8B-Instruct-unsloth-bnb-4bit
  • Method: LoRA (r=16, alpha=16), language-model layers only, vision encoder frozen
  • Trainable parameters: 43,646,976 (0.50% of 8.81B total)
  • Hardware: single NVIDIA T4 (Kaggle, free tier). Unsloth automatically offloaded the embedding layer to CPU RAM to fit comfortably within the T4's 14.6GB VRAM (peak usage during training: 7.50GB)
  • Precision: fp16
  • Epochs: 3, ~57 minutes total training time - the eval loss was still improving at the final step (no overfitting observed, unlike the Gemma 4 run), so the full 3 epochs were used

Evaluation

Evaluated on the dataset's official 30-page test set (never seen during training or validation), CER/WER computed after normalizing diacritics and unifying ى/ي; digits evaluated separately and exactly.

Subset Pages CER WER
All test pages 30 1.97% 3.76%
Pages without a table 28 0.61% 2.90%
Pages with a table 2 21.01% 15.82%

Digit-run error rate on non-table pages: 1.73%.

Comparison with the companion models (non-table pages)

Model Size CER WER Digit error
Qwen3-VL-4B 4.47B 2.32% 8.09% 11.69%
Gemma 4 E4B 8.03B 2.89% 7.87% 9.98%
Qwen3-VL-8B (this model) 8.81B 0.61% 2.90% 1.73%

This 8B model is a clear step up on every metric, not just a marginal improvement. Even its weakest test pages scored better than the average of the two smaller models. Table performance also improved over both companion models, though it remains the weakest area relative to this model's own plain-text performance.

Known limitations

  • Dataset size. The training set is 280 pages from 19 documents, and only 18 of 336 training pages (5.4%) contained a table - the 2 table pages in the test set (CER 21.0%) remain a small sample and the weakest area of this model's otherwise strong performance.
  • Occasional non-deterministic output. On one development run, a single page produced a garbled word merge (two adjacent short words fused together) that did not reproduce when the same page was re-run as part of the full 30-page batch (which scored near-perfect on that page). This is consistent with ordinary floating-point non-determinism in GPU inference rather than a systematic model weakness, but it means greedy decoding on this setup is not perfectly reproducible run-to-run.
  • Labels in the training dataset are LLM-generated with spot checks, not fully human-verified - see the dataset card for details.

Usage

import torch
from unsloth import FastVisionModel
from PIL import Image

BASE_MODEL = "unsloth/Qwen3-VL-8B-Instruct-unsloth-bnb-4bit"
ADAPTER = "mohamedwasef/Docora"

model, processor = FastVisionModel.from_pretrained(
    ADAPTER,
    load_in_4bit=True,
    max_seq_length=2048,
    dtype=None,
)
FastVisionModel.for_inference(model)

image = Image.open("your_page.png").convert("RGB")

instruction = (
    "اقرأ هذه الصفحة ونسخ نصها كاملاً بالضبط كما هو مكتوب، "
    "بنفس ترتيب القراءة وبدون أي إضافة أو تلخيص."
)

messages = [
    {"role": "user", "content": [
        {"type": "image", "image": image},
        {"type": "text", "text": instruction},
    ]}
]

input_text = processor.apply_chat_template(messages, add_generation_prompt=True)
inputs = processor(
    image, input_text, add_special_tokens=False, return_tensors="pt"
).to("cuda")

with torch.inference_mode():
    output_ids = model.generate(
        **inputs,
        max_new_tokens=1024,
        use_cache=True,
        temperature=0.0,
        do_sample=False,
    )

generated_ids = output_ids[:, inputs["input_ids"].shape[1]:]
prediction = processor.batch_decode(generated_ids, skip_special_tokens=True)[0].strip()
print(prediction)

Note: this model needs roughly 7-8GB VRAM for inference on a T4 (Unsloth automatically offloads the embedding layer to CPU RAM to fit comfortably).

License

Apache 2.0, matching the base model. The training data annotations are CC BY 4.0 (see the dataset card); source documents are official Egyptian government publications.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mohamedwasef/Docora