Abjad 0.2

Abjad 0.2 is a standalone MLX-VLM model for Arabic OCR and text extraction from images and scanned documents. It is the Qwen3-VL 4B Instruct 4-bit base model with the final Abjad LoRA training fused into the model weights.

The training mix included reviewed Arabic OCR examples, Arabic document replay examples, and a conservative Quran auxiliary set for Arabic letter and diacritic recognition. The Quran auxiliary set retained only exact normalized matches against the Tanzil Quran text reference; its manifest is included in the project training artifacts.

Training data summary

The current Abjad 0.2 training lineage contains 8,111 cumulative training rows across sequential fine-tuning stages:

  • 3,912 initial Arabic OCR rows
  • 3,005 selected Arabic OCR/title and replay rows
  • 6 manually corrected page-12 rows
  • 1,188 auxiliary rows for Arabic letters and diacritics (588 verified Quran rows plus 600 general OCR replay rows)

Local MLX usage

git lfs install
git clone https://huggingface.co/Alsamir/Abjad_0.2
cd Abjad_0.2

python /path/to/Abjad_0.1/mlx_ocr_pipeline.py input.pdf \
  -o output.txt \
  --model . \
  --no-adapter

For scanned pages, use a prompt that requests exact transcription, preserves right-to-left Arabic, keeps Arabic-Indic digits, excludes watermarks, and returns only visible document text. Tables may require a layout-aware post-processing step.

License and attribution

The base model remains subject to the Qwen model license and its upstream terms. Users must also comply with the licenses of the training sources. The Quran reference text used for the auxiliary data was obtained from Tanzil and should retain the required attribution when redistributed.

Repository guide

This model repository follows a simple layout so it can be used both as a downloadable checkpoint and as a small reproducible OCR project:

Abjad_0.2/
├── model.safetensors       # Fused standalone MLX weights
├── config.json              # Model configuration
├── tokenizer*               # Tokenizer files
├── preprocessor_config.json # Image processor configuration
├── fused_manifest.json      # Fusion and provenance summary
├── docs/                    # Inference, data, training, and evaluation guides
├── configs/                 # Portable configuration examples
└── scripts/                 # Convenience entry points

The model files at the repository root are unchanged. The documentation and helper files are additive and do not replace the fused weights or require an adapter for inference.

Recommended inference settings

For scanned Arabic pages, use a deterministic prompt that requests exact visible transcription, preserves right-to- left order and printed digit shapes, ignores watermarks and signatures, and forbids repetition. For single-column scans, horizontal regions can improve reading order. Tables and legal documents should always be reviewed against the source image.

See docs/inference.md, docs/data_format.md, docs/training.md, and docs/evaluation.md.

Downloads last month
89
Safetensors
Model size
4B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Alsamir/Abjad_0.2

Quantized
(1)
this model