Abjad — checkpoint 9750

Standalone BF16 Transformers export of Qwen/Qwen3-VL-4B-Instruct with the Abjad stage-2 LoRA checkpoint-9750 merged into its weights. No separate adapter is required. This uses Transformers rather than the older MLX release format.

Training status

This checkpoint stores step 9,750 of 11,000 planned steps. Training was interrupted; this is not a completed 11,000-step run. No verified total of unique training images or accuracy score is claimed. Earlier-stage lineage has not been independently audited.

Pre-merge and post-merge outputs matched on one dataset image. This checks export behavior only, not OCR accuracy. The image may have been used for training. Repetition, omissions, incorrect letters/digits, and layout errors remain possible. Review output against the source, especially names and numbers.

Usage

Install PyTorch, Accelerate, Pillow, and a Transformers version supporting Qwen3-VL. Exact export versions are recorded in release_manifest.json.

import torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForImageTextToText
repo = "Alsamir/Abjad-checkpoint-9750"
processor = AutoProcessor.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto").eval()
image = Image.open("page.png").convert("RGB")
messages = [{"role":"user", "content":[{"type":"image"}, {"type":"text", "text":"Transcribe all visible Arabic text exactly. Preserve Arabic letters and digits. Return only the transcription."}]}]
prompt = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[prompt], images=[image], return_tensors="pt").to(model.device)
with torch.inference_mode():
    result = model.generate(**inputs, max_new_tokens=2048, do_sample=False)
print(processor.batch_decode(result[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])

Convert PDF pages to images before inference. Long pages may hit the token limit. Word conversion and exact layout reconstruction need a separate document pipeline. Use hardware supporting the selected dtype; CPU inference may be slow.

Contents and provenance

Includes full model shards, tokenizer, processor, configuration and a checksum/provenance manifest. Training documents, optimizer state, credentials and machine-specific paths are excluded. The original checkpoint is retained separately. The base model is licensed under Apache 2.0. No dataset is redistributed.

Windows and OCI

The published repository is a standalone Transformers model. It contains no adapter and no machine-specific paths. Windows users can run it with NVIDIA CUDA, Python 3.10 or 3.11, and a CUDA-enabled PyTorch build. OCI users can use the same files in a CUDA Linux image. A GPU with enough VRAM for the BF16 model is recommended.

py -3.11 -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
# Install the CUDA-enabled PyTorch wheel matching your driver from pytorch.org
pip install -r requirements-windows.txt
python run_ocr_windows.py page.png --output result.txt

The runner selects CUDA when available and falls back to CPU. Rasterize PDFs into PNG or TIFF pages first. Long pages may require cropping. Review names, numbers, tables, stamps, and signatures against the source document.

Downloads last month
17
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Alsamir/Abjad-checkpoint-9750

Finetuned
(456)
this model