Chandra OCR 2 BF16 GGUF — one-GPU RDNA 4 validation

Unofficial, reproducible BF16 GGUF conversion of datalab-to/chandra-ocr-2, validated end to end with llama.cpp on one AMD Radeon RX 9070 XT (gfx1201).

This is a container conversion, not a fine-tune, new model, or claim of improved OCR quality. The main GGUF preserves BF16 precision while omitting Chandra's unused multi-token-prediction (MTP) draft head. The vision projector is BF16. The original model, architecture, training, and authorship belong to Datalab.

Files

File Bytes SHA-256
chandra-ocr-2.BF16.gguf 9,695,791,648 4e9d5fa9854cf820d4425d28034df31ec1221a7f9d1082b0c4359d79f318cb56
chandra-ocr-2.mmproj-bf16.gguf 675,568,864 54ddb8285933512cdbf1c84238aa0435b473a6efef2caeda8ca802c2899e87b3
chat_template.jinja 7,622 0d158f349ca965f7eea9db0eb45cd177b85bb0e4ae05dcdd0f060da8f7d41812

The complete machine-readable provenance is in manifest.json.

Pinned provenance

  • Source model: datalab-to/chandra-ocr-2
  • Source revision: af93b47dba1b47b6640c86ccf487ed2260ab9a09
  • Source model.safetensors SHA-256: 0804568be9f099d6479fad9ed77a4da4611f3c1e7bc6e009af7dce45e8aa3847
  • Converter/runtime: ggml-org/llama.cpp
  • Converter revision: 8f5ab832ca7d8a7b4f23687693fb8b0ecbc227e7
  • chandra-ocr: 0.2.0

See CONVERSION.md for the exact commands and integrity checks.

Verified configuration

Component Verified value
GPU AMD Radeon RX 9070 XT, 16 GB (gfx1201)
GPU allocation One isolated GPU; full layer offload
OS / kernel Ubuntu 24.04.4 / 6.17.0-40-generic
ROCm 7.2.1
llama.cpp 8f5ab832ca7d8a7b4f23687693fb8b0ecbc227e7
Context 24,576 tokens
Maximum tested output envelope 12,384 tokens
Observed model-process VRAM peak Approximately 11.2 GB

The GGUF format is not RDNA4-specific. gfx1201 is the hardware on which this exact pair completed image, PDF, financial-table, handwritten-document, deterministic-output, and lifecycle tests. Other llama.cpp-supported hardware may work but is not validated here.

Run with llama.cpp and the Chandra CLI

Build llama.cpp for your accelerator, then start a loopback server. These are the validated inference settings; replace the device selector as appropriate for your machine.

ROCR_VISIBLE_DEVICES=GPU-YOUR-STABLE-UUID HIP_VISIBLE_DEVICES=0 \
llama-server \
  --model chandra-ocr-2.BF16.gguf \
  --mmproj chandra-ocr-2.mmproj-bf16.gguf \
  --alias chandra \
  --host 127.0.0.1 \
  --port 18100 \
  --ctx-size 24576 \
  --n-gpu-layers 999 \
  --split-mode none \
  --main-gpu 0 \
  --flash-attn on \
  --fit off \
  --parallel 1 \
  --batch-size 2048 \
  --ubatch-size 512 \
  --jinja \
  --chat-template-file chat_template.jinja \
  --image-min-tokens 1024

In another shell:

python3 -m venv .venv
.venv/bin/pip install 'chandra-ocr==0.2.0'

VLLM_API_BASE=http://127.0.0.1:18100/v1 \
VLLM_API_KEY=EMPTY \
VLLM_MODEL_NAME=chandra \
.venv/bin/python -m chandra.scripts.cli input.pdf output \
  --method vllm \
  --batch-size 1 \
  --max-workers 1 \
  --max-retries 2 \
  --max-output-tokens 12384 \
  --no-images \
  --no-html

Do not expose an unauthenticated llama.cpp server to a public network. Generative OCR can omit or hallucinate content; verify consequential documents against their source.

Validation result and limits

The exact-pinned conversion matched the public comparison BF16 tensor schema. All projector tensor payloads matched. In the main file, every tensor payload matched except twelve scalar F32 values across ten ssm_a tensors; the maximum absolute difference was 9.5367431640625e-07. Accepted OCR outputs matched. There is no evidence that this conversion improves OCR quality over other correct BF16 conversions.

The useful contribution is the pinned, MTP-free artifact pair; complete provenance; one-16-GB-GPU validation; and a reproducible launch envelope.

License and attribution

Chandra OCR 2 weights use Datalab's AI Pubs Open RAIL-M License (Modified). It contains usage, redistribution, commercial, competitive-use, attribution, and share-alike conditions. Read LICENSE in full before using or redistributing these files. The license and its restrictions apply to this derivative conversion.

Modified-file notice: the original Chandra OCR 2 checkpoint was converted to GGUF at the pinned llama.cpp revision; the unused MTP draft head was omitted from the main GGUF; model tensor precision otherwise remains BF16. No Datalab endorsement is claimed.

Downloads last month
280
GGUF
Model size
5B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dougvk/chandra-ocr-2-BF16-GGUF-RDNA4

Quantized
(35)
this model