Image-Text-to-Text
PEFT
Safetensors
Marathi
lora
ocr
marathi

Chitrapathak Marathi v7 LoRA

LoRA adapter for krutrim-ai-labs/Chitrapathak-2, fine-tuned for Marathi line OCR on real Maharashtra government resolutions.

This repository contains the adapter only (about 228 MB). Load it on top of Chitrapathak-2. The base weights are not included.

The adapter is a non-commercial research derivative of Chitrapathak-2 and is distributed under the Krutrim Community License Agreement Version 1.0. Attribution: Krutrim.

Evaluation

Same frozen 100 Marathi line images for every model (marathi_ocr_eval_v2). This set was not used to train this adapter.

Overall winner: IndicOCR v7 LoRA πŸ†

πŸ† marks the best score on that row. Lower is better for CER, WER, and boundary errors. Higher is better for accuracy, F1, and exact match. Tied winners each get a trophy.

Metric PARSeq Base PARSeq V3 PARSeq All Real D Chitrapathak Base Chitrapathak old LoRA Chitrapathak v7 LoRA (this model) IndicOCR v7 LoRA
CER ↓ 28.04% 26.70% 24.56% 6.63% 3.90% 3.60% 3.52% πŸ†
Content-no-WS CER ↓ 21.76% 16.15% 15.56% 3.65% 3.27% 3.18% πŸ† 3.22%
Strict WER ↓ 96.40% 98.93% 96.49% 37.20% 13.83% 11.78% 10.52% πŸ†
Normalized WER ↓ 96.40% 98.93% 96.49% 37.20% 13.83% 11.78% 10.52% πŸ†
Content-aware WER ↓ 56.57% 46.45% 46.45% 7.21% 5.94% 5.84% πŸ† 6.04%
Lexical WER ↓ 91.57% 58.17% 55.52% 5.52% 3.86% πŸ† 4.11% 4.66%
Punctuation F1 ↑ 0.00% 76.80% 76.20% 97.45% 97.99% πŸ† 97.89% 96.98%
Matra accuracy ↑ 86.53% 86.73% 86.82% 96.58% πŸ† 96.24% 96.19% 96.58% πŸ†
Conjunct accuracy ↑ 82.16% 89.67% 84.98% 98.12% πŸ† 98.12% πŸ† 98.12% πŸ† 98.12% πŸ†
Numeral accuracy ↑ 83.68% 86.56% 86.36% 98.10% 98.86% 98.71% 99.00% πŸ†
Character accuracy ↑ 71.96% 73.30% 75.44% 93.37% 96.10% 96.40% 96.48% πŸ†
Exact match ↑ 0/100 0/100 0/100 38/100 52/100 60/100 63/100 πŸ†
Boundary merges ↓ 125 223 217 79 36 25 17 πŸ†
Boundary splits ↓ 6 3 3 0 πŸ† 0 πŸ† 0 πŸ† 1

This adapter leads content-no-whitespace CER and content-aware WER, and ties the best conjunct accuracy and boundary-split count.

Training

  • Base: Chitrapathak-2 (Qwen2.5-VL, vision tower frozen)
  • Data: 9,200 real Marathi line crops from official government resolutions, at most 10 words per line. 800 lines from other documents were held out of training.
  • LoRA rank 32, alpha 64, dropout 0.05, on language-model attention and MLP projections
  • 2 epochs, learning rate 1e-4, effective batch 16, bf16, AdamW, cosine schedule, 3% warmup
  • Prompt: system β€œYou are a helpful assistant.” plus β€œPerform OCR on this image and transcribe all visible text exactly as it appears.”
  • Image budget during training: max_pixels=300000, min_pixels=3136

Usage

import torch
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor

base = "krutrim-ai-labs/Chitrapathak-2"
adapter = "snxtyle/chitrapathak-marathi-v7-lora"

processor = AutoProcessor.from_pretrained(base, max_pixels=300_000, min_pixels=56 * 56)
model = AutoModelForImageTextToText.from_pretrained(base, dtype=torch.bfloat16)
model = PeftModel.from_pretrained(model, adapter)
model.eval()

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "line.png"},
            {"type": "text", "text": "Perform OCR on this image and transcribe all visible text exactly as it appears."},
        ],
    },
]
Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for snxtyle/chitrapathak-marathi

Adapter
(1)
this model