Instructions to use snxtyle/chitrapathak-marathi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use snxtyle/chitrapathak-marathi with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("krutrim-ai-labs/Chitrapathak-2") model = PeftModel.from_pretrained(base_model, "snxtyle/chitrapathak-marathi") - Notebooks
- Google Colab
- Kaggle
Chitrapathak Marathi v7 LoRA
LoRA adapter for krutrim-ai-labs/Chitrapathak-2, fine-tuned for Marathi line OCR on real Maharashtra government resolutions.
This repository contains the adapter only (about 228 MB). Load it on top of Chitrapathak-2. The base weights are not included.
The adapter is a non-commercial research derivative of Chitrapathak-2 and is distributed under the Krutrim Community License Agreement Version 1.0. Attribution: Krutrim.
Evaluation
Same frozen 100 Marathi line images for every model (marathi_ocr_eval_v2). This set was not used to train this adapter.
Overall winner: IndicOCR v7 LoRA π
π marks the best score on that row. Lower is better for CER, WER, and boundary errors. Higher is better for accuracy, F1, and exact match. Tied winners each get a trophy.
| Metric | PARSeq Base | PARSeq V3 | PARSeq All Real D | Chitrapathak Base | Chitrapathak old LoRA | Chitrapathak v7 LoRA (this model) | IndicOCR v7 LoRA |
|---|---|---|---|---|---|---|---|
| CER β | 28.04% | 26.70% | 24.56% | 6.63% | 3.90% | 3.60% | 3.52% π |
| Content-no-WS CER β | 21.76% | 16.15% | 15.56% | 3.65% | 3.27% | 3.18% π | 3.22% |
| Strict WER β | 96.40% | 98.93% | 96.49% | 37.20% | 13.83% | 11.78% | 10.52% π |
| Normalized WER β | 96.40% | 98.93% | 96.49% | 37.20% | 13.83% | 11.78% | 10.52% π |
| Content-aware WER β | 56.57% | 46.45% | 46.45% | 7.21% | 5.94% | 5.84% π | 6.04% |
| Lexical WER β | 91.57% | 58.17% | 55.52% | 5.52% | 3.86% π | 4.11% | 4.66% |
| Punctuation F1 β | 0.00% | 76.80% | 76.20% | 97.45% | 97.99% π | 97.89% | 96.98% |
| Matra accuracy β | 86.53% | 86.73% | 86.82% | 96.58% π | 96.24% | 96.19% | 96.58% π |
| Conjunct accuracy β | 82.16% | 89.67% | 84.98% | 98.12% π | 98.12% π | 98.12% π | 98.12% π |
| Numeral accuracy β | 83.68% | 86.56% | 86.36% | 98.10% | 98.86% | 98.71% | 99.00% π |
| Character accuracy β | 71.96% | 73.30% | 75.44% | 93.37% | 96.10% | 96.40% | 96.48% π |
| Exact match β | 0/100 | 0/100 | 0/100 | 38/100 | 52/100 | 60/100 | 63/100 π |
| Boundary merges β | 125 | 223 | 217 | 79 | 36 | 25 | 17 π |
| Boundary splits β | 6 | 3 | 3 | 0 π | 0 π | 0 π | 1 |
This adapter leads content-no-whitespace CER and content-aware WER, and ties the best conjunct accuracy and boundary-split count.
Training
- Base: Chitrapathak-2 (Qwen2.5-VL, vision tower frozen)
- Data: 9,200 real Marathi line crops from official government resolutions, at most 10 words per line. 800 lines from other documents were held out of training.
- LoRA rank 32, alpha 64, dropout 0.05, on language-model attention and MLP projections
- 2 epochs, learning rate 1e-4, effective batch 16, bf16, AdamW, cosine schedule, 3% warmup
- Prompt: system βYou are a helpful assistant.β plus βPerform OCR on this image and transcribe all visible text exactly as it appears.β
- Image budget during training:
max_pixels=300000,min_pixels=3136
Usage
import torch
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor
base = "krutrim-ai-labs/Chitrapathak-2"
adapter = "snxtyle/chitrapathak-marathi-v7-lora"
processor = AutoProcessor.from_pretrained(base, max_pixels=300_000, min_pixels=56 * 56)
model = AutoModelForImageTextToText.from_pretrained(base, dtype=torch.bfloat16)
model = PeftModel.from_pretrained(model, adapter)
model.eval()
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{
"role": "user",
"content": [
{"type": "image", "image": "line.png"},
{"type": "text", "text": "Perform OCR on this image and transcribe all visible text exactly as it appears."},
],
},
]
- Downloads last month
- 17
Model tree for snxtyle/chitrapathak-marathi
Base model
Qwen/Qwen2.5-VL-3B-Instruct