Request access to Lipikar's UnifiedIndicOCR Model
This model is released under CC BY-NC-ND 4.0. Weights are made available for inference and evaluation purposes. Fine-tuning, further training, or creation of derivative models is not permitted under this license.
By requesting access, you confirm that you have read and agree to the CC BY-NC-ND 4.0 license terms, and that you will use this model for non-commercial inference/testing purposes only — not for fine-tuning, retraining, or building derivative models.
Log in or Sign Up to review the conditions and access this model content.
UnifiedIndicOCR (PARSeq)
Multilingual Indic OCR model based on PARSeq (permuted autoregressive
sequence model), with a ViT-L encoder. This repo is self-contained for
inference — only torch, torchvision, timm and Pillow are required,
no pytorch_lightning/hydra/training-repo dependencies.
Repo layout
├── infer.py # run this: OCR on one or more images
├── requirements.txt # torch, torchvision, timm, Pillow — nothing else
├── weights/
│ └── parseq_indic.pt # model weights + architecture/charset config
├── samples/ # a couple of sample images to try
│ ├── adhyaksh_new_1.png
│ └── bn_test_4.jpg
└── strhub/ # trimmed copy of the original strhub package —
# just the PARSeq nn.Module + Tokenizer, no
# pytorch_lightning/hydra/nltk training code
Model
- Architecture: PARSeq, ViT-L encoder —
embed_dim=1024,enc_depth=24,enc_num_heads=16, 3-layer decoder withdec_num_heads=16. Matches theparseq-Lmodel config combined withconfigs/10lang.yamlin the training repo (verified against the checkpoint's own saved hyperparameters). - Input:
32×128(H×W) RGB crops,patch_size=[4, 8]. - Charset: 985 characters (multilingual Indic + Latin + punctuation) + 3
special tokens (
[E]/[B]/[P]) = 988 tokens. Note: this is the exact charset stored inside the source checkpoint's own hyperparameters, which has one extra character (ZWJ, U+200D) compared to the checked-inconfigs/10lang.yamlin the training repo — that config file is stale by one character, so the checkpoint's own charset is used here instead of re-deriving it from the yaml. - Weights: converted from the training repo's
psmr/train-saves/old-a-final-pt.pt(355M params, fp32, ~1.4 GB).
Setup
conda create --name indicocr python=3.10 -y
conda activate indicocr
pip install -r requirements.txt
Run
# Uses the bundled sample images by default
python infer.py
# Or point it at your own images
python infer.py --images path/to/img1.png path/to/img2.jpg --device cuda
Output:
samples/adhyaksh_new_1.png: '...recognized text...' (confidence: 0.98xx)
samples/bn_test_4.jpg: '...recognized text...' (confidence: 0.97xx)