Request access to Lipikar's UnifiedIndicOCR Model

This model is released under CC BY-NC-ND 4.0. Weights are made available for inference and evaluation purposes. Fine-tuning, further training, or creation of derivative models is not permitted under this license.

By requesting access, you confirm that you have read and agree to the CC BY-NC-ND 4.0 license terms, and that you will use this model for non-commercial inference/testing purposes only — not for fine-tuning, retraining, or building derivative models.

Log in or Sign Up to review the conditions and access this model content.

UnifiedIndicOCR (PARSeq)

Multilingual Indic OCR model based on PARSeq (permuted autoregressive sequence model), with a ViT-L encoder. This repo is self-contained for inference — only torch, torchvision, timm and Pillow are required, no pytorch_lightning/hydra/training-repo dependencies.

Repo layout

├── infer.py                 # run this: OCR on one or more images
├── requirements.txt         # torch, torchvision, timm, Pillow — nothing else
├── weights/
│   └── parseq_indic.pt      # model weights + architecture/charset config
├── samples/                 # a couple of sample images to try
│   ├── adhyaksh_new_1.png
│   └── bn_test_4.jpg
└── strhub/                  # trimmed copy of the original strhub package —
                              # just the PARSeq nn.Module + Tokenizer, no
                              # pytorch_lightning/hydra/nltk training code

Model

  • Architecture: PARSeq, ViT-L encoder — embed_dim=1024, enc_depth=24, enc_num_heads=16, 3-layer decoder with dec_num_heads=16. Matches the parseq-L model config combined with configs/10lang.yaml in the training repo (verified against the checkpoint's own saved hyperparameters).
  • Input: 32×128 (H×W) RGB crops, patch_size=[4, 8].
  • Charset: 985 characters (multilingual Indic + Latin + punctuation) + 3 special tokens ([E]/[B]/[P]) = 988 tokens. Note: this is the exact charset stored inside the source checkpoint's own hyperparameters, which has one extra character (ZWJ, U+200D) compared to the checked-in configs/10lang.yaml in the training repo — that config file is stale by one character, so the checkpoint's own charset is used here instead of re-deriving it from the yaml.
  • Weights: converted from the training repo's psmr/train-saves/old-a-final-pt.pt (355M params, fp32, ~1.4 GB).

Setup

conda create --name indicocr python=3.10 -y
conda activate indicocr
pip install -r requirements.txt

Run

# Uses the bundled sample images by default
python infer.py

# Or point it at your own images
python infer.py --images path/to/img1.png path/to/img2.jpg --device cuda

Output:

samples/adhyaksh_new_1.png: '...recognized text...'  (confidence: 0.98xx)
samples/bn_test_4.jpg: '...recognized text...'  (confidence: 0.97xx)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support