YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Bina 0.2 - RizehPizeh

Persian full-page OCR pipeline based on PP-OCRv6:

  1. Detect ordered Persian text regions.
  2. Recognize CTC-safe RTL word groups.
  3. Reconstruct lines and complete pages in reading order.

Current training state

The initial PP-OCRv6 medium recognizer pilot reached a held-out normalized edit similarity of 0.1752708888 before its ephemeral Colab runtime reset. No model checkpoint survived that reset.

The files under training-kit/ reproduce data preparation, RTL segmentation, and evaluation-driven early stopping. Future best checkpoints are uploaded under checkpoints/ so runtime resets do not erase training progress.

Data

  • Reza2kn/persian-handwriting-pages-3.69m
  • Reza2kn/persian-printed-ocr-3.5m
  • Reza2kn/visualears-hardword-sentences for rare-character coverage

The Persian recognizer uses a character dictionary rather than a language-model tokenizer.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support