PiotrSty/trocr-pl-mixed-v3 (experimental)

Longest fine-tune of PiotrSty/trocr-pl-base on synthetic Polish print + real EHRI typewritten Polish lines (CC-BY 4.0, ehri-pl-lines). Same recipe as trocr-pl-mixed-v2 (run 4) but 30 epochs instead of 15, because run 4's val CER was still improving at epoch 15.

Training

  • Base: PiotrSty/trocr-pl-base
  • Method: QLoRA on decoder attention (q/k/v/out_proj), rank 16, alpha 32
  • Train: 2000 synthetic lines + 349 real EHRI lines (3 docs)
  • Val: 38 EHRI lines (held-out doc ZIH3010905)
  • Epochs: 30, batch 8, lr 1e-4, T4 x2
  • Best checkpoint: /kaggle/working/trocr-pl-run5/checkpoint-4116 (val CER 0.2976, WER 0.6684)
  • Document-level split, no line leakage.

Evaluation on frozen held-out sets

Model EHRI test (81, typewriter) real-lines-v1 (75, print)
trocr-pl-base (run2) CER 47.30% / WER 90.82% CER 11.11% / WER 35.84%
trocr-pl-mixed-v1 (run3, 5 ep) CER 33.95% / WER 85.69% CER 7.09% / WER 29.44%
trocr-pl-mixed-v2 (run4, 15 ep) CER 30.42% / WER 78.85% CER 5.47% / WER 23.20%
trocr-pl-mixed-v3 (run5, 30 ep) (fill from eval cell) (fill from eval cell)

Limitations

  • Line recognizer only; page segmentation on faded typewriter is unreliable.
  • Only 349 real typewritten training lines.
  • Do NOT use as drop-in replacement without your own eval.

Provenance

See run.json, selection.json, best_metrics.json in this repo. Source: https://github.com/PiotrStyla/OCR_engine (commit c39baf4) EHRI dataset: https://huggingface.co/datasets/PiotrSty/ehri-pl-lines

Downloads last month
82
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PiotrSty/trocr-pl-mixed-v3

Finetuned
(4)
this model