YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

DocEng_Tessarect_mlt

A fine-tuned Tesseract OCR model for Maltese-language news articles.

Overview

This model was fine-tuned on a synthetic dataset of ~20k PNG image "snippets," each depicting a short excerpt of Maltese text paired with its ground-truth transcription. Training ran for 2000 iterations on top of the base Tesseract Maltese trained data.

Dataset Generation

The synthetic training images were produced in two stages:

  1. Text generation โ€” Maltese news-style text snippets were generated by querying Gemini 3.1 Flash-Lite to generate 120 pdfs covering 10 different topics. The snippets were based on the character set provided in the competition assets and generated predominately in Maltese with some English covering different fonts, text and background colours.

  2. Snippet Extraction - The snippet extraction process scans each individual PDF, identifies text coordinates, and crops specific areas into individual images. The resulting images are saved alongside a CSV index that links each snippet to its corresponding text, and word count. Gaussian blur is applied to 10% of these snippets.

  3. Manual Noising โ€” Randomised manual degradation was implemented.

Usage

Download the trained data file from the Hugging Face Hub and point Tesseract at its containing directory:

import os
from huggingface_hub import hf_hub_download

traineddata_path = hf_hub_download(
repo_id="AlanaBusu/DocEng_Tessarect_mlt",
filename="finetune_tess_extended_2000.traineddata",
)
tessdata_dir = os.path.dirname(traineddata_path)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support