grc-macronizer-char (seed 13, original) -- experimental variant

This is one of three random-seed replicates of Ericu950/oga-macronizer-char, provided for experimentation and reproducibility only. Use the main repo unless you specifically need this seed.

This was the original single-seed checkpoint (default --seed 13, no explicit override) previously published as the main repo; it has since been superseded there by a better-performing seed. Same architecture, training data, and hyperparameters as the main model (see that repo's README for the full description) -- only the random seed differs.

Evaluation

Full Norma Syllabarum Graecarum benchmark (17 works, 2,957 gold-annotated positions), scored with macron_model/eval_norma.py:

system accuracy (unmarked defaults to short) raw accuracy unmarked / total
rule-based grc-macronizer 86.84% 53.77% 1,300 / 2,957
this checkpoint (seed 13) 88.47% 88.37% 9 / 2,957
main repo (seed 42) 89.89% 89.79% 9 / 2,957

Across three seeds, defaults-to-short accuracy averages 89.36±0.78% (88.47% / 89.72% / 89.89%), with every seed beating the rule-based teacher -- this checkpoint is the weakest of the three, though still ahead of the rule-based system.

Usage

import sys
sys.path.insert(0, "path/to/grc-macronizer/macron_model")  # for predict.py
from predict import MacronPredictor

predictor = MacronPredictor("path/to/downloaded/checkpoint")
macronized = predictor.macronize("ανθρωπος ανηρ")

Citation

Thörn Cleland, Albin and Eric Cullhed (forthcoming). Automatic Annotation of Ancient Greek Vowel Length.

License

GNU GPL v3, matching grc-macronizer.

Downloads last month
-
Safetensors
Model size
867k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support