LFM2.5 TitleGen SFT v2

This model is the supervised fine-tuning stage of the LFM2.5 TitleGen experiment. Model repository: ManhHoDinh/lfm25-titlegen-sft. Benchmark publication timestamp: 2026-08-16T00:00:00+07:00.

Preliminary benchmark

The current model passed 291 of 300 evaluated English and Vietnamese examples (97.0%).

Model Passed Overall English Vietnamese Evidence
DPO 293 / 300 97.7% 99.5% 94.1% Preliminary automated EN/VI
SFT v2 291 / 300 97.0% 99.5% 92.2% Preliminary automated EN/VI
Curriculum 289 / 300 96.3% 98.5% 92.2% Preliminary automated EN/VI

Language coverage

Language Samples Passed Rate Status
German 0 0 0.0% NOT_EVALUATED
English 198 197 99.5% EVALUATED
Spanish 0 0 0.0% NOT_EVALUATED
Filipino 0 0 0.0% NOT_EVALUATED
French 0 0 0.0% NOT_EVALUATED
Indonesian 0 0 0.0% NOT_EVALUATED
Japanese 0 0 0.0% NOT_EVALUATED
Korean 0 0 0.0% NOT_EVALUATED
Lao 0 0 0.0% NOT_EVALUATED
Malay 0 0 0.0% NOT_EVALUATED
Burmese 0 0 0.0% NOT_EVALUATED
Portuguese 0 0 0.0% NOT_EVALUATED
Russian 0 0 0.0% NOT_EVALUATED
Tamil 0 0 0.0% NOT_EVALUATED
Thai 0 0 0.0% NOT_EVALUATED
Vietnamese 102 94 92.2% EVALUATED
Chinese 0 0 0.0% NOT_EVALUATED

0 means no evaluated examples when the status is NOT_EVALUATED; it is not a measured zero score.

Methodology

This preliminary benchmark evaluates aggregate English and Vietnamese results with deterministic decoding (do_sample: false, max_new_tokens: 32). Automated rubric identifiers: 3-8_tu, khong_cham_cuoi, mot_dong, dung_ngon_ngu, khong_chep. The remaining contracted languages are shown explicitly as not evaluated.

Limitations

  • Only English and Vietnamese have evaluated examples.
  • Language correctness uses an automated heuristic.
  • No native review or blind preference evidence is included.
  • These aggregate automated results do not establish production readiness, causal improvement, or statistical significance.

Machine-readable results

See benchmark-report.json for the validated aggregate report.

Downloads last month
227
Safetensors
Model size
0.2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ManhHoDinh/lfm25-titlegen-sft

Finetuned
(36)
this model