IndicF5 Haryanvi (Bangru) — public preview

This repository is the release destination for an experimental Bangru-focused Haryanvi adaptation of AI4Bharat/IndicF5.

Current status

The released IndicF5 checkpoint has been converted into the upstream F5 training format with full intended parameter coverage. A 22-prompt zero-shot baseline, a four-clip full-precision smoke test, a 200-update tiny overfit test, and checkpoint inference have passed.

A finished Haryanvi checkpoint is not published yet. Native-speaker listening approval and the pilot fine-tune remain required before release.

The reproducible source code is available at sauravtom/haryanvi-indicf5. The accompanying Space is kumarx/indicf5-haryanvi-demo.

Dataset scope

The public dataset identifies itself as the Bangru dialect of Haryanvi. This project does not claim coverage of every Haryanvi dialect. At the pinned revision, 2,766 metadata-linked rows passed automatic filtering, representing approximately 4.12 hours of audio. Native-speaker review is still a release gate.

Responsible use

Only use reference voices that you own or have explicit permission to clone. Unauthorized impersonation and voice cloning are prohibited.

Licenses and attribution

  • Project release: Apache-2.0
  • Base model: AI4Bharat IndicF5, MIT
  • Fine-tuning dataset: ankitdhiman/haryanvi-tts, Apache-2.0

The final model card will include training configuration, hardware, duration, evaluation methodology, listening results, known failures, and samples.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kumarx/indicf5-haryanvi

Finetuned
(12)
this model

Dataset used to train kumarx/indicf5-haryanvi

Space using kumarx/indicf5-haryanvi 1