finetune to other language

#1
by mikinko - opened

hi, plz where or how it can be trained to finetune to other lang ?

need russian

need for Portuguese Brazil

@mikinko @rekillkos @lailton I couldn't find a Whistle audio fine-tuning recipe in the sources checked here: the needle repo at af2654d, and this checkpoint at b358dda.

  • needle finetune trains Needle, not Whistle. finetune.py turns query/answers JSONL into token ids and trains LoRA adapters on them. That training module operates on token ids, not audio.
  • The language set is fixed. whistle.py has LANGUAGES = ("en", "de", "fr", "es", "it", "nl", "pl"), and the vocabulary inside whistle.cact holds exactly those seven language tags. Neither Russian nor Portuguese has a listed language code or vocabulary tag. That alone does not establish what training changes would be needed.
  • The shipped vocabulary has 8,199 entries, including 256 byte tokens, and no Cyrillic pieces. This is a vocabulary check; it does not show whether the model can transcribe Russian or Portuguese, or whether a particular fine-tuning method would work.

The open question is already on GitHub: needle#169 is adding Bulgarian from whistle.safetensors, lists what the published code leaves out (log-mel settings, the conv stem, the encoder's conv-module order) and asks whether a Whistle fine-tuning path is planned. No reply there yet. That issue is an existing place to follow the request for a documented training and export path.

To reproduce the counts, save the code as check_vocab.py and run python check_vocab.py whistle.cact on this exact file. It reads the vocabulary table at the end of that file:

import struct, sys
d = open(sys.argv[1], "rb").read()                 # whistle.cact
p = d.find(b"\x05\x00<pad>") - 5                    # entry: f32 score, u8 type, u16 length, UTF-8
pieces = []
while p < len(d):
    n, = struct.unpack("<H", d[p + 5:p + 7])
    pieces.append(d[p + 7:p + 7 + n].decode()); p += 7 + n
print(len(pieces), [s for s in pieces if len(s) == 6 and s[:2] == "<|"])
print("byte tokens:", sum(s.startswith("<0x") for s in pieces))
print("Cyrillic:", sum(any("Ѐ" <= c <= "ӿ" for c in s) for s in pieces))

I haven't trained or run Whistle; this comes from the package source, the issue and the vocabulary in the published file.

Sign up or log in to comment