Training-data provenance: could you enumerate the training datasets?

#2
by hudsonmolthan - opened

Hi, and thank you for releasing these weights.

We're evaluating MossFormer2_SS_16K (via ClearerVoice-Studio) for a commercial application and need to confirm the provenance of its training data. The model card states it was "trained on large scale datasets including open-sourced and private data" but does not enumerate them.

Could you help us with two questions:

  1. Could you enumerate the specific datasets used to train this 16 kHz separation checkpoint (both the open-sourced and, to the extent shareable, the private portions)?
  2. In particular, do any LDC-distributed corpora β€” e.g. WSJ / WSJ0-mix, Fisher, CALLHOME, Switchboard, or NIST-SRE β€” appear in the training data for THIS checkpoint? (We note the sibling mossformer2-wsj0mix-3spk is WSJ-derived, and want to confirm the provenance of this particular model.)

We ask because commercial deployment requires the training corpora to be free of redistribution/commercial-use restrictions. A short confirmation either way would let us proceed. Thank you for your time.

Sign up or log in to comment