Training-data provenance: could you enumerate the training corpora?

#20
by hudsonmolthan - opened

Hi, and thank you for releasing this model.

We're evaluating MOSS-Transcribe-Diarize for a commercial application and, as part of due diligence, need to confirm the provenance of the training data. The technical report (arXiv 2601.01554, §3 Data Composition) names AISHELL-4 explicitly, and additionally refers to sampling from "public corpora" collected from the Internet (§3.1) and an unspecified "in-house corpus" used to construct the simulated mixtures (§3.2).

Could you help us with two questions:

  1. Could you enumerate the specific public corpora used for training beyond AISHELL-4?
  2. In particular, do any LDC-distributed corpora — e.g. Fisher, CALLHOME, Switchboard, WSJ, NIST-SRE, DIHARD, or RT-03 — appear anywhere in the training data (including the in-house simulated-mixture pool)?

We ask because commercial use requires the training corpora to be free of redistribution/commercial-use restrictions. A short confirmation either way would let us proceed. Thanks very much for your time.

Sign up or log in to comment