Unsupervised training data inquiry

#3
by majfu - opened

Hello,
I was wondering if there's any more detail available on the datasets used in stages 1-3 training, the report paper mentions "12-million bilingual unsupervised paragraph dataset, with a roughly 1:1 ratio between Chinese and English texts". Is there any information on which specific datasets were used? If not, then maybe some general information about domain coverage, where the texts were sourced or how long the paragraphs were?
I would be very grateful for any help πŸ˜„

Sign up or log in to comment