facebook/xnli
Viewer • Updated • 6.4M • 20.2k • 72
MorpheL MI-guided tokenizer trained on facebook/xnli subset sw, using all train/validation/test premise and hypothesis text under the same corpus protocol as the project baselines.
Downstream notebooks must load morphel_vocab.json, mi_index.pkl, segmentation_cache.pkl, and morphel_config.json; tokenizer.json alone is only a WordLevel wrapper.
{
"fertility": 1.259849551558451,
"tokens_per_char": 0.232687970096,
"avg_seq_len": 17.329933333333333,
"vocab_coverage": 1.0,
"oov_rate": 0.0,
"fallback_event_rate": 0.1561573736720706,
"char_shatter_rate": 0.0,
"n_sentences": 15000,
"n_word_types": 18059
}