Hindi BPE Tokenizer

This tokenizer was trained on a mixed Hindi dataset using Byte-Pair Encoding (BPE).

  • Vocabulary size: 6000
  • Base tokens: UTF-8 bytes (256)
  • Trained in Python using a custom implementation
  • Original length: 7053637
  • Compressed length: 869484
  • Compression ratio: 8.11X

Files

  • vocab.json: Token IDs
  • merges.json: Merge rules
  • metadata.json: Tokenizer configuration
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support