Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

ChrisMcCormick
/
decoderstack-d24

English
decoderstack
nanochat
pretraining
Model card Files Files and versions
xet
Community
decoderstack-d24 / tokenizer
677 kB
Ctrl+K
Ctrl+K
  • 1 contributor
History: 3 commits
ChrisMcCormick's picture
ChrisMcCormick
tokenizer: keep the pre-fix token_bytes so published bpb numbers stay reproducible
4aad0ae verified about 11 hours ago
  • token_bytes.pt
    133 kB
    xet
    tokenizer: regenerate token_bytes.pt with upstream 2ce972a raw-bytes fix (191/32768 ids corrected) about 11 hours ago
  • token_bytes_legacy.pt
    133 kB
    xet
    tokenizer: keep the pre-fix token_bytes so published bpb numbers stay reproducible about 11 hours ago
  • tokenizer.pkl

    Detected Pickle imports (1)

    • "tiktoken.core.Encoding"

    How to fix it?

    412 kB
    xet
    Add model card, meta.json, logs, tokenizer, training source, and nanochat converter about 17 hours ago