MiniCPM5 Stage1 SFT โ€” Transformers checkpoints

This public repository contains model weights only, converted to Hugging Face Transformers format from a 192-NPU tensor-parallel training run. It does not include optimizer, scheduler, or RNG state. Each checkpoint-{step}/ directory is a separate model snapshot with its own config, tokenizer assets, and safetensors weight files.

Available steps: 1112 (original stage1), 1113, 2224, 3336, 4448, 5560, 6672, 7784, and 8900 (final completed step). The earlier snapshots use five weight shards each; step 8900 uses one model-00001.safetensors file and an index. The repo name is an experiment label, not a parameter count: the converted model has 2,516,756,480 parameters in bfloat16.

For the final step, 381 converted tensors were compared against the tensor-parallel training weights without a mismatch, and a CPU-only Transformers load and short generation test passed. No GPU was used for this check. These model-only snapshots can be used for inference or a weight-only training resume; they cannot restore an optimizer or reproduce a full-state distributed resume.

Example:

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "eigentom/nanocode_sft_60b"
model = AutoModelForCausalLM.from_pretrained(repo, subfolder="checkpoint-8900")
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder="checkpoint-8900")
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for eigentom/nanocode_sft_60b

Finetunes
1 model