MiMo-V2.6-Pro-EXL3-FP4

Partial research checkpoint: layer 1 initial fit only. This is not a complete or ready-to-serve model.

Contains all 384 layer-1 cold expert artifacts and SHA-256 receipts. The original 5% hot-expert allocation is recorded in allocation.json; layer 1 has no hot experts. Other layers, the model backbone, and hot weight tensors are not included in this partial upload.

Each expert uses mixed-rate EXL3 trellis coding with fitted scales and an actual packed budget at or below 2 bits per weight, including the artifact header. All neurons are retained. These artifacts already include per-expert PV-style refinement: fixed trellis codes with joint expert-output scale fitting for up to 1,500 steps, selected on held-out validation. Full-corpus joint-layer PV, which optimizes the combined routed output across experts, remains pending. The subsequent layer-level scalar-refinement trial is not included.

On cached held-out routed layer-1 inputs, decoded BF16 arithmetic measured 2.4034% validation relative L2 and 4.7751% audit relative L2. These are whole-layer output reconstruction errors, not perplexity, KL divergence, or full-model accuracy. See reports/layer1_routed_eval.json.

Runtime status

The EXL3-FP4 name describes the intended experimental SM120 serving path. The uploaded packed artifacts are EXL3 trellis weights. FP4 execution is not qualified, and this upload does not provide a production vLLM loader or a completed FP4 MMA kernel. Full-corpus PV and sequential propagation remain pending.

decode.py reconstructs an expert with the official EXL3 1.5.1 runtime, using a compatible CUDA/PyTorch wheel:

python decode.py layers/layer_00001/expert_00000/selected.bin expert_00000.pt

The resulting file contains gate_proj, up_proj, and down_proj tensors. manifest.json records every artifact hash and byte count. Source: XiaomiMiMo/MiMo-V2.6-Pro-RL, revision 73875d00b30a89ef8cc353a0b60b0e9f9561952d.

Calibration coverage

The saved training corpus contains 18,006,461 tokens, comprising reasoning, code, agentic, instruction, medical, and prose text tokenized with the pinned MiMo tokenizer. This initial fit used 65,536 cached training token rows plus 65,536 allocation-calibration rows, filtered by each expert’s routes; overlap between these sets has not been measured. It did not optimize across all 18 million tokens. Reported validation and audit each use 16,384 held-out tokens. Full-corpus activation capture is complete on the A100 host, but full-corpus joint-layer optimization is not complete.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jarrelscy/MiMo-V2.6-Pro-EXL3-FP4

Finetuned
(3)
this model