SmolLM-1.7B-base-mla-topk2-rank640

This repository contains the attention-only incremental weights for an MLAfication conversion of HuggingFaceTB/SmolLM-1.7B. It is not a standalone full-model checkpoint.

Variant

Field Value
Base model HuggingFaceTB/SmolLM-1.7B
Base revision d7449ff7241c863f3e8accc475155f0f97afa011
MLA latent rank 640
RoPE dimensions per KV head 2
Training checkpoint step 5600
Training variant stage1+2-distill-QKV
Delta tensors 168
Delta size 0.56 GiB
SHA-256 aba8fbaf9ef079a5ca7695c305346c7f1c764dc79435153b6f10352c6594780d

The training run froze every parameter outside attention and also froze each attention output projection. The uploaded delta therefore contains exactly the parameters matching "attn" in name and "o_proj" not in name; embeddings, MLP, normalization outside attention, output projections, and LM head are omitted. Frozen tensors from the historical full checkpoint were verified exactly against the pinned base weights after casting the base tensors to the checkpoint dtype.

Loading

Use the MHA2MLA-V2 implementation to instantiate the MLAfication architecture from the pinned base model, then load model.safetensors with strict=False:

from safetensors.torch import load_file

delta = load_file("model.safetensors")
load_result = mla_model.load_state_dict(delta, strict=False)

The architecture must be patched before loading the delta. Loading this repository directly with vanilla AutoModelForCausalLM.from_pretrained(...) is not supported. config.json contains the layer-wise RoPE indices and sanitized MLAfication metadata.

Files

  • model.safetensors: attention-only delta weights
  • config.json: base architecture, RoPE indices, and MLAfication metadata
  • manifest.json: provenance, byte size, tensor count, and checksum

License

The weights follow the Apache 2.0 license of the SmolLM base model.

Downloads last month
14
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenMOSS-Team/SmolLM-1.7B-base-mla-topk2-rank640

Finetuned
(15)
this model