SmolLM-1.7B-base-mla-topk2-rank384

This repository contains the attention-only incremental weights for an MLAfication conversion of HuggingFaceTB/SmolLM-1.7B. It is not a standalone full-model checkpoint.

Variant

Field Value
Base model HuggingFaceTB/SmolLM-1.7B
Base revision d7449ff7241c863f3e8accc475155f0f97afa011
MLA latent rank 384
RoPE dimensions per KV head 2
Training checkpoint step 8000
Training variant stage1+2-distill-QKV
Delta tensors 168
Delta size 0.49 GiB
SHA-256 2a2e1812d02e35cfddbcd915b515dffe8226e1ec9454d3a59f507c8c16e77f09

The training run froze every parameter outside attention and also froze each attention output projection. The uploaded delta therefore contains exactly the parameters matching "attn" in name and "o_proj" not in name; embeddings, MLP, normalization outside attention, output projections, and LM head are omitted. Frozen tensors from the historical full checkpoint were verified exactly against the pinned base weights after casting the base tensors to the checkpoint dtype.

Loading

Use the MHA2MLA-V2 implementation to instantiate the MLAfication architecture from the pinned base model, then load model.safetensors with strict=False:

from safetensors.torch import load_file

delta = load_file("model.safetensors")
load_result = mla_model.load_state_dict(delta, strict=False)

The architecture must be patched before loading the delta. Loading this repository directly with vanilla AutoModelForCausalLM.from_pretrained(...) is not supported. config.json contains the layer-wise RoPE indices and sanitized MLAfication metadata.

Files

  • model.safetensors: attention-only delta weights
  • config.json: base architecture, RoPE indices, and MLAfication metadata
  • manifest.json: provenance, byte size, tensor count, and checksum

License

The weights follow the Apache 2.0 license of the SmolLM base model.

Downloads last month
14
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenMOSS-Team/SmolLM-1.7B-base-mla-topk2-rank384

Finetuned
(15)
this model