SmolLM-1.7B-base-mla-topk2-rank640
This repository contains the attention-only incremental weights for an MLAfication conversion of HuggingFaceTB/SmolLM-1.7B. It is not a standalone full-model checkpoint.
Variant
| Field | Value |
|---|---|
| Base model | HuggingFaceTB/SmolLM-1.7B |
| Base revision | d7449ff7241c863f3e8accc475155f0f97afa011 |
| MLA latent rank | 640 |
| RoPE dimensions per KV head | 2 |
| Training checkpoint | step 5600 |
| Training variant | stage1+2-distill-QKV |
| Delta tensors | 168 |
| Delta size | 0.56 GiB |
| SHA-256 | aba8fbaf9ef079a5ca7695c305346c7f1c764dc79435153b6f10352c6594780d |
The training run froze every parameter outside attention and also froze each attention output projection. The uploaded delta therefore contains exactly the parameters matching "attn" in name and "o_proj" not in name; embeddings, MLP, normalization outside attention, output projections, and LM head are omitted. Frozen tensors from the historical full checkpoint were verified exactly against the pinned base weights after casting the base tensors to the checkpoint dtype.
Loading
Use the MHA2MLA-V2 implementation to instantiate the MLAfication architecture from the pinned base model, then load model.safetensors with strict=False:
from safetensors.torch import load_file
delta = load_file("model.safetensors")
load_result = mla_model.load_state_dict(delta, strict=False)
The architecture must be patched before loading the delta. Loading this repository directly with vanilla AutoModelForCausalLM.from_pretrained(...) is not supported. config.json contains the layer-wise RoPE indices and sanitized MLAfication metadata.
Files
model.safetensors: attention-only delta weightsconfig.json: base architecture, RoPE indices, and MLAfication metadatamanifest.json: provenance, byte size, tensor count, and checksum
License
The weights follow the Apache 2.0 license of the SmolLM base model.
- Downloads last month
- 14
Model tree for OpenMOSS-Team/SmolLM-1.7B-base-mla-topk2-rank640
Base model
HuggingFaceTB/SmolLM-1.7B