Configuration Parsing Warning:Invalid JSON for config file config.json

m1b-sft-lc-1k

A checkpoint from MambaRL: does outcome-only RL extend the effective context-retrieval horizon of a pure Mamba-2 LM, without changing the architecture? No attention is added and the recurrent state is not grown โ€” only what the policy learns to keep in it changes.

  • Base model: AntonV/mamba2-2.7b-hf
  • LoRA adapter merged in from checkpoints/M1b_sft_lc_1k (targets: in_proj only โ€” PEFT refuses out_proj/conv1d on mamba2 because the fused kernels read those weights directly)
  • Weights here are merged, so AutoModelForCausalLM.from_pretrained is enough.

Phase 1 (proposal 2.4): generic CoT LoRA SFT on 1024 examples, final train loss 0.3550, 0.02 h on one L40.

Answers follow a <think>...</think><answer>...</answer> contract.

Downloads last month
14
Safetensors
Model size
3B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Lixing-Li/m1b-sft-lc-1k

Adapter
(4)
this model