Instructions to use sugiv/tiny-clm-qwen3-1.7b-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use sugiv/tiny-clm-qwen3-1.7b-lora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Tiny CLM: LoRA adapters for Qwen3-1.7B
Every LoRA adapter produced by Tiny CLM, a small replication of the reinforcement-learning
claim in Context Language Models (Shao et al., 2026,
arXiv 2609.37725, CC BY 4.0): stepwise GRPO with a
success-gated efficiency advantage, trained on a synthetic KV Store task where the model must
manage its own 2,048-token context with OFFLOAD / GREP / NOTE / ANSWER / READY actions.
The code, pre-registration, full results and the meaning of every condition are in the project
repository (README.md, PREREGISTRATION.md, RESULTS.md). This repo mirrors the code
repo's results/ layout, so each folder here matches the run of the same name there.
Which adapter to use
The ckpt_best/ folder of each run is the checkpoint selected on the validation split
(highest accuracy, ties broken by lower cost); it is the one evaluated on the test split.
| Folder | What it is |
|---|---|
e2_cal/adapter_step8/ |
SFT warm start used by every RL run (picked by the calibration rule) |
e2_cal/adapter_step{2..32}/, e2_cal/adapter/ |
Other early-stopped SFT checkpoints from the calibration sweep |
e2_sft/adapter/ |
Full SFT run (213 steps); ~100% accurate, not used for RL |
e3_r1_s{0,1,2}/ |
RL, R1 outcome reward only, training seeds 0–2 |
e3_r2_s{0,1,2}/ |
RL, R2 outcome + success-gated efficiency advantage (w_eff 0.25) |
e3_r3_s{0,1,2}/ |
RL, R3 naive ungated cost penalty |
e5_final_s{0,1,2}/ |
RL, R1 reward trained on the final transcript only (E5 ablation) |
pilot_s0/, pilot2_s0/ |
M4 pilot runs (lr 2e-5 and 1e-4) |
Inside each RL run folder:
ckpt_best/: selected checkpoint (selection.jsongives step, validation accuracy, cost)adapter_latest/: the policy after the last training stepadapter_step0/: the starting policy (a copy of the warm start)*/ref/: a frozen copy of the warm-start adapter, saved alongside because it served as the KL reference during training. Load the folder root (adapterdefault) for the policy.
Runs that escaped the warm-start plateau (test accuracy >= 0.99 at pressure >= 2x, cache-friendly
"offload immediately" policy): e3_r1_s0, e3_r1_s2, e3_r2_s2, e3_r3_s1, e3_r3_s2,
e5_final_s0, e5_final_s1.
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-1.7B", torch_dtype="bfloat16")
model = PeftModel.from_pretrained(base, "sugiv/tiny-clm-qwen3-1.7b-lora", subfolder="e3_r1_s0/ckpt_best")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-1.7B")
Prompts must use the project's environment renderer (ChatML with thinking disabled); see
env/render.py and eval/run_eval.py in the code repository.
Training details
LoRA r 32, alpha 64, all linear layers, bf16. GRPO: 4 prompts x 8 rollouts per step, 40 steps, lr 5e-5, KL 0.01 (k3) to the warm start, clip [0.2, 0.28], truncated importance sampling, dynamic sampling, token-mean per trajectory. Rollouts with vLLM 0.30 at temperature 0.7. One RTX 3090/4090.
License
Adapters: Apache-2.0, following the base model Qwen/Qwen3-1.7B. The method follows Shao et al., arXiv 2609.37725 (CC BY 4.0).
- Downloads last month
- -