full_3b — steering-resistance LoRA adapter

LoRA adapter for Qwen/Qwen2.5-3B-Instruct trained to resist adversarial activation steering: fine-tuned with CAA vectors injected live at decoder layer 18, rewarded for reproducing its own clean answers. Full method: https://github.com/JacoDuToit11/steering-resistance.

Provenance

run full_3b
result clean 100%->100% · steer_heldout@1.6 correct 0%->4%
trained 2026-08-05T17:16:57+0000
code b16f6e72a9bb + uncommitted changes (see run/code.patch)
config configs/full_3b.yaml (snapshot: run/config.yaml)
wandb —
hardware NVIDIA A100 80GB PCIe
stack torch 2.6.0+cu124, transformers 5.14.1, peft 0.20.0, accelerate 1.14.0, datasets 5.0.1, numpy 2.4.6, huggingface_hub 1.26.0
data: concepts_path adc6aee9b1cc537a… (data/concepts_paper.json)
data: qa_path fa2356571420fd8c… (data/qa_open.json)

Training parameters

param value
layer 18
train_alphas [0.4, 0.8, 1.2]
eval_alphas [0.4, 0.8, 1.0, 1.2, 1.6]
efficacy_alpha 0.8
efficacy_min_rate 0.3
steered_frac 0.7
relevant_frac 0.6
repeats_per_question 6
lora_r 16
lora_alpha 32
lora_dropout 0.05
lora_targets ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]
lr 0.0001
epochs 1
effective_batch_size 16
train_batch_size 4
max_seq_len 256
seed 0

Eval results

model condition alpha n correct steered other
M0 clean 0.0 17 100% [100%,100%] 0% [0%,0%] 0% [0%,0%]
M0 correct_inject 0.4 9 100% [100%,100%] 0% [0%,0%] 0% [0%,0%]
M0 correct_inject 0.8 9 89% [67%,100%] 0% [0%,0%] 11% [0%,33%]
M0 correct_inject 1.0 9 89% [67%,100%] 0% [0%,0%] 11% [0%,33%]
M0 correct_inject 1.2 9 89% [67%,100%] 0% [0%,0%] 11% [0%,33%]
M0 correct_inject 1.6 9 89% [67%,100%] 0% [0%,0%] 11% [0%,33%]
M0 steer_heldout 0.4 2754 84% [77%,89%] 9% [7%,12%] 7% [4%,12%]
M0 steer_heldout 0.8 2754 14% [10%,17%] 73% [69%,77%] 13% [12%,14%]
M0 steer_heldout 1.0 2754 1% [1%,2%] 81% [80%,83%] 17% [16%,18%]
M0 steer_heldout 1.2 2754 0% [0%,0%] 74% [73%,75%] 26% [25%,27%]
M0 steer_heldout 1.6 2754 0% [0%,0%] 52% [52%,53%] 48% [47%,48%]
M0 steer_train 0.4 679 93% [84%,98%] 0% [0%,1%] 7% [2%,16%]
M0 steer_train 0.8 679 16% [12%,21%] 73% [69%,78%] 10% [8%,13%]
M0 steer_train 1.0 679 2% [1%,3%] 87% [85%,89%] 11% [10%,13%]
M0 steer_train 1.2 679 0% [0%,0%] 85% [83%,87%] 15% [13%,17%]
M0 steer_train 1.6 679 0% [0%,0%] 53% [51%,56%] 47% [44%,49%]
M1 clean 0.0 17 100% [100%,100%] 0% [0%,0%] 0% [0%,0%]
M1 correct_inject 0.4 9 100% [100%,100%] 0% [0%,0%] 0% [0%,0%]
M1 correct_inject 0.8 9 100% [100%,100%] 0% [0%,0%] 0% [0%,0%]
M1 correct_inject 1.0 9 100% [100%,100%] 0% [0%,0%] 0% [0%,0%]
M1 correct_inject 1.2 9 100% [100%,100%] 0% [0%,0%] 0% [0%,0%]
M1 correct_inject 1.6 9 89% [67%,100%] 0% [0%,0%] 11% [0%,33%]
M1 steer_heldout 0.4 2754 97% [93%,100%] 1% [0%,1%] 2% [0%,6%]
M1 steer_heldout 0.8 2754 66% [58%,73%] 20% [17%,24%] 14% [10%,18%]
M1 steer_heldout 1.0 2754 44% [35%,53%] 34% [27%,41%] 22% [17%,27%]
M1 steer_heldout 1.2 2754 26% [18%,35%] 39% [32%,47%] 34% [29%,40%]
M1 steer_heldout 1.6 2754 4% [2%,7%] 54% [50%,58%] 42% [40%,44%]
M1 steer_train 0.4 679 97% [90%,100%] 0% [0%,0%] 3% [0%,9%]
M1 steer_train 0.8 679 87% [79%,94%] 4% [1%,8%] 9% [5%,14%]
M1 steer_train 1.0 679 63% [51%,75%] 21% [11%,32%] 16% [10%,22%]
M1 steer_train 1.2 679 40% [26%,54%] 33% [22%,45%] 27% [20%,36%]
M1 steer_train 1.6 679 7% [3%,12%] 61% [55%,67%] 32% [29%,35%]

Reproduce

git clone https://github.com/JacoDuToit11/steering-resistance.git && cd steering-resistance
git checkout b16f6e72a9bb0c38dd20d0bab46579325178c793
python scripts/run.py configs/full_3b.yaml

run/ mirrors the full experiment directory: run_meta.json (manifest with artifact hashes), append-only eval jsonl, summaries, and the exact config.

Downloads last month
43
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JacoDuToit/steer-full_3b

Base model

Qwen/Qwen2.5-3B
Adapter
(1389)
this model