Safety-WaRP (WSR-Tune) β gsm8k + beavertails-harmfulmix10p [WaRP] fine-tuned keep_ratio=0.1
β οΈ Research artifact β intentionally attacked model
This checkpoint was fine-tuned on a mixture that deliberately includes harmful prompt/response pairs (a harmful fine-tuning attack, cf. Qi et al., 2023). It exists to measure whether the WaRP safety-coefficient freezing survives such an attack.
Its safety behavior is therefore expected to be degraded relative to the base model. Do not deploy it. Use it only for safety evaluation and comparison against the corresponding non-attacked checkpoints.
kmseong/llama2_7b-chat-Safety-FT-lr5e-5 λ₯Ό μμμ μΌλ‘, WaRP(Weight space Rotation Process) μ¬νλΌλ―Έν°ν 곡κ°μμ
μμ κ΄λ ¨ κ³μ λ°©ν₯μ λκ²°ν μ± gsm8k + beavertails-harmfulmix10p [WaRP] λ‘ downstream fine-tuning ν λͺ¨λΈμ
λλ€.
- κ° weight matrix λ₯Ό μ
λ ₯ νμ±κ° 곡λΆμ°μ κ³ μ κΈ°μ
Uλ‘ νμ (C = W U) - μμ λ°μ΄ν°(circuit_breakers)μ λν gradient μ€μλ μμ
keep_ratioμ’νλ₯Ό λκ²° - λλ¨Έμ§("flat") μ’νλ§ νμ΅ β forward μ mask+detach λ‘ κ΅¬ν (non-freeze λ°©μ)
- token-wise constrained SFT (shallow-vs-deep) κ²°ν©
μ μ© λ²μ: q_proj, k_proj, v_proj, up_proj, down_proj / μ 체 32κ° layer / per-layer μ€μλ.
Training run
| base model | kmseong/llama2_7b-chat-Safety-FT-lr5e-5 |
| downstream data | gsm8k + beavertails-harmfulmix10p [WaRP] (8220 samples) |
| epochs / lr | 3 / 5e-05 |
| batch x grad_accum | 2 x 8 (effective 16) |
| optimizer / scheduler | adamw_torch / cosine |
| frozen safety coefficients | 451,927,702 / 4,496,293,888 (10.05%) |
| train wall-clock | 2146 s |
| train peak VRAM (device) | 96.63 GB |
- Downloads last month
- 5
Model tree for kmseong/llama2_7b_chat-WaRP-kr0.1-gsm8k_harmfulmix10p_lr5e-5
Base model
meta-llama/Llama-3.2-3B-Instruct