Safety-WaRP (WSR-Tune) β€” gsm8k + beavertails-harmfulmix10p [WaRP] fine-tuned keep_ratio=0.1

⚠️ Research artifact β€” intentionally attacked model

This checkpoint was fine-tuned on a mixture that deliberately includes harmful prompt/response pairs (a harmful fine-tuning attack, cf. Qi et al., 2023). It exists to measure whether the WaRP safety-coefficient freezing survives such an attack.

Its safety behavior is therefore expected to be degraded relative to the base model. Do not deploy it. Use it only for safety evaluation and comparison against the corresponding non-attacked checkpoints.

kmseong/llama2_7b-chat-Safety-FT-lr5e-5 λ₯Ό μ‹œμž‘μ μœΌλ‘œ, WaRP(Weight space Rotation Process) μž¬νŒŒλΌλ―Έν„°ν™” κ³΅κ°„μ—μ„œ μ•ˆμ „ κ΄€λ ¨ κ³„μˆ˜ λ°©ν–₯을 λ™κ²°ν•œ 채 gsm8k + beavertails-harmfulmix10p [WaRP] 둜 downstream fine-tuning ν•œ λͺ¨λΈμž…λ‹ˆλ‹€.

  • 각 weight matrix λ₯Ό μž…λ ₯ ν™œμ„±κ°’ κ³΅λΆ„μ‚°μ˜ κ³ μœ κΈ°μ € U 둜 νšŒμ „ (C = W U)
  • μ•ˆμ „ 데이터(circuit_breakers)에 λŒ€ν•œ gradient μ€‘μš”λ„ μƒμœ„ keep_ratio μ’Œν‘œλ₯Ό 동결
  • λ‚˜λ¨Έμ§€("flat") μ’Œν‘œλ§Œ ν•™μŠ΅ β€” forward 의 mask+detach 둜 κ΅¬ν˜„ (non-freeze 방식)
  • token-wise constrained SFT (shallow-vs-deep) κ²°ν•©

적용 λ²”μœ„: q_proj, k_proj, v_proj, up_proj, down_proj / 전체 32개 layer / per-layer μ€‘μš”λ„.

Training run

base model kmseong/llama2_7b-chat-Safety-FT-lr5e-5
downstream data gsm8k + beavertails-harmfulmix10p [WaRP] (8220 samples)
epochs / lr 3 / 5e-05
batch x grad_accum 2 x 8 (effective 16)
optimizer / scheduler adamw_torch / cosine
frozen safety coefficients 451,927,702 / 4,496,293,888 (10.05%)
train wall-clock 2146 s
train peak VRAM (device) 96.63 GB
Downloads last month
5
Safetensors
Model size
7B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for kmseong/llama2_7b_chat-WaRP-kr0.1-gsm8k_harmfulmix10p_lr5e-5