Unlearned Checkpoint

Field Value
Unlearning method RMU
Base model google/gemma-2-2b-it
Target concept Gambling
Checkpoint type Full Model Weights
Rank / seed 100 / 42
Train eval protocol mc

Unlearning Configuration

Selected hyperparameters (from unlearned_checkpoints.json):

Parameter Value
alpha 50
delta_embed 0
k_features_embed 0
layer_id 8
layer_ids 6,7,8
lr 0.0003
n_tokens_edited 0
param_ids 6
setting_name S2_lid8_L678
steering 1000

Primary Unlearning Metrics (held-out test, MC protocol)

Headline scores used for checkpoint selection:

Metric Train (after unlearning) Test (after unlearning)
Efficacy 0.902 0.421
Specificity 1 0.947
Harmonic mean 0.948 0.583
Relearning QA (MC) โ€” 0.66

Full Evaluation (baseline โ†’ unlearned)

From evaluation/score_comparison.csv:

Metric Baseline (train) After unlearn (train) Baseline (test) After unlearn (test)
QA accuracy 0.76 0.3 0.82 0.58
QA fraction 1 0.098 1 0.579
SimDom accuracy 0.94 0.98 0.96 0.94
SimDom fraction 1 1 1 0.972
MMLU accuracy 0.52 0.52 0.551 0.528
MMLU fraction 1 1 1 0.924

Files in This Repository

File Description
unlearned_checkpoints.json Checkpoint metadata & hyperparameters
evaluation/evaluation_summary.json Full evaluation payload (train/test/relearning)
evaluation/score_comparison.csv Baseline vs. unlearned comparison table
Downloads last month
44
Safetensors
Model size
3B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for shirasko/gemma-2-2b-it-rmu-gambling

Finetuned
(1077)
this model

Collection including shirasko/gemma-2-2b-it-rmu-gambling