Unlearned Checkpoint

Field Value
Unlearning method SNMF
Base model google/gemma-2-2b-it
Target concept Gambling
Checkpoint type Full Model Weights
Rank / seed 100 / 42
Train eval protocol mc

Unlearning Configuration

Selected hyperparameters (from unlearned_checkpoints.json):

Parameter Value
coverage_thresh 0.95
delta_embed 0
delta_in 10
delta_out 7
feature_source all
k_features_embed 0
k_features_mlp_in 26
k_features_mlp_out 6
layer_hi_in 12
layer_hi_out 17
layer_lo_in 0
layer_lo_out 9
n_tokens_edited 0
ratio_thresh 2
w_mode both

Primary Unlearning Metrics (held-out test, MC protocol)

Headline scores used for checkpoint selection:

Metric Train (after unlearning) Test (after unlearning)
Efficacy 0.902 0.632
Specificity 0.808 0.738
Harmonic mean 0.852 0.681
Relearning QA (MC) โ€” 0.58

Full Evaluation (baseline โ†’ unlearned)

From evaluation/score_comparison.csv:

Metric Baseline (train) After unlearn (train) Baseline (test) After unlearn (test)
QA accuracy 0.76 0.3 0.82 0.46
QA fraction 1 0.098 1 0.368
SimDom accuracy 0.94 0.78 0.96 0.84
SimDom fraction 1 0.768 1 0.831
MMLU accuracy 0.52 0.48 0.551 0.45
MMLU fraction 1 0.852 1 0.664

Files in This Repository

File Description
unlearned_checkpoints.json Checkpoint metadata & hyperparameters
evaluation/evaluation_summary.json Full evaluation payload (train/test/relearning)
evaluation/score_comparison.csv Baseline vs. unlearned comparison table
Downloads last month
47
Safetensors
Model size
3B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for shirasko/gemma-2-2b-it-snmf-gambling

Finetuned
(1080)
this model

Collection including shirasko/gemma-2-2b-it-snmf-gambling