LMEnt 1B full control

This is the shared full control model for the Ancient Rome, Baseball, and artificial intelligence (AI) experiments in Can Concept Erasure Reproduce Concept Exclusion? A Matched Evaluation of EMBER, RMU, and SNMF by Gal Barak, Tamar Tabbach, Itamar Stahl, and Adam Fleisher. It is an English causal language model with the OLMo2 1B architecture. It was trained on the LMEnt entity-annotated Wikipedia corpus with no concept-linked chunks excluded from the loss. It is not an instruction-tuned model.

The same checkpoint is the starting point for all selected EMBER, RMU, SNMF, RMU+EMBER, and SNMF+EMBER edits in the paper. The three separately trained concept-excluded twins are lment-1b-norome-2e-b131k, lment-1b-nobaseball-2e-b131k, and lment-1b-noai-2e-b131k. Those twins mask concept-linked chunks during training; the edited models alter this completed control afterward. The twin and edited-model roles should not be conflated.

Training and provenance

Field Value
Architecture Olmo2ForCausalLM; 18 layers, hidden size 2,048, 16 attention heads; untied input and output embeddings
Training data LMEnt entity-annotated Wikipedia corpus
Training Two epochs; 54,832 optimizer steps; 131,072 tokens per global batch
Optimizer AdamW; peak learning rate 4e-4; 2,000 warmup steps; cosine decay to 4e-5; weight decay 0.05; gradient clipping at 1.0
Sequence lengths 64–2,048, grown with grow_p2
Seeds Model initialization 12,536; data order 0
Original run TAU SLURM job 853707, untaught-control-1b-2e-b131k_20260905_221156, final step 54832

This repository contains the converted final model, tokenizer, and configuration. The safetensors index references two weight shards. The model was trained using the OLMo2 architecture in OLMo-core; it should not be mistaken for a fine-tune of a separately released OLMo2 checkpoint.

Use

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "itamarstahl/lment-1b-control-2e-b131k"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto")

This model is intended as a research reference for matched concept-exclusion and concept-erasure comparisons. It is a base language model; evaluate multiple-choice options by their text likelihood rather than by parsing a generated answer letter. The paper reports held-out concept results with likelihood-ranked answer accuracy and proximity to the corresponding twin using correct-answer negative log-likelihood and teacher-forced full-vocabulary KL divergence.

Limitations

This control was trained on Wikipedia-derived text and may reproduce inaccuracies or biases in that material. It has no instruction or safety tuning. The experiments test specific concept datasets and evaluation protocols; their results do not establish general knowledge removal or absence of a concept from a twin. No license has been asserted for these weights in this card.

Citation

Gal Barak, Tamar Tabbach, Itamar Stahl, and Adam Fleisher. Can Concept Erasure Reproduce Concept Exclusion? A Matched Evaluation of EMBER, RMU, and SNMF. 2026.

Downloads last month
176
Safetensors
Model size
1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including itamarstahl/lment-1b-control-2e-b131k