EngramEdit: MQuAKE · 3K

Overview

Conditional memory expands an LLM's capacity through learned n-gram embeddings that participate in its computation. EngramEdit enables decoupled knowledge updates by updating these embeddings while keeping the Transformer backbone fixed. Revised facts remain usable across different expressions and in multi-hop reasoning, while unrelated knowledge and general capabilities are largely preserved.

This repository provides the EngramEdit checkpoint for the factual edits specified by 3,000 MQuAKE cases on LongCat-Flash-Lite. The checkpoint contains the learned conditional memory updates. Load it together with the base model using the EngramEdit code below.

Links

📄 Paper · 🤗 Hugging Face Paper · 🌐 Project Page · 💻 GitHub · 🧪 Usage Example · 📊 Evaluation

Checkpoints: 🤗 CounterFact · 2K · 🤗 ZsRE · 2K · 🤗 MQuAKE · 3K

Usage

Clone the code repository, follow its installation instructions, and run the example from that checkout. Download the base model once:

hf download meituan-longcat/LongCat-Flash-Lite --local-dir data/LongCat-Flash-Lite

The example compares the model's response before and after loading the checkpoint. No editing, expression generation, or frequency-cache preparation is needed.

Each MQuAKE case may specify multiple factual edits. This example queries one edited fact; the repository also provides multi-hop evaluation. Counterfactual targets are benchmark edits, not claims about real-world facts.

from pathlib import Path

import torch
from huggingface_hub import hf_hub_download

from experiments.utils import load_model
from EngramEdit import prepare_engramedit_model

# Load the base model.
model, tokenizer = load_model("data/LongCat-Flash-Lite", "bfloat16")
prompt = "Fernando Santos is a citizen of"
target = "United Kingdom"


@torch.inference_mode()
def answer(prompt):
    inputs = tokenizer(prompt, return_tensors="pt").to(
        model.get_input_embeddings().weight.device
    )
    output = model.generate(
        **inputs, do_sample=False, max_new_tokens=32,
        pad_token_id=tokenizer.pad_token_id,
    )
    return tokenizer.decode(
        output[0, inputs.input_ids.shape[1]:], skip_special_tokens=True
    )


# Inspect the query and the target fact.
print("Prompt:", prompt)
# Prompt: Fernando Santos is a citizen of
print("Target answer:", target)
# Target answer: United Kingdom

# Responses below are illustrative, not recorded checkpoint outputs.
print("Before loading:", answer(prompt))
# Before loading: Portugal

# Download and load the learned conditional memory updates.
state_path = hf_hub_download(
    repo_id="ModalityDance/EngramEdit-LongCat-Flash-Lite-MQuAKE-3K",
    filename="engramedit_state.pt",
    local_dir="checkpoints/mquake",
)
model = prepare_engramedit_model(model, state_dir=Path(state_path).parent)

# Ask the same question with the checkpoint loaded.
print("After loading:", answer(prompt))
# After loading: United Kingdom

Citation

@misc{cai2026engramedit,
  title  = {EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory},
  author = {Hongru Cai and Ran Wei and Wenjie Wang and Chengfa Wu and Ning Song and Yongqi Li and Wenjie Li},
  year   = {2026},
  url    = {https://github.com/ModalityDance/EngramEdit}
}

License

This checkpoint is released under the MIT License. The base model is distributed separately under its own license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ModalityDance/EngramEdit-LongCat-Flash-Lite-MQuAKE-3K

Finetuned
(9)
this model

Collection including ModalityDance/EngramEdit-LongCat-Flash-Lite-MQuAKE-3K