GenMem
GenMem provides MemRetriever and MemEvolver checkpoints and a frozen four-layer RQ-KMeans codebook. The encoding utility maps your own memory text to a SID using Qwen3-Embedding-0.6B; it does not require loading the 8B memory agents.
Released artifacts
| Directory | Artifact |
|---|---|
memr/ |
Qwen3-8B MemRetriever, v11 RL step50 |
meme/ |
Qwen3-8B MemEvolver, v39 RL step80 |
codebook/ |
Frozen 48 × 16 × 8 × 8 codebooks, each center a 1024-dimensional float32 vector |
prompts.json |
Experiment inference system prompts and memory_lookup tool schema |
The memory agents include 80 added SID tokens. Use the tokenizer shipped with
each checkpoint. This repository is a multi-component bundle: specify memr or
meme as the Transformers subfolder, not the repository root as a standalone model.
No server model, private memory bank, training examples, trajectories, optimizer
states, evaluation results, or credentials are included.
Memory to SID
Download the small encoder files, then install dependencies:
hf download chuchuxwx/GenMem --local-dir GenMem --include 'codebook/*' 'encode_memory.py' 'requirements.txt'
cd GenMem
pip install -r requirements.txt
python encode_memory.py --memory 'When a tool call fails, inspect the error, fix its arguments, and verify the returned result.'
CPU is the default; pass --device cuda:0 to use an available GPU. The embedding
model downloads separately from Qwen/Qwen3-Embedding-0.6B at the revision pinned
in codebook/config.json. A local matching snapshot can be supplied with
--embedding-model. Output contains sid_codes and the concatenated four SID
tokens, for example the illustrative format
<SID_L1_5><SID_L2_3><SID_L3_2><SID_L4_1> (not a claimed output for the text above).
from encode_memory import SIDEncoder
encoder = SIDEncoder()
result = encoder.encode("Your reusable experience text")
print(result[0])
The exact instruction prefix is stored in codebook/config.json. Input is
instruction plus the supplied text, without a chat template. The encoder uses
last-EOS pooling when EOS is present, otherwise the last token, followed by L2
normalization. It processes texts individually to avoid the historical
padded-batch last-token ambiguity. Inputs exceeding 8192 tokens are truncated.
CPU float32 and CUDA float16 can differ near quantization boundaries.
At each level the encoder minimizes
0.5 * squared_euclidean(residual, center) + 0.5 * (1 - cosine(residual, center)),
then subtracts the selected center. Do not normalize each residual, independently
cluster all four levels, renumber centers, or substitute another embedding model
or prompt. The four codebooks address up to 49,152 slots, not unique identifiers
for individual texts. Multiple memories can share a SID.
The NumPy encoder is checked against the original weighted-distance implementation
on synthetic normalized vectors; see validation.json. This verifies quantizer
parity, not exact reproduction of every historical text-to-SID assignment.
Historical input construction, padding and precision can affect those assignments.
Supplying only a rewritten summary instead of its original memory text can also
produce a different SID.
Loading the memory agents
import json
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "chuchuxwx/GenMem"
role = "memr" # or "meme"
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=role, extra_special_tokens={})
model = AutoModelForCausalLM.from_pretrained(
repo, subfolder=role, torch_dtype="auto", device_map="auto"
)
with open(hf_hub_download(repo, "prompts.json")) as f:
prompts = json.load(f)
system = prompts["MEMR_SYSTEM_PROMPT" if role == "memr" else "MEME_SYSTEM_PROMPT"]
content = "<query> Your retrieval query </query>" if role == "memr" else "<trajectory> Your execution trajectory </trajectory>"
prompt = tokenizer.apply_chat_template(
[{"role": "system", "content": system}, {"role": "user", "content": content}],
tools=prompts["MEMORY_LOOKUP_TOOLS"], tokenize=False,
add_generation_prompt=True, enable_thinking=True,
)
This prepares the initial prompt, not a complete agent runtime. Your application
must parse memory_lookup, validate the four SID indices, fetch the corresponding
memory from your own bank, and supply the tool response before requesting the final
answer. MemE also reasons after the tool response. Do not decode SID output with
skip_special_tokens=True. The MemE v39 prompt is specialized for failed trajectories.
The experiment used bounded generation and tool-call handling; a bare unlimited
generate call is not a reproduction of that pipeline.
Responsible use and limitations
Generated memories may contain mistakes, repetitive content, or task-specific answers. Validate updates before applying them and keep recoverable bank snapshots. SID encoding does not guarantee useful retrieval, unique assignment, or exact reconstruction of the source text. No benchmark performance is promised by the standalone examples. Do not submit secrets or personal data as memory content.
The checkpoints derive from Qwen3-8B, and the text encoder uses Qwen3-Embedding-0.6B.
License
Code, codebooks, and fine-tuned weights are released under Apache-2.0. See LICENSE
and NOTICE. The Qwen base models retain their upstream attribution and license.