GenMem

GenMem provides MemRetriever and MemEvolver checkpoints and a frozen four-layer RQ-KMeans codebook. The encoding utility maps your own memory text to a SID using Qwen3-Embedding-0.6B; it does not require loading the 8B memory agents.

Released artifacts

Directory Artifact
memr/ Qwen3-8B MemRetriever, v11 RL step50
meme/ Qwen3-8B MemEvolver, v39 RL step80
codebook/ Frozen 48 × 16 × 8 × 8 codebooks, each center a 1024-dimensional float32 vector
prompts.json Experiment inference system prompts and memory_lookup tool schema

The memory agents include 80 added SID tokens. Use the tokenizer shipped with each checkpoint. This repository is a multi-component bundle: specify memr or meme as the Transformers subfolder, not the repository root as a standalone model. No server model, private memory bank, training examples, trajectories, optimizer states, evaluation results, or credentials are included.

Memory to SID

Download the small encoder files, then install dependencies:

hf download chuchuxwx/GenMem --local-dir GenMem --include 'codebook/*' 'encode_memory.py' 'requirements.txt'
cd GenMem
pip install -r requirements.txt
python encode_memory.py --memory 'When a tool call fails, inspect the error, fix its arguments, and verify the returned result.'

CPU is the default; pass --device cuda:0 to use an available GPU. The embedding model downloads separately from Qwen/Qwen3-Embedding-0.6B at the revision pinned in codebook/config.json. A local matching snapshot can be supplied with --embedding-model. Output contains sid_codes and the concatenated four SID tokens, for example the illustrative format <SID_L1_5><SID_L2_3><SID_L3_2><SID_L4_1> (not a claimed output for the text above).

from encode_memory import SIDEncoder
encoder = SIDEncoder()
result = encoder.encode("Your reusable experience text")
print(result[0])

The exact instruction prefix is stored in codebook/config.json. Input is instruction plus the supplied text, without a chat template. The encoder uses last-EOS pooling when EOS is present, otherwise the last token, followed by L2 normalization. It processes texts individually to avoid the historical padded-batch last-token ambiguity. Inputs exceeding 8192 tokens are truncated. CPU float32 and CUDA float16 can differ near quantization boundaries.

At each level the encoder minimizes 0.5 * squared_euclidean(residual, center) + 0.5 * (1 - cosine(residual, center)), then subtracts the selected center. Do not normalize each residual, independently cluster all four levels, renumber centers, or substitute another embedding model or prompt. The four codebooks address up to 49,152 slots, not unique identifiers for individual texts. Multiple memories can share a SID.

The NumPy encoder is checked against the original weighted-distance implementation on synthetic normalized vectors; see validation.json. This verifies quantizer parity, not exact reproduction of every historical text-to-SID assignment. Historical input construction, padding and precision can affect those assignments. Supplying only a rewritten summary instead of its original memory text can also produce a different SID.

Loading the memory agents

import json
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "chuchuxwx/GenMem"
role = "memr"  # or "meme"
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=role, extra_special_tokens={})
model = AutoModelForCausalLM.from_pretrained(
    repo, subfolder=role, torch_dtype="auto", device_map="auto"
)
with open(hf_hub_download(repo, "prompts.json")) as f:
    prompts = json.load(f)
system = prompts["MEMR_SYSTEM_PROMPT" if role == "memr" else "MEME_SYSTEM_PROMPT"]
content = "<query> Your retrieval query </query>" if role == "memr" else "<trajectory> Your execution trajectory </trajectory>"
prompt = tokenizer.apply_chat_template(
    [{"role": "system", "content": system}, {"role": "user", "content": content}],
    tools=prompts["MEMORY_LOOKUP_TOOLS"], tokenize=False,
    add_generation_prompt=True, enable_thinking=True,
)

This prepares the initial prompt, not a complete agent runtime. Your application must parse memory_lookup, validate the four SID indices, fetch the corresponding memory from your own bank, and supply the tool response before requesting the final answer. MemE also reasons after the tool response. Do not decode SID output with skip_special_tokens=True. The MemE v39 prompt is specialized for failed trajectories. The experiment used bounded generation and tool-call handling; a bare unlimited generate call is not a reproduction of that pipeline.

Responsible use and limitations

Generated memories may contain mistakes, repetitive content, or task-specific answers. Validate updates before applying them and keep recoverable bank snapshots. SID encoding does not guarantee useful retrieval, unique assignment, or exact reconstruction of the source text. No benchmark performance is promised by the standalone examples. Do not submit secrets or personal data as memory content.

The checkpoints derive from Qwen3-8B, and the text encoder uses Qwen3-Embedding-0.6B.

License

Code, codebooks, and fine-tuned weights are released under Apache-2.0. See LICENSE and NOTICE. The Qwen base models retain their upstream attribution and license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for chuchuxwx/GenMem

Finetuned
Qwen/Qwen3-8B
Finetuned
(2161)
this model