llmae-qwen-0.5b-s2

LLMAE checkpoint: Qwen2.5-0.5B, stage 2, + KL + additive codec: the paper's Qwen model.

From the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code). LLMAE turns a pretrained decoder-only LLM into a text autoencoder by reading the activations of K latent tokens at an intermediate layer through a structured attention mask; the same LLM decodes the latent back into the text.

Recipe

Backbone Qwen/Qwen2.5-0.5B
Bottleneck layer 15
Latent 256 x 896
LoRA r = 32, alpha = 64, q/k/v/o + latent-token embedding rows
Data 1M Pile-uncopyrighted + 400K C4-RealNewsLike (arkanathp/1M-stratified-pile-uncopyrighted, arkanathp/400K-stratified-c4-realnewslike), shuffled, 2 epochs
Latent KL / reference KL 0 / 1e-4
Codec yes (codec.pt, KL 1e-3 on the codec embedding)
Input / latent noise 1024 tokens / sigma 0.1

Warm-started from arkanathp/llmae-qwen-0.5b-s1 and trained for 2 further epochs with the codec (codec KL 1e-3, codec LR 1e-4). Ships codec.pt.

Reconstruction (C4-News-Stratified, 500 documents of up to 1024 tokens, greedy)

BLEU-4 Exact match Word edit distance
1.000 97.4% 0.005%

Usage

from llmae.loading import load_autoencoder   # https://github.com/arkanath/LLMAE
ae, tokenizer, config = load_autoencoder("arkanathp/llmae-qwen-0.5b-s2", device="cuda")
z = ae.encode("some text")["latent_hidden_states"]   # (1, 256, 896)
print(ae.decode(z)[0])

The tokenizer contains the historical <GIST_i> tokens, which are the paper's Embed tokens (the K learnable tokens whose activations at the bottleneck layer form the latent), and config.json carries an llmae_config block; llmae/compat.py in the code release bridges the historical naming.

Downloads last month
141
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arkanathp/llmae-qwen-0.5b-s2

Adapter
(466)
this model