Instructions to use arkanathp/llmae-qwen-0.5b-s2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use arkanathp/llmae-qwen-0.5b-s2 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B") model = PeftModel.from_pretrained(base_model, "arkanathp/llmae-qwen-0.5b-s2") - Notebooks
- Google Colab
- Kaggle
llmae-qwen-0.5b-s2
LLMAE checkpoint: Qwen2.5-0.5B, stage 2, + KL + additive codec: the paper's Qwen model.
From the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code). LLMAE turns a pretrained decoder-only LLM into a text autoencoder by reading the activations of K latent tokens at an intermediate layer through a structured attention mask; the same LLM decodes the latent back into the text.
Recipe
| Backbone | Qwen/Qwen2.5-0.5B |
| Bottleneck layer | 15 |
| Latent | 256 x 896 |
| LoRA | r = 32, alpha = 64, q/k/v/o + latent-token embedding rows |
| Data | 1M Pile-uncopyrighted + 400K C4-RealNewsLike (arkanathp/1M-stratified-pile-uncopyrighted, arkanathp/400K-stratified-c4-realnewslike), shuffled, 2 epochs |
| Latent KL / reference KL | 0 / 1e-4 |
| Codec | yes (codec.pt, KL 1e-3 on the codec embedding) |
| Input / latent noise | 1024 tokens / sigma 0.1 |
Warm-started from arkanathp/llmae-qwen-0.5b-s1 and trained for 2 further epochs with the codec (codec KL 1e-3, codec LR 1e-4). Ships codec.pt.
Reconstruction (C4-News-Stratified, 500 documents of up to 1024 tokens, greedy)
| BLEU-4 | Exact match | Word edit distance |
|---|---|---|
| 1.000 | 97.4% | 0.005% |
Usage
from llmae.loading import load_autoencoder # https://github.com/arkanath/LLMAE
ae, tokenizer, config = load_autoencoder("arkanathp/llmae-qwen-0.5b-s2", device="cuda")
z = ae.encode("some text")["latent_hidden_states"] # (1, 256, 896)
print(ae.decode(z)[0])
The tokenizer contains the historical <GIST_i> tokens, which are the paper's
Embed tokens (the K learnable tokens whose activations at the bottleneck layer
form the latent), and config.json carries an llmae_config block;
llmae/compat.py in the code release bridges the historical naming.
- Downloads last month
- 141
Model tree for arkanathp/llmae-qwen-0.5b-s2
Base model
Qwen/Qwen2.5-0.5B