Instructions to use Jonnester/Purrence-3M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Jonnester/Purrence-3M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Jonnester/Purrence-3M", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Jonnester/Purrence-3M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Jonnester/Purrence-3M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Jonnester/Purrence-3M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Jonnester/Purrence-3M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Jonnester/Purrence-3M
- SGLang
How to use Jonnester/Purrence-3M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Jonnester/Purrence-3M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Jonnester/Purrence-3M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Jonnester/Purrence-3M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Jonnester/Purrence-3M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Jonnester/Purrence-3M with Docker Model Runner:
docker model run hf.co/Jonnester/Purrence-3M
Purrence-3M
I built Purrence as a base language model with 2,987,712 inference parameters and an 8,192 token context. I reuse one shared transformer block through six recurrent loops in a single pass, with Attention Residuals, half RoPE, document-aware key offsets, per-head XSA and sigmoid attention gates. I tie the embedding and output weights.
I include my FP32 inference weights, original tokenizer, native PyTorch loader and Transformers inference interface here. I use the model as a base language model without a chat template.
My training blog
I share some of my training thoughts and other considerations in my blog, if you are interested in reading!
Model architecture
I spend most of my parameter budget on one wide transformer block and reuse it six times in a single pass. I use LR AttnRes to read earlier residual sources with learned depth queries. The block weights are shared across loops, while the queries let different reads select different mixtures of earlier representations.
| Field | Value |
|---|---|
| Inference parameters | 2,987,712 |
| Hidden size | 384 |
| Blocks | One shared block applied six times |
| Blocks stored / applied | 1 / 6 |
| Sequential passes | 1 |
| Intermediate size | 2,096 |
| Attention heads | 6 |
| Key and value heads | 6 |
| Head dimension | 64 |
| Attention | Causal multihead attention with QK normalization, per head XSA and sigmoid gates |
| Residual routing | LR AttnRes with rank 128 |
| Depth query bank | 12 learned queries of width 128, totaling 1,536 parameters |
| Activation | Squared ReLU in the backbone |
| Normalization | RMS normalization with epsilon 1e-6. QK and routing normalization use 2^-23 |
| Positional encoding | Half RoPE on 32 of each head's 64 dimensions, with frequencies from 1 to 1/1,024 |
| Key offset | The unrotated key channels use the previous position within each document |
| Context length | 8,192 tokens |
| Vocabulary | 2,048 IDs, comprising 2,047 byte level BPE tokens and one structural boundary |
| Embeddings | Tied input and output, with 786,432 parameters |
| Logit cap | None |
| Released weights | FP32 Safetensors |
I normalize queries and keys before applying half RoPE. After causal attention, I use a learned XSA coefficient to adjust the component along the current token's value vector, then apply a learned sigmoid gate to each head. I use the tied embedding matrix for the final vocabulary projection.
During training, I used CE with NITP and a separate SwiGLU projector. That projector added 1,769,472 parameters, bringing the training total to 4,757,184. I discard it for inference, and it is absent from the weights in this repository. I trained with BF16 and evaluate the released weights in FP32.
How I trained it
I based my training data on the Pulvis v2 mixture. I trained this checkpoint from scratch for 5B tokens, with a 4B token plateau followed by a 1B token cooldown. I kept this release at 5B because my follow-up benchmark experiments showed little benefit from training beyond 5B tokens. I treat that as a finding from my experiments, rather than a general limit on what longer training can achieve.
What I tested
I ran these experiments while developing Purrence, across several training budgets and model revisions. I compare each result with its own control. I report cross entropy (CE) in nats per token, and use the length normalized local intelligence index where I report downstream results. These earlier experimental scores are separate from the final model's 8.4873 result.
Residual routing
I compared LR AttnRes with learned scalar residual weights and ordinary pre norm residual addition. In a three loop run, LR AttnRes reached held out cross-entropy (CE) of 2.6239 nats per token, compared with 2.7152 for the scalar baseline and 2.7336 for pre norm. It took about 11% longer than the scalar baseline and 17% longer than pre norm. The routing queries added 768 parameters in this three loop model.
Loop count and training length
In an approximately 1B token loop count sweep, held out CE was 2.6239 at three loops, 2.6116 at six, 2.6156 at nine, and 2.6101 at twelve. Twelve loops were marginally best at equal token count, while three processed tokens about 3.9 times faster. Over a longer horizon, the six loop run reached held out CE 2.2260 at 12B tokens versus 2.2993 for three loops. A twelve loop run reached 2.1979, but its length normalized local intelligence index was 5.758, below the six loop CE only run's 6.039. In my longest run, held out CE kept improving without sustained downstream gains. These results show that lower CE did not guarantee a higher downstream index.
The NITP projector
I tested CE with Next Implicit Token Prediction (NITP), using a SwiGLU projection head with auxiliary weight 1. At six loops and 12B tokens, held out CE was 2.2172 with NITP versus 2.2260 with CE alone. At three loops, NITP was slightly worse: 2.3037 versus 2.2993. A six loop linear NITP projector reached 2.2320, worse than the CE only control in that comparison, so I kept the SwiGLU projector.
NITP during cooldown
For a cooldown comparison, I restarted four continuations from the same six loop, linear-projector checkpoint at 10B tokens. Removing NITP gave the lowest held out CE at 12B, 2.2244 versus 2.2320 for the original cooldown, but its downstream index was lower, 6.029 versus 6.128. Freezing the projector barely changed CE, to 2.2322, while its index was 6.346. These older runs were evaluated in BF16, and deterministic execution settings changed between the original run and continuations. They used one seed, so I treated the downstream differences as suggestive rather than decisive and kept the later SwiGLU projector trainable.
DeepCrossAttention
I adapted DeepCrossAttention as another way to read earlier representations. In this short, roughly 1B token test, held out CE rose from 2.5885 for the matched LR AttnRes control to 2.6230, while counted training took 32.6% longer. I used one seed and my own unfused adaptation. I dropped this version from my design, without drawing a general conclusion about the method.
One wide block or two narrower blocks
I compared one width 384 block applied six times with two distinct width 256 blocks applied three times each, keeping six total block applications. Around 4.24B tokens, smoothed training CE was 2.2902 for one shared block and 2.3170 for two blocks. At twelve block applications the ordering reversed slightly, 2.3122 versus 2.2996 around 2.61B tokens. I did not find that one layout was universally better, though twelve applications were too slow for my budget.
Precision and feedforward choices
In a short precision screen, moving the FFN and output projection from the tested FP8 paths to BF16 reduced held out loss from 2.5889 to 2.5013, while recorded training time changed from 297.2 to 292.2 seconds. The output projection change accounted for the largest loss difference. I also compared backbone activations at the same parameter count: fused ReLU² reached held out loss 2.5085 versus 2.5137 for SwiGLU, with SwiGLU taking 6% longer. FFN width and optimizer settings also changed to fit the budget, so I cannot attribute that result to activation alone. In three paired BF16 seeds, width 2,096 beat 2,112 on held out loss in every pair, with means of 2.5003 and 2.5064 respectively. I kept width 2,096.
The Pulvis v2 architecture comparison
I compared the complete six loop model with a narrower Pulvis v2 shaped adaptation over the same 5B token stream. This changed several parts of the architecture at once, including block layout, attention, FFN, and context length, so it tests two architecture packages and does not isolate weight sharing. With one seed and full held out and downstream evaluation in FP32, the six loop model had lower held out CE, 2.0232 versus 2.1071, while the Pulvis shaped model had a higher length normalized index, 6.673 versus 6.172. Task scores were mixed. I read this as a comparison of the two complete configurations. I did not isolate any single component.
Why I abandoned two pass training
I also tried a two pass latent token variant. Even with zero Jacobi iterations, it processed three times as many positions as the corresponding single pass for the same number of real training tokens. In an earlier pilot, updates with three Jacobi rounds took a median 0.955 seconds versus 0.334 seconds with zero rounds. The added work was too costly for my budget, so I abandoned the two pass run and returned to the single pass six loop model.
My CPU benchmark results
I reran all six pinned benchmark tasks on my CPU on October 5, 2026, covering 18,928 examples in about 8.9 minutes. I used FP32 math, zero-shot prompts and the full vocabulary, with autocast and TF32 disabled.
I obtained a normalized Intelligence Index of 8.4873. My unrounded result was 8.487306383044526, and my raw-accuracy index was 6.997659441758839.
| Benchmark | Examples I evaluated | My accuracy | My normalized accuracy |
|---|---|---|---|
| ARC Easy | 2,376 | 34.18% | 34.55% |
| ARC Challenge | 1,172 | 20.39% | 23.63% |
| HellaSwag | 10,042 | 27.50% | 27.81% |
| PIQA | 1,838 | 55.33% | 56.64% |
| ArithMark 3 | 1,000 | 34.80% | 34.80% |
| ArithMark 2 | 2,500 | 26.44% | 25.92% |
I independently recalculated the predictions, accuracies and both index variants from the saved per-example outputs. I reproduced every task accuracy and normalized accuracy from my original native evaluation. I use ArithMark 3 in the index and report ArithMark 2 separately.
I used the pinned local evaluation protocol with harness revision d6de81643928d653435c431bae19945d41d32520. I am reporting local measurements and have not obtained independent leaderboard verification. I selected this checkpoint through repeated benchmark evaluations.
Reproducing the evaluation
I reproduce the five tasks used in my index with the zero shot protocol in evaluation_config.json: ARC Easy, ARC Challenge, HellaSwag, PIQA, and the official ArithMark 3 evaluation. I use CPU FP32 continuation scoring without a chat template. I pin the task datasets and evaluator revisions, check hashes for the weights, configuration, tokenizer and model code, and independently recalculate both accuracy measures from the saved predictions.
I use Python 3.11. From the downloaded model directory, I install the evaluation dependencies and run:
python -m pip install -r requirements-eval.txt
python reproduce_evaluation.py --model . --output results
I use an empty or new output directory. For a Hub copy, I sign in with hf auth login if access requires it, then run:
python reproduce_evaluation.py --model Jonnester/Purrence-3M --output results
I record the resolved Hub commit automatically. I can also pass --revision COMMIT_SHA to select a particular snapshot. I require the same verified model file hashes in either case.
I save full per example outputs in results/lm_eval_results.json and results/arithmark3/. In results/reaggregation.json, I record the model and evaluator hashes, dataset revisions, settings, example counts, recalculated metrics and index. My complete CPU reproduction covered 16,428 examples and returned 8.487306383044526. I report ArithMark 2 separately from my earlier six task run.
How I run it
I use Python 3.11 and install the pinned dependencies from the repository directory:
python -m pip install -r requirements.txt
python load_model.py --prompt "Once upon a time" --max-new-tokens 32
I use CPU inference and greedy generation by default. I load the Safetensors weights with strict parameter checks and verify the tokenizer hash. I also check for nonfinite weights and an unexpected parameter count.
I can generate text or score continuations through the Python API:
from load_model import load_model
from purrence_runtime.scorer import Scorer
model, tokenizer = load_model()
scorer = Scorer(model, tokenizer)
print(scorer.generate("Once upon a time", max_new_tokens=32))
print(scorer.score([(tokenizer.encode("The sky is"), tokenizer.encode(" blue."))]))
How I use Transformers
I use the same weights and tokenizer through the Transformers Auto classes. I enable trust_remote_code because I include the inference implementation for my custom architecture. I use FP32 and do not use a KV cache.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "Jonnester/Purrence-3M"
# I can set revision to a commit hash to pin this snapshot.
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo, trust_remote_code=True, torch_dtype=torch.float32
).eval()
inputs = tokenizer("Once upon a time", return_tensors="pt")
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=32, do_sample=False, use_cache=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
I preserve the original tokenizer behavior: literal <|endoftext|> in a prompt is ordinary text. I reserve token ID 2047 for structural boundaries, end of generation and padding.
My generation example
I ran the command above on my CPU with the prompt Once upon a time and a limit of 32 new tokens. I received this exact, unedited continuation:
, had a far apartment border. The border was 100 million years older than the next
I show only the generated continuation because that is what the command prints. I reached the token limit before completing the sentence.
My license
I have not selected a release license yet. I include the third-party license notices for their respective source components. I do not present those notices as a license for my model weights or the entire package.
- Downloads last month
- 173