QELVRA
A 1.85M-parameter byte-level Transformer language model β trained from scratch
Every token is a raw byte. No BPE, no sentencepiece, no third-party tokenizer:
256 byte IDs in, 256 byte IDs out, and a 4-block causal Transformer in between β
written in plain PyTorch and released as raw .pt state dicts you can load in
five lines.
Built as an experimental from-scratch language model β architecture, training loop and tokenizer implemented directly, then measured and documented honestly.
Contents
- Highlights
- Architecture
- Training
- Dataset
- Tokenizer
- Download
- Quick start
- Model files
- Honest scope
- License
- Links
Highlights
- π§ Built from scratch β model, training loop and tokenizer written directly in
PyTorch: no Hugging Face
transformers, no pretrained weights, no copied config. - π€ True byte-level tokens β UTF-8 bytes mapped to IDs 0β255. The whole
tokenizer is
list(text.encode("utf-8")); decoding isbytes(ids).decode(). - π¦ Tiny and inspectable β 1,853,568 parameters, 4 layers Γ 4 heads, 192-dim embeddings, 128-token context. Both checkpoints are ~22.6 MB and load on CPU in well under a second.
- π§ͺ Measured, not claimed β best validation loss 0.0341 at step 5,000, trained in 2.66 minutes on a GPU; the exact hyperparameters and token counts are published below and embedded in every checkpoint.
- π₯ CPU and CUDA β the reference inference script picks CUDA when
torch.cuda.is_available()is true and falls back to CPU automatically. - π Published with checksums β every file ships with a SHA-256 entry so you can verify exactly what you downloaded.
Architecture
| Component | Value |
|---|---|
| Architecture | Decoder-only causal Transformer (pre-norm blocks) |
| Vocabulary | 256 (one token = one UTF-8 byte) |
| Context length | 128 tokens (= 128 bytes) |
| Embedding dimension | 192 |
| Attention heads | 4 (head dim 48) |
| Transformer layers | 4 |
| Feed-forward dimension | 768 |
| Dropout | 0.05 |
| QKV projection | fused Linear(192 β 576) with bias |
| Final layer norm | yes |
| Weight tying | lm_head tied to token_embedding |
| Parameters | 1,853,568 |
token bytes ββΆ token_embedding (256Γ192) ββ
βββΆ + ββΆ 4 Γ [ LayerNorm ββΆ Causal MHA ββΆ + ββΆ LayerNorm ββΆ 192β768β192 FFN ββΆ + ]
position ids ββΆ position_embedding (128Γ192) β β
final LayerNorm ββΆ lm_head ββΆ byte logits
Training
| Setting | Value |
|---|---|
| Training steps | 5,000 |
| Batch size | 32 |
| Context length | 128 |
| Learning rate | 0.0003 |
| Weight decay | 0.01 |
| Gradient clipping | 1.0 |
| Training time | 2.66 minutes (GPU) |
| Best validation loss | 0.0341 |
| Final training loss | 0.0348 |
Dataset
| Split | Tokens |
|---|---|
| Total | 3,402,000 |
| Training | 3,061,800 |
| Validation | 170,100 |
| Test | 170,100 |
Tokenizer
QELVRA uses UTF-8 byte-level tokenization.
- Encoding rule: each UTF-8 byte is represented by its integer value (0β255).
- Decoding rule: token IDs are converted back to bytes and decoded as UTF-8.
- Token ID range:
[0, 255]β every ID is always valid, so no unknown-token handling is ever needed.
ids = list("Hello".encode("utf-8")) # [72, 101, 108, 108, 111]
text = bytes(ids).decode("utf-8") # "Hello"
Download
Weights β Hugging Face: jojin1709/qelvra
# with the Hugging Face CLI
pip install -U huggingface_hub
hf download jojin1709/qelvra
# or in Python
# from huggingface_hub import hf_hub_download
# path = hf_hub_download("jojin1709/qelvra", "qelvra_best.pt", repo_type="model")
Code β GitHub: jojin1709/qelvra
git clone https://github.com/jojin1709/qelvra.git
Quick start
The reference inference package (model.py, tokenizer.py, generate.py) lives in
inference/ on GitHub.
It only reads the checkpoints β nothing is retrained or rewritten.
pip install -r inference/requirements.txt
# 1 β loading verification test (prints config, param count, PASS/FAIL)
python inference/generate.py --verify
# 2 β generate
python inference/generate.py --prompt "Hello" --max-new-tokens 100
# 3 β interactive mode
python inference/generate.py --interactive
Generation controls:
python inference/generate.py --prompt "The future of AI" `
--max-new-tokens 150 --temperature 0.8 --top-k 40 --top-p 0.9 --seed 42
| Option | Default | Meaning |
|---|---|---|
--max-new-tokens |
100 | bytes to generate |
--temperature |
0.8 | sampling temperature; 0 = greedy |
--top-k |
40 | top-k sampling, 0 disables |
--top-p |
0.9 | nucleus sampling |
--seed |
none | reproducible output |
--device |
auto | auto / cpu / cuda |
--checkpoint |
auto-found | path to qelvra_best.pt or qelvra_final.pt |
The script looks for the checkpoint next to itself, in the working directory, in
QELVRA_MODEL_DIR, and in the usual download folders β the resolved path is printed
on every run.
Example output (seed 42):
PROMPT : The future of AI
CONTINUATION :
enstral twithe keeping the model computationally efficient.
Qelvra focuses on e
Model files
| File | Size | Purpose |
|---|---|---|
qelvra_best.pt |
22.6 MB | checkpoint with the best validation loss |
qelvra_final.pt |
22.6 MB | final checkpoint (step 5,000) |
config.json |
793 B | model configuration and training summary |
metadata.json |
436 B | short machine-readable model card |
tokenizer.json |
310 B | tokenizer specification |
README.md |
β | this model card |
SHA256SUMS.txt |
β | SHA-256 checksums for every file above |
Each .pt file contains model_state_dict, optimizer_state_dict,
model_config, step, and the recorded losses β so the architecture always comes
from the checkpoint itself, never from guesswork.
Verify what you downloaded:
Get-FileHash -Algorithm SHA256 .\qelvra_best.pt # 6f1f53aa7da5ea44a8a02113fa2c142fc552ab6cb65a11641b2dc43dea97cb7e
Honest scope
| Claim | What it actually means |
|---|---|
| Val loss 0.0341 | Very low for a language model β the model largely memorised a small 3.4M-byte corpus. Low loss here is not a claim of broad generalization. |
| 1.85M parameters | Toy scale. It produces locally coherent English about its own training text and can drift into malformed bytes on prompts it never saw. |
| Byte-level tokens | 1 token = 1 byte, so the 128-token context covers only 128 bytes of text. |
Raw .pt format |
Not a Hugging Face transformers model β there is no config.json in transformers format and no AutoModelForCausalLM support. Load it with the reference scripts. |
| Feed-forward activation | Not stored in the checkpoint (activations have no weights). Measured empirically across gelu/relu/silu β gelu reproduces the recorded validation loss most closely, and --activation lets you compare. |
| Status | experimental (see metadata.json). |
License
MIT β free to use, modify and redistribute with the copyright notice retained.
Links
- Model weights: https://huggingface.co/jojin1709/qelvra
- Source code: https://github.com/jojin1709/qelvra
- Issues: https://github.com/jojin1709/qelvra/issues
Developed by JOJIN JOHN
- Downloads last month
- -