QELVRA

A 1.85M-parameter byte-level Transformer language model β€” trained from scratch

Every token is a raw byte. No BPE, no sentencepiece, no third-party tokenizer: 256 byte IDs in, 256 byte IDs out, and a 4-block causal Transformer in between β€” written in plain PyTorch and released as raw .pt state dicts you can load in five lines.

Built as an experimental from-scratch language model β€” architecture, training loop and tokenizer implemented directly, then measured and documented honestly.


License PyTorch Hugging Face Model GitHub Stars


         


Download weights from Hugging Face  Get the code on GitHub  Run it locally  Files and checksums


Contents


Highlights

  • πŸ”§ Built from scratch β€” model, training loop and tokenizer written directly in PyTorch: no Hugging Face transformers, no pretrained weights, no copied config.
  • πŸ”€ True byte-level tokens β€” UTF-8 bytes mapped to IDs 0–255. The whole tokenizer is list(text.encode("utf-8")); decoding is bytes(ids).decode().
  • πŸ“¦ Tiny and inspectable β€” 1,853,568 parameters, 4 layers Γ— 4 heads, 192-dim embeddings, 128-token context. Both checkpoints are ~22.6 MB and load on CPU in well under a second.
  • πŸ§ͺ Measured, not claimed β€” best validation loss 0.0341 at step 5,000, trained in 2.66 minutes on a GPU; the exact hyperparameters and token counts are published below and embedded in every checkpoint.
  • πŸ–₯ CPU and CUDA β€” the reference inference script picks CUDA when torch.cuda.is_available() is true and falls back to CPU automatically.
  • πŸ“œ Published with checksums β€” every file ships with a SHA-256 entry so you can verify exactly what you downloaded.

Architecture

Component Value
Architecture Decoder-only causal Transformer (pre-norm blocks)
Vocabulary 256 (one token = one UTF-8 byte)
Context length 128 tokens (= 128 bytes)
Embedding dimension 192
Attention heads 4 (head dim 48)
Transformer layers 4
Feed-forward dimension 768
Dropout 0.05
QKV projection fused Linear(192 β†’ 576) with bias
Final layer norm yes
Weight tying lm_head tied to token_embedding
Parameters 1,853,568
token bytes ─▢ token_embedding (256Γ—192) ─┐
                                          β”œβ”€β–Ά + ─▢ 4 Γ— [ LayerNorm ─▢ Causal MHA ─▢ + ─▢ LayerNorm ─▢ 192β†’768β†’192 FFN ─▢ + ]
position ids ─▢ position_embedding (128Γ—192) β”˜                                                    β”‚
                                                                                          final LayerNorm ─▢ lm_head ─▢ byte logits

Training

Setting Value
Training steps 5,000
Batch size 32
Context length 128
Learning rate 0.0003
Weight decay 0.01
Gradient clipping 1.0
Training time 2.66 minutes (GPU)
Best validation loss 0.0341
Final training loss 0.0348

Dataset

Split Tokens
Total 3,402,000
Training 3,061,800
Validation 170,100
Test 170,100

Tokenizer

QELVRA uses UTF-8 byte-level tokenization.

  • Encoding rule: each UTF-8 byte is represented by its integer value (0–255).
  • Decoding rule: token IDs are converted back to bytes and decoded as UTF-8.
  • Token ID range: [0, 255] β€” every ID is always valid, so no unknown-token handling is ever needed.
ids = list("Hello".encode("utf-8"))   # [72, 101, 108, 108, 111]
text = bytes(ids).decode("utf-8")     # "Hello"

Download

Weights β€” Hugging Face: jojin1709/qelvra

# with the Hugging Face CLI
pip install -U huggingface_hub
hf download jojin1709/qelvra

# or in Python
# from huggingface_hub import hf_hub_download
# path = hf_hub_download("jojin1709/qelvra", "qelvra_best.pt", repo_type="model")

Code β€” GitHub: jojin1709/qelvra

git clone https://github.com/jojin1709/qelvra.git

Quick start

The reference inference package (model.py, tokenizer.py, generate.py) lives in inference/ on GitHub. It only reads the checkpoints β€” nothing is retrained or rewritten.

pip install -r inference/requirements.txt

# 1 β€” loading verification test (prints config, param count, PASS/FAIL)
python inference/generate.py --verify

# 2 β€” generate
python inference/generate.py --prompt "Hello" --max-new-tokens 100

# 3 β€” interactive mode
python inference/generate.py --interactive

Generation controls:

python inference/generate.py --prompt "The future of AI" `
  --max-new-tokens 150 --temperature 0.8 --top-k 40 --top-p 0.9 --seed 42
Option Default Meaning
--max-new-tokens 100 bytes to generate
--temperature 0.8 sampling temperature; 0 = greedy
--top-k 40 top-k sampling, 0 disables
--top-p 0.9 nucleus sampling
--seed none reproducible output
--device auto auto / cpu / cuda
--checkpoint auto-found path to qelvra_best.pt or qelvra_final.pt

The script looks for the checkpoint next to itself, in the working directory, in QELVRA_MODEL_DIR, and in the usual download folders β€” the resolved path is printed on every run.

Example output (seed 42):

PROMPT       : The future of AI
CONTINUATION :
enstral twithe keeping the model computationally efficient.

Qelvra focuses on e

Model files

File Size Purpose
qelvra_best.pt 22.6 MB checkpoint with the best validation loss
qelvra_final.pt 22.6 MB final checkpoint (step 5,000)
config.json 793 B model configuration and training summary
metadata.json 436 B short machine-readable model card
tokenizer.json 310 B tokenizer specification
README.md β€” this model card
SHA256SUMS.txt β€” SHA-256 checksums for every file above

Each .pt file contains model_state_dict, optimizer_state_dict, model_config, step, and the recorded losses β€” so the architecture always comes from the checkpoint itself, never from guesswork.

Verify what you downloaded:

Get-FileHash -Algorithm SHA256 .\qelvra_best.pt   # 6f1f53aa7da5ea44a8a02113fa2c142fc552ab6cb65a11641b2dc43dea97cb7e

Honest scope

Claim What it actually means
Val loss 0.0341 Very low for a language model β€” the model largely memorised a small 3.4M-byte corpus. Low loss here is not a claim of broad generalization.
1.85M parameters Toy scale. It produces locally coherent English about its own training text and can drift into malformed bytes on prompts it never saw.
Byte-level tokens 1 token = 1 byte, so the 128-token context covers only 128 bytes of text.
Raw .pt format Not a Hugging Face transformers model β€” there is no config.json in transformers format and no AutoModelForCausalLM support. Load it with the reference scripts.
Feed-forward activation Not stored in the checkpoint (activations have no weights). Measured empirically across gelu/relu/silu β€” gelu reproduces the recorded validation loss most closely, and --activation lets you compare.
Status experimental (see metadata.json).

License

MIT β€” free to use, modify and redistribute with the copyright notice retained.


Links

Developed by JOJIN JOHN

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support