YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Tiny GPT-2 2-Layer
A tiny GPT-2 model designed for learning and debugging Transformer inference and vLLM.
Model Architecture
- Architecture: GPT-2
- Number of layers: 2
- Hidden size: 128
- Attention heads: 4
- Head dimension: 32
- Vocabulary size: 50,257
- Maximum sequence length: 128
- Parameters: ~6.6M
- Data type: float32
Purpose
This model is randomly initialized and is not intended to produce meaningful text.
The purpose of this model is to make it easier to study the internals of Transformer inference and vLLM without having to step through a large model with dozens of Transformer layers.
It can be used to study:
- Transformer forward pass
- Self-Attention
- Q / K / V projections
- Causal attention
- KV Cache
- Prefill and Decode
- vLLM Scheduler
- vLLM ModelRunner
- vLLM Attention backend
- KV Cache block management
Model Structure
Input IDs
β
βΌ
Token Embedding
β
βΌ
βββββββββββββββββββ
β Transformer β
β Layer 0 β
β β
β Self-Attention β
β MLP β
ββββββββββ¬βββββββββ
β
βΌ
βββββββββββββββββββ
β Transformer β
β Layer 1 β
β β
β Self-Attention β
β MLP β
ββββββββββ¬βββββββββ
β
βΌ
LM Head
β
βΌ
Logits
Using with Transformers
from transformers import AutoTokenizer, AutoModelForCausalLM
model_name = "YOUR_USERNAME/tiny-gpt2-2layer"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
inputs = tokenizer("Hello", return_tensors="pt")
outputs = model(**inputs)
print(outputs.logits.shape)
Using with vLLM
vllm serve YOUR_USERNAME/tiny-gpt2-2layer \
--served-model-name tiny-gpt2 \
--gpu-memory-utilization 0.5 \
--max-model-len 128
Important
This model was created for educational and debugging purposes.
The weights are randomly initialized, so generated text is not expected to be meaningful.
- Downloads last month
- 60
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support