Tiny CED
Tiny CED is a small custom PyTorch encoder-decoder language model trained from scratch on the English TinyStories dataset. This package contains inference-only weights in Safetensors format.
This is not a Transformers PreTrainedModel. Load it with the included
model.py and safetensors.torch.load_model as shown below.
Model details
| Item | Value |
|---|---|
| Independent parameters | 19,667,712 |
| Vocabulary size | 8,192 |
| Hidden size | 384 |
| Encoder layers | 4 |
| Decoder layers | 4 |
| Attention heads | 6 |
| Feed-forward size | 1,024 |
| Local attention window | 64 |
| Trained context length | 512 |
| Weight dtype | FP32 |
The token embedding and output head are tied. model.safetensors was exported
with safetensors.torch.save_model so the shared tensor is stored only once.
Usage
Install the runtime dependencies:
pip install -r requirements.txt
Run generation on CPU:
python generate.py \
--device cpu \
--text "Once upon a time, a little rabbit" \
--max-new-tokens 100
Use --device cuda when CUDA is available. Set --temperature 0 for greedy
decoding.
Training and evaluation
The exported weights come from the existing best.pt checkpoint; no retraining
was performed during export.
- Optimizer: AdamW
- Peak learning rate: 3e-4
- Best-checkpoint training tokens: 40,004,782
- Best-checkpoint step: 5,856
- Validation loss: 1.896527
- Validation perplexity: 6.662715
Validation used the TinyStories validation split and the same 8,192-token BPE tokenizer included in this repository.
Files
model.safetensors: inference weightsmodel.py: exact custom PyTorch architectureconfig.json: architecture and token IDstokenizer.json: BPE tokenizergenerate.py: minimal generation example
Limitations
This is a small research model trained only on synthetic English children's stories. It may generate incorrect, repetitive, biased, unsafe, or otherwise inappropriate text. It is not suitable for factual, safety-critical, or production use without further evaluation.
The source project did not specify a license for its code or weights. Choose an appropriate license and confirm the training-data terms before publishing this package publicly. TinyStories is listed as CDLA-Sharing-1.0 on its dataset page.
- Downloads last month
- 376