Talos 10M
Talos 10M is a custom causal language model developed by Promethean Studios.
Model
- Parameters: 9,952,320
- Hidden size: 448
- Layers: 3
- Attention heads: 28
- KV heads: 14
- Head dimension: 16
- FFN intermediate size: 1,792
- Context length: 512 tokens
- Model vocabulary: 1,024
- Training tokenizer vocabulary: 512
- Architecture: TalosGPT
- Attention: Full attention
- RoPE theta: 10,000
Training
- Training steps: 100,000
- Validation loss: 1.1523912596702575
- Batch size: 8
- Sequence length: 512
- Maximum learning rate: 3e-4
- Minimum learning rate: 3e-5
- Warmup steps: 1,000
- Weight decay: 0.1
- Seed: 1337
Files
talos_10m_step_100000.pt
- PyTorch training checkpoint containing the trained Talos 10M weights.
tokenizer.json
- Talos-native ByteLevelBPE tokenizer used to produce the training tokens.
config.json
- Model and training configuration.
Notes
Talos is a custom architecture and is not a Transformers-native model.
The model was trained with a 512-vocabulary tokenizer while the model architecture contains 1,024 vocabulary slots. Training data used token IDs through 509.
- Downloads last month
- 20
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support