Talos 10M

Talos 10M is a custom causal language model developed by Promethean Studios.

Model

  • Parameters: 9,952,320
  • Hidden size: 448
  • Layers: 3
  • Attention heads: 28
  • KV heads: 14
  • Head dimension: 16
  • FFN intermediate size: 1,792
  • Context length: 512 tokens
  • Model vocabulary: 1,024
  • Training tokenizer vocabulary: 512
  • Architecture: TalosGPT
  • Attention: Full attention
  • RoPE theta: 10,000

Training

  • Training steps: 100,000
  • Validation loss: 1.1523912596702575
  • Batch size: 8
  • Sequence length: 512
  • Maximum learning rate: 3e-4
  • Minimum learning rate: 3e-5
  • Warmup steps: 1,000
  • Weight decay: 0.1
  • Seed: 1337

Files

talos_10m_step_100000.pt

  • PyTorch training checkpoint containing the trained Talos 10M weights.

tokenizer.json

  • Talos-native ByteLevelBPE tokenizer used to produce the training tokens.

config.json

  • Model and training configuration.

Notes

Talos is a custom architecture and is not a Transformers-native model.

The model was trained with a 512-vocabulary tokenizer while the model architecture contains 1,024 vocabulary slots. Training data used token IDs through 509.

Downloads last month
20
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support