Styx 100M

Styx 100M is an experimental language model from Promethean Studios and part of the Talos model family.

This release contains the model trained through 100,000 training steps and has 96,482,304 parameters.

Styx is not being presented as a finished or production-ready language model.

It is a research checkpoint and a step in the development of the Talos model family. The goal is to experiment with larger model sizes, training infrastructure, data efficiency, and practical local inference.

Model Details

Property Value Parameters 96,482,304 Vocabulary 1,024 Hidden size 1,024 Layers 6 Attention heads 64 KV heads 32 Feed-forward network SwiGLU Maximum sequence length 512 Training steps 100,000 Architecture Causal Transformer

Styx uses grouped-query attention with 32 KV heads and a SwiGLU feed-forward network.

Styx is a relatively small model by modern language-model standards. Its capabilities should be evaluated accordingly. More parameters do not automatically produce a better model, and this release is primarily useful as a development and research milestone.

What to Expect

Styx has known limitations.

You may encounter:

  • Incorrect or fabricated information
  • Repetitive generations
  • Incomplete responses
  • Weak reasoning
  • Inconsistent instruction following
  • Poor generalization
  • Unsafe or undesirable generations
  • Outputs that simply make very little sense

These limitations are intentionally documented rather than hidden.

Styx is not a production assistant.

It has not been trained or evaluated to provide the reliability, accuracy, instruction following, or safety expected from larger production language models.

The purpose of this release is to make the model and its progress available for experimentation.

Why Release It?

Styx is part of an ongoing development process.

Releasing an imperfect model provides a reference point for evaluating what changes between generations.

This includes experimentation with:

  • Model scale
  • Training efficiency
  • Data efficiency
  • Tokenization
  • Attention architectures
  • Training infrastructure
  • Local inference
  • Benchmarking
  • Checkpoint formats
  • Larger future models

If you’re interested in small language models, Styx can be used as a lightweight model to download, inspect, benchmark, and experiment with locally.

Files

  • model.safetensors — model weights in SafeTensors format
  • config.json — model configuration
  • tokenizer/tokenizer.json — Styx tokenizer
  • checkpoints/step-100000.pt — original training checkpoint

model.safetensors contains the model weights for inference.

The original .pt checkpoint contains additional training state and metadata from the training run.

Training

The released model corresponds to:

100,000 training steps

The checkpoint was produced using the Talos training infrastructure developed by Promethean Studios.

Training metadata, including recorded training and validation information, is preserved in the original training checkpoint.

The Talos Model Family

Styx is one stage in a larger model-development effort.

The project began with significantly smaller models and is progressively moving toward larger and more capable architectures.

The purpose of each release is not simply to produce another model checkpoint. It is to learn what works, what doesn’t, and what needs to change before scaling further.

Styx has known problems. That’s expected.

This release is a snapshot of the project at this stage of development, not the final destination.

More Versions Are Coming

Styx is not the end of the project.

Additional models and revisions are planned as development continues. Future releases will incorporate lessons learned from the current generation, including its failures.

Future development may include:

  • Improved training procedures
  • Better data preparation
  • Architectural changes
  • Improved tokenizer and data pipelines
  • More extensive evaluation
  • Larger parameter counts
  • Better inference support

Future releases may differ substantially from Styx.

Compatibility, architecture, training methods, and model behavior should not be assumed to remain identical between generations.

Intended Use

Styx 100M is intended primarily for:

  • Small language-model research
  • Local inference experimentation
  • Architecture research
  • Benchmarking
  • Training research
  • Data-efficiency experiments
  • Educational experimentation
  • Development within the Talos ecosystem

It is particularly suited to people interested in seeing what can be done with relatively small language models.

Limitations

Styx has not undergone the level of evaluation associated with large production models.

Do not rely on its outputs for:

  • Medical decisions
  • Legal decisions
  • Financial decisions
  • Safety-critical systems
  • High-stakes automation
  • Unsupervised production applications

Treat model output as untrusted.

Styx can generate incorrect, misleading, or undesirable content. Verify anything important independently.

License

Apache License 2.0.

Organization

Promethean Studios

Styx 100M is released as part of the Talos model family.

⸻

Status

Experimental • 100M class • 100K training steps • Actively evolving

Styx is a checkpoint, not the finish line.

Downloads last month
174
Safetensors
Model size
96.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including PrometheanStudio/styx-100m