alephn-intro-post

Alephn 1 (111M)

Alephn 1 is a 111M-parameter language model, trained from scratch by following the Cerebras-GPT 111M recipe on a single NVIDIA L40S GPU. It is a base model: it continues text and has not been instruction-tuned or aligned, so it will not follow instructions or hold a real conversation.

This project is an independent replication and is not affiliated with Cerebras. Its sole purpose is to test the data quality of Elebush's first corpus, KnowSpread.

Model details

Architecture GPT-2-style decoder, dense attention in every block, sequential (non-parallel) blocks
Parameters ~111M (embeddings tied with the output head)
Layers / hidden / heads 10 / 768 / 12 (head dim 64)
FFN size 3072, exact GELU
Positions Learned absolute, context length 2048
Vocabulary GPT-2 BPE, 50,257 tokens
Dropout None
Weights format safetensors, fp32

Training

Data KnowSpread
Tokens ~2.2B (about 20 tokens per parameter, compute-optimal)
Batch size ~246K tokens (120 sequences x 2048)
Optimizer AdamW, betas (0.9, 0.95), eps 1e-8
Weight decay 0.1 (2D weight matrices only; biases and LayerNorm excluded)
Peak LR 6e-4, linear warmup over 375M tokens, then linear decay to 10% of peak
Gradient clipping 1.0 (global norm)
Precision bfloat16 autocast
Init Truncated normal (std 0.02); residual output projections scaled by 1/sqrt(2 * n_layer)
Hardware 1x NVIDIA L40S
Framework PyTorch, torch.compile

Evaluation

Zero-shot accuracy using EleutherAI's lm-evaluation-harness.

image

Alephn 1 scores higher than Cerebras-GPT 111M on 5 of the 7 tasks reported in the Cerebras-GPT paper.

Limitations and intended use

  • This is a small base model trained on about 2.2B tokens. Expect fluent-looking but frequently incoherent or factually wrong text.
  • Scores on commonsense and reasoning benchmarks are near chance level, as is typical at this scale.
  • It has had no safety training or alignment, and it can produce biased, offensive, or inappropriate text reflecting its training data.
  • Intended for research, education, and experimentation (for example as a starting point for fine-tuning). It is not suitable for production use or for any decision-making.

Our Article on Alephn 1: Introducing Alephn 1

Downloads last month
501
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for elebush/alephn-1

Quantizations
1 model

Paper for elebush/alephn-1