Parallax 8B lambda: mid-training checkpoints
These are research checkpoints from lambda, a training run of Parallax that is still in progress. Parallax is a mixture-of-experts language model trained on consumer GPUs spread over several countries. The hosts have no direct connections to each other and synchronize over the public internet.
- Base model. It is pretrained on a mixture of web, document, math, code, knowledge-focused and book text. It is not instruction-tuned, chat-tuned or safety-tuned, and it will continue text rather than follow instructions.
- Mid-training. Each export is a snapshot of a run that has not finished. Later exports are expected to differ, and quality between exports is not monotonic.
- Research artifact. It is published so the training can be followed and inspected. It is not intended for production use.
Live training dashboard: parallax.chutes.ai.
Model
| Parameters | ~7.8B total, ~1.25B active per token |
| Layers | 64: 18 gated delta-rule (GDN2) recurrent, 8 sliding-window attention (window 2048), 6 sparse attention, 32 mixture-of-experts |
| Experts | 4096 routed experts (128 per MoE layer), 12 routed + 1 shared expert per token |
| Expert weights | ternary (-1, 0, +1) with per-row scales; at most two nonzero pairs in every group of eight |
| Width | d_model 1152; ReLU² expert activation |
| Tokenizer | Llama 3 tokenizer (vocabulary padded to 128,384); tied input/output embeddings |
| Training context | 4096 tokens |
| Output logit scale | learnable, capped at 3.0 |
Training
- Data: the same data mix as the earlier Parallax runs epsilon and kappa, not FineWeb-Edu. The main phase draws from seven sources totalling about 1.09T tokens: two general pretraining mixtures of web, document, math and code text (about 78% of the tokens), an additional web set, two math sets, a knowledge-focused set and a books set. At about 899.6B tokens the data switches to two further sources for the rest of the run; the learning rate does not change at that point. The run is planned as one pass of about 973B tokens.
- Optimization: the dense trunk is synchronized with a decoupled DiLoCo scheme over the internet. Routed experts are trained through low-rank adapters on the GPUs that use them, and the adapter updates are folded into full-precision masters that publish new ternary versions. The tech report describes the details.
- Learning rate: linear warm-up from 0 to 6e-4 over the first 19.66B tokens, then 6e-4 until 400B tokens, 3e-4 until 700B and 1.5e-4 through the end of training. There is no final decay.
- Batch: 10 gradient-accumulation micro-batches per step (about 8.5M tokens per fleet step) for about the first 100B tokens, then 20 (about 17M tokens per step) after a switch made while the run continued. The fleet is 26 machines with 208 RTX 5090 GPUs.
- Output logits: the learnable output logit scale is capped at 3.0 in training and inference (see Files and format).
Exports
One folder per export under exports/, named by training tokens (rounded down to whole billions) and
fleet step. A new export is added about once an hour while the run continues. Published exports are
never modified or removed.
| Export | Tokens | Step | Size | val_mix (nats/token) | 0-shot macro | 5-shot macro | Exported (UTC) |
|---|---|---|---|---|---|---|---|
194B-tokens_step17828 (latest) |
194.3B | 17828 | 2.45 GB | - | - | - | 2026-10-05 11:11 |
191B-tokens_step17631 |
191.5B | 17631 | 2.45 GB | - | - | - | 2026-10-05 10:44 |
187B-tokens_step17334 |
187.0B | 17334 | 2.45 GB | - | - | - | 2026-10-05 10:12 |
184B-tokens_step17145 |
184.2B | 17145 | 2.45 GB | - | - | - | 2026-10-05 09:45 |
179B-tokens_step16851 |
179.8B | 16851 | 2.45 GB | - | - | - | 2026-10-05 09:11 |
176B-tokens_step16645 |
176.9B | 16645 | 2.45 GB | - | - | - | 2026-10-05 08:45 |
172B-tokens_step16359 |
172.6B | 16359 | 2.45 GB | - | - | - | 2026-10-05 08:12 |
169B-tokens_step16164 |
169.7B | 16164 | 2.45 GB | - | - | - | 2026-10-05 07:45 |
165B-tokens_step15861 |
165.2B | 15861 | 2.45 GB | - | - | - | 2026-10-05 07:10 |
162B-tokens_step15667 |
162.4B | 15667 | 2.46 GB | - | - | - | 2026-10-05 06:45 |
157B-tokens_step15384 |
157.9B | 15384 | 2.46 GB | - | - | - | 2026-10-05 06:13 |
155B-tokens_step15179 |
155.1B | 15179 | 2.46 GB | - | - | - | 2026-10-05 05:46 |
150B-tokens_step14880 |
150.6B | 14880 | 2.46 GB | - | - | - | 2026-10-05 05:12 |
147B-tokens_step14693 |
147.8B | 14693 | 2.46 GB | - | - | - | 2026-10-05 04:45 |
143B-tokens_step14118 |
143.4B | 14118 | 2.47 GB | - | - | - | 2026-10-05 04:13 |
140B-tokens_step13940 |
140.5B | 13940 | 2.47 GB | - | - | - | 2026-10-05 03:44 |
135B-tokens_step13714 |
135.8B | 13714 | 2.47 GB | - | - | - | 2026-10-05 03:14 |
132B-tokens_step13524 |
132.5B | 13524 | 2.47 GB | - | - | - | 2026-10-05 02:48 |
127B-tokens_step13242 |
127.7B | 13242 | 2.48 GB | - | - | - | 2026-10-05 02:12 |
123B-tokens_step13016 |
123.6B | 13016 | 2.48 GB | - | - | - | 2026-10-05 01:43 |
119B-tokens_step12790 |
119.5B | 12790 | 2.48 GB | - | - | - | 2026-10-05 01:13 |
115B-tokens_step12532 |
115.4B | 12532 | 2.48 GB | - | - | - | 2026-10-05 00:38 |
111B-tokens_step12302 |
111.4B | 12302 | 2.48 GB | - | - | - | 2026-10-05 00:10 |
107B-tokens_step12081 |
107.2B | 12081 | 2.48 GB | - | - | - | 2026-10-04 23:41 |
103B-tokens_step11842 |
103.2B | 11842 | 2.48 GB | - | - | - | 2026-10-04 23:14 |
99B-tokens_step11533 |
99.2B | 11533 | 2.48 GB | - | - | - | 2026-10-04 22:39 |
95B-tokens_step11160 |
96.0B | 11160 | 2.47 GB | - | - | - | 2026-10-04 22:12 |
92B-tokens_step10766 |
92.6B | 10766 | 2.47 GB | - | - | - | 2026-10-04 21:43 |
89B-tokens_step10388 |
89.3B | 10388 | 2.47 GB | - | - | - | 2026-10-04 21:12 |
85B-tokens_step9986 |
86.0B | 9986 | 2.47 GB | - | - | - | 2026-10-04 20:41 |
82B-tokens_step9591 |
82.7B | 9591 | 2.47 GB | - | - | - | 2026-10-04 20:09 |
79B-tokens_step9195 |
79.3B | 9195 | 2.47 GB | - | - | - | 2026-10-04 19:39 |
76B-tokens_step8835 |
76.1B | 8835 | 2.48 GB | - | - | - | 2026-10-04 19:13 |
72B-tokens_step8443 |
72.8B | 8443 | 2.48 GB | - | - | - | 2026-10-04 18:40 |
69B-tokens_step8039 |
69.5B | 8039 | 2.48 GB | - | - | - | 2026-10-04 18:11 |
66B-tokens_step7652 |
66.1B | 7652 | 2.48 GB | - | - | - | 2026-10-04 17:42 |
62B-tokens_step7275 |
62.9B | 7275 | 2.48 GB | - | - | - | 2026-10-04 17:10 |
59B-tokens_step6711 |
59.6B | 6711 | 2.48 GB | - | - | - | 2026-10-04 16:40 |
56B-tokens_step6444 |
56.3B | 6444 | 2.48 GB | - | - | - | 2026-10-04 16:09 |
53B-tokens_step6096 |
53.0B | 6096 | 2.48 GB | - | - | - | 2026-10-04 15:40 |
49B-tokens_step5723 |
49.8B | 5723 | 2.48 GB | - | - | - | 2026-10-04 15:12 |
46B-tokens_step5320 |
46.4B | 5320 | 2.48 GB | - | - | - | 2026-10-04 14:40 |
43B-tokens_step4946 |
43.1B | 4946 | 2.48 GB | - | - | - | 2026-10-04 14:10 |
39B-tokens_step4549 |
39.8B | 4549 | 2.48 GB | - | - | - | 2026-10-04 13:38 |
36B-tokens_step4166 |
36.5B | 4166 | 2.48 GB | - | - | - | 2026-10-04 13:15 |
33B-tokens_step3768 |
33.2B | 3768 | 2.48 GB | - | - | - | 2026-10-04 12:40 |
29B-tokens_step3391 |
29.8B | 3391 | 2.49 GB | - | - | - | 2026-10-04 12:16 |
26B-tokens_step3002 |
26.5B | 3002 | 2.49 GB | - | - | - | 2026-10-04 11:43 |
23B-tokens_step2604 |
23.1B | 2604 | 2.50 GB | - | - | - | 2026-10-04 11:15 |
19B-tokens_step2198 |
19.9B | 2198 | 2.50 GB | - | - | - | 2026-10-04 10:40 |
16B-tokens_step1817 |
16.7B | 1817 | 2.50 GB | - | - | - | 2026-10-04 10:11 |
13B-tokens_step1419 |
13.3B | 1419 | 2.50 GB | - | - | - | 2026-10-04 09:42 |
6B-tokens_step635 |
6.6B | 635 | 2.50 GB | - | - | - | 2026-10-04 08:41 |
- val_mix: mean cross-entropy (nats per token, lower is better) on a fixed held-out validation set stratified over the nine sources of the data mix (2048 windows of 4096 tokens), the same set used for kappa. It is not comparable with the FineWeb-Edu validation number published for theta.
- 0-shot macro: mean over 11 tasks (ARC-Challenge, ARC-Easy, BoolQ, COPA, HellaSwag, LAMBADA, OpenBookQA, PIQA, SciQ, SIQA, WinoGrande). 5-shot macro: the same tasks without LAMBADA (10 tasks). Per task, acc_norm is used for ARC, HellaSwag, OpenBookQA, PIQA and SciQ, and acc for the rest. Values are percentages.
- The scores come from the project's own scorer, which runs the native ternary experts. They track progress within this run. Compare them with numbers from other evaluation harnesses with care.
- Benchmark prompts are scored with every run of two or more newlines collapsed to a single newline (including the few-shot separator).
- A
-means the export has not been scored yet. The table fills in as scores arrive.
Latest export
exports/194B-tokens_step17828: 194.3B training tokens, step 17828.
hf download chutesai/parallax-8b-lambda --include "exports/194B-tokens_step17828/*" --local-dir parallax-8b-lambda
cd parallax-8b-lambda/exports/194B-tokens_step17828
tar -xf packed_experts.tar
sha256sum -c --quiet SHA256SUMS # every file of the original export, byte for byte
Files and format
The files are in Parallax's native compact export format (inference only, no optimizer state). They are byte-identical to the export the training system produced:
| File | Contents |
|---|---|
manifest.json |
export manifest: tensor inventory, per-file sha256 digests, token clock |
model_config.json |
model configuration |
coverage.json, layouts.json |
tensor coverage and expert frame layouts |
indexer.bundle |
sparse-attention indexer weights |
relay_pack/ |
trunk (non-expert) weights in bf16, with their own manifest |
packed_experts.tar |
the 4096 routed experts (packed_experts/*.t24p, packed ternary codes and scales) |
SHA256SUMS |
sha256 of every file of the original export |
export_info.json |
step, tokens, time, sizes and digests of this export |
The only change from the original export is packaging. The 4096 expert files are stored in one
uncompressed tar to keep the repository's file count manageable. Extract it and check SHA256SUMS
as shown above. Every upload was checked against the training system's own digests before and
after it was published.
Logit-scale bound. The model's output logit scale is bounded: the forward pass uses
exp(min(logit_scale_log, log 3.0)). The stored trunk tensor is the raw training parameter (it can sit
slightly above the bound, e.g. from bf16 rounding), and each export records the bound in manifest.json
(logit_scale: max, raw, effective) and in coverage.json (inference_policy.logit_scale_max).
A loader must apply the recorded bound; the files themselves are left byte-identical.
Running it
This is a base model only. It is not chat- or instruction-tuned, so it does plain text completion: give it the start of a text and it continues it. It will not follow instructions or hold a conversation.
Standard transformers cannot load this format. Use our llama.cpp fork, https://github.com/chutesai/llama.cpp, which adds a
dedicated runtime, llama-parallax, for CPU (x86-64, ARM64) and Apple GPUs (Metal):
git clone https://github.com/chutesai/llama.cpp && cd llama.cpp
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON # -DGGML_METAL=OFF for CPU only
cmake --build build --target llama-parallax -j 8
python tools/parallax/run.py --binary build/bin/llama-parallax --model parallax-t9.gguf \
--tokenizer tokenizer.json --backend metal --experts lut9 \
--prompt 'The capital of France is' --predict 64
The runtime reads GGUF files converted from these exports. Ready-made GGUF files are not published
yet; tools/parallax/README.md in the fork covers conversion, options and the tokenizer.
Tech report
The Parallax tech report: https://parallax.chutes.ai/tech-report.pdf. It is AI-generated from the project's measurements and logs, and it is a living document that changes as the run progresses. It covers the predecessor run (kappa); lambda differs as described above.
Limitations
This is an early base model. It can produce incorrect, biased or nonsensical output. It has no alignment or safety tuning, and it is small and far from converged.
License
MIT.
Table updated 2026-10-05 11:22 UTC.