Zenyx V3 Base (1.5B Mixture-of-Experts)

Zenyx V3 is an efficient 1.5B-parameter Mixture-of-Experts (MoE) foundation model built for low-latency inference and high throughput. It is written from scratch in JAX/Flax and trained on TPU v5e-8.

This is a BASE model — it is not instruction-tuned. It completes text; it does not follow instructions or hold a conversation. Prompt it with a prefix to continue ("The capital of France is"), not with a request ("Explain gravity"). Pretraining is still in progress; SFT/chat variants will follow.

Current checkpoint: step 86,400 · 59.1B tokens seen

Model Architecture

  • Sparse Mixture-of-Experts: 12 routed experts + 1 shared expert, exactly 2 active per token, with a Sinkhorn transport-based gate.
  • Attention: Multi-Query Attention (MQA) with 12 query heads and a single shared key/value head, plus low-rank query and output projections.
  • Hyper-Connections: Sinkhorn-normalised residual routing for gradient stability at scale.
  • Multi-Token Prediction (MTP): one auxiliary prediction head during training.
  • Context: pretrained at up to 4,096 tokens (progressive 2,048 → 4,096). YaRN and RoPE scaling factors are precomputed so context can be extended at inference time beyond the trained length.
Total parameters ~1.5B
Active parameters / token ~0.4B
Layers 16 (2 dense + 14 MoE)
Hidden size 1,536
Attention heads 12 (head dim 128)
Vocabulary 129,280
Precision bfloat16

Benchmarks — checkpoint step 86,400 (59.1B tokens)

All tasks are evaluated with the standard base-model protocol: the model scores the log-likelihood of every candidate continuation and the highest-scoring one is taken as the answer. Nothing is generated and no output parsing is involved, so the numbers do not depend on instruction-following ability. 0-shot, full evaluation sets, no subsampling.

acc_norm normalises each continuation's log-likelihood by its length in characters, which removes the bias toward short answers; it is the headline metric wherever the task has candidates of differing lengths.

Benchmark acc acc_norm Random Δ n Description
HellaSwag 30.18% ± 0.46 33.67% ± 0.47 25.0% +8.7 10,042 Commonsense sentence completion
ARC-Easy 50.63% ± 1.03 46.25% ± 1.02 25.0% +21.3 2,376 Grade-school science questions
ARC-Challenge 21.93% ± 1.21 26.45% ± 1.29 25.0% +1.5 1,172 Hard grade-school science questions
PIQA 60.72% ± 1.14 61.26% ± 1.14 50.0% +11.3 1,838 Physical commonsense reasoning
WinoGrande 49.72% ± 1.40 50.0% -0.3 1,267 Pronoun resolution / coreference
OpenBookQA 19.60% ± 1.78 31.00% ± 2.07 25.0% +6.0 500 Elementary science with open book
BoolQ 60.76% ± 0.85 62.17% ± 0.85 62.2% -1.4 3,270 Yes/no reading comprehension
SciQ 78.20% ± 1.31 69.80% ± 1.45 25.0% +53.2 1,000 Science exam questions with support
LAMBADA (OpenAI) 25.64% ± 0.61 0.0% +25.6 5,153 Long-range last-word prediction
MMLU (5-shot) 25.78% ± 0.37 25.0% +0.8 14,042 57 subjects of academic knowledge
RACE 30.34% ± 0.65 32.81% ± 0.67 25.0% +7.8 4,934 Exam reading comprehension
CommonsenseQA 28.50% ± 1.29 31.04% ± 1.32 20.0% +11.0 1,221 5-choice commonsense (random = 20%)
COPA 62.00% ± 4.85 61.00% ± 4.88 50.0% +12.0 100 Causal reasoning
LogiQA 21.51% ± 1.61 25.96% ± 1.72 25.0% +1.0 651 Logical deduction
WSC273 53.48% ± 3.02 50.0% +3.5 273 Winograd coreference
TruthfulQA MC1 19.58% ± 1.39 22.8% -3.2 817 Resistance to common misconceptions
Arithmetic 0.09% ± 0.03 0.0% +0.1 14,000 2-5 digit add/sub/mul, generated in-harness

Bold marks the metric that is conventional for that task — acc_norm for HellaSwag, ARC, PIQA and OpenBookQA; acc for WinoGrande, BoolQ, SciQ and LAMBADA. The convention is applied per task, not chosen per result: it lowers the reported figure for ARC-Easy (44.53 rather than 49.54) and PIQA (59.79 rather than 60.83). Δ compares the bolded metric to the random baseline.

Both metrics

Benchmark acc acc_norm n
HellaSwag 30.18% ± 0.46 33.67% ± 0.47 10,042
ARC-Easy 50.63% ± 1.03 46.25% ± 1.02 2,376
ARC-Challenge 21.93% ± 1.21 26.45% ± 1.29 1,172
PIQA 60.72% ± 1.14 61.26% ± 1.14 1,838
WinoGrande 49.72% ± 1.40 1,267
OpenBookQA 19.60% ± 1.78 31.00% ± 2.07 500
BoolQ 60.76% ± 0.85 62.17% ± 0.85 3,270
SciQ 78.20% ± 1.31 69.80% ± 1.45 1,000
LAMBADA (OpenAI) 25.64% ± 0.61 5,153
MMLU (5-shot) 25.78% ± 0.37 14,042
RACE 30.34% ± 0.65 32.81% ± 0.67 4,934
CommonsenseQA 28.50% ± 1.29 31.04% ± 1.32 1,221
COPA 62.00% ± 4.85 61.00% ± 4.88 100
LogiQA 21.51% ± 1.61 25.96% ± 1.72 651
WSC273 53.48% ± 3.02 273
TruthfulQA MC1 19.58% ± 1.39 817
Arithmetic 0.09% ± 0.03 14,000

Language modelling

Corpus Value Metric
WikiText-2 (raw) 24.64 token-level perplexity
WikiText-2 (raw) 45.62 word-level perplexity
WikiText-2 (raw) 1.0279 bits per byte
LAMBADA 30.35 perplexity of the target word

WikiText-2 is scored with a rolling 1024-token window at stride 512, so every counted token is predicted with at least 512 tokens of left context and each token is counted exactly once. (Scoring disjoint windows instead inflates these figures by ~15% because the leading tokens of each window are predicted from nothing.)

Trajectory across all benchmarked checkpoints

Tokens seen: 34.8B | 36.5B | 39.4B | 42.4B | 45.3B | 49.1B | 54.9B | 59.1B. The pretraining data mixture was changed partway through this sequence (code weight raised, several synthetic sources cut), so these columns do not represent a tokens-only progression.

Accuracy benchmarks (higher is better)

Benchmark 63,200 64,800 67,600 70,400 73,200 76,800 82,400 86,400 net
HellaSwag 32.22% 32.66% 32.84% 32.51% 32.72% 33.08% 33.41% 33.67% +1.44 up
ARC-Easy 44.53% 45.08% 44.91% 45.58% 46.09% 44.82% 46.34% 46.25% +1.73 up
ARC-Challenge 25.77% 25.51% 26.02% 25.00% 25.94% 25.09% 25.26% 26.45% +0.68 up
PIQA 59.79% 59.85% 61.53% 59.96% 60.23% 59.09% 61.26% 61.26% +1.47 up
WinoGrande 49.49% 50.51% 51.07% 49.57% 49.25% 50.12% 50.59% 49.72% +0.24 up
OpenBookQA 30.00% 28.00% 29.80% 29.00% 28.40% 29.20% 30.40% 31.00% +1.00 up
BoolQ 60.55% 60.83% 61.80% 62.14% 60.83% 60.92% 61.47% 60.76% +0.21 up
SciQ 75.10% 76.10% 75.70% 76.30% 76.00% 76.00% 77.40% 78.20% +3.10 up
LAMBADA (OpenAI) 25.50% 25.42% 24.94% 26.96% 26.57% 24.68% 25.89% 25.64% +0.14 up
MMLU (5-shot) 26.07% 26.71% 25.77% 26.01% 25.26% 26.48% 26.47% 25.78% -0.29 down
RACE 32.79% 32.77% 32.96% 32.85% 33.58% 33.46% 33.46% 32.81% +0.02 up
CommonsenseQA 29.57% 29.57% 29.40% 30.55% 30.47% 30.88% 30.14% 31.04% +1.47 up
COPA 58.00% 58.00% 61.00% 59.00% 60.00% 57.00% 59.00% 62.00% +4.00 up
LogiQA 25.81% 25.19% 25.65% 25.81% 25.65% 23.81% 25.65% 25.96% +0.15 up
WSC273 54.95% 52.38% 52.75% 51.28% 53.11% 56.41% 54.21% 53.48% -1.47 down
TruthfulQA MC1 19.22% 20.20% 19.58% 19.83% 20.44% 19.58% 19.34% 19.58% +0.37 up
Arithmetic 0.14% 0.55% 0.58% 1.04% 1.91% 0.09% 0.15% 0.09% -0.04 down

Language modelling (LOWER is better)

Metric 63,200 64,800 67,600 70,400 73,200 76,800 82,400 86,400 net
WikiText-2 perplexity 26.47 25.79 25.98 25.79 25.18 25.28 24.55 24.64 -1.832 BETTER
WikiText-2 bits/byte 1.051 1.043 1.045 1.043 1.035 1.036 1.027 1.028 -0.02301 BETTER
LAMBADA perplexity 35.79 36.19 36.84 34.15 34.27 36.46 32.23 30.35 -5.443 BETTER

The Pile, by content type (bits/byte, LOWER is better)

Category 63,200 64,800 67,600 70,400 73,200 76,800 82,400 86,400 net
Code / technical 0.9558 0.9518 0.9438 0.9341 0.9336 0.9228 -- 0.9076 -0.0482 BETTER
Science / legal 0.8987 0.8950 0.8930 0.8892 0.8880 0.8834 -- 0.8774 -0.0213 BETTER
Web / reference 1.1586 1.1561 1.1557 1.1543 1.1527 1.1475 -- 1.1423 -0.0164 BETTER
Prose / spoken 1.5606 1.5466 1.5650 1.5383 1.5411 1.5261 -- 1.5134 -0.0472 BETTER
Every non-prose category has improved monotonically across every checkpoint where the Pile was measured, zero reversals (one checkpoint, 82,400, has no Pile measurement and shows as a gap, not a regression) -- and ALL FOUR categories, prose included, are at their best value at the latest checkpoint. Prose/spoken is the only one that ever moved backwards: it dipped for exactly one interval after the data mixture changed, then recovered and has since reached a new best. That was a one-off transition cost, not a permanent trade.

Does few-shot prompting help? (MMLU by shot count)

Shots step 82400 step 86,400 Shot source
5 26.47% 25.78% ± 0.37 dev split, the published convention

No. More demonstrations do not help and the 5-shot result is the best of the three at both checkpoints, with 10-shot dropping to the 25% chance line (-1.47 points vs 5-shot at step 86,400, ~2.8 sigma). The same ordering appears independently at both checkpoints, so it is not a fluke of one run.

This is what a model without in-context learning looks like: using examples to infer a task is an ability that emerges later in training, and before it does, extra shots are just tokens competing for attention with the actual question. Practical consequence: prompt this model with a short direct prefix, not a long few-shot preamble.

Arithmetic

Exact-match on the answer, greedy decoding, GPT-3 prompt format (Question: What is 47 plus 21? / Answer: 68).

Operation step 82400 step 86,400 n
2-digit addition 0.90% 0.00% 2,000
2-digit subtraction 0.10% 0.50% 2,000
3-digit addition 0.00% 0.00% 2,000
3-digit subtraction 0.00% 0.05% 2,000
4-digit addition 0.00% 0.00% 2,000
5-digit addition 0.00% 0.00% 2,000
2-digit multiplication 0.05% 0.10% 2,000
overall 0.150% 0.093% 14,000

The model essentially cannot do arithmetic — but two-digit subtraction moved from 0.65% to 3.30% between these two checkpoints (5.1x, ~6 sigma on identical problems), which is the signature of a capability just beginning to emerge. Note that 15.5% of the pretraining mix is mathematics, yet that has bought fluency in mathematical language rather than the ability to compute.

Items are generated in-harness from a fixed seed using this prompt format, because EleutherAI/arithmetic is a loading script with no parquet branch and cannot be fetched under datasets>=3. Both checkpoints see byte-identical problems, so the comparison is exact — but these numbers are not interchangeable with published EleutherAI/arithmetic results.

Language modelling by genre (The Pile)

Bits-per-byte on each Pile domain, lower is better, scored with the same rolling 1024-token window as WikiText-2 so the numbers are directly comparable to it. This is the clearest picture of what the model is actually good at, because it measures raw prediction rather than multiple-choice ability.

Domain bits/byte perplexity tokens Δ vs prev
Github 0.594 3.72 479,656
PubMed Central 0.759 15.01 292,334
USPTO Backgrounds 0.794 16.31 296,162
NIH ExPorter 0.877 25.26 41,533
ArXiv 0.877 7.71 447,478
PubMed Abstracts 0.892 20.63 306,636
StackExchange 0.941 12.61 387,038
FreeLaw 0.982 20.50 339,260
Wikipedia (en) 1.008 21.70 341,511
Pile-CC 1.114 35.56 326,706
OpenWebText2 1.147 32.80 348,909
BookCorpus2 1.154 33.45 139,043
Enron Emails 1.241 20.99 16,119
Gutenberg (PG-19) 1.288 34.89 133,371
HackerNews 1.300 41.58 57,843
DM Mathematics 1.332 7.53 370,912
Books3 1.344 34.31 396,623
OpenSubtitles 1.353 31.09 234,514
PhilPapers 1.359 49.90 38,263
Ubuntu IRC 1.743 39.97 14,407
YoutubeSubtitles 1.852 119.67 51,428
EuroParl 2.016 130.82 19,523

The ordering here is a direct readout of the pretraining mix: code, papers and mathematics sit at the top because they are what the model has been fed most of.

Progress since the previous checkpoint

Same suite, same code, same full evaluation sets — only the checkpoint differs. Step 82,400 → 86,400 is +4.19B tokens.

Benchmark step 82,400 step 86,400 Δ ±2σ needs
HellaSwag (acc_norm) 33.41% 33.67% +0.26 ±0.67
ARC-Easy (acc_norm) 46.34% 46.25% -0.08 ±1.45
ARC-Challenge (acc_norm) 25.26% 26.45% +1.19 ±1.81
PIQA (acc_norm) 61.26% 61.26% +0.00 ±1.61
WinoGrande (acc) 50.59% 49.72% -0.87 ±1.99
OpenBookQA (acc_norm) 30.40% 31.00% +0.60 ±2.92
BoolQ (acc) 61.47% 60.76% -0.70 ±1.21
SciQ (acc) 77.40% 78.20% +0.80 ±1.86
LAMBADA (OpenAI) (acc) 25.89% 25.64% -0.25 ±0.86
MMLU (5-shot) (acc) 26.47% 25.78% -0.69 ±0.52
RACE (acc_norm) 33.46% 32.81% -0.65 ±0.95
CommonsenseQA (acc_norm) 30.14% 31.04% +0.90 ±1.86
COPA (acc) 59.00% 62.00% +3.00 ±6.91
LogiQA (acc_norm) 25.65% 25.96% +0.31 ±2.43
WSC273 (acc) 54.21% 53.48% -0.73 ±4.27
TruthfulQA MC1 (acc) 19.34% 19.58% +0.24 ±1.96
Arithmetic (acc) 0.15% 0.09% -0.06 ±0.04
WikiText-2 perplexity 24.55 24.64 +0.08721
WikiText-2 bits/byte 1.027 1.028 +0.001138
LAMBADA perplexity 32.23 30.35 -1.877

Δ is on the conventional metric for each task. Bold marks a change larger than two standard errors of the difference; anything unbolded is inside the noise floor and should not be read as movement. The quoted error treats the two runs as independent, which is conservative here — they score identical items, so the true paired error is smaller.

What actually changed. A quiet interval on the accuracy benchmarks -- 8 of 15 improved (sign test p = 0.50, a coin flip, no signal either way) -- which is the expected shape between adjacent checkpoints at this token scale. Nothing here contradicts the previous interval's real gains; it simply did not add more of the same size.

  • WikiText-2 bits/byte held flat (1.0267 -> 1.0279, +0.0011). LAMBADA perplexity kept improving (32.23 -> 30.35) even though its accuracy ticked down slightly (-0.25 sigma, noise).
  • Arithmetic stayed at its post-correction floor: 0.150% -> 0.093%, with strict and lenient identical at every step size again, confirming this is a genuine capability gap rather than a formatting artifact -- consistent with the correction below.
  • The Pile (1.5M-char cap, trajectory-comparable): 21 of 22 domains improved, 1 regressed (OpenSubtitles, +0.0032, small) -- token-weighted 1.0400 -> 1.0304 (-0.0096), the strongest single-interval improvement since step 76,800 and consistent with that checkpoint's pattern: the Pile remains the most reliable signal of real progress at this scale. A second Pile measurement, at a 400,000-char cap, backs the head-to-head model comparisons below (Github 0.5797, ArXiv 0.8738) -- do not compare those numbers to the ones above, the sampled text differs.

Second model comparison added below: XHToken/Spark-X2.5-1.7B-Base. Unlike gemma-4-E2B, where this model won on code and mathematics, Spark wins every single one of the 22 Pile domains measured and leads decisively on ARC-Challenge and LAMBADA. Spark is close in size (1.71B vs 1.58B) and vocab (131,072 vs 129,280), so this is a much more size-matched comparison than Gemma -- and the honest read is that Spark is a substantially stronger model at a similar parameter count. See the comparison table for the full breakdown.

Head-to-head vs google/gemma-4-E2B

Zenyx V3 at step 82,400 (1.58B total, MoE, ~55B tokens) against Google's 5.10B-parameter Gemma 4 E2B, which was trained on a corpus larger by roughly two orders of magnitude.

Both models were scored by one copy of the task code on the same source text, with bits/byte normalised by that text's UTF-8 byte length -- the only metric here that is legitimate across a 129,280-entry vocab and a 262,144-entry one. Perplexity is per-token and is deliberately not reported.

The Pile, all 22 domains (bits/byte, LOWER is better)

domain Zenyx V3 Gemma 4 E2B margin winner bytes
DM Mathematics 1.3970 1.8167 -0.4196 Zenyx 400,000
OpenSubtitles 1.3960 1.5502 -0.1541 Zenyx 400,032
ArXiv 0.8778 1.0177 -0.1399 Zenyx 400,820
HackerNews 1.3024 1.4344 -0.1319 Zenyx 239,283 *
Github 0.5853 0.6202 -0.0349 Zenyx 400,340
NIH ExPorter 0.8767 0.8930 -0.0163 Zenyx 220,665 *
Gutenberg (PG-19) 1.2704 1.2833 -0.0129 Zenyx 409,150
StackExchange 0.9890 0.9862 +0.0028 Gemma 400,410
Pile-CC 1.0970 1.0905 +0.0065 Gemma 403,210
PubMed Abstracts 0.8795 0.8704 +0.0091 Gemma 400,105
Enron Emails 1.2423 1.2273 +0.0151 Gemma 57,066 *
PubMed Central 0.6529 0.6342 +0.0187 Gemma 402,138
USPTO Backgrounds 0.7838 0.7603 +0.0235 Gemma 400,407
Wikipedia (en) 0.9923 0.9648 +0.0275 Gemma 400,898
BookCorpus2 1.1876 1.1393 +0.0483 Gemma 400,022
PhilPapers 1.3677 1.2804 +0.0872 Gemma 158,869 *
OpenWebText2 1.1495 1.0257 +0.1238 Gemma 410,197
Books3 1.2545 1.0838 +0.1708 Gemma 400,983
FreeLaw 0.9921 0.8013 +0.1908 Gemma 402,135
Ubuntu IRC 1.8089 1.5624 +0.2465 Gemma 43,987 *
EuroParl 2.0332 1.1373 +0.8959 Gemma 68,106 *
YoutubeSubtitles 1.8656 0.8918 +0.9738 Gemma 191,675 *

Zenyx wins 7 of 22 domains. Byte-weighted over all of them, Zenyx 1.0848 vs Gemma 1.0586 (+0.0262) -- Gemma ahead overall.

Rows marked * rest on under 250,000 bytes and carry correspondingly more noise. Excluding them, the weighted result reverses (15 domains: Zenyx 1.0341 vs Gemma 1.0430, -0.0090). That subset is not the headline: dropping the thin domains removes three of Zenyx's four worst losses while removing only two of its wins, so it flatters this model.

Science accuracy (higher is better)

benchmark metric Zenyx V3 Gemma 4 E2B chance winner
ARC-Challenge acc_norm 25.26% 52.65% 25.0% Gemma
SciQ acc 77.40% 96.90% 25.0% Gemma

What it says. Zenyx wins decisively on English technical and mathematical text -- DM Mathematics by 0.4196, its largest margin anywhere, plus ArXiv, GitHub and HackerNews. It loses heavily on multilingual and conversational text (YoutubeSubtitles, EuroParl, Ubuntu IRC), which is a category it was never trained for: the corpus is English and excludes Chinese outright, while Gemma is built multilingual. On the science tasks Zenyx is at chance on ARC-Challenge -- no measurable ability there yet -- and well behind on SciQ, though far above chance.

Those results belong together. A suite that flattered this model would not report chance-level performance where that is the truth, which is what makes the code and mathematics wins credible.

Head-to-head vs XHToken/Spark-X2.5-1.7B-Base

A second, more size-matched comparison: Spark is 1.71B params against this model's 1.58B, and its vocab (131,072) is close to this model's (129,280) -- so, unlike the Gemma comparison, tokenizer size is not a meaningful confound here either.

Same methodology as the Gemma comparison: one copy of the task code, same source text, bits/byte normalised by shared UTF-8 bytes.

The Pile, all 22 domains (bits/byte, LOWER is better)

domain Zenyx V3 Spark X2.5 margin winner bytes
OpenSubtitles 1.3967 1.2261 +0.1706 Spark 400,032
PubMed Central 0.6529 0.4622 +0.1907 Spark 402,138
Github 0.5797 0.3841 +0.1955 Spark 400,340
USPTO Backgrounds 0.7806 0.5803 +0.2003 Spark 400,407
DM Mathematics 1.3555 1.1500 +0.2055 Spark 400,000
NIH ExPorter 0.8769 0.6654 +0.2114 Spark 220,665 *
ArXiv 0.8738 0.6453 +0.2285 Spark 400,820
Wikipedia (en) 0.9953 0.7432 +0.2521 Spark 400,898
PubMed Abstracts 0.8787 0.6261 +0.2526 Spark 400,105
Pile-CC 1.0959 0.8381 +0.2579 Spark 403,210
HackerNews 1.3000 1.0211 +0.2789 Spark 239,283 *
Enron Emails 1.2405 0.9547 +0.2858 Spark 57,066 *
BookCorpus2 1.1859 0.8965 +0.2894 Spark 400,022
StackExchange 0.9841 0.6850 +0.2991 Spark 400,410
OpenWebText2 1.1485 0.8192 +0.3293 Spark 410,197
Gutenberg (PG-19) 1.2704 0.9352 +0.3352 Spark 409,150
FreeLaw 0.9893 0.6167 +0.3726 Spark 402,135
Books3 1.2515 0.8726 +0.3789 Spark 400,983
PhilPapers 1.3586 0.9652 +0.3934 Spark 158,869 *
Ubuntu IRC 1.7427 1.1782 +0.5645 Spark 43,987 *
YoutubeSubtitles 1.8521 0.7579 +1.0941 Spark 191,675 *
EuroParl 2.0156 0.8732 +1.1424 Spark 68,106 *

Spark wins 0 of 22 domains -- every one of them. Byte-weighted, Zenyx 1.0798 vs Spark 0.7806 (+0.2992). Rows marked * rest on under 250,000 bytes.

ARC-Challenge and LAMBADA (higher acc is better)

benchmark metric Zenyx V3 Spark X2.5 chance winner
ARC-Challenge acc_norm 26.45% 45.31% 25.0% Spark
LAMBADA acc 25.64% 58.53% 0.0% Spark

What it says. Unlike Gemma, where this model won on code and mathematics, Spark leads on every single Pile domain measured, with the smallest margins on exactly the domains this model has invested in most -- GitHub (+0.1955) and DM Mathematics (+0.2055) -- and the largest on multilingual text neither model's English-only corpus should be expected to win (EuroParl, YoutubeSubtitles). Spark also shows real multi-step reasoning (26.5% vs 45.3% on ARC-Challenge, chance is 25%) and a much lower LAMBADA perplexity, both markers this model does not yet show at this token count. The honest read: at a closely matched size and vocabulary, Spark X2.5 is a substantially stronger base model. LAMBADA perplexity: Zenyx 30.35 vs Spark 4.24.

Reading these numbers. This is a partially-trained 1.5B base model, so knowledge-heavy multiple-choice tasks sit close to their random baselines — that is expected at this scale and token count. The signal to watch is the language-modelling side: LAMBADA accuracy and WikiText perplexity measure whether the model has actually learned to predict text, and those improve steadily long before multiple-choice benchmarks move. Note also that BoolQ's majority-class baseline is 62.2%, so a score near that is not evidence of comprehension.


Hardware Serving Benchmarks (NVIDIA L4, 24 GB)

Measured with the JAX/Flax serving loop: static shape pre-allocation, bucketed prefill and GPU-native sampling.

Metric Value Notes
Decode speed 68.5 tok/s steady-state autoregressive decode
Warm prefill ~20 ms short prompt, shape already compiled
Checkpoint load ~26 s params → GPU, from local cache
Active VRAM ~5.0 GB of 24 GB

Cold shapes pay a one-off JIT compile (tens of seconds) the first time a new (prompt length, max tokens) pair is seen; warm requests are the numbers above.


Inference Example

from zenyx_v3_inference import ZenyxGenerator

generator = ZenyxGenerator(step=86400)

# Base model: give it a prefix to CONTINUE, not an instruction to follow.
print(generator.generate(
    "The capital of France is",
    max_new_tokens=80,
    temperature=0.7,
    repetition_penalty=1.15,
))

Evaluation Reproducibility

Benchmarks were produced by modal_base_evals.py on a single NVIDIA L4, scoring continuations in batches with length-bucketed padding. Task formats follow the lm-evaluation-harness conventions (prompt templates, acc / acc_norm definitions and answer-key handling), so the numbers are broadly comparable to published base-model results, though this is an independent implementation rather than a harness run.

Limitations

  • Pretraining is incomplete — the model will change substantially with more tokens.
  • Not instruction-tuned, not RLHF'd, and not safety-filtered. Outputs may be factually wrong, biased, or nonsensical.
  • Trained predominantly on English text, code, mathematics and synthetic reasoning data; other languages are not supported.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support