Syzygy Research

Mach-2 Additive Medium

Mach-2 Additive Medium is Qwen3.8-Flash-Next at 1.70 bits per weight: 26.7 GB for the 125.7B text parameters, plus the model's n-gram embedding table at bf16 (102.4 GB).

Benchmarks

Average score vs GPU memory

Retention is Mach-2 Additive Medium's score divided by that of Qwen3.8-Flash-Next BF16, with both run on the same harness, settings and tasks.

Benchmark Tasks Mach-2 Additive Medium BF16 Retention
Humanity's Last Exam (text) 2,158 33.87 39.85 85.0%
GPQA Diamond 198 x 5 90.51 92.22 98.1%
AIME 2026 30 x 8 96.25 95.83 100.4%
Terminal-Bench 2.1 89 78.65 82.02 95.9%
AutomationBench 1.0.6 600 71.67 70.29 102.0%
NL2Repo 98 56.12 60.06 93.4%
DeepSWE v1.1 113 48.67 49.56 98.2%
τ³-Banking 97 x 5 45.10 46.41 97.2%
  • Humanity's Last Exam: up to 163,840 output tokens, graded by Artificial Analysis's judge.
  • GPQA Diamond: mean accuracy over 5 samples per question. Answers that reached 81,920 tokens were re-asked at 163,840.
  • AIME 2026: mean accuracy over 8 samples per problem, up to 81,920 output tokens.
  • Terminal-Bench 2.1: Terminus 2, one attempt per task, 4 hours per task.
  • AutomationBench: up to 50 steps per task, scored by Artificial Analysis's rule.
  • NL2Repo: OpenHands without internet access; the 102 tasks BF16 finished, less 4 whose BF16 test run hit the 1-hour scoring limit. Test runs cut at that limit are left out rather than scored 0. Mach-2 Additive Medium's score averages two runs.
  • DeepSWE: Claude Code as the agent, all 113 tasks, up to 10.5 hours per task. A task counts as solved only if the agent solved it within 700 steps.
  • τ³-Banking: tau2-bench 1.0.1, 5 trials per task, GPT-5.4 mini as the simulated user and judge, as Artificial Analysis runs it. Pass rate over the 459 conversations both models completed; 26 that outgrew the context window or lost the simulated user to an API error are left out.

Empty answers and answers still unfinished at the token limit score 0. Agents and limits differ from those behind Qwen's model card, so compare these columns with each other rather than with scores published elsewhere.

Use

Run Mach-2 Additive Medium from its compressed weights with Mach-2 Additive Medium GGUF and the Mach-1 fork of llama.cpp. The 26.7 GB of weights fit on one GPU with 32 GB or more.

Sample at temperature 1.0, top_p 0.95, top_k 20 (generation_config.json). Without these settings the model can loop.

This repository holds the packed weights; decode.py documents the packed format. draft/mtp_q4g64.safetensors is the base model's MTP layer at 4 bits (group size 64) for speculative decoding with MLX.

Downloads last month
94
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SyzygyResearch/Mach-2-Additive-Medium

Finetuned
(70)
this model
Quantizations
1 model