YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Per-Layer E2E Cos: Verify MLX Quantization Fidelity Layer-by-Layer
Does your 4-bit MLX model compute what the FP32 weights say it should?
This tool runs a single token through every transformer layer of an MLX 4-bit model, executing the EXACT same architecture in numpy FP32 as a ground-truth reference. It computes cosine similarity per layer β telling you precisely where quantization drift creeps in.
Results (Qwen3.5-35B-A3B-4bit, 40 layers)
- 24/40 layers match perfectly (T-, cos > 0.99)
- 15/40 layers show minor drift (H-class, 0.80 < cos < 0.99)
- 1/40 layers diverge (T+, cos < 0.80)
- Mean cosine: 0.966
Key Findings
| Finding | Result |
|---|---|
| Per-component matmul fidelity | 160/160 at cos > 0.99999 β every operation correct |
| Self-determinism | 40/40 L-inf=0 β MLX 4-bit is bit-identical across runs |
| Logit self-match | 10/10 top-1 token match β output deterministic |
| bf16 vs 4-bit (same framework) | Semantically equivalent β 54% word overlap, identical answers |
| bf16 vs 4-bit (cross-framework vs numpy) | 0/10 top-1 match β measures framework, not model |
| Production quality | 25/25 probes correct on OptiQ |
| Benchmarks (5 tasks, both models) | Tied within error bars β per-layer cos β benchmark accuracy |
What This Tells Us
Cross-framework per-layer E2E cosine similarity measures the GAP between MLX GPU fused-kernel fp16 and numpy CPU FP32 β not model quality. The 4-bit model is proven correct by self-determinism, logit match, bf16 comparison, and output quality probes. Always run a self-determinism baseline before interpreting cross-framework results.
Usage
python3 per_layer_e2e_cos.py --model mlx-community/Qwen3.5-35B-A3B-4bit --token-id 151644
How It Works
- Loads the MLX model and does a forward pass β extracts per-layer output
- Loads the raw safetensors weights, dequantizes to FP32, reimplements the full layer in numpy
- Computes cosine similarity between MLX output and numpy reference at each layer
- Persists results as JSON
Architecture Support
- Full attention: QKV projections, Q-norm/K-norm, interleaved Q-gate per head, sigmoid gating, GQA
- Gated DeltaNet: Conv1d, QKVZ projections, delta rule recurrence, silu gating
- Mixture of Experts: Top-k routing, shared expert gate, switch MLP dispatch
- 4-bit + 8-bit quantization: group_size=64 affine dequant with scales/biases
Requirements
- Apple Silicon (MLX runtime)
- mlx >= 0.10.0, safetensors >= 0.4.0, numpy >= 1.24.0
What Not to Measure
After 12 days and 109 verified measurements, the lessons are clear:
- Measure: Per-component matmul cos, self-determinism, output-level logit match, standard benchmarks
- Skip: Cross-framework per-layer E2E without self-baseline, adversary scores as quality metric, single-token measurements at low norms
- Always: Persist to non-/tmp storage, link every claim to a measurement file, self-compare first
License
MIT