YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Per-Layer E2E Cos: Verify MLX Quantization Fidelity Layer-by-Layer

Does your 4-bit MLX model compute what the FP32 weights say it should?

This tool runs a single token through every transformer layer of an MLX 4-bit model, executing the EXACT same architecture in numpy FP32 as a ground-truth reference. It computes cosine similarity per layer β€” telling you precisely where quantization drift creeps in.

Results (Qwen3.5-35B-A3B-4bit, 40 layers)

  • 24/40 layers match perfectly (T-, cos > 0.99)
  • 15/40 layers show minor drift (H-class, 0.80 < cos < 0.99)
  • 1/40 layers diverge (T+, cos < 0.80)
  • Mean cosine: 0.966

Key Findings

Finding Result
Per-component matmul fidelity 160/160 at cos > 0.99999 β€” every operation correct
Self-determinism 40/40 L-inf=0 β€” MLX 4-bit is bit-identical across runs
Logit self-match 10/10 top-1 token match β€” output deterministic
bf16 vs 4-bit (same framework) Semantically equivalent β€” 54% word overlap, identical answers
bf16 vs 4-bit (cross-framework vs numpy) 0/10 top-1 match β€” measures framework, not model
Production quality 25/25 probes correct on OptiQ
Benchmarks (5 tasks, both models) Tied within error bars β€” per-layer cos β‰  benchmark accuracy

What This Tells Us

Cross-framework per-layer E2E cosine similarity measures the GAP between MLX GPU fused-kernel fp16 and numpy CPU FP32 β€” not model quality. The 4-bit model is proven correct by self-determinism, logit match, bf16 comparison, and output quality probes. Always run a self-determinism baseline before interpreting cross-framework results.

Usage

python3 per_layer_e2e_cos.py --model mlx-community/Qwen3.5-35B-A3B-4bit --token-id 151644

How It Works

  1. Loads the MLX model and does a forward pass β€” extracts per-layer output
  2. Loads the raw safetensors weights, dequantizes to FP32, reimplements the full layer in numpy
  3. Computes cosine similarity between MLX output and numpy reference at each layer
  4. Persists results as JSON

Architecture Support

  • Full attention: QKV projections, Q-norm/K-norm, interleaved Q-gate per head, sigmoid gating, GQA
  • Gated DeltaNet: Conv1d, QKVZ projections, delta rule recurrence, silu gating
  • Mixture of Experts: Top-k routing, shared expert gate, switch MLP dispatch
  • 4-bit + 8-bit quantization: group_size=64 affine dequant with scales/biases

Requirements

  • Apple Silicon (MLX runtime)
  • mlx >= 0.10.0, safetensors >= 0.4.0, numpy >= 1.24.0

What Not to Measure

After 12 days and 109 verified measurements, the lessons are clear:

  • Measure: Per-component matmul cos, self-determinism, output-level logit match, standard benchmarks
  • Skip: Cross-framework per-layer E2E without self-baseline, adversary scores as quality metric, single-token measurements at low norms
  • Always: Persist to non-/tmp storage, link every claim to a measurement file, self-compare first

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support