MiniMax M3: a tiny toolchain fixture

A small, independently initialized random-weight reference model for learning, debugging, and reproducing QFS capture workflows. It is a base/root fixture, not a quantized child.

These weights are untrained, not an assistant, and not a language-quality benchmark. Generated text has no useful semantic quality. No upstream trained weights or training data are implied by an architecture name.

Weight lineage

This checkpoint was initialized directly for an architecture test. It was not fine-tuned from a production model and does not inherit that model's trained weights. Architecture/configuration lineage is documented separately from weight lineage.

Browse the Random Architecture Fixtures collection. A random base may also serve as the common source for an aligned quantization family.

At a glance

Property Observed value
Artifact role Base/reference model; capture role root
Serialized top-level weight files 156,272 bytes (0.149 MiB); metadata, tokenizer and evidence excluded
Generated parameters 72,000
Included vision parameters 8,336; inclusion is not vision-quality evidence
Saved semantic tensors 76 (parameters and buffers are not interchangeable)
Fixture vocabulary 272 tokens; independent byte tokenizer, not upstream vocabulary
CPU stack Python 3.12; Torch 2.11.0+cpu; Transformers 5.16.1; two Torch threads in the recorded capture workflow
Remote model code Not required: native Transformers class behind QFS

Size is serialized artifact size, not runtime RAM or a capacity/performance guarantee. Packed-array element counts are not model parameter counts.

What is actually included

Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and untied LM head. Dense-to-MoE transition, shared experts, Gemma norms, per-head Q/K norms, clipped SwiGLU and learned block-sparse indexer preserved. Text panel and separate raw-patch vision smoke only; no processor, image quality, video or MTP claim. Native sparse mask materializes densely on CPU; no production sparse-kernel performance claim.

The repository file inventory includes config.json, generation_config.json, tokenizer.json, tokenizer_config.json, build-manifest.json, support.json, requirements-cpu.txt. Exact build/runtime/license inventories are linked below. A file may describe historical provenance without being an executable entry point.

Good community uses—and boundaries

  • Learn how to download pinned artifacts, seal a small synthetic token panel, capture hidden states and replay a full vocabulary head.
  • Debug model-family adapters, strict tensor loading, storage decoders, or reproducibility tooling without downloading a production-sized checkpoint.
  • Reproduce a narrowly scoped result, report an adapter/reader regression, and retain the source, panel and runtime identities needed to explain it.

Not established: trained-model accuracy; useful instruction following; quantizer optimization quality; original production-weight compatibility; GPU/NPU/serving-kernel parity; cross-hardware determinism; long-context behavior outside the recorded panel; throughput or paid-compute admission.

Original-release limitations

Official checkpoint is BF16 (no quantization_config); full model is approximately 428B parameters and is not a tiny-CPU workload. Native eager/SDPA implements learned sparse selection using a dense materialized mask, not the production MSA sparse kernel. Official text config advertises seven MTP modules/one next-token predictor; native class intentionally omits MTP and has a narrow upstream unexpected-key rule (^|.)mtp..*. Tiny checkpoint contains no MTP keys. Existing QFS layer_outer mapping handles source language_model.model.layers.N -> native model.language_model.layers.N; this fixture does not duplicate that conversion. No original-scale, MSA-kernel, video, left-padding equivalence or original-quality measurement is claimed.

Model family, root dataset and evidence

This model is the base/root of its own random fixture family, not a reproduction of the trained upstream model. The fidelity-root dataset repository is the family reference location. A link is not a claim that registration or publication has completed.

The historical CPU evidence bundle retains two independent captures at first/ and repeat/, plus comparison/. Its repository root is a receipt bundle, not a single canonical QFS dataset. Those recorded same-machine/same-stack forced comparisons reported 0.0 nats on 252 synthetic scored positions; that is not a trained-quality result or a transferable hardware floor. The original detailed receipt and caveats remain authoritative.

Runtime requirements and safe local reproduction

Replay the published evidence without loading a model

This uses QFS's existing NumPy FP64 comparator in a Torch-free environment. It reads the stored hidden states and each side's own head. No model forward, remote model code, GPU, upload or registry mutation is involved. Runtime receipts name the actual backend; last-bit differences from another FP64 implementation are not a new quality claim. Synthetic exact controls remain zero.

Set QFS to a reviewed Quant Fidelity Suite checkout and use Bash:

: "${QFS:?Set QFS to your reviewed QFS checkout}"
WORK=$(mktemp -d)
export OMP_NUM_THREADS=2 MKL_NUM_THREADS=2 OPENBLAS_NUM_THREADS=2
python3.12 -m venv "$WORK/replay-env"
"$WORK/replay-env/bin/pip" install 'numpy==2.5.3' 'huggingface-hub==1.30.0'
"$WORK/replay-env/bin/hf" download malaiwah/qfs-fixture-root-captures-v1 --repo-type dataset \
  --revision f53b204091c988ce4a2161af81886f3018745559 --include 'roots/minimax-m3/**' --local-dir "$WORK/evidence"
"$WORK/replay-env/bin/python" "$QFS/bin/fidelity_dataset.py" verify \
  "$WORK/evidence/roots/minimax-m3/first" --verify-tensors
"$WORK/replay-env/bin/python" "$QFS/bin/fidelity_dataset.py" verify \
  "$WORK/evidence/roots/minimax-m3/repeat" --verify-tensors
"$WORK/replay-env/bin/python" "$QFS/bin/fidelity_dataset.py" compare \
  --reference "$WORK/evidence/roots/minimax-m3/first" \
  --candidate "$WORK/evidence/roots/minimax-m3/repeat" \
  --out "$WORK/replayed" --device cpu --replay-device numpy --replay-dtype float32 \
  --vocab-chunk 8192 --verify-tensors --self-compare --force-compute

Reconstructed format comparisons are intentionally advisory and normally return exit code 2 while writing a valid receipt. Inspect that receipt; do not silence refusals or interpret an advisory result as a production-quality ranking.

Capture the actual checkpoint

Capture uses a separate pinned Torch CPU environment. The input below is the original token-panel format, not a sealed capture's internal panel/ folder. The native source, tokenizer and model revisions remain explicit. Set AUTHOR to your own HF handle; the dataset repository argument is attribution only and nothing is uploaded by these commands.

: "${AUTHOR:?Set AUTHOR to your Hugging Face handle}"
"$WORK/replay-env/bin/hf" download malaiwah/minimax-m3-tiny-random-bf16 --revision ca904dff2941004a3f5bd5af62064fcfb34da8b9 --local-dir "$WORK/model"
"$WORK/replay-env/bin/hf" download malaiwah/minimax-m3-tiny-random-bf16 --revision ca904dff2941004a3f5bd5af62064fcfb34da8b9 --local-dir "$WORK/source"
"$WORK/replay-env/bin/hf" download malaiwah/qfs-fixture-root-captures-v1 --repo-type dataset \
  --revision f53b204091c988ce4a2161af81886f3018745559 --include 'requirements-capture.txt' --include 'roots/minimax-m3/input-panel/**' \
  --local-dir "$WORK/inputs"
python3.12 -m venv "$WORK/capture-env"
"$WORK/capture-env/bin/pip" install -r "$WORK/inputs/requirements-capture.txt"
"$WORK/capture-env/bin/python" "$QFS/bin/fidelity_dataset.py" architectures prepare \
  --architecture minimax-m3 --model-dir "$WORK/model" \
  --model-repository malaiwah/minimax-m3-tiny-random-bf16 --model-revision ca904dff2941004a3f5bd5af62064fcfb34da8b9 \
  --panel "$WORK/inputs/roots/minimax-m3/input-panel" --tokenizer-root "$WORK/source" \
  --author "$AUTHOR" --dataset-repository "$AUTHOR/minimax-m3-tiny-random-bf16-capture" \
  --dataset-id "fidelity--$AUTHOR.minimax-m3-tiny-random-bf16" --out "$WORK/workflow"

Inspect workflow.json. This launcher executes its exact capture/verification commands and selects the Torch-free interpreter only for the final comparison:

export OMP_NUM_THREADS=2 MKL_NUM_THREADS=2 OPENBLAS_NUM_THREADS=2
"$WORK/replay-env/bin/python" - "$WORK/workflow/workflow.json" <<'PY'
import json, subprocess, sys
workflow = json.load(open(sys.argv[1]))
for step in workflow["commands"]:
    argv = list(step["argv"])
    if step["step"] == "compare":
        argv[0] = sys.executable
    result = subprocess.run(argv)
    if result.returncode:
        raise SystemExit(result.returncode)
PY

Use the actual root at malaiwah/minimax-m3-tiny-fidelity-root-v1@4bdcedf286fa2de29bc11102e90cac47cd77cd14. The community collection groups models and captures. Custom code, where required above, is explicitly pinned and executed only after your consent; hash verification is provenance, not a sandbox.

Licensing and detailed provenance

  • Fixture repository license: Exact upstream LICENSE retained conservatively; independently generated random weights, no upstream weights copied. Native Transformers runtime separately Apache-2.0. No vendored runtime Python.
  • Summary: MiniMax-M3 pinned MINIMAX COMMUNITY LICENSE: commercial use requires Built with MiniMax M3 attribution and one-time email notice; products/services exceeding US$20M annual revenue require prior written authorization. Prohibited-use appendix applies. Exact upstream terms retained; no claim that upstream is MIT.
Immutable provenance, historical cards, source inventories and full caveats

Selected original provenance fields (full tensor/component evidence remains in the linked inventories):

{
  "generation": "Complete seeded native FP32 initialization rounded to BF16; exact semantic native save/reload including dtype required. No loader guard overrides.",
  "generator_sha256": "3eb685652c96791a1a7da0e9da453426ade42e50f5a5bdeb9ef3abc4a9523299",
  "lineage": {
    "relationship": "architecture lineage only; no source weights copied",
    "repository": "MiniMaxAI/MiniMax-M3",
    "revision": "f0e1c1e04d40177e4673a22097036854f536e9c0"
  },
  "native_fp32_exceptions": [],
  "parameter_count": 72000,
  "scope": "Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and untied LM head. Dense-to-MoE transition, shared experts, Gemma norms, per-head Q/K norms, clipped SwiGLU and learned block-sparse indexer preserved. Text panel and separate raw-patch vision smoke only; no processor, image quality, video or MTP claim. Native sparse mask materializes densely on CPU; no production sparse-kernel performance claim.",
  "seed": 20260907,
  "versions": {
    "safetensors": "0.8.0",
    "tokenizers": "0.23.2",
    "torch": "2.11.0+cpu",
    "transformers": "5.16.1"
  },
  "vision_parameter_count": 8336
}
Downloads last month
304
Safetensors
Model size
72k params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collections including malaiwah/minimax-m3-tiny-random-bf16