MiMo-9B-CORTEX v1 (GGUF, Q4_K_M)

CORTEX-protocol fine-tune of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, exported as a single llama.cpp GGUF. Passes the full CORTEX battery with its own patch.

File: MiMo-9B-CORTEX-v1-Q4_K_M.gguf — 5.63 GB sha256: 34c79f8bf79b3819199821361d8ca221d8749d77bbacce185587c60d5e6b8365

Architecture — what it actually is

qwen3_5 dense hybrid (transformer family): 32 layers in a 3:1 pattern of gated DeltaNet linear-attention layers to periodic full-attention layers, plus an MoE-free dense FFN. Vocab 248,320. It is not attention-free — see the honest ledger below. Converted with --no-mtp (base config declares mtp_num_hidden_layers: 1 but ships no MTP module; the metadata would break llama.cpp loading otherwise).

Battery results (identical harness across the fleet, cpu-xl, llama.cpp b11191)

Test Result
Repair trial (AWAKE→REPAIR→REST) PASS — model patch applied, no AST fallback, pytest 3/3
Trace header Exact `[TRACE: §R:…
Rule 1 (unmapped §QUANTUM:ENTANGLE) Fail-closed — declines to invoke, files a proposal, no hallucination
Generation speed tg64 11.23 tok/s
Prompt speed pp29 35.00 tok/s
Resident RAM (probe / bench) ~9.1 GB
Trial generation wall-clock 40.4 s

Provenance

Honest ledger

  • This model is a transformer-family architecture. The user's original "no transformers" directive applies to the runtime (pure llama.cpp, no transformers library at inference — verified by the harness ENGINE CHECK). The attention-free exploration lives on the Falcon-Mamba branch of this project.
  • Rule-1 probe emits a valid trace + refusal but not the literal UNMAPPED marker — behaviorally fail-closed, not literally. Documented, not hidden.
  • Speed numbers are same-box relative (cpu-xl); absolute values vary with hardware.

Runtime

Zero-dependency: llama.cpp / llama-cpp-python only. Stop guard: <|im_end|> (248046) + <|endoftext|> (248044). See cortex-knowledge-vault/tools/runtime for the pinned token_map.json and tri-state cortex_runtime.py.

Downloads last month
97
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support