Mistral Small 4 119B-A6B GGUF β€” Quantized by BatiAI

Mistral's unified open-weight model β€” reasoning + multimodal + agentic coding in one. 119B Mixture-of-Experts with only 6B active per token, so it runs at small-model speed while keeping frontier quality. Apache 2.0 (fully commercial-friendly, no gating).

Quantized directly from official Mistral weights β€” not a re-quant of someone else's GGUF. Signed with BatiAI metadata for BatiFlow.

Quick Start

ollama run batiai/mistral-small-4:q4

Available Quantizations

Quant Size RAM target Recommended For
Q3_K_M 54GB 64GB Mac Compact
Q4_K_M 68GB 96GB Mac Recommended (balance)
Q5_K_M 79GB 128GB Mac Max quality

119B total params β†’ these are for 64GB+ Macs (M-series Max/Ultra). The 6B active means inference is fast despite the size. IQ3/IQ4 (smaller, imatrix) can be added on request.

RAM Requirements

Your Mac RAM Q3 (54GB) Q4 (68GB) Q5 (79GB)
64GB βœ… tight ❌ ❌
96GB βœ… βœ… ❌ tight
128GB βœ… βœ… βœ…
192GB+ βœ… βœ… βœ… comfortable

Why Mistral Small 4?

  • One model, three jobs β€” reasoning, multimodal understanding, and agentic coding unified (no model-switching).
  • 6B active / 119B total MoE β€” frontier-class capability at the inference speed/cost of a small model.
  • Apache 2.0 β€” no license friction, no gating. Build commercial products freely.
  • Native llama.cpp support β€” Mistral ships official GGUF tooling; arch (mistral3) is mainstream.

Why BatiAI Quantization?

  • Original-source β€” quantized straight from Mistral's official weights, not a copy of a third-party GGUF.
  • BatiAI-signed β€” general.author: BatiAI, general.url: https://flow.bati.ai.
  • Mac-tuned selection β€” quant sizes chosen for real Apple Silicon RAM tiers.

Technical Details

  • Original Model: mistralai/Mistral-Small-4-119B-2603
  • Architecture: mistral3 MoE, 119B total / ~6B active per token
  • License: Apache 2.0
  • Quantized with: llama.cpp (Q8_0 intermediate β†’ K-quants, --allow-requantize)
  • Quantized by: BatiAI
  • Note: YaRN config verified clean (no yarn_log_multiplier bug, unlike the earlier Medium 3.5 release).

About BatiFlow

BatiFlow β€” free, on-device AI automation for Mac. 5MB app, 100% local, unlimited. 60+ tools.

License

Quantized from mistralai/Mistral-Small-4-119B-2603. License: Apache 2.0.

Benchmarks

Mac ν•˜λ“œμ›¨μ–΄ μ‹€μΈ‘ 벀치 λŒ€κΈ° 쀑 (bench.sh). μΈ‘μ • ν›„ μžλ™ κ°±μ‹ .

Downloads last month
184
GGUF
Model size
119B params
Architecture
mistral4
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for batiai/Mistral-Small-4-119B-GGUF

Quantized
(31)
this model