K2-Horizon-7B-Uno MLX 4-bit

Uniform 4-bit quantization of the merged IFM/K2-Horizon-7B-Uno model — a diffusion-augmented LLM based on K2-Horizon-7B. The LoRA adapter is baked into the base weights, so it runs as a standard autoregressive model.

Upstream model: IFM/K2-Horizon-7B-Uno by Institute of Foundation Models, released under Apache 2.0.

Conversion: Merged and quantized to MLX format using Hermes Agent with mlx-lm.

Quickstart

pip install -U mlx-lm

python3 -m mlx_lm.generate \
  --model hermitdave/K2-Horizon-7B-Uno-MLX-4bit \
  --prompt "Explain step by step." \
  --max-tokens 512 --temp 1.0 --top-p 0.95

Reasoning

K2-Horizon-7B is a reasoning model. Always use reasoning_effort="high":

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="hermitdave/K2-Horizon-7B-Uno-MLX-4bit",
    messages=[{"role": "user", "content": "Explain step by step."}],
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print("Reasoning:", getattr(response.choices[0].message, "reasoning_content", None))
print("Answer:", response.choices[0].message.content)

oMLX Patch

K2-Horizon requires oMLX v0.6.4+ with the K2-Horizon support patch.

Citation

@misc{k2_horizon_7b_uno,
  title        = {K2-Horizon-7B-Uno},
  author       = {Institute of Foundation Models},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/IFM/K2-Horizon-7B-Uno}},
}

License

Apache 2.0 (same as upstream).

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hermitdave/K2-Horizon-7B-Uno-MLX-4bit

Finetuned
(2)
this model