Instructions to use hermitdave/K2-Horizon-7B-Uno-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use hermitdave/K2-Horizon-7B-Uno-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("hermitdave/K2-Horizon-7B-Uno-MLX-4bit") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use hermitdave/K2-Horizon-7B-Uno-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "hermitdave/K2-Horizon-7B-Uno-MLX-4bit" --prompt "Once upon a time"
- Atomic Chat
K2-Horizon-7B-Uno MLX 4-bit
Uniform 4-bit quantization of the merged IFM/K2-Horizon-7B-Uno model — a diffusion-augmented LLM based on K2-Horizon-7B. The LoRA adapter is baked into the base weights, so it runs as a standard autoregressive model.
Upstream model: IFM/K2-Horizon-7B-Uno by Institute of Foundation Models, released under Apache 2.0.
Conversion: Merged and quantized to MLX format using Hermes Agent with mlx-lm.
Quickstart
pip install -U mlx-lm
python3 -m mlx_lm.generate \
--model hermitdave/K2-Horizon-7B-Uno-MLX-4bit \
--prompt "Explain step by step." \
--max-tokens 512 --temp 1.0 --top-p 0.95
Reasoning
K2-Horizon-7B is a reasoning model. Always use reasoning_effort="high":
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="hermitdave/K2-Horizon-7B-Uno-MLX-4bit",
messages=[{"role": "user", "content": "Explain step by step."}],
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print("Reasoning:", getattr(response.choices[0].message, "reasoning_content", None))
print("Answer:", response.choices[0].message.content)
oMLX Patch
K2-Horizon requires oMLX v0.6.4+ with the K2-Horizon support patch.
Citation
@misc{k2_horizon_7b_uno,
title = {K2-Horizon-7B-Uno},
author = {Institute of Foundation Models},
year = {2026},
howpublished = {\url{https://huggingface.co/IFM/K2-Horizon-7B-Uno}},
}
License
Apache 2.0 (same as upstream).
- Downloads last month
- -
Model size
9B params
Tensor type
BF16
·
F32 ·
Hardware compatibility
Log In to add your hardware
4-bit