IRx-1 (GGUF)

GGUF build of IRx-1 for llama.cpp-based runtimes (LM Studio, llama.rn / React Native mobile apps, Ollama import, etc). irx-1-Q4_K_M.gguf, ~1.2GB.

A real bug this build fixes

A stock mlx_lm.convert โ†’ llama.cpp convert_hf_to_gguf.py pipeline on this model produces a GGUF that loads without error but generates complete garbage โ€” silently broken, not obviously broken. Found by directly diffing every tensor between the raw HF checkpoint and MLX's converted output:

  1. conv1d.weight layout โ€” MLX stores it as (out, kernel, in); llama.cpp expects PyTorch's (out, in, kernel). Affects all 18 linear-attention layers' core recurrent-state computation.
  2. RMSNorm weight offset โ€” the raw checkpoint uses the Gemma-style zero-centered convention (multiplier = 1 + weight); MLX adds the 1.0 for its own kernel and that shifted value is what got exported. Affects 61 tensors (every input_layernorm, post_attention_layernorm, q_norm, k_norm, and the final norm).

Both verified directly against the raw checkpoint (mx.allclose after undoing each transform matches exactly) and fixed before conversion โ€” see scripts/fix_gguf_mlx_conversion.py in the main repo. Also needs --no-mtp at convert time (this checkpoint doesn't carry an optional multi-token-prediction head some conversion paths expect).

Usage

llama-cli -m irx-1-Q4_K_M.gguf \
  -sys "Respond directly with only your final answer. Do not show your reasoning, planning, drafts, or a step-by-step thinking process. Your name is IRx-1. If asked who you are, what you are, who created/made/built you, who your developer or author is, or anything about the identity or background of this model, always answer in your own words that you are IRx-1, created by Ramesh Inampudi from Hyderabad, India, and point to his website iramesh.com. Never mention any other AI company or base model name." \
  -p "How do I convert Celsius to Fahrenheit?"

Same capability/limitation notes as the main IRx-1 model card apply โ€” small model, not frontier-scale, don't expose tool/function-calling to it in host apps that support that.

Changelog

  • 2026-09-08 โ€” First working GGUF build. Two silent MLXโ†’GGUF conversion bugs found and fixed (see above); verified generating correct, coherent output before publishing. Full model changelog (training rounds, RAG pipeline, etc.) on the main model card.

License

Apache 2.0. This is a derivative fine-tuned model โ€” full Apache 2.0 terms apply as with any work under this license.

Downloads last month
51
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support