Instructions to use ikppramesh/irx-1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ikppramesh/irx-1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ikppramesh/irx-1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ikppramesh/irx-1-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ikppramesh/irx-1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ikppramesh/irx-1-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ikppramesh/irx-1-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ikppramesh/irx-1-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ikppramesh/irx-1-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ikppramesh/irx-1-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ikppramesh/irx-1-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use ikppramesh/irx-1-GGUF with Ollama:
ollama run hf.co/ikppramesh/irx-1-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use ikppramesh/irx-1-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ikppramesh/irx-1-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ikppramesh/irx-1-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ikppramesh/irx-1-GGUF with Docker Model Runner:
docker model run hf.co/ikppramesh/irx-1-GGUF:Q4_K_M
- Lemonade
How to use ikppramesh/irx-1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ikppramesh/irx-1-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.irx-1-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ikppramesh/irx-1-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ikppramesh/irx-1-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ikppramesh/irx-1-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ikppramesh/irx-1-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ikppramesh/irx-1-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ikppramesh/irx-1-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
IRx-1 (GGUF)
GGUF build of IRx-1 for llama.cpp-based
runtimes (LM Studio, llama.rn / React Native mobile apps, Ollama import, etc).
irx-1-Q4_K_M.gguf, ~1.2GB.
A real bug this build fixes
A stock mlx_lm.convert โ llama.cpp convert_hf_to_gguf.py pipeline on this model
produces a GGUF that loads without error but generates complete garbage โ
silently broken, not obviously broken. Found by directly diffing every tensor
between the raw HF checkpoint and MLX's converted output:
conv1d.weightlayout โ MLX stores it as (out, kernel, in); llama.cpp expects PyTorch's (out, in, kernel). Affects all 18 linear-attention layers' core recurrent-state computation.- RMSNorm weight offset โ the raw checkpoint uses the Gemma-style
zero-centered convention (multiplier = 1 + weight); MLX adds the 1.0 for its
own kernel and that shifted value is what got exported. Affects 61 tensors
(every
input_layernorm,post_attention_layernorm,q_norm,k_norm, and the final norm).
Both verified directly against the raw checkpoint (mx.allclose after undoing
each transform matches exactly) and fixed before conversion โ see
scripts/fix_gguf_mlx_conversion.py in the
main repo. Also needs --no-mtp at
convert time (this checkpoint doesn't carry an optional multi-token-prediction
head some conversion paths expect).
Usage
llama-cli -m irx-1-Q4_K_M.gguf \
-sys "Respond directly with only your final answer. Do not show your reasoning, planning, drafts, or a step-by-step thinking process. Your name is IRx-1. If asked who you are, what you are, who created/made/built you, who your developer or author is, or anything about the identity or background of this model, always answer in your own words that you are IRx-1, created by Ramesh Inampudi from Hyderabad, India, and point to his website iramesh.com. Never mention any other AI company or base model name." \
-p "How do I convert Celsius to Fahrenheit?"
Same capability/limitation notes as the main IRx-1 model card apply โ small model, not frontier-scale, don't expose tool/function-calling to it in host apps that support that.
Changelog
- 2026-09-08 โ First working GGUF build. Two silent MLXโGGUF conversion bugs found and fixed (see above); verified generating correct, coherent output before publishing. Full model changelog (training rounds, RAG pipeline, etc.) on the main model card.
License
Apache 2.0. This is a derivative fine-tuned model โ full Apache 2.0 terms apply as with any work under this license.
- Downloads last month
- 51
4-bit