Instructions to use ikppramesh/irx-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ikppramesh/irx-2 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ikppramesh/irx-2") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ikppramesh/irx-2 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ikppramesh/irx-2"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ikppramesh/irx-2" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ikppramesh/irx-2 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ikppramesh/irx-2"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ikppramesh/irx-2" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ikppramesh/irx-2", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ikppramesh/irx-2 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ikppramesh/irx-2"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ikppramesh/irx-2
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ikppramesh/irx-2 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ikppramesh/irx-2"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ikppramesh/irx-2" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
IRx-2 — ⭐ Best Overall
The recommended model of the IRx family: a private, offline AI chat assistant with the best balance of answer quality and speed. Everything runs on your device — no internet connection, no account, and nothing you type ever leaves your phone, tablet or laptop.
The IRx family
All four run fully offline — nothing you type ever leaves your device.
| Model | Download | Best for | Phones & tablets (GGUF) | Mac (MLX) |
|---|---|---|---|---|
| IRx-mini | ~0.8GB | The smallest. Simple, quick answers on phones with just 4GB of RAM | irx-mini-GGUF | irx-mini |
| IRx-1 | ~1.2GB | Fast, quick everyday answers on almost any phone | irx-1-GGUF | irx-1 |
| ⭐ IRx-2 — Best Overall | ~2.7GB | The best balance of answer quality and speed | irx-2-GGUF | irx-2 |
| IRx-2 Pro | ~5.3GB | The most detailed, long-form answers and long conversations | irx-2-pro-GGUF | irx-2-pro |
Not sure which to pick? Start with ⭐ IRx-2. Choose IRx-mini for a 4GB-RAM phone, IRx-1 for fast everyday answers on most phones, and IRx-2 Pro for the richest answers on a device with 12GB+ RAM.
See them side by side: the IRx board compares all four models, shows their measured speed live, and has an animated step-by-step setup guide for PocketPal on iPhone and Android.
What IRx-2 does best — ⭐ Best Overall
IRx-2 is the recommended model for most people. It is about twice the size of IRx-1, which shows in clearer reasoning, better-structured answers, more reliable writing and stronger coding help — while still running comfortably on phones and tablets with 8GB+ RAM.
| Who | Example things to ask |
|---|---|
| Software developers | "Explain this error and how to fix it: …" · "Write a function that validates an email address, with tests." · "Why is this SQL query slow?" |
| Farmers | "Plan a monthly budget for a 2-acre vegetable farm." · "Compare drip and flood irrigation: pros and cons." · "Write a loan application letter to my bank." |
| Students | "Make a one-week study plan for my exams." · "Explain Newton's laws with everyday examples." |
| Teachers | "Create a 40-minute lesson plan on the water cycle, with a short quiz." |
| Shop & small business owners | "Outline a simple business plan for a tea stall." · "Write a polite reply to a customer complaint." |
| Writers & creators | "Outline a 5-minute YouTube script about saving money." · "Make this paragraph sound more professional." |
| Job seekers | "Improve these resume bullet points." · "Give me 10 likely interview questions for a sales role." |
| Everyday life | "Plan a family weekend on a budget." · "Help me write a birthday message for my father." |
Measured on the same Mac CPU: IRx-2 generates about 26 tokens/s vs IRx-1's 54 — roughly half the speed, for noticeably better answers.
Usage (Mac, MLX)
pip install mlx-lm
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load("ikppramesh/irx-2")
messages = [{"role": "user", "content": "How do I convert Celsius to Fahrenheit?"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=1024,
sampler=make_sampler(temp=0.7, top_p=0.95)))
No system prompt is required: the built-in chat template supplies the IRx-2 one when
none is given, and always keeps thinking mode off. Sample at a non-zero temperature
(e.g. temp=0.7). On a phone or tablet, use the
GGUF build.
Limitations
- Not a frontier-scale model — it won't match large hosted AI services on very hard multi-step reasoning, deep coding problems or breadth of world knowledge.
- Facts can be wrong. Like any small offline model it can state things confidently that aren't true — double-check anything important, such as prices, medicines and doses, laws, and local farming advice (seed varieties, chemical quantities, weather).
- Not a replacement for professionals — medical, legal and financial answers are general information only.
- News up to 3 October 2026 only — it can't look anything up, and news details can be imprecise.
- Don't enable native tool/function-calling in chat apps — never trained; plain chat is reliable.
Changelog
2026-10-03 — News refresh and quality update
- Built-in news digest (up to 3 October 2026). The model's built-in prompt now carries 15 of the day's top headlines across India, world, business, science and technology, AI and sports, so it can answer questions about them offline, in any app, with no internet. In testing, answers about digest headlines contained the real facts 47% of the time, against 0% for news learned through training alone.
- Knows its cutoff. Asked how current it is, it says its news goes up to 3 October 2026 and suggests checking a news source for anything later.
- News sources: 18. Times of India, The Hindu (National, Business, Sport, Sci-Tech), Indian Express, NDTV, LiveMint, BBC (World, Business, Science, Technology), Al Jazeera, ESPNcricinfo, NASA, TechCrunch, The Verge and MIT Technology Review. Recent headlines were also added to training.
- Cleaner training data. Removed answers that repeat themselves, refusals, answers copied from earlier in a chat and leftover chat-export text. Greetings no longer trigger a self-introduction, and identity answers never credit another AI company.
- Passed the pre-publish check: repeated 8-turn conversations with the repeat penalty off and the app's thinking request on — no loops, no copied answers, no hidden reasoning, correct identity.
- Keep in mind: news outside the digest is not reliably known, and small models can mix up names, numbers and dates. Double-check anything important. The digest adds a few seconds before the first reply of a new chat.
2026-09-30 — First release. Built with the IRx pipeline's fixes: fine-tune merged into the full-precision base and quantized once with an importance matrix (no repeating/looping replies), thinking mode always off, identity built in. Passed the automated pre-publish check: repeated 8-turn conversations with the repeat penalty off and the app's thinking request on — no loops, no copied answers, no hidden reasoning, correct identity.
License
Apache 2.0. IRx-2 is a derivative fine-tuned model — full Apache 2.0 terms apply as with any work under this license.
- Downloads last month
- 411
4-bit