Instructions to use ritiksuman/rudra with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ritiksuman/rudra with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ritiksuman/rudra") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ritiksuman/rudra") model = AutoModelForMultimodalLM.from_pretrained("ritiksuman/rudra", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ritiksuman/rudra with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ritiksuman/rudra:BF16 # Run inference directly in the terminal: llama cli -hf ritiksuman/rudra:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ritiksuman/rudra:BF16 # Run inference directly in the terminal: llama cli -hf ritiksuman/rudra:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ritiksuman/rudra:BF16 # Run inference directly in the terminal: ./llama-cli -hf ritiksuman/rudra:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ritiksuman/rudra:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ritiksuman/rudra:BF16
Use Docker
docker model run hf.co/ritiksuman/rudra:BF16
- LM Studio
- Jan
- vLLM
How to use ritiksuman/rudra with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ritiksuman/rudra" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ritiksuman/rudra", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ritiksuman/rudra:BF16
- SGLang
How to use ritiksuman/rudra with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ritiksuman/rudra" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ritiksuman/rudra", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ritiksuman/rudra" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ritiksuman/rudra", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use ritiksuman/rudra with Ollama:
ollama run hf.co/ritiksuman/rudra:BF16
- Unsloth Desktop
- Pi
How to use ritiksuman/rudra with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ritiksuman/rudra:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ritiksuman/rudra:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ritiksuman/rudra with Docker Model Runner:
docker model run hf.co/ritiksuman/rudra:BF16
- Lemonade
How to use ritiksuman/rudra with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ritiksuman/rudra:BF16
Run and chat with the model
lemonade run user.rudra-BF16
List all available models
lemonade list
- Hermes Agent
How to use ritiksuman/rudra with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ritiksuman/rudra:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ritiksuman/rudra:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ritiksuman/rudra with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ritiksuman/rudra:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ritiksuman/rudra:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
🦥 Rudra — 4B Coding Agent, Function-Calling & Reasoning LLM (GGUF + Transformers)
Rudra is a 4B coding-agent / reasoning language model built by Vyomma Intelligence. The brain behind Rudra is Ritik Suman from Vyomma Intelligence.
It is a fine-tune of unsloth/Qwen3.5-4B
with QLoRA on Unsloth, trained for coding, agentic tool use, and
reasoning, with an anti-hallucination instruction set.
Runs locally on CPU or a small GPU via Ollama / llama.cpp, or with Transformers. Great for a local coding assistant, function-calling experiments, and agent prototypes.
⚡ TL;DR
ollama pull ... # not hosted on Ollama registry yet
ollama create rudra -f Modelfile
ollama run rudra "Plan how to fix a failing pytest suite."
- 🧩 Function calling (JSON
<tool_call>) - 🛠️ Agentic loop: plan → inspect → edit → run → verify → summarize
- 🧠 Reasoning with
[VERIFY]+CONFIDENCE:markers - 🖥️ GGUF Q4_K_M (~2.8 GB) — CPU friendly
- 🔓 Apache-2.0 (base license)
🚀 Quickstart
Ollama
FROM ./Qwen3.5-4B.Q4_K_M.gguf
SYSTEM """Built by vyomma intelligence. Brain behind rudra is ritik suman from vyomma intelligence.
You are Rudra, a coding, agentic and reasoning assistant. Never fabricate facts or outputs; if unsure, say so; end with CONFIDENCE: <high|medium|low>."""
PARAMETER temperature 0.4
PARAMETER num_ctx 8192
ollama create rudra -f Modelfile
ollama run rudra "Write a Python LRU cache."
llama.cpp
llama-cli -m Qwen3.5-4B.Q4_K_M.gguf -p "Explain what a list comprehension does."
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("ritiksuman/rudra")
model = AutoModelForCausalLM.from_pretrained("ritiksuman/rudra", device_map="auto")
msgs = [{"role": "system", "content": "Built by vyomma intelligence. Brain behind rudra is ritik suman from vyomma intelligence."},
{"role": "user", "content": "Write a Python function to check for primes."}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")
print(tok.decode(model.generate(ids.to(model.device), max_new_tokens=256)[0], skip_special_tokens=True))
🧰 Tools (function calling)
| Tool | Purpose |
|---|---|
list_dir(path) |
List files/directories |
read_file(path) |
Read a file |
write_file(path, content) |
Create/overwrite a file |
edit_file(path, old, new) |
Replace an exact substring |
run_command(cmd) |
Run a shell command |
search_code(pattern) |
Search the codebase |
Calls are emitted as:
<tool_call>
{"name": "read_file", "arguments": {"path": "calc.py"}}
</tool_call>
📊 Evaluation (honest, reproducible)
Measured with the harness in the repo (HumanEval+ with real execution; BFCL-style function calling on a held-out xLAM slice; a private set including unanswerable prompts). Small samples — treat as directional.
| Metric | Base Qwen3.5-4B | Rudra (this model) |
|---|---|---|
| HumanEval+ pass@1 (n=15) | 40% | 20% |
| Function-calling name acc (n=30) | 96.7% | 56.7% |
| Answerable accuracy (n=6) | 6/6 | 6/6 |
| Decline rate on unanswerable (n=8) | 88% | 62% |
Overconfident (CONFIDENCE: high on unanswerable) |
0% | 50% |
⚠️ New Stable Fine‑tune (2026‑10‑10) – A new fine‑tune using the original rudra_agent_train.jsonl dataset (1.1k examples) is currently training on a local RTX 3050 (rank 16, seq 1024, GA 4, LR 2e‑5, 1 epoch). Preference optimization (ORPO) with the existing rudra_prefs.jsonl will follow after this run completes. Evaluation results will be updated here.
⚠️ Honest interpretation
This fine-tune currently does NOT beat its base model. On this evaluation it
regressed coding and function-calling, and its CONFIDENCE: tag is not
calibrated (it can say “high” when it is wrong). Part of the function-calling
gap is caused by the custom chat template's JSON tool format vs. the base
model's native format.
Recommendation: for production use today, prefer the base model, or re-train Rudra with the larger recipe below (more data, native tool format, higher rank) before relying on it.
🏋️ Training
| Setting | Value |
|---|---|
| Base | unsloth/Qwen3.5-4B |
| Method | QLoRA (4-bit), Unsloth |
| LoRA rank | 16 (alpha 16) |
| Epochs | 2 |
| LR | 5e-5 |
| Max seq | 1024 |
| Hardware | NVIDIA RTX 3050 Laptop (6 GB) |
| Data | coding-agent tool trajectories + CodeAlpaca + GSM8K + curated identity/anti-hallucination |
⚠️ Current Run (2026‑10‑10) – A new fine‑tune using the improved 3136‑example dataset (Magicoder, xLAM, Hermes, Orca‑Math, local agentic data) is in progress on a local RTX 3050 (rank 16, seq 1024, GA 4, LR 2e‑5, 1 epoch). Preference optimization (ORPO) with the existing rudra_prefs.jsonl will follow after this run completes.
🔬 Reproduce a stronger run
A ready-to-run Kaggle T4 notebook is included in the source project (~32k examples, rank 32, LR 2e-4, 4096 tokens, better datasets: Magicoder-OSS-Instruct, OpenCodeInstruct, xLAM, Hermes, Orca-Math). This is the recommended path to a model that actually beats base.
❓ FAQ
Is it free to run? Yes — GGUF Q4_K_M runs on CPU; no API needed.
Does it need a GPU? No; a small GPU helps.
Context length? Trained at 1024; base supports up to 40,960 — raise num_ctx
in Ollama on bigger hardware.
License? Apache-2.0 (inherited).
🙏 Credits
Fine-tune & tooling by Vyomma Intelligence. Brain behind Rudra: Ritik Suman. Base model: Qwen3.5-4B (Alibaba). Training: Unsloth + TRL.
- Downloads last month
- 270