🦥 Rudra — 4B Coding Agent, Function-Calling & Reasoning LLM (GGUF + Transformers)

Rudra is a 4B coding-agent / reasoning language model built by Vyomma Intelligence. The brain behind Rudra is Ritik Suman from Vyomma Intelligence.

It is a fine-tune of unsloth/Qwen3.5-4B with QLoRA on Unsloth, trained for coding, agentic tool use, and reasoning, with an anti-hallucination instruction set.

Runs locally on CPU or a small GPU via Ollama / llama.cpp, or with Transformers. Great for a local coding assistant, function-calling experiments, and agent prototypes.


⚡ TL;DR

ollama pull ...           # not hosted on Ollama registry yet
ollama create rudra -f Modelfile
ollama run rudra "Plan how to fix a failing pytest suite."
  • 🧩 Function calling (JSON <tool_call>)
  • 🛠️ Agentic loop: plan → inspect → edit → run → verify → summarize
  • 🧠 Reasoning with [VERIFY] + CONFIDENCE: markers
  • 🖥️ GGUF Q4_K_M (~2.8 GB) — CPU friendly
  • 🔓 Apache-2.0 (base license)

🚀 Quickstart

Ollama

FROM ./Qwen3.5-4B.Q4_K_M.gguf
SYSTEM """Built by vyomma intelligence. Brain behind rudra is ritik suman from vyomma intelligence.
You are Rudra, a coding, agentic and reasoning assistant. Never fabricate facts or outputs; if unsure, say so; end with CONFIDENCE: <high|medium|low>."""
PARAMETER temperature 0.4
PARAMETER num_ctx 8192
ollama create rudra -f Modelfile
ollama run rudra "Write a Python LRU cache."

llama.cpp

llama-cli -m Qwen3.5-4B.Q4_K_M.gguf -p "Explain what a list comprehension does."

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("ritiksuman/rudra")
model = AutoModelForCausalLM.from_pretrained("ritiksuman/rudra", device_map="auto")
msgs = [{"role": "system", "content": "Built by vyomma intelligence. Brain behind rudra is ritik suman from vyomma intelligence."},
        {"role": "user", "content": "Write a Python function to check for primes."}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")
print(tok.decode(model.generate(ids.to(model.device), max_new_tokens=256)[0], skip_special_tokens=True))

🧰 Tools (function calling)

Tool Purpose
list_dir(path) List files/directories
read_file(path) Read a file
write_file(path, content) Create/overwrite a file
edit_file(path, old, new) Replace an exact substring
run_command(cmd) Run a shell command
search_code(pattern) Search the codebase

Calls are emitted as:

<tool_call>
{"name": "read_file", "arguments": {"path": "calc.py"}}
</tool_call>

📊 Evaluation (honest, reproducible)

Measured with the harness in the repo (HumanEval+ with real execution; BFCL-style function calling on a held-out xLAM slice; a private set including unanswerable prompts). Small samples — treat as directional.

Metric Base Qwen3.5-4B Rudra (this model)
HumanEval+ pass@1 (n=15) 40% 20%
Function-calling name acc (n=30) 96.7% 56.7%
Answerable accuracy (n=6) 6/6 6/6
Decline rate on unanswerable (n=8) 88% 62%
Overconfident (CONFIDENCE: high on unanswerable) 0% 50%

⚠️ New Stable Fine‑tune (2026‑10‑10) – A new fine‑tune using the original rudra_agent_train.jsonl dataset (1.1k examples) is currently training on a local RTX 3050 (rank 16, seq 1024, GA 4, LR 2e‑5, 1 epoch). Preference optimization (ORPO) with the existing rudra_prefs.jsonl will follow after this run completes. Evaluation results will be updated here.

⚠️ Honest interpretation

This fine-tune currently does NOT beat its base model. On this evaluation it regressed coding and function-calling, and its CONFIDENCE: tag is not calibrated (it can say “high” when it is wrong). Part of the function-calling gap is caused by the custom chat template's JSON tool format vs. the base model's native format.

Recommendation: for production use today, prefer the base model, or re-train Rudra with the larger recipe below (more data, native tool format, higher rank) before relying on it.


🏋️ Training

Setting Value
Base unsloth/Qwen3.5-4B
Method QLoRA (4-bit), Unsloth
LoRA rank 16 (alpha 16)
Epochs 2
LR 5e-5
Max seq 1024
Hardware NVIDIA RTX 3050 Laptop (6 GB)
Data coding-agent tool trajectories + CodeAlpaca + GSM8K + curated identity/anti-hallucination

⚠️ Current Run (2026‑10‑10) – A new fine‑tune using the improved 3136‑example dataset (Magicoder, xLAM, Hermes, Orca‑Math, local agentic data) is in progress on a local RTX 3050 (rank 16, seq 1024, GA 4, LR 2e‑5, 1 epoch). Preference optimization (ORPO) with the existing rudra_prefs.jsonl will follow after this run completes.

🔬 Reproduce a stronger run

A ready-to-run Kaggle T4 notebook is included in the source project (~32k examples, rank 32, LR 2e-4, 4096 tokens, better datasets: Magicoder-OSS-Instruct, OpenCodeInstruct, xLAM, Hermes, Orca-Math). This is the recommended path to a model that actually beats base.


❓ FAQ

Is it free to run? Yes — GGUF Q4_K_M runs on CPU; no API needed. Does it need a GPU? No; a small GPU helps. Context length? Trained at 1024; base supports up to 40,960 — raise num_ctx in Ollama on bigger hardware. License? Apache-2.0 (inherited).

🙏 Credits

Fine-tune & tooling by Vyomma Intelligence. Brain behind Rudra: Ritik Suman. Base model: Qwen3.5-4B (Alibaba). Training: Unsloth + TRL.

Downloads last month
270
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ritiksuman/rudra

Finetuned
Qwen/Qwen3.5-4B
Adapter
(209)
this model