smeagle-4b — a small agentic helper you can run at home

smeagle is a compact (4B) agentic model — a terminal / software-engineering helper fine-tuned by Hob Forge from Qwen/Qwen3.5-4B-Base. It's built for the people big models leave behind: laptops and modest PCs, low RAM, no datacentre GPU. It calls tools, edits files, writes and checks code, and drives an agent loop — and it fits in a few GB.

It works methodically: read before you write, verify before you claim, report honestly. Small, focused, and unreasonably effective at the work it was trained for.

Why it's worth your VRAM

Two honest measurements, run against the raw Qwen3.5-4B base as the baseline (same data, same load — no cherry-picking):

  • −26.9% loss on the assistant tokens it was trained to produce (masked held-out cross-entropy: base 0.568 → smeagle 0.415). That's a large, real shift on the objective that matters.
  • 5 / 5 on an agentic task benchmark, zero tool-protocol failures — build-a-plugin (writes + syntax-checks
    • reads back valid Python), multi-tool chaining, file read/modify/write, multi-step calc with verification, and tool-based word counting. Scored by task completion on a real harness, not by proxy loss.

It's native tool-calling tuned — hand it tool schemas (Ollama tools, or an OpenAI-compatible /v1 endpoint) and it emits proper tool_calls. No prompt-wrestling to make it act.

What smeagle can do for you

Small, but genuinely agentic. Give it tools and it:

  • Uses tools correctly — picks the right tool, fills the arguments, no coaxing (native function-calling).
  • Chains steps — computes a value, writes it to a file, reads it back, uses it in the next step; holds the thread across a multi-step plan.
  • Works with files — creates, reads, modifies, and lists files; builds multi-file layouts.
  • Writes and verifies code — generates Python functions/modules and syntax-checks them before claiming done, the way a careful engineer would.
  • Recovers from errors — when a tool errors (a syntax bug, a bad path), it reads the message and fixes it instead of looping or giving up.
  • Extends its own runtime — it can write a working plugin for an agent harness (it built one for the very harness it was benchmarked in).
  • Respects constraints — when you need exact output, it gives you exact output, no chatter.

And with no tools at all, it's still a handy little coding/terminal assistant: write a function, explain a snippet, draft a shell command, reason through a small task.

Measured: 16/16 on our agentic task suite (scored by real task completion — tool chains, file ops, code-gen+verify, error-recovery, constrained output), and −27% loss vs the base model on its training objective.

Try this — getting started

ollama run hf.co/Hob-forge/smeagle-4b:Q8_0

Feel it out with a few prompts (no tools needed):

  • Code: Write a Python function is_prime(n), then walk me through why it's correct for 17 and 18.
  • Terminal: I want to find every .log file over 10MB and delete it. Give me the one-liner and explain each part.
  • Reason: A file has 12,000 lines and I need to process it in chunks of 500 without loading it all into memory. Sketch the approach.

Then give it tools (Ollama tools or an OpenAI-compatible /v1 endpoint) and let the agent out:

  • List the files in this project, read the main entry point, and tell me what it does.
  • Create three files a.txt, b.txt, c.txt — each containing its own name — then list them to confirm.
  • Write a factorial(n) function to factorial.py and syntax-check it before you tell me it's done.

The small-but-mighty part shows when it's driving — that's what it was trained for.

Choose your quant (does it run on YOUR machine?)

File Size Fits comfortably in Use it when
smeagle-4b-v0.1-Q4_K_M.gguf 2.6 GB ~4 GB RAM/VRAM smallest — older laptops, tight memory
smeagle-4b-v0.1-Q5_K_M.gguf 3.0 GB ~5 GB a little more headroom
smeagle-4b-v0.1-Q6_K.gguf 3.4 GB ~5–6 GB near-lossless, still small
smeagle-4b-v0.1-Q8_0.gguf 4.3 GB ~6 GB best quality; the recommended default

262K context. Runs on CPU alone, or a sliver of GPU.

Run it with llama.cpp

git clone https://github.com/ggml-org/llama.cpp.git && cd llama.cpp && cmake -B build && cmake --build build -j
./build/bin/llama-cli -hf Hob-forge/smeagle-4b:Q8_0 -p "list the files here, then tell me what this project is"

Give it your own system prompt and tools and it stays out of the way. Ask it "who are you" with no system prompt and it'll introduce itself as smeagle — the identity is injected only in that bare case, never over your system message or tool calls.

Honest limits — where it actually breaks

We pushed it with a hard trap-suite (10 tasks built specifically to make a strong 4B fail — 8-hop dependent chains, an RPN evaluator, byte-exact files, self-referential puzzles) and re-ran the flaky ones several times so we'd report reliable behaviour, not a lucky single draw. The genuine limits — double-check it on these:

  • Byte-exact / multi-line formatting. Told to write four exact lines with trailing spaces and no final newline, it collapsed them onto one line. It normalizes formatting — if you need bytes exactly, verify.
  • Operand-order-sensitive code. In a generated RPN evaluator it got the subtract/divide order wrong (a op b vs b op a) — the classic bug. Review order-sensitive logic it writes.
  • Long single-file running state. Overwriting one file with a running value across many steps, it loses the thread (0/3 in our reruns). Give it distinct filenames or fewer hops when precision matters.
  • Temperature-sensitive on one-shot precision. Some exact/self-referential tasks it gets right most of the time but not every time (e.g. "write a sentence stating its own word count" — ~2 in 3 at default temp 0.6). If you need it reliable in one shot, lower the temperature or let it retry.

What it does do well: the full agentic suite (16/16 — single-tool → multi-tool chains → file ops → code-gen+verify → error recovery → constrained output), plus long dependent tool-chains, building larger working modules with correct logic, and recognizing a tool's limits.

  • It's 4B — a frontier model beats it on hard long-horizon work; smeagle's job is to be the capable little one that fits, not to out-muscle models 100× its size. And it's a specialist (agentic/terminal/ SWE), not a generalist oracle.
  • Why v0.1 and not v0.2? We built a v0.2 and benchmarked both, in the open, on all 26 tasks with reruns. v0.2's apparent edge turned out to be temperature noise once repeated — no reliable agentic gain — and it regressed slightly on held-out SWE loss. So v0.1 stays the recommended release. That's the honest call, and we'll only ship a v0.2 when it clearly earns it.

License & rules — free for almost everyone

smeagle is released under the Hob Forge Community License v1.0 (LICENSE.md). In plain English:

  • Free to use, run, modify, quantize, and ship products with — commercially or not — for individuals, hobbyists, startups, research, schools, and small-to-mid-size companies. Run it on your laptop; build on it.
  • Attribution always — keep a visible "Hob Forge — smeagle" credit when you redistribute it.
  • The billboard clause — if you're a frontier AI lab or serve smeagle to 10M+ monthly users, you must publicly disclose that your product uses a Hob Forge fine-tune. If it helps you serve millions, say so.
  • Guardrails (non-negotiable): no child-sexualizing/endangering use; don't abliterate-and-relabel as Hob Forge; don't pass its output off as human-authored to deceive.

The base model (Qwen/Qwen3.5-4B-Base, Apache-2.0) keeps its Apache-2.0 terms, unaffected.

Provenance

Fine-tuned from Qwen/Qwen3.5-4B-Base by Hob Forge — a tiny sovereign AI lab shipping small, honest specialists for people who run AI on their own hardware. Not abliterated. Trained with anti-doom-loop and long-horizon reasoning data. Full catalogue: huggingface.co/Hob-forge.

Downloads last month
28
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Hob-forge/smeagle-4b

Finetuned
(133)
this model

Collection including Hob-forge/smeagle-4b