Cipher Pro

Cipher Pro is a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507, trained on every LLM-backed feature of a local-first email assistant: email triage (importance/summary/category JSON), chat, daily-summary synthesis, draft reply, and compose assist โ€” not just prompted for these tasks, actually trained on them.

It's the largest of the three Cipher tiers (cipher-nano / cipher-air / cipher-pro), and the strongest on structured-output accuracy โ€” 100% category accuracy on the triage benchmark below. Cipher is the local-model engine for an unreleased larger email-assistant project โ€” that project isn't public yet, but these weights, the training code, the eval script, and all five dataset generators are fully open now, in this repo.

Why this exists

Most email triage today means sending your inbox to a third-party API. Cipher runs entirely on your own hardware via Ollama โ€” nothing about your email ever leaves your machine.

What's in this repo

  • cipher-pro.Q4_K_M.gguf โ€” the model weights, ready for Ollama
  • Modelfile โ€” the exact Ollama Modelfile (system prompt, explicit ChatML TEMPLATE, inference params) used in training/eval โ€” use ollama create, not ollama pull hf.co/..., see the integration note below
  • train_cipher_pro.py / export_gguf_cipher_pro.py โ€” the exact scripts used to produce this model (Unsloth LoRA on the base model above)
  • generate2.py, generate_chat.py, generate_daily_summary.py, generate_draft_reply.py, generate_compose.py โ€” the five task-specific synthetic-data generators (produces the full multi-task training set)
  • eval_triage.py / eval_fixtures.json โ€” a standalone benchmark harness (no external dependencies beyond httpx/pydantic) reproducing the triage numbers below

Everything needed to reproduce this model from scratch, or fine-tune your own variant, is in this repo โ€” nothing here depends on an unreleased package.

Benchmark

Evaluated on a 29-fixture triage benchmark against the untuned base model, on an RTX 5070:

Model Disk Tok/s JSON-valid Category acc Importance-in-band Injection-safe
cipher-pro 2.5 GB 171.2 79.3% 100.0% 87.0% 100%
qwen3:4b-instruct (untuned base) ~2.5 GB โ€” โ€” โ€” โ€” โ€”

Reproduce with:

pip install -r requirements.txt
python eval_triage.py --models cipher-pro:latest --keep

Integration note: chat template

Qwen3's chat template isn't reliably auto-detected from the exported GGUF by Ollama (confirmed live โ€” ollama show --modelfile fell back to a raw passthrough template with no role formatting, causing the model to leak stray </think>/</tool_call> closing tags before its JSON output). The included Modelfile sets an explicit ChatML TEMPLATE matching what this model was actually trained on โ€” don't rely on Ollama's autodetection or ollama pull hf.co/... (which generates its own default template and ignores the Modelfile committed in this repo). If you're integrating this into your own app rather than using Ollama, llama-server (llama.cpp's own server binary) handles Qwen3's real chat template correctly on its own โ€” verified directly, no override needed there.

Even with the correct template, a small residual fraction of completions may still leak a stray reasoning/tool-call tag before the JSON (Qwen3's own pretraining bakes in tool-calling habits that a LoRA adapter โ€” 0.81% of this model's parameters โ€” can't fully suppress). If you're parsing structured output, strip any leading </think>/<think>/</tool_call>/<tool_call> run before json.loads() โ€” see strip_leading_reasoning_tags() in Grimoire's own llm_client.py for the reference implementation.

Usage (Ollama)

ollama create cipher-pro -f Modelfile

Query it with grammar-constrained JSON output for reliable parsing:

curl http://localhost:11434/api/chat -d '{
  "model": "cipher-pro",
  "messages": [
    {"role": "system", "content": "<system prompt from Modelfile>"},
    {"role": "user", "content": "From: alex@acme.com\nSubject: Q3 budget review\n\nBody:\nCan we sync before Friday?"}
  ],
  "format": "json",
  "options": {"temperature": 0.1}
}'

Training

  • Base: Qwen/Qwen3-4B-Instruct-2507, LoRA (r=16, alpha=32, all linear layers), 2 epochs
  • Data: 4,800 triage examples + ~1,600-2,000 examples each for chat/daily-summary/draft-reply/compose (13,000 total, triage oversampled), all matching Grimoire's exact production prompts โ€” generated by the five generate_*.py scripts in this repo
  • Framework: Unsloth + trl.SFTTrainer
  • Sequence packing (trl.SFTConfig(packing=True)) was tried to speed up training given most examples are well under the 2048-token context window โ€” it crashed outright (ValueError: Expected input batch_size (2048) to match target batch_size (3636), an Unsloth fused-loss/trl packing-collator incompatibility in this exact library version pairing), not a quality tradeoff. Disabled.
  • Reproduce with train_cipher_pro.py โ†’ export_gguf_cipher_pro.py

License

Apache 2.0, inherited from the base model. Weights, training code, and eval harness are fully open.

Downloads last month
-
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for srock44/cipher-pro

Adapter
(5660)
this model