Instructions to use OculusMindAI/OculusMind-ToolCall-8B-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="OculusMindAI/OculusMind-ToolCall-8B-v2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("OculusMindAI/OculusMind-ToolCall-8B-v2") model = AutoModelForCausalLM.from_pretrained("OculusMindAI/OculusMind-ToolCall-8B-v2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0 # Run inference directly in the terminal: llama cli -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0 # Run inference directly in the terminal: llama cli -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
Use Docker
docker model run hf.co/OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
- LM Studio
- Jan
- vLLM
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OculusMindAI/OculusMind-ToolCall-8B-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OculusMindAI/OculusMind-ToolCall-8B-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
- SGLang
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OculusMindAI/OculusMind-ToolCall-8B-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OculusMindAI/OculusMind-ToolCall-8B-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OculusMindAI/OculusMind-ToolCall-8B-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OculusMindAI/OculusMind-ToolCall-8B-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with Ollama:
ollama run hf.co/OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
- Unsloth Desktop
- Pi
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with Docker Model Runner:
docker model run hf.co/OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
- Lemonade
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
Run and chat with the model
lemonade run user.OculusMind-ToolCall-8B-v2-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use OculusMindAI/OculusMind-ToolCall-8B-v2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "OculusMindAI/OculusMind-ToolCall-8B-v2:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
OculusMind-ToolCall-8B-v2 ("Pico V2")
Uno, one of the two OculusMind.AI agents whose tool calls trained this model.
Pico V2 is an 8-billion-parameter tool-calling model that runs on a laptop and
picks the right tool, with valid arguments, on 95 % of the held-out production
turns it was built for — 9 points above its base. It is Mistral's
Ministral-3-8B-Instruct-2512 fine-tuned on production data from Uno and Vera, the
agents at OculusMind.AI — the tool calls working
assistants make for people all day — and selected
from a ten-seed sweep under a hard rule: a seed that ever looped was out,
whatever it scored. It is built for, and deploying to, the OculusMind.AI agents
at oculusmind.ai. It is the
second generation of OculusMind.AI's tool-calling fine-tune series; the
first was
OculusMindAI/OculusMind-ToolCall-8B-v1.
What it does well:
- Chooses the right tool, with valid arguments, first time. On OculusMind.AI's held-out production evals it agrees with production's tool choice on 95.45 % of tool-turns — +9.1 points over the base model on the same items.
- Does not loop. Degenerate repeated tool calls — the failure that makes small models unusable as agents — measured 0.00 % on every population for the published Q8_0 build, including a slice built to provoke them.
- Keeps its general ability. Ahead of the base on GSM8K (87.6 vs 79.7), IFEval (72.2 vs 68.7) and BFCL V4 (69.3 vs 65.2 over all 4,441 items; ahead on 15 of its 17 categories).
- Reads long context. Planted-fact recall and a correct tool call at every step from 2K to 250K tokens, identical to the base.
- Runs anywhere. Q8_0 GGUF (9.0 GB) for llama.cpp, Ollama and LM Studio; BF16 safetensors for transformers, vLLM and MLX; and the 29 MB LoRA adapter to apply to the base yourself. Apache-2.0.
Text only: this release does not include the base's vision encoder, so image input is not supported; a vision-capable build is planned for V3.
Headline results
On OculusMind.AI's production-derived tool-calling evals, Pico V2 agrees with production's tool choice more often than its untrained base on every population, with no degenerate tool-call loops — and it is ahead of the base on the general-capability tasks that were run. Same scorer, same cap (4,096), Q8_0 GGUF through llama.cpp; every figure is per artifact and reproducible from the files listed in §8 and the appendix.
| eval population | base (untrained) | Pico V2 | Δ vs base |
|---|---|---|---|
report-vera-sealed — held out, 66 tool-turns, the reported figure |
86.36 † | 95.45 | 🟢 +9.09 pts (+10.5 %) |
select-vera — 146 tool-turns, temperature 0 |
80.82 † | 96.58 | 🟢 +15.76 pts (+19.5 %) |
select-vera at production temperature 0.35 |
82.88 | 97.26 | 🟢 +14.38 pts (+17.4 %) |
| loop-provocation slice — 12 turns built to provoke loops | 83.33 | 100.0 | 🟢 +16.67 pts (+20.0 %) |
| degenerate loops on the provocation slice | 🔴 8.33 % | 0.00 % | 🟢 −8.33 pts — eliminated |
† Every figure in these tables is a direct collection on that population.
The base and Pico V2 figures on report-vera-sealed and select-vera were each
collected two or three times with identical results (zero spread; §1); the
production-temperature and loop-slice rows are single collections. Δ is in
percentage points, with the relative change in parentheses.
| general capability (vs base) | base | Pico V2 | Δ vs base |
|---|---|---|---|
| GSM8K exact-match, flexible-extract (1,319 items) | 79.68 | 87.57 | 🟢 +7.89 pts (+9.9 %) |
| IFEval instruction-level strict (541 prompts) | 68.71 | 72.18 | 🟢 +3.47 pts (+5.1 %) |
| BFCL V4, all 17 categories, 4,441 items (item-weighted accuracy) | 65.17 | 69.29 | 🟢 +4.12 pts (+6.3 %) |
| long-context smoke, planted-fact recall + tool call, 50K → 250K tokens | all ✓ | all ✓ (§9) | parity |
BFCL V4 by category — 17 categories, Pico V2 against the base
| category (n) | base | Pico V2 | Δ vs base |
|---|---|---|---|
simple_python (400) |
88.50 | 92.25 | 🟢 +3.75 pts (+4.2 %) |
simple_java (100) |
20.00 | 44.00 | 🟢 +24.00 pts (+120 %) |
simple_javascript (50) |
32.00 | 48.00 | 🟢 +16.00 pts (+50 %) |
multiple (200) |
88.50 | 93.00 | 🟢 +4.50 pts (+5.1 %) |
parallel (200) |
83.50 | 87.50 | 🟢 +4.00 pts (+4.8 %) |
parallel_multiple (200) |
76.00 | 81.50 | 🟢 +5.50 pts (+7.2 %) |
irrelevance (240) |
89.58 | 86.25 | 🔴 −3.33 pts (−3.7 %) |
live_simple (258) |
59.69 | 65.50 | 🟢 +5.81 pts (+9.7 %) |
live_multiple (1,053) |
68.19 | 69.99 | 🟢 +1.80 pts (+2.6 %) |
live_parallel (16) |
50.00 | 68.75 | 🟢 +18.75 pts (+37.5 %) |
live_parallel_multiple (24) |
62.50 | 75.00 | 🟢 +12.50 pts (+20.0 %) |
live_irrelevance (884) |
85.29 | 83.03 | 🔴 −2.26 pts (−2.6 %) |
live_relevance (16) |
81.25 | 87.50 | 🟢 +6.25 pts (+7.7 %) |
multi_turn_base (200) |
23.50 | 37.50 | 🟢 +14.00 pts (+59.6 %) |
multi_turn_miss_func (200) |
9.50 | 14.00 | 🟢 +4.50 pts (+47.4 %) |
multi_turn_miss_param (200) |
14.00 | 26.50 | 🟢 +12.50 pts (+89.3 %) |
multi_turn_long_context (200) |
18.50 | 35.00 | 🟢 +16.50 pts (+89.2 %) |
Pico V2 is ahead on 15 of the 17 categories and trails only on the two
irrelevance categories — the ones that reward declining to call a tool — by
2–3 points: it was trained on a corpus in which reaching for a tool is the norm
and restraint was not a target. Multi-turn categories were served at a
131,072-token window; two miss_func items and two long_context items ran
away across turns until that window was exhausted and are scored wrong, as BFCL
scores them (§3). The base column is the base's published unseeded BFCL run,
scored with the same harness revision (NOTICE).
The harness's own category-group accuracies, from its leaderboard CSV:
| BFCL V4 group | base | Pico V2 | Δ vs base |
|---|---|---|---|
| Non-Live AST (single, multiple, parallel, parallel-multiple) | 73.71 | 80.85 | 🟢 +7.14 pts (+9.7 %) |
| Live AST | 66.25 | 69.21 | 🟢 +2.96 pts (+4.5 %) |
| Multi-Turn | 16.38 | 28.25 | 🟢 +11.87 pts (+72.5 %) |
| Relevance detection | 81.25 | 87.50 | 🟢 +6.25 pts (+7.7 %) |
| Irrelevance detection | 87.44 | 84.64 | 🔴 −2.80 pts (−3.2 %) |
The "BFCL V4 overall" figure in the headline table is the item-weighted accuracy over these 17 categories (4,441 items), computed identically for both models. BFCL's own scoring script also prints a single "Overall Acc" number for each model (31.95 for Pico V2, 27.65 for the base); it is not shown in this card because its formula weights in the agentic web-search and memory categories, which neither model was run on and which it therefore counts as zero — a leaderboard convention, not a comparable measurement of these two models.
Read these as what they measure: agreement with OculusMind.AI's production tool
choice on turns held out from training — the behaviour this model was fine-tuned
to reproduce — not a general ranking. §1 states the claim precisely and what it is
not; §3 is the loop measurement; §5 the limitations. Comparisons with Claude Haiku
4.5 and 5.5 on the same populations are in
APPENDIX-claude-comparison.md.
1. Summary
The claim, precisely stated
On report-vera-sealed — 138 production-derived items, 66 of them tool-turns,
from conversations not in the training data, held back from the seed sweep that
chose this model and scored once — this model chooses the
same first tool production chose, with valid arguments and no degenerate repeat,
on 95.45 % of tool-turns, against the untrained base's 86.36 %.
This figure, and the loop-free claim throughout, describe the Q8_0 GGUF. The same weights at BF16 and Q4_K_M were measured separately and do loop; see §5.4.
That is +9.09 points on the reported population. It is measured at
temperature 0, cap 4096, on Q8_0 GGUF through llama.cpp c812c543f.
How the numbers were collected. Every figure is a direct collection on the population it is quoted for. The two main populations were collected at least twice for the base and for Pico V2, in separate runs — same artifact, same cap, same temperature:
| population | base | Pico V2 | Δ vs base |
|---|---|---|---|
report-vera-sealed (66 tool-turns) |
86.36 · 86.36 | 95.45 · 95.45 · 95.45 | 🟢 +9.09 pts (+10.5 %) |
select-vera (146 tool-turns) |
80.82 · 80.82 · 80.82 | 96.58 · 96.58 · 96.58 | 🟢 +15.76 pts (+19.5 %) |
Zero spread across repeats for both models on both populations — the base's two sealed repeats are byte-identical in every response field — so the headline deltas carry no run-to-run band on this hardware. What they do carry is the population dependence stated next.
What this claim is not
- Not +21.92, and not +15.76. Both figures are from
select-vera, the population used to select this model; the reported population is the held-out one. Onselect-verathe base scored 80.82 in each of three direct collections, so the selection-population gap is +15.76 — and it is the selection score, not the claim. - Not a broad capability claim. The general-capability evidence is three public suites — BFCL V4 (above), and IFEval and GSM8K through lm-evaluation-harness 0.4.13 on the full task sets (IFEval 541 prompts, GSM8K 1,319 items), chat endpoint with the chat template applied, measured 2026-10-02 — shown next.
| metric | base | Pico V2 (seed 46) | Δ vs base | seed 42 |
|---|---|---|---|---|
gsm8k exact_match, flexible-extract |
0.7968 | 0.8757 | 🟢 +0.0789 (+9.9 %) | 0.8666 |
gsm8k exact_match, strict-match |
0.4367 | 0.8347 | 🟢 +0.3980 (+91.1 %) | 0.8173 |
ifeval inst_level_strict |
0.6871 | 0.7218 | 🟢 +0.0347 (+5.1 %) | 0.7242 |
ifeval prompt_level_strict |
0.5638 | 0.6137 | 🟢 +0.0499 (+8.9 %) | 0.6174 |
ifeval inst_level_loose |
0.7494 | 0.7614 | 🟢 +0.0120 (+1.6 %) | 0.7650 |
ifeval prompt_level_loose |
0.6414 | 0.6636 | 🟢 +0.0222 (+3.5 %) | 0.6673 |
The two GSM8K rows answer different questions. Flexible-extract takes the
last number in the response and is the reasoning comparison: +7.9 points
over the base. Strict-match requires the #### <answer> line the task's
few-shot examples use; the base emits it less than half the time, so the +39.8
there is mostly answer formatting, not arithmetic, and is not a reasoning
claim.
This model is ahead of the base on all six, on tasks with no relationship to tool calling. That is evidence of no damage outside the trained domain — the classic failure of narrow fine-tuning — which is what it was run to detect. It is not evidence the model is broadly better: two task families is a narrow instrument, and the +7.9 flexible-extract gain is larger than a tool-calling fine-tune ought to produce and is not explained. Treat it as "no damage found, plus an unexplained gain."
Seed 42 edges this model on IFEval by 0.4 points — two prompts of 541, inside noise, and it does not disturb the selection.
- Not a converged model. This is checkpoint 150 of a run configured for 1,005 iterations, stopped deliberately.
- Not statistically separated from the runner-up on the reported population. Seed 42 also scores 95.45 — the same 63 of 66 tool-turns. See the appendix.
Why the multiple-choice benchmarks are absent
MMLU, ARC, HellaSwag and Winogrande are not scored by generating an answer. The harness supplies each candidate answer and asks the model to score it, and the harness could not run those probes against our serving setup — the scoring interface it needs is not available through the endpoint the model is served on. This was established by running it, not assumed.
So the general-capability evidence here is IFEval and GSM8K, both generation tasks. That is a narrow instrument and the card does not imply broader coverage.
Comparisons with Claude Haiku
Claude Haiku 4.5 and Claude Haiku 5.5 were run on the same populations with the same
scorer and cap. The results, how each was run, and what they do and do not show are
in APPENDIX-claude-comparison.md.
2. Selection, in one paragraph
Ten runs were made, seeds 42–51, identical in dataset, recipe, checkpoint and scoring population — the seed was the only difference. Five of the ten exhibit a tool-call loop defect, so the sweep was not a search for a better score but for a model that does not exhibit it. Looping is a gate: any non-zero rate excludes the seed, and survivors are ranked on score afterwards. The gate cost something real — one excluded seed scored identically to this model on the selection population.
This model was then scored once on report-vera-sealed, which informed no seed
or checkpoint choice. One earlier decision did see it: before the sweep, training
recipes were compared on a 404-turn screen that included these 138 items, so the
recipe — though not the model — was chosen with them in view. That held-out figure
is the claim in §1; the selection score is not quoted as a performance claim.
Full detail — the defect, the ten-seed table, the detector thresholds, the
selection-bias objection and what the numbers do not establish:
APPENDIX-looping.md.
3. Looping, measured where it actually matters
| population | temp | Pico V2 loop rate | base loop rate | Δ vs base | worst response, Pico V2 (base) |
|---|---|---|---|---|---|
select-vera (146 tool-turns) |
0 | 0.00 % | 🔴 0.68 % | 🟢 −0.68 pts | 13 calls (272) |
report-vera-sealed (66) |
0 | 0.00 % | 0.00 % | ±0 | 6 calls (9) |
| loop-provocation slice (12) | 0 | 0.00 % | 🔴 8.33 % | 🟢 −8.33 pts | 13 calls (272) |
select-vera (146) |
0.35 | 0.00 % | 0.00 % | ±0 | 19 calls (14) |
Zero looping turns on every population measured, including at production temperature.
The base loops, measured over the whole corpus. A dedicated scan of the full faithful item set — 1,965 items, 1,133 tool-turns, Q8_0, temperature 0, cap 4096, completed 2026-10-02 — puts the base rate at 0.09 %: one turn, which ran 111 tool calls and repeated a single tool name 108 times. The name-repeat histogram is otherwise flat (1,059 turns at 1, nothing between 11 and 108), so the defect is rare and catastrophic rather than a gradient.
That is the number to quote for the base. The 0.68 % figure elsewhere in this
repo is the same defect measured on select-vera's 146 tool-turns, where a single
looping turn is 0.68 % — a small-sample artifact of the same one event, not a
different finding.
So this is a defect removed, not one avoided — on the Q8_0 GGUF through llama.cpp, which is where every row above was measured.
Through other engines the picture is different, and not in the way the word
"loop" suggests. The BF16 safetensors served through mlx_lm (the path MLX
users take) returned at most one tool call per response on every one of 224
turns across three populations, so the loop metric cannot fire there — and
the scores are lower for the same reason: 89.39 on the sealed set, 86.3 at
production temperature, and 25.0 on the loop-provocation slice, whose turns
need several calls in sequence. That is a serving-path limitation (how
mlx_lm renders or parses Mistral's [TOOL_CALLS] list), not a property of the
weights, and it is unresolved as of this release: MLX users should expect
single-call responses until it is. In BFCL's agentic multi-turn harness the Q8_0
model itself showed the defect's other face: on 2 of 200 miss_func items it
kept calling tools turn after turn until a 131K context was exhausted, which
BFCL scores as failures and this card reports as such.
A caution that applies to any model selected this way: two seeds that scored
0.00 % on select-vera loop on report-vera-sealed, one of them on two turns of
124 and 119 calls. A gate verdict belongs to the population it was measured on.
This model was checked on four.
4. Training data
The training data is OculusMind.AI production data, and production data
re-run through additional models: 462 unique turns — 240 from Uno and 222 from
Vera — each repeated 3 times, giving 1,386 training rows, plus 50 validation rows.
No turn comes from a conversation that appears in the evaluation populations. Six
external corpora were screened and contributed zero rows. Every row carries an
independent reviewer verdict. Detail and attribution: NOTICE.
5. Limitations
Checkpoint 150 of 1,005 — a partial-training checkpoint.
Training data is from the same two agents as the evaluation — 240 Uno and 222 Vera turns. The reported population is Vera turns from conversations that are not in the training data, so the headline number measures generalisation to new conversations with a familiar agent, not transfer to a new one.
General-capability evidence is three public suites — BFCL V4, IFEval and GSM8K — not a broad benchmark battery. The multiple-choice probes (MMLU, ARC, HellaSwag, Winogrande) could not run through the chat endpoint here (§1).
Scores are per precision, and only Q8_0 is published. The same fused weights were also built as BF16 and Q4_K_M GGUFs and measured identically; the loop-free property holds only for Q8_0, and the other two are therefore not distributed (they exist, and are withheld for this reason):
artifact sealed Δ vs Q8_0 loops worst response loop slice loops Q8_0 (published) 95.45 — 0.00 % 6 calls 100.0 0.00 % BF16 (not published) 95.45 ±0.00 pts 🔴 1.52 % (+1.52 pts) 124 calls 100.0 0.00 % Q4_K_M (not published) 93.94 🔴 −1.51 pts (−1.6 %) 🔴 3.03 % (+3.03 pts) 124 calls† 🔴 91.67 (−8.33 pts) 🔴 8.33 % (+8.33 pts) † "loops" and "worst response" are defined over the 66 tool-turns of the sealed set. Read over all 138 items, Q4_K_M's worst response was 226 calls (
read_url× 225, to the generation cap) on a turn where production answered without a tool. Q8_0's worst over all 138 is still 6.BF16 scores identically to Q8_0 and loops anyway. Quoting the score alone would present them as equivalent while one exhibits the defect this model was selected for not having. BF16 is the higher precision, so the honest reading is that Q8_0's quantisation suppresses a tendency the full-precision weights retain — not that Q8_0 is better than its own weights. We have not established why.
Q4_K_M loops at 8.33 % on the provocation slice — the untrained base's rate. It would have been the smallest and most tempting download, and it is the one that gives the benefit back; that is why it is not offered.
The screen is not byte-deterministic — 14/20 identical sequences on re-collection — though the loop verdict and first-tool choice agreed 20/20.
Aggregate run-to-run variance is unbounded on current evidence: 20 items were re-collected, not 266.
An inherited upstream tokenizer issue —
transformerswarns the base checkpoint ships an incorrect regex pattern. SeeNOTICE§6.Scores are per serving path. Every headline figure is the Q8_0 GGUF through llama.cpp's completion endpoint with our own template rendering and tool parser. The same weights through other front ends, on the same held-out set, same cap and temperature:
serving path artifact sealed Δ vs llama.cpp loops tool calls per response llama.cpp /completion(our harness)Q8_0 GGUF 95.45 — 0.00 % up to 6 Ollama 0.35 ( Modelfile, its own template and parser)Q8_0 GGUF 93.94 🔴 −1.51 pts (−1.6 %) 0.00 % up to 6 mlx_lm.serverBF16 safetensors 89.39 🔴 −6.06 pts (−6.3 %) 0.00 % never more than 1 (§3) Each front end renders Mistral's
[TOOL_CALLS]grammar its own way, and the differences are the front ends', not the weights'. transformers and vLLM were not measured.
6. Intended use
Tool-calling assistants and agent workloads. Pico V2 is built for choosing a tool and constructing its arguments inside an agent loop — the production workload of the OculusMind.AI agents it was trained on and is deploying to — and it is validated as a general tool-calling assistant model on public suites: BFCL V4 (ahead of the base on 15 of its 17 categories, across single, parallel, live and multi-turn function calling), IFEval (instruction following, +3.5 points over the base) and GSM8K (reasoning, +7.9 points). Use it where an 8B model that reliably picks and calls tools is wanted: local agents, tool routers, structured-output assistants, and assistants that drive external APIs.
Where the evidence stops. It declines to call a tool slightly less often than the base (BFCL's two irrelevance categories, −2 to −3 points), so a deployment that depends on restraint should keep its own guard. Safety and refusal behaviour were not evaluated beyond what the base provides; multilingual use, long free-form writing and code generation were not evaluated. Text only — the base model's vision encoder is not part of this release and image inputs are not supported; a vision-capable build is planned for V3.
7. Licence and attribution
Apache-2.0, as is the base. The base's licence was verified live against the
upstream repository on 2026-09-29 and the pinned revision still matched upstream
main. See NOTICE for the base-model attribution and the statement
of modifications.
Lineage. Pico V2 is the second generation of the series that began with
OculusMindAI/OculusMind-ToolCall-8B-v1:
the same aim — an 8B model that calls tools the way our production agents do —
and the same base model at the same revision, now with a production-only
training corpus, a seed sweep under a loop gate, and held-out measurement.
If you build on Pico V2, the attribution travels with it. Under Apache-2.0
§4, any redistribution or derivative work — a further fine-tune, a merge, a
re-quantisation, a conversion to another format — must include this repository's
NOTICE file, which names OculusMind.AI and this repository as the
source, together with a copy of the licence and a prominent note of what was
changed. Please also cite it:
@misc{oculusmind2026picov2,
title = {Pico V2: an 8B tool-calling model fine-tuned on production agent traffic},
author = {{OculusMind.AI (Andrew Paris LLC)}},
year = {2026},
url = {https://huggingface.co/OculusMindAI/OculusMind-ToolCall-8B-v2}
}
8. Files
Pick the group for the software you run; you need one group, not all of them.
A. llama.cpp, Ollama, LM Studio → the Q8_0 GGUF. This is the measured
artifact: every figure in this card is from this file, through llama.cpp.
Ollama users: USAGE-OLLAMA.md walks through create, run,
a tool call and the tool-result round trip, every command executed against these
files on Ollama 0.35.1; ollama create oculusmind-ai-pico-v2:q8_0 -f Modelfile
builds it with the right stop tokens and an identity-only system prompt. To
install it in one line instead, with the same settings:
ollama run oculusmindai/toolcall-8b-v2 (from the Ollama catalog) or
ollama run hf.co/OculusMindAI/OculusMind-ToolCall-8B-v2 (from this repository,
which reads system and params below).
| file | bytes | sha256 |
|---|---|---|
gguf/Q8_0/model-q8_0.gguf |
9,028,871,552 | fe76fb07e5ca7351171bcd94838129514f8ad9bfc0a4e84ac6d9c247b33e1c56 |
Modelfile — Ollama definition for the GGUF above |
2,458 | af1404948cc6758687de0851efb0786f819a088139792e19b26681e471b62a1c |
system — the Modelfile's system prompt, for ollama run hf.co/... |
463 | 3172c95e8a831428f1a1085d1f141b045bce0769d5341c775b0a8e3f98305323 |
params — the Modelfile's stop tokens, temperature and context, for ollama run hf.co/... |
122 | 86cb31f9330417900cdcd032702beb1d0af4aeb24896872827c0b7d2f8fd6316 |
B. transformers, vLLM, MLX → the BF16 safetensors. The same merged weights at
full precision, in the standard Hugging Face layout (Ministral3ForCausalLM,
text only), with the tokenizer and chat template. Loaded from this repository,
they produce a correct structured tool call through transformers 5.19 and
through mlx_lm 0.32; neither path was scored on the evals above, and vLLM was
not run. Through MLX they return at most one tool call per response (§3, §5).
| file | bytes | sha256 |
|---|---|---|
model-00001-of-00004.safetensors |
5,318,535,144 | 627cbaa0706299fd1e695a9151da514ac05dd34bb978137c24fb2aff11e8b990 |
model-00002-of-00004.safetensors |
5,352,157,840 | 97a1d4852fe20c89411cec0610521aa52ae0c2d066f8b6935e45b3429b4d31c1 |
model-00003-of-00004.safetensors |
5,234,708,880 | da958797386cfdb2281a908d1c0d4caa5938d44a337f31a3e4f70c32a4209c02 |
model-00004-of-00004.safetensors |
1,073,741,952 | bf02e982fd7a72202ab02eb837a9f7c3cddea4e87db6d0def674689de38fba11 |
model.safetensors.index.json |
25,472 | 047bcb4995569e1d935ad5710cc76a237981987a64e71fdfa2cd17d6856e30ba |
config.json |
844 | 89573a75764b477492b4a789dc665c36c0125f16c2c1bf0ea7d29f695b2f7f81 |
generation_config.json |
131 | e0923390059f84a9180b00e5501778acc45ea9856cd7f2fd68208b360927c677 |
tokenizer.json |
17,078,128 | d5f6046775b112f0e2d456ee9dba450684ab964fe5c4e231599bdc6773028135 |
tokenizer_config.json |
198,094 | f59f7294e4f26383d0ea93840fe21cf197784be0842a8301a0343e8c34ed0d6d |
special_tokens_map.json |
21,436 | 8233d50ad79471ccb483a869703b940f5db1763db7ffd7c0134e0cec1a00a56e |
chat_template.jinja |
11,912 | 74eeb55fd3341286ec3fd44e902b7120721acc81cd394e96b431f85e93a1ea56 |
C. The LoRA adapter → apply the fine-tune to the base yourself. Rank 8, on
the attention projections only, in mlx_lm's adapter format. Fused into
mistralai/Ministral-3-8B-Instruct-2512-BF16 at revision f6fae979… it
reproduces the BF16 safetensors above tensor for tensor; applied at serve time
through mlx_lm (--adapter-path, or the adapters field of a server request;
mlx_lm.server before 0.32.0 ignores its --adapter-path flag and serves the
plain base, so on those versions use the request field) it scores 89.39 on
the held-out set and 25.0 on the loop-provocation slice — the merged weights' figures on that path to the hundredth, with
identical answer text on 136 of 138 held-out items and identical tool calls on
137 — against 80.3 and 16.67 for the base alone there. The fine-tune is the
adapter; the weights are one way of carrying it.
| file | bytes | sha256 |
|---|---|---|
adapter/adapters.safetensors |
29,000,642 | b3e2672fc0781ec8c5c9f604ed09fd0cb6958b75420ed6fe3e8aee0816ee4d94 |
adapter/adapter_config.json |
1,208 | 8ee3dff9c1e3ef47313cfdb7c7b973edee3b695bf5c2afc705e033afc0f51b7d |
D. Documentation. README.md (this card), USAGE-OLLAMA.md (verified
Ollama walkthrough), NOTICE (licence, attribution and the statement of
modifications — derivative works must carry it, §7), LICENSE (Apache-2.0),
APPENDIX-looping.md (the loop measurement and the seed sweep),
APPENDIX-claude-comparison.md (Claude Haiku 4.5 and 5.5 on the same populations,
with its charts in assets/),
CHECKSUMS.sha256 (every file above), and assets/uno.png (the portrait above;
OculusMind.AI's own artwork, 1024 × 1024).
Every sha256 above was recorded by the build that produced the file, and
model-q8_0.gguf was re-hashed from disk on 2026-10-05 and matched — a
manifest that agrees with itself proves nothing about the bytes on disk.
There is deliberately no BF16 or Q4_K_M download. Both were built from the same weights and measured (§5.4); both loop on the sealed set where Q8_0 does not, and the decision was to publish no artifact that exhibits the defect this model was selected against. Quantising the published GGUF further yourself will not reproduce the Q8_0 numbers — §5.4 is the measurement of exactly that.
9. Context window — smoke-tested to 250K tokens
The GGUF declares a 262,144-token window (YaRN ×16 over the base's original
16,384). The fine-tune trained at ≤ 20,480 tokens and every score above was
measured at a 65,536-token serving context, so the window was tested separately
rather than inferred from metadata. On 2026-10-06, through llama.cpp
c812c543f at --ctx-size 262144 with a q8_0 KV cache, a deterministic
synthetic document with three facts planted at 10 / 50 / 90 % depth was sent at
each size with a three-tool schema attached, asking for each fact and for one
web_search call about the mid-depth fact:
| document tokens | facts recalled (depth .1 / .5 / .9) | tool call formed | time to first token (prefill, cold) | response once the cache is populated: fact recall · tool call |
|---|---|---|---|---|
| 2,048 | 3 / 3 | yes | 5 s | 0.2 s · 0.4 s |
| 51,200 | 3 / 3 | yes | 130 s | 0.5 s · 1.5 s |
| 102,400 | 3 / 3 | yes | 459 s | 0.8 s · 1.7 s |
| 153,600 | 3 / 3 | yes | 988 s | 1.2 s · 2.5 s |
| 204,800 | 3 / 3 | yes | 1,712 s (28.5 min) | 1.6 s · 4.1 s |
| 250,880 | 3 / 3 | yes | 2,559 s (42.7 min) | 2.2 s · 5.1 s |
The prefill column is the one-time cost of reading the document into the KV cache; every further question against the same document — the second and third fact probes and the tool call — answered in the last column's times, because llama.cpp reuses the cached prefix. The untrained base produced the identical table to within a tenth of a second. This is a smoke test, not a long-context benchmark: it shows the window serves and the model reads from deep inside it and still forms a correct tool call; it does not rank models or measure degradation with length. Two practical facts follow from it: prefill cost grows roughly quadratically with context on this engine (half an hour to first token for a 200K prompt on an Apple M3 Ultra at Q8_0), and a 262,144 window costs about 32 GB of memory here including a ~18 GB KV cache.
- Downloads last month
- 321
Model tree for OculusMindAI/OculusMind-ToolCall-8B-v2
Base model
mistralai/Ministral-3-8B-Base-2512