OculusMind-ToolCall-8B-v2 ("Pico V2")

Uno, the OculusMind.AI agent whose production tool-calling traffic this model was trained on

Uno, one of the two OculusMind.AI agents whose tool calls trained this model.

Pico V2 is an 8-billion-parameter tool-calling model that runs on a laptop and picks the right tool, with valid arguments, on 95 % of the held-out production turns it was built for — 9 points above its base. It is Mistral's Ministral-3-8B-Instruct-2512 fine-tuned on production data from Uno and Vera, the agents at OculusMind.AI — the tool calls working assistants make for people all day — and selected from a ten-seed sweep under a hard rule: a seed that ever looped was out, whatever it scored. It is built for, and deploying to, the OculusMind.AI agents at oculusmind.ai. It is the second generation of OculusMind.AI's tool-calling fine-tune series; the first was OculusMindAI/OculusMind-ToolCall-8B-v1.

What it does well:

  • Chooses the right tool, with valid arguments, first time. On OculusMind.AI's held-out production evals it agrees with production's tool choice on 95.45 % of tool-turns — +9.1 points over the base model on the same items.
  • Does not loop. Degenerate repeated tool calls — the failure that makes small models unusable as agents — measured 0.00 % on every population for the published Q8_0 build, including a slice built to provoke them.
  • Keeps its general ability. Ahead of the base on GSM8K (87.6 vs 79.7), IFEval (72.2 vs 68.7) and BFCL V4 (69.3 vs 65.2 over all 4,441 items; ahead on 15 of its 17 categories).
  • Reads long context. Planted-fact recall and a correct tool call at every step from 2K to 250K tokens, identical to the base.
  • Runs anywhere. Q8_0 GGUF (9.0 GB) for llama.cpp, Ollama and LM Studio; BF16 safetensors for transformers, vLLM and MLX; and the 29 MB LoRA adapter to apply to the base yourself. Apache-2.0.

Text only: this release does not include the base's vision encoder, so image input is not supported; a vision-capable build is planned for V3.

Headline results

On OculusMind.AI's production-derived tool-calling evals, Pico V2 agrees with production's tool choice more often than its untrained base on every population, with no degenerate tool-call loops — and it is ahead of the base on the general-capability tasks that were run. Same scorer, same cap (4,096), Q8_0 GGUF through llama.cpp; every figure is per artifact and reproducible from the files listed in §8 and the appendix.

eval population base (untrained) Pico V2 Δ vs base
report-vera-sealed — held out, 66 tool-turns, the reported figure 86.36 † 95.45 🟢 +9.09 pts (+10.5 %)
select-vera — 146 tool-turns, temperature 0 80.82 † 96.58 🟢 +15.76 pts (+19.5 %)
select-vera at production temperature 0.35 82.88 97.26 🟢 +14.38 pts (+17.4 %)
loop-provocation slice — 12 turns built to provoke loops 83.33 100.0 🟢 +16.67 pts (+20.0 %)
degenerate loops on the provocation slice 🔴 8.33 % 0.00 % 🟢 −8.33 pts — eliminated

† Every figure in these tables is a direct collection on that population. The base and Pico V2 figures on report-vera-sealed and select-vera were each collected two or three times with identical results (zero spread; §1); the production-temperature and loop-slice rows are single collections. Δ is in percentage points, with the relative change in parentheses.

general capability (vs base) base Pico V2 Δ vs base
GSM8K exact-match, flexible-extract (1,319 items) 79.68 87.57 🟢 +7.89 pts (+9.9 %)
IFEval instruction-level strict (541 prompts) 68.71 72.18 🟢 +3.47 pts (+5.1 %)
BFCL V4, all 17 categories, 4,441 items (item-weighted accuracy) 65.17 69.29 🟢 +4.12 pts (+6.3 %)
long-context smoke, planted-fact recall + tool call, 50K → 250K tokens all ✓ all ✓ (§9) parity

BFCL V4 by category — 17 categories, Pico V2 against the base

category (n) base Pico V2 Δ vs base
simple_python (400) 88.50 92.25 🟢 +3.75 pts (+4.2 %)
simple_java (100) 20.00 44.00 🟢 +24.00 pts (+120 %)
simple_javascript (50) 32.00 48.00 🟢 +16.00 pts (+50 %)
multiple (200) 88.50 93.00 🟢 +4.50 pts (+5.1 %)
parallel (200) 83.50 87.50 🟢 +4.00 pts (+4.8 %)
parallel_multiple (200) 76.00 81.50 🟢 +5.50 pts (+7.2 %)
irrelevance (240) 89.58 86.25 🔴 −3.33 pts (−3.7 %)
live_simple (258) 59.69 65.50 🟢 +5.81 pts (+9.7 %)
live_multiple (1,053) 68.19 69.99 🟢 +1.80 pts (+2.6 %)
live_parallel (16) 50.00 68.75 🟢 +18.75 pts (+37.5 %)
live_parallel_multiple (24) 62.50 75.00 🟢 +12.50 pts (+20.0 %)
live_irrelevance (884) 85.29 83.03 🔴 −2.26 pts (−2.6 %)
live_relevance (16) 81.25 87.50 🟢 +6.25 pts (+7.7 %)
multi_turn_base (200) 23.50 37.50 🟢 +14.00 pts (+59.6 %)
multi_turn_miss_func (200) 9.50 14.00 🟢 +4.50 pts (+47.4 %)
multi_turn_miss_param (200) 14.00 26.50 🟢 +12.50 pts (+89.3 %)
multi_turn_long_context (200) 18.50 35.00 🟢 +16.50 pts (+89.2 %)

Pico V2 is ahead on 15 of the 17 categories and trails only on the two irrelevance categories — the ones that reward declining to call a tool — by 2–3 points: it was trained on a corpus in which reaching for a tool is the norm and restraint was not a target. Multi-turn categories were served at a 131,072-token window; two miss_func items and two long_context items ran away across turns until that window was exhausted and are scored wrong, as BFCL scores them (§3). The base column is the base's published unseeded BFCL run, scored with the same harness revision (NOTICE).

The harness's own category-group accuracies, from its leaderboard CSV:

BFCL V4 group base Pico V2 Δ vs base
Non-Live AST (single, multiple, parallel, parallel-multiple) 73.71 80.85 🟢 +7.14 pts (+9.7 %)
Live AST 66.25 69.21 🟢 +2.96 pts (+4.5 %)
Multi-Turn 16.38 28.25 🟢 +11.87 pts (+72.5 %)
Relevance detection 81.25 87.50 🟢 +6.25 pts (+7.7 %)
Irrelevance detection 87.44 84.64 🔴 −2.80 pts (−3.2 %)

The "BFCL V4 overall" figure in the headline table is the item-weighted accuracy over these 17 categories (4,441 items), computed identically for both models. BFCL's own scoring script also prints a single "Overall Acc" number for each model (31.95 for Pico V2, 27.65 for the base); it is not shown in this card because its formula weights in the agentic web-search and memory categories, which neither model was run on and which it therefore counts as zero — a leaderboard convention, not a comparable measurement of these two models.

Read these as what they measure: agreement with OculusMind.AI's production tool choice on turns held out from training — the behaviour this model was fine-tuned to reproduce — not a general ranking. §1 states the claim precisely and what it is not; §3 is the loop measurement; §5 the limitations. Comparisons with Claude Haiku 4.5 and 5.5 on the same populations are in APPENDIX-claude-comparison.md.

1. Summary

The claim, precisely stated

On report-vera-sealed — 138 production-derived items, 66 of them tool-turns, from conversations not in the training data, held back from the seed sweep that chose this model and scored once — this model chooses the same first tool production chose, with valid arguments and no degenerate repeat, on 95.45 % of tool-turns, against the untrained base's 86.36 %.

This figure, and the loop-free claim throughout, describe the Q8_0 GGUF. The same weights at BF16 and Q4_K_M were measured separately and do loop; see §5.4.

That is +9.09 points on the reported population. It is measured at temperature 0, cap 4096, on Q8_0 GGUF through llama.cpp c812c543f.

How the numbers were collected. Every figure is a direct collection on the population it is quoted for. The two main populations were collected at least twice for the base and for Pico V2, in separate runs — same artifact, same cap, same temperature:

population base Pico V2 Δ vs base
report-vera-sealed (66 tool-turns) 86.36 · 86.36 95.45 · 95.45 · 95.45 🟢 +9.09 pts (+10.5 %)
select-vera (146 tool-turns) 80.82 · 80.82 · 80.82 96.58 · 96.58 · 96.58 🟢 +15.76 pts (+19.5 %)

Zero spread across repeats for both models on both populations — the base's two sealed repeats are byte-identical in every response field — so the headline deltas carry no run-to-run band on this hardware. What they do carry is the population dependence stated next.

What this claim is not

  • Not +21.92, and not +15.76. Both figures are from select-vera, the population used to select this model; the reported population is the held-out one. On select-vera the base scored 80.82 in each of three direct collections, so the selection-population gap is +15.76 — and it is the selection score, not the claim.
  • Not a broad capability claim. The general-capability evidence is three public suites — BFCL V4 (above), and IFEval and GSM8K through lm-evaluation-harness 0.4.13 on the full task sets (IFEval 541 prompts, GSM8K 1,319 items), chat endpoint with the chat template applied, measured 2026-10-02 — shown next.
metric base Pico V2 (seed 46) Δ vs base seed 42
gsm8k exact_match, flexible-extract 0.7968 0.8757 🟢 +0.0789 (+9.9 %) 0.8666
gsm8k exact_match, strict-match 0.4367 0.8347 🟢 +0.3980 (+91.1 %) 0.8173
ifeval inst_level_strict 0.6871 0.7218 🟢 +0.0347 (+5.1 %) 0.7242
ifeval prompt_level_strict 0.5638 0.6137 🟢 +0.0499 (+8.9 %) 0.6174
ifeval inst_level_loose 0.7494 0.7614 🟢 +0.0120 (+1.6 %) 0.7650
ifeval prompt_level_loose 0.6414 0.6636 🟢 +0.0222 (+3.5 %) 0.6673

The two GSM8K rows answer different questions. Flexible-extract takes the last number in the response and is the reasoning comparison: +7.9 points over the base. Strict-match requires the #### <answer> line the task's few-shot examples use; the base emits it less than half the time, so the +39.8 there is mostly answer formatting, not arithmetic, and is not a reasoning claim.

This model is ahead of the base on all six, on tasks with no relationship to tool calling. That is evidence of no damage outside the trained domain — the classic failure of narrow fine-tuning — which is what it was run to detect. It is not evidence the model is broadly better: two task families is a narrow instrument, and the +7.9 flexible-extract gain is larger than a tool-calling fine-tune ought to produce and is not explained. Treat it as "no damage found, plus an unexplained gain."

Seed 42 edges this model on IFEval by 0.4 points — two prompts of 541, inside noise, and it does not disturb the selection.

  • Not a converged model. This is checkpoint 150 of a run configured for 1,005 iterations, stopped deliberately.
  • Not statistically separated from the runner-up on the reported population. Seed 42 also scores 95.45 — the same 63 of 66 tool-turns. See the appendix.

Why the multiple-choice benchmarks are absent

MMLU, ARC, HellaSwag and Winogrande are not scored by generating an answer. The harness supplies each candidate answer and asks the model to score it, and the harness could not run those probes against our serving setup — the scoring interface it needs is not available through the endpoint the model is served on. This was established by running it, not assumed.

So the general-capability evidence here is IFEval and GSM8K, both generation tasks. That is a narrow instrument and the card does not imply broader coverage.

Comparisons with Claude Haiku

Claude Haiku 4.5 and Claude Haiku 5.5 were run on the same populations with the same scorer and cap. The results, how each was run, and what they do and do not show are in APPENDIX-claude-comparison.md.

2. Selection, in one paragraph

Ten runs were made, seeds 42–51, identical in dataset, recipe, checkpoint and scoring population — the seed was the only difference. Five of the ten exhibit a tool-call loop defect, so the sweep was not a search for a better score but for a model that does not exhibit it. Looping is a gate: any non-zero rate excludes the seed, and survivors are ranked on score afterwards. The gate cost something real — one excluded seed scored identically to this model on the selection population.

This model was then scored once on report-vera-sealed, which informed no seed or checkpoint choice. One earlier decision did see it: before the sweep, training recipes were compared on a 404-turn screen that included these 138 items, so the recipe — though not the model — was chosen with them in view. That held-out figure is the claim in §1; the selection score is not quoted as a performance claim.

Full detail — the defect, the ten-seed table, the detector thresholds, the selection-bias objection and what the numbers do not establish: APPENDIX-looping.md.

3. Looping, measured where it actually matters

population temp Pico V2 loop rate base loop rate Δ vs base worst response, Pico V2 (base)
select-vera (146 tool-turns) 0 0.00 % 🔴 0.68 % 🟢 −0.68 pts 13 calls (272)
report-vera-sealed (66) 0 0.00 % 0.00 % ±0 6 calls (9)
loop-provocation slice (12) 0 0.00 % 🔴 8.33 % 🟢 −8.33 pts 13 calls (272)
select-vera (146) 0.35 0.00 % 0.00 % ±0 19 calls (14)

Zero looping turns on every population measured, including at production temperature.

The base loops, measured over the whole corpus. A dedicated scan of the full faithful item set — 1,965 items, 1,133 tool-turns, Q8_0, temperature 0, cap 4096, completed 2026-10-02 — puts the base rate at 0.09 %: one turn, which ran 111 tool calls and repeated a single tool name 108 times. The name-repeat histogram is otherwise flat (1,059 turns at 1, nothing between 11 and 108), so the defect is rare and catastrophic rather than a gradient.

That is the number to quote for the base. The 0.68 % figure elsewhere in this repo is the same defect measured on select-vera's 146 tool-turns, where a single looping turn is 0.68 % — a small-sample artifact of the same one event, not a different finding.

So this is a defect removed, not one avoided — on the Q8_0 GGUF through llama.cpp, which is where every row above was measured.

Through other engines the picture is different, and not in the way the word "loop" suggests. The BF16 safetensors served through mlx_lm (the path MLX users take) returned at most one tool call per response on every one of 224 turns across three populations, so the loop metric cannot fire there — and the scores are lower for the same reason: 89.39 on the sealed set, 86.3 at production temperature, and 25.0 on the loop-provocation slice, whose turns need several calls in sequence. That is a serving-path limitation (how mlx_lm renders or parses Mistral's [TOOL_CALLS] list), not a property of the weights, and it is unresolved as of this release: MLX users should expect single-call responses until it is. In BFCL's agentic multi-turn harness the Q8_0 model itself showed the defect's other face: on 2 of 200 miss_func items it kept calling tools turn after turn until a 131K context was exhausted, which BFCL scores as failures and this card reports as such.

A caution that applies to any model selected this way: two seeds that scored 0.00 % on select-vera loop on report-vera-sealed, one of them on two turns of 124 and 119 calls. A gate verdict belongs to the population it was measured on. This model was checked on four.

4. Training data

The training data is OculusMind.AI production data, and production data re-run through additional models: 462 unique turns — 240 from Uno and 222 from Vera — each repeated 3 times, giving 1,386 training rows, plus 50 validation rows. No turn comes from a conversation that appears in the evaluation populations. Six external corpora were screened and contributed zero rows. Every row carries an independent reviewer verdict. Detail and attribution: NOTICE.

5. Limitations

  1. Checkpoint 150 of 1,005 — a partial-training checkpoint.

  2. Training data is from the same two agents as the evaluation — 240 Uno and 222 Vera turns. The reported population is Vera turns from conversations that are not in the training data, so the headline number measures generalisation to new conversations with a familiar agent, not transfer to a new one.

  3. General-capability evidence is three public suites — BFCL V4, IFEval and GSM8K — not a broad benchmark battery. The multiple-choice probes (MMLU, ARC, HellaSwag, Winogrande) could not run through the chat endpoint here (§1).

  4. Scores are per precision, and only Q8_0 is published. The same fused weights were also built as BF16 and Q4_K_M GGUFs and measured identically; the loop-free property holds only for Q8_0, and the other two are therefore not distributed (they exist, and are withheld for this reason):

    artifact sealed Δ vs Q8_0 loops worst response loop slice loops
    Q8_0 (published) 95.45 — 0.00 % 6 calls 100.0 0.00 %
    BF16 (not published) 95.45 ±0.00 pts 🔴 1.52 % (+1.52 pts) 124 calls 100.0 0.00 %
    Q4_K_M (not published) 93.94 🔴 −1.51 pts (−1.6 %) 🔴 3.03 % (+3.03 pts) 124 calls† 🔴 91.67 (−8.33 pts) 🔴 8.33 % (+8.33 pts)

    † "loops" and "worst response" are defined over the 66 tool-turns of the sealed set. Read over all 138 items, Q4_K_M's worst response was 226 calls (read_url × 225, to the generation cap) on a turn where production answered without a tool. Q8_0's worst over all 138 is still 6.

    BF16 scores identically to Q8_0 and loops anyway. Quoting the score alone would present them as equivalent while one exhibits the defect this model was selected for not having. BF16 is the higher precision, so the honest reading is that Q8_0's quantisation suppresses a tendency the full-precision weights retain — not that Q8_0 is better than its own weights. We have not established why.

    Q4_K_M loops at 8.33 % on the provocation slice — the untrained base's rate. It would have been the smallest and most tempting download, and it is the one that gives the benefit back; that is why it is not offered.

  5. The screen is not byte-deterministic — 14/20 identical sequences on re-collection — though the loop verdict and first-tool choice agreed 20/20.

  6. Aggregate run-to-run variance is unbounded on current evidence: 20 items were re-collected, not 266.

  7. An inherited upstream tokenizer issue — transformers warns the base checkpoint ships an incorrect regex pattern. See NOTICE §6.

  8. Scores are per serving path. Every headline figure is the Q8_0 GGUF through llama.cpp's completion endpoint with our own template rendering and tool parser. The same weights through other front ends, on the same held-out set, same cap and temperature:

    serving path artifact sealed Δ vs llama.cpp loops tool calls per response
    llama.cpp /completion (our harness) Q8_0 GGUF 95.45 — 0.00 % up to 6
    Ollama 0.35 (Modelfile, its own template and parser) Q8_0 GGUF 93.94 🔴 −1.51 pts (−1.6 %) 0.00 % up to 6
    mlx_lm.server BF16 safetensors 89.39 🔴 −6.06 pts (−6.3 %) 0.00 % never more than 1 (§3)

    Each front end renders Mistral's [TOOL_CALLS] grammar its own way, and the differences are the front ends', not the weights'. transformers and vLLM were not measured.

6. Intended use

Tool-calling assistants and agent workloads. Pico V2 is built for choosing a tool and constructing its arguments inside an agent loop — the production workload of the OculusMind.AI agents it was trained on and is deploying to — and it is validated as a general tool-calling assistant model on public suites: BFCL V4 (ahead of the base on 15 of its 17 categories, across single, parallel, live and multi-turn function calling), IFEval (instruction following, +3.5 points over the base) and GSM8K (reasoning, +7.9 points). Use it where an 8B model that reliably picks and calls tools is wanted: local agents, tool routers, structured-output assistants, and assistants that drive external APIs.

Where the evidence stops. It declines to call a tool slightly less often than the base (BFCL's two irrelevance categories, −2 to −3 points), so a deployment that depends on restraint should keep its own guard. Safety and refusal behaviour were not evaluated beyond what the base provides; multilingual use, long free-form writing and code generation were not evaluated. Text only — the base model's vision encoder is not part of this release and image inputs are not supported; a vision-capable build is planned for V3.

7. Licence and attribution

Apache-2.0, as is the base. The base's licence was verified live against the upstream repository on 2026-09-29 and the pinned revision still matched upstream main. See NOTICE for the base-model attribution and the statement of modifications.

Lineage. Pico V2 is the second generation of the series that began with OculusMindAI/OculusMind-ToolCall-8B-v1: the same aim — an 8B model that calls tools the way our production agents do — and the same base model at the same revision, now with a production-only training corpus, a seed sweep under a loop gate, and held-out measurement.

If you build on Pico V2, the attribution travels with it. Under Apache-2.0 §4, any redistribution or derivative work — a further fine-tune, a merge, a re-quantisation, a conversion to another format — must include this repository's NOTICE file, which names OculusMind.AI and this repository as the source, together with a copy of the licence and a prominent note of what was changed. Please also cite it:

@misc{oculusmind2026picov2,
  title  = {Pico V2: an 8B tool-calling model fine-tuned on production agent traffic},
  author = {{OculusMind.AI (Andrew Paris LLC)}},
  year   = {2026},
  url    = {https://huggingface.co/OculusMindAI/OculusMind-ToolCall-8B-v2}
}

8. Files

Pick the group for the software you run; you need one group, not all of them.

A. llama.cpp, Ollama, LM Studio → the Q8_0 GGUF. This is the measured artifact: every figure in this card is from this file, through llama.cpp. Ollama users: USAGE-OLLAMA.md walks through create, run, a tool call and the tool-result round trip, every command executed against these files on Ollama 0.35.1; ollama create oculusmind-ai-pico-v2:q8_0 -f Modelfile builds it with the right stop tokens and an identity-only system prompt. To install it in one line instead, with the same settings: ollama run oculusmindai/toolcall-8b-v2 (from the Ollama catalog) or ollama run hf.co/OculusMindAI/OculusMind-ToolCall-8B-v2 (from this repository, which reads system and params below).

file bytes sha256
gguf/Q8_0/model-q8_0.gguf 9,028,871,552 fe76fb07e5ca7351171bcd94838129514f8ad9bfc0a4e84ac6d9c247b33e1c56
Modelfile — Ollama definition for the GGUF above 2,458 af1404948cc6758687de0851efb0786f819a088139792e19b26681e471b62a1c
system — the Modelfile's system prompt, for ollama run hf.co/... 463 3172c95e8a831428f1a1085d1f141b045bce0769d5341c775b0a8e3f98305323
params — the Modelfile's stop tokens, temperature and context, for ollama run hf.co/... 122 86cb31f9330417900cdcd032702beb1d0af4aeb24896872827c0b7d2f8fd6316

B. transformers, vLLM, MLX → the BF16 safetensors. The same merged weights at full precision, in the standard Hugging Face layout (Ministral3ForCausalLM, text only), with the tokenizer and chat template. Loaded from this repository, they produce a correct structured tool call through transformers 5.19 and through mlx_lm 0.32; neither path was scored on the evals above, and vLLM was not run. Through MLX they return at most one tool call per response (§3, §5).

file bytes sha256
model-00001-of-00004.safetensors 5,318,535,144 627cbaa0706299fd1e695a9151da514ac05dd34bb978137c24fb2aff11e8b990
model-00002-of-00004.safetensors 5,352,157,840 97a1d4852fe20c89411cec0610521aa52ae0c2d066f8b6935e45b3429b4d31c1
model-00003-of-00004.safetensors 5,234,708,880 da958797386cfdb2281a908d1c0d4caa5938d44a337f31a3e4f70c32a4209c02
model-00004-of-00004.safetensors 1,073,741,952 bf02e982fd7a72202ab02eb837a9f7c3cddea4e87db6d0def674689de38fba11
model.safetensors.index.json 25,472 047bcb4995569e1d935ad5710cc76a237981987a64e71fdfa2cd17d6856e30ba
config.json 844 89573a75764b477492b4a789dc665c36c0125f16c2c1bf0ea7d29f695b2f7f81
generation_config.json 131 e0923390059f84a9180b00e5501778acc45ea9856cd7f2fd68208b360927c677
tokenizer.json 17,078,128 d5f6046775b112f0e2d456ee9dba450684ab964fe5c4e231599bdc6773028135
tokenizer_config.json 198,094 f59f7294e4f26383d0ea93840fe21cf197784be0842a8301a0343e8c34ed0d6d
special_tokens_map.json 21,436 8233d50ad79471ccb483a869703b940f5db1763db7ffd7c0134e0cec1a00a56e
chat_template.jinja 11,912 74eeb55fd3341286ec3fd44e902b7120721acc81cd394e96b431f85e93a1ea56

C. The LoRA adapter → apply the fine-tune to the base yourself. Rank 8, on the attention projections only, in mlx_lm's adapter format. Fused into mistralai/Ministral-3-8B-Instruct-2512-BF16 at revision f6fae979… it reproduces the BF16 safetensors above tensor for tensor; applied at serve time through mlx_lm (--adapter-path, or the adapters field of a server request; mlx_lm.server before 0.32.0 ignores its --adapter-path flag and serves the plain base, so on those versions use the request field) it scores 89.39 on the held-out set and 25.0 on the loop-provocation slice — the merged weights' figures on that path to the hundredth, with identical answer text on 136 of 138 held-out items and identical tool calls on 137 — against 80.3 and 16.67 for the base alone there. The fine-tune is the adapter; the weights are one way of carrying it.

file bytes sha256
adapter/adapters.safetensors 29,000,642 b3e2672fc0781ec8c5c9f604ed09fd0cb6958b75420ed6fe3e8aee0816ee4d94
adapter/adapter_config.json 1,208 8ee3dff9c1e3ef47313cfdb7c7b973edee3b695bf5c2afc705e033afc0f51b7d

D. Documentation. README.md (this card), USAGE-OLLAMA.md (verified Ollama walkthrough), NOTICE (licence, attribution and the statement of modifications — derivative works must carry it, §7), LICENSE (Apache-2.0), APPENDIX-looping.md (the loop measurement and the seed sweep), APPENDIX-claude-comparison.md (Claude Haiku 4.5 and 5.5 on the same populations, with its charts in assets/), CHECKSUMS.sha256 (every file above), and assets/uno.png (the portrait above; OculusMind.AI's own artwork, 1024 × 1024).

Every sha256 above was recorded by the build that produced the file, and model-q8_0.gguf was re-hashed from disk on 2026-10-05 and matched — a manifest that agrees with itself proves nothing about the bytes on disk.

There is deliberately no BF16 or Q4_K_M download. Both were built from the same weights and measured (§5.4); both loop on the sealed set where Q8_0 does not, and the decision was to publish no artifact that exhibits the defect this model was selected against. Quantising the published GGUF further yourself will not reproduce the Q8_0 numbers — §5.4 is the measurement of exactly that.

9. Context window — smoke-tested to 250K tokens

The GGUF declares a 262,144-token window (YaRN ×16 over the base's original 16,384). The fine-tune trained at ≤ 20,480 tokens and every score above was measured at a 65,536-token serving context, so the window was tested separately rather than inferred from metadata. On 2026-10-06, through llama.cpp c812c543f at --ctx-size 262144 with a q8_0 KV cache, a deterministic synthetic document with three facts planted at 10 / 50 / 90 % depth was sent at each size with a three-tool schema attached, asking for each fact and for one web_search call about the mid-depth fact:

document tokens facts recalled (depth .1 / .5 / .9) tool call formed time to first token (prefill, cold) response once the cache is populated: fact recall · tool call
2,048 3 / 3 yes 5 s 0.2 s · 0.4 s
51,200 3 / 3 yes 130 s 0.5 s · 1.5 s
102,400 3 / 3 yes 459 s 0.8 s · 1.7 s
153,600 3 / 3 yes 988 s 1.2 s · 2.5 s
204,800 3 / 3 yes 1,712 s (28.5 min) 1.6 s · 4.1 s
250,880 3 / 3 yes 2,559 s (42.7 min) 2.2 s · 5.1 s

The prefill column is the one-time cost of reading the document into the KV cache; every further question against the same document — the second and third fact probes and the tool call — answered in the last column's times, because llama.cpp reuses the cached prefix. The untrained base produced the identical table to within a tenth of a second. This is a smoke test, not a long-context benchmark: it shows the window serves and the model reads from deep inside it and still forms a correct tool call; it does not rank models or measure degradation with length. Two practical facts follow from it: prefill cost grows roughly quadratically with context on this engine (half an hour to first token for a 200K prompt on an Apple M3 Ultra at Q8_0), and a 262,144 window costs about 32 GB of memory here including a ~18 GB KV cache.

Downloads last month
321
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OculusMindAI/OculusMind-ToolCall-8B-v2

Finetuned
(11)
this model
Quantizations
2 models