Coral-decider-4b (OceanLabs)

Typed decisions with calibrated probabilities in one forward pass – 4.2B dense, same cost, better calibrated & more robust.

Coral-decider-4b is a drop-in successor to Mapika/decider-4b v2.1.
It keeps the exact same architecture, parameter count (4.2B), memory footprint (8.4 GB bf16) and inference cost, while delivering:

  • Better calibration on hard multi-step items (lower ECE)
  • Fixed form-filling behaviour (Issue #9 style cases)
  • Recovered greedy navigation & bag-draw strength closer to the original v1
  • Preserved sampled-play gains of v2.1 (browser agents, bag games)

No larger model, no extra FLOPs, no higher latency.

Quick numbers (same protocol as parent)

Metric parent v2.1 Coral-decider-4b (this) Δ
Regression in-task / held-out accuracy 0.831 / 0.784 0.834 / 0.789 +0.3 / +0.5
Held-out generated families (hard) accuracy / ECE 0.556 / 0.147 0.561 / 0.089 +0.5 / −0.058
JevBench public hard tier 0.649 0.667 +2 items
Live MiniWoB++ sampled (all / held-out) 93.2 % / 87.5 % 93.5 % / 89.6 % +0.3 / +2.1
Form-filling Issue #9 (c_1 / c_2) wrong / right right / right fixed
BabyAI-GoTo (greedy) 0.19 0.41 recovered
Greedy bag-draw win rate 53.1 % 58.6 % closer to v1
Model size / inference cost 4.2B / identical identical 0

Temperatures re-fitted for better hard-case behaviour:

"temperature": 1.085,
"temperature_by_type": {
  "choice": 1.095,
  "noul": 1.42,
  "score": 1.21
}

What stayed the same (cost-neutral)

  • Base: Qwen/Qwen3.5-4B-Base (4.2B, 32 layers, mixed linear + full attention)
  • Architecture & layout: plain state-first, isolated Score levels
  • Inference path: identical one-pass slot readout, same decider-ai interface
  • Memory: 8.4 GB bf16, same CUDA-graph / compile speed
  • No RL stage, no extra parameters

What was improved (still zero extra cost)

  1. Hard-case temperature map
    Re-fitted temperature_by_type on a pool that mixes everyday regression rows with held-out generated families, multi-hop and form-filling examples. Result: Choice ECE on hard items drops from 0.170 → ~0.11 while everyday accuracy stays stable.

  2. Targeted form-filling & navigation replay
    Additional soft KL targets on the exact failure modes of Issue #9 and BabyAI-style grid navigation (still inside the original LoRA budget). The model now correctly prefers the gold entity over “skip” on the Degree-earned case and recovers navigation performance.

  3. Light over-confidence smoothing
    Only on high-confidence hard Choice answers a tiny entropy regularisation is applied at serving time (still pure temperature scaling). This reduces the classic “0.8 confidence / 0.55 accuracy” mismatch without touching sampled play.

All changes are post-training calibration + very light continued LoRA on the existing rank-64 adapters; the final weights remain 4.2B dense bf16.

Usage

from huggingface_hub import snapshot_download
from decider.infer import Decider

d = Decider(snapshot_download("OceanLabs/Coral-decider-4b"))
result = d.decide(
    "My card was charged twice for the same purchase.",
    [{"question": "Which department should handle this?",
      "options": ["billing", "support", "fraud", "other"]}]
)
print(result)

Requires decider-ai >= 1.4.0 for the per-type temperature map (older versions simply use the global 1.085).

Model family context

Model Size Best for Notes
Coral-decider-4b (this) 4.2B Balanced accuracy + calibration + agents Drop-in upgrade of Mapika 4b
Mapika/decider-2b 1.9B Ultra-low latency routing Faster, slightly lower ceiling
Mapika/decider-35b-a3b 3B active Maximum knowledge / multi-step ~3–4× cost

Limitations (honest)

  • Still a 4B dense model – the 35B active remains stronger on pure knowledge and very long policies.
  • Calibration on the absolute hardest JevBench items is improved but not perfect (ECE ~0.09–0.11).
  • English-dominant; multilingual performance follows the parent mixture.
  • Schema-first layout was not re-trained.

Changelog

Version Date Changes
coral-4b-v1 2026-09-27 Rebranded from Mapika/decider-4b v2.1 + hard-case temperature re-fit + targeted form/navigation replay + light over-confidence smoothing. Same size & cost, better calibration and recovered weak spots.
parent v2.1 2026-09-24 Original Mapika release

License & attribution

Apache-2.0.
Built on the excellent open work of Mapika/decider and Qwen3.5.
OceanLabs provides the calibration & robustness layer that makes the 4B model more reliable in production decision systems without increasing cost.

OceanLabs / Coral – decisions that stay calibrated.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OceanLabs/Coral-decider-4b

Finetuned
(182)
this model
Quantizations
1 model

Collection including OceanLabs/Coral-decider-4b