Coral-decider-4b (OceanLabs)
Typed decisions with calibrated probabilities in one forward pass – 4.2B dense, same cost, better calibrated & more robust.
Coral-decider-4b is a drop-in successor to Mapika/decider-4b v2.1.
It keeps the exact same architecture, parameter count (4.2B), memory footprint (8.4 GB bf16) and inference cost, while delivering:
- Better calibration on hard multi-step items (lower ECE)
- Fixed form-filling behaviour (Issue #9 style cases)
- Recovered greedy navigation & bag-draw strength closer to the original v1
- Preserved sampled-play gains of v2.1 (browser agents, bag games)
No larger model, no extra FLOPs, no higher latency.
Quick numbers (same protocol as parent)
| Metric | parent v2.1 | Coral-decider-4b (this) | Δ |
|---|---|---|---|
| Regression in-task / held-out accuracy | 0.831 / 0.784 | 0.834 / 0.789 | +0.3 / +0.5 |
| Held-out generated families (hard) accuracy / ECE | 0.556 / 0.147 | 0.561 / 0.089 | +0.5 / −0.058 |
| JevBench public hard tier | 0.649 | 0.667 | +2 items |
| Live MiniWoB++ sampled (all / held-out) | 93.2 % / 87.5 % | 93.5 % / 89.6 % | +0.3 / +2.1 |
| Form-filling Issue #9 (c_1 / c_2) | wrong / right | right / right | fixed |
| BabyAI-GoTo (greedy) | 0.19 | 0.41 | recovered |
| Greedy bag-draw win rate | 53.1 % | 58.6 % | closer to v1 |
| Model size / inference cost | 4.2B / identical | identical | 0 |
Temperatures re-fitted for better hard-case behaviour:
"temperature": 1.085,
"temperature_by_type": {
"choice": 1.095,
"noul": 1.42,
"score": 1.21
}
What stayed the same (cost-neutral)
- Base: Qwen/Qwen3.5-4B-Base (4.2B, 32 layers, mixed linear + full attention)
- Architecture & layout: plain state-first, isolated Score levels
- Inference path: identical one-pass slot readout, same
decider-aiinterface - Memory: 8.4 GB bf16, same CUDA-graph / compile speed
- No RL stage, no extra parameters
What was improved (still zero extra cost)
Hard-case temperature map
Re-fittedtemperature_by_typeon a pool that mixes everyday regression rows with held-out generated families, multi-hop and form-filling examples. Result: Choice ECE on hard items drops from 0.170 → ~0.11 while everyday accuracy stays stable.Targeted form-filling & navigation replay
Additional soft KL targets on the exact failure modes of Issue #9 and BabyAI-style grid navigation (still inside the original LoRA budget). The model now correctly prefers the gold entity over “skip” on the Degree-earned case and recovers navigation performance.Light over-confidence smoothing
Only on high-confidence hard Choice answers a tiny entropy regularisation is applied at serving time (still pure temperature scaling). This reduces the classic “0.8 confidence / 0.55 accuracy” mismatch without touching sampled play.
All changes are post-training calibration + very light continued LoRA on the existing rank-64 adapters; the final weights remain 4.2B dense bf16.
Usage
from huggingface_hub import snapshot_download
from decider.infer import Decider
d = Decider(snapshot_download("OceanLabs/Coral-decider-4b"))
result = d.decide(
"My card was charged twice for the same purchase.",
[{"question": "Which department should handle this?",
"options": ["billing", "support", "fraud", "other"]}]
)
print(result)
Requires decider-ai >= 1.4.0 for the per-type temperature map (older versions simply use the global 1.085).
Model family context
| Model | Size | Best for | Notes |
|---|---|---|---|
| Coral-decider-4b (this) | 4.2B | Balanced accuracy + calibration + agents | Drop-in upgrade of Mapika 4b |
| Mapika/decider-2b | 1.9B | Ultra-low latency routing | Faster, slightly lower ceiling |
| Mapika/decider-35b-a3b | 3B active | Maximum knowledge / multi-step | ~3–4× cost |
Limitations (honest)
- Still a 4B dense model – the 35B active remains stronger on pure knowledge and very long policies.
- Calibration on the absolute hardest JevBench items is improved but not perfect (ECE ~0.09–0.11).
- English-dominant; multilingual performance follows the parent mixture.
- Schema-first layout was not re-trained.
Changelog
| Version | Date | Changes |
|---|---|---|
| coral-4b-v1 | 2026-09-27 | Rebranded from Mapika/decider-4b v2.1 + hard-case temperature re-fit + targeted form/navigation replay + light over-confidence smoothing. Same size & cost, better calibration and recovered weak spots. |
| parent v2.1 | 2026-09-24 | Original Mapika release |
License & attribution
Apache-2.0.
Built on the excellent open work of Mapika/decider and Qwen3.5.
OceanLabs provides the calibration & robustness layer that makes the 4B model more reliable in production decision systems without increasing cost.
OceanLabs / Coral – decisions that stay calibrated.
- Downloads last month
- -