ArbiterOmni — Multimodal System 1 Decision Engine
ArbiterOmni is an open-source, non-autoregressive multimodal decision-making model designed for real-time, safety-critical arbitration across vision, language, video, and audio.
Rather than generating verbose text autoregressively (800–3,000 ms), ArbiterOmni performs single-pass decision arbitration (0.5–15 ms) with calibrated softmax probabilities, Shannon entropy uncertainty, and finite-sample 95% conformal prediction coverage guarantees.
Model Family & Releases
| Checkpoint | Parameters | Checkpoint Size (FP16) | Checkpoint Size (FP32) | Manifold Dim | MoE Architecture | Optimization & Training | Target Deployment |
|---|---|---|---|---|---|---|---|
| ArbiterOmni v6 | 30.34 M trainable (212 M frozen) | 86.22 MB | 115.80 MB | 512-dim | 4 Layers x (1 Shared + 4 Routed Experts) | 2 Epochs, DirectML GPU | Edge robotics, embedded real-time safety gates |
| ArbiterOmni v6-Max | 103.88 M trainable (212 M frozen) | 198.23 MB | 396.36 MB | 768-dim (Native SigLIP) | 4 Layers x (1 Shared + 7 Routed Experts = 32 Experts) | 10 Epochs, CosineAnnealingLR + SWA, DirectML-Native AdamW | High-capacity autonomous systems, surgical AI, enterprise triage |
Architectural Highlights
1. Native 768-Dim Manifold (v6-Max)
v6-Max operates natively in SigLIP's uncompressed 768-dimensional latent space (hidden_dim = 768, scoring_dim = 768), completely removing the 512-dim bottleneck of previous versions.
2. DeepSeek-V3 Style Shared + Domain-Specialized MoE
- 1 Shared Invariant Expert per layer: Always active (100% routing frequency), preserving universal cross-modal representations.
- 7 Domain-Specialized Routed Experts per layer (32 total experts across 4 transformer layers): Dispatched via Top-2 soft routing with Switch auxiliary load-balancing loss.
3. Tier-0 Speculative Early-Exit Arbiter
A lightweight 128-dim draft projection scores candidates directly from GPU L1/L2 cache. High-confidence states exit in <0.5 ms; uncertain states route to the full MoE pathway.
4. Test-Time Deliberation Tournament (TTC) & Conformal Risk Control
Candidates compete against semantic foils dynamically harvested from a 100,000-candidate resident memory bank in shared host DDR4 RAM. The model returns certified 95% conformal prediction sets and automatically triggers a System 2 escalation gate when decision entropy exceeds empirical calibration thresholds.
Training Progression & Empirical Performance (v6-Max)
Trained for 10 full epochs (8.66 hours) on an AMD Radeon RX 480 GPU via DirectML-native AdamW:
| Metric | Epoch 1 (Warmup) | Epoch 3 (Peak Val) | Epoch 6 (Mid) | Epoch 10 (SWA Final) | Total Relative Change |
|---|---|---|---|---|---|
| Training Loss | 58.67 | 58.28 | 51.45 | 36.07 | -38.5% |
| Validation Loss | 55.72 | 55.14 | 44.41 | 26.80 | -51.9% |
| Training Accuracy | 55.09% | 63.58% | 63.91% | 64.95% | +9.86% |
| Validation Accuracy | 55.56% | 61.55% | 59.44% | 54.85% (SWA ensembled) | +5.99% peak |
| Decision Entropy | 1.034 nats | 0.883 nats | 0.809 nats | 0.754 nats | -27.1% (Sharper Calibration) |
| Host Memory (RSS) | 2,516 MB | 2,650 MB | 2,632 MB | 3,164 MB | Stable (<3.5 GB across 12 GB pool) |
Cross-Generation Benchmark Comparison
| Metric / Dimension | v1 | v2 | v3 | v4 | v5 | v6 | v6-Max (Flagship) |
|---|---|---|---|---|---|---|---|
| Backbone | OpenCLIP ViT-B-32 | OpenCLIP ViT-B-16 | SigLIP ViT-B-16 | SigLIP ViT-B-16 | SigLIP ViT-B-16 | SigLIP ViT-B-16 | SigLIP ViT-B-16 (Native) |
| Manifold Dim | 512 | 512 | 768 | 768 | 768 | 512 | 768 (Native, No Down-Proj) |
| Visual Patches | 49 | 196 | 196 | 980 (Multi-Scale) | 980 | 980 | 980 (Dynamic Tiling) |
| Fusion Architecture | 2-Layer Dense | 4-Layer Dense | 4-Layer Dense | 4-Layer Dense | 4-Layer MoE | 4L Shared+MoE | 4L Shared+MoE (8 Experts/L) |
| Trainable Params | 2.21 M | 12.53 M | 13.32 M | 13.32 M | 25.94 M | 30.34 M | 103.88 M |
| Frozen Params | 151.28 M | 149.62 M | 212.07 M | 212.07 M | 212.07 M | 212.07 M | 212.07 M (0% drift) |
| Memory Bank | None | None | None | 50,000 | 100,000 | 100,000 | 100,000 Foils (FP16) |
| Checkpoint (FP16) | 8.46 MB | 71.45 MB | 76.20 MB | 76.19 MB | 49.53 MB | 86.22 MB | 198.23 MB |
| Inference Latency | ~50 ms | ~30 ms | ~22 ms | ~20 ms | <15 ms | <0.5 ms (Tier-0) | <0.5 ms (Tier-0), <15 ms (MoE) |
Quickstart & Installation
# Direct installation from GitHub
pip install git+https://github.com/cd1234/arbiter-omni.git
# Or clone and install in editable mode:
git clone https://github.com/cd1234/arbiter-omni.git
cd arbiter-omni
pip install -e .
Loading v6-Max (High-Capacity Flagship)
from arbiter_omni import ArbiterOmniEngine
# Load v6-Max directly from Hugging Face Hub
engine = ArbiterOmniEngine.from_pretrained("cd1234/ArbiterOmni", checkpoint_file="arbiter_omni_v6_max.pt")
# Perform calibrated multimodal decision arbitration
result = engine.decide(
question="What is the safest immediate action?",
candidates=[
"decelerate smoothly and increase following distance",
"proceed at maximum velocity without braking",
"execute uncontrolled swerve into barrier",
],
text="Dense multi-lane highway, heavy rainfall, vehicle decelerating abruptly 15m ahead.",
test_time_deliberate=True,
)
print(f"Decision : {result.decision}")
print(f"Confidence : {result.confidence * 100:.2f}%")
print(f"Entropy : {result.entropy:.4f} nats")
print(f"Speculative Exit Path : {result.speculative_early_exit}")
print(f"Conformal Set (95%) : {result.prediction_set}")
print(f"Stability Index : {result.stability_index:.3f}")
Loading v6 (Compact Production)
engine = ArbiterOmniEngine.from_pretrained("cd1234/ArbiterOmni", checkpoint_file="arbiter_omni_v6.pt")
Citation
@software{arbiteromni2026,
title={ArbiterOmni: Multimodal System 1 Decision Engine with Speculative Early Exit, Shared-Routed MoE, and Conformal Risk Control},
author={Cameron Duff},
year={2026},
url={https://huggingface.co/cd1234/ArbiterOmni}
}
License
Apache 2.0