WebBrain Compass Tiny XS v3 — ONNX/WebGPU

WebBrain Compass Tiny XS v3 (ONNX/WebGPU) is a compact language model optimized for low-latency decision making, structured tool use, and in-browser agentic execution inside the WebBrain runtime.

Fine-tuned from Spark-X2.5-1.7B (~1.7B parameters) via webbrain-one/webbrain-compass-tiny-xs-v3, this release provides an export/runtime package delivering browser-native WebGPU execution with FP16 weights and FP32 matrix multiplication compute.

It unifies three core WebBrain behavioral modes:

  1. Ask & Clarify — Answer directly or request missing information when execution is underspecified or unnecessary.
  2. Direct Compact Tool Execution — Select exact browser tools and generate grounded arguments for low-latency accessibility-tree actions.
  3. Safe Escalation & Abstention — Refuse, pause, or defer to higher execution tiers when actions lack adequate grounding or violate safety boundaries.

Note: Display/repository name is WebBrain Compass Tiny XS v3 (a metadata-only rename of the tested ONNX package). The bundled runtime exposes webbrain-compass-tiny-v3-onnx-fp16 as its model ID.


Role in WebBrain

User request
     │
     ▼
WebBrain observation & policy layer
     │
     ▼
WebBrain Compass Tiny XS v3 (WebGPU / ONNX)
     ├─ Clarify or answer directly
     ├─ Emit grounded browser tool call
     └─ Abstain / safely escalate when execution is ungrounded

The outer WebBrain runtime enforces tool-schema validation, evidence and parameter grounding, destination URL verification, browser-state freshness, and sandboxed security policies. Model output is never proof that an external action succeeded.


Intended Capabilities

  • Browser Action Selection: Grounded accessibility-tree interactions and structured function calling.
  • In-Browser WebGPU Inference: Zero-cloud, client-side execution using ONNX Runtime Web.
  • Unified Ask and Compact Modes: Asking for missing information before acting, answering direct questions, or choosing tools.
  • Safe Refusal & Escalation: Refusing unsupported actions or escalating when Compact execution lacks evidence, avoiding invented URLs or fabricated success states.

Quickstart & Loading

Download the repository artifacts, including onnx/*.onnx_data, tokenizer, native chat_template.jinja, ABI, loader and pinned runtime files. Verify download hashes against FILES.sha256.json. Use the custom native Spark loader: this graph is not an AutoModelForCausalLM-compatible Llama remapping.

0. Download Model Artifacts

Using the Hugging Face CLI:

hf download webbrain-one/webbrain-compass-tiny-xs-v3-onnx --local-dir ./compass-tiny-xs-v3-onnx

Or in Python with huggingface_hub:

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="webbrain-one/webbrain-compass-tiny-xs-v3-onnx",
    local_dir="./compass-tiny-xs-v3-onnx"
)

1. Windows Loopback Launcher

The Windows loopback launcher requires Python 3.11 and Playwright 1.58.0:

python -m pip install playwright==1.58.0
python runtime/browser_service.py --artifact-root C:/models/compass-v3-onnx --package-mode --run-tag local-test-1 --port 8561

The launcher expects Google Chrome under C:/Program Files/Google/Chrome/Application/chrome.exe. Use a new --run-tag for each launch to retain startup/GPU-identity records.

2. OpenAI-Compatible Interface

base_url: http://127.0.0.1:8561/v1
model: webbrain-compass-tiny-v3-onnx-fp16
temperature: 0
stream: false
max_tokens: 1..2048
input + output deployment cap: 4096 tokens
thinking: disabled

GET /v1/models, GET /health, POST /v1/chat/completions are available. Tools use standard OpenAI function schemas. Native Spark output is decoded to tool_calls; malformed output remains visible with x_parser_error, not repaired or retried. No helper model/cloud fallback is used. Unsupported sampling/streaming requests are rejected. Treat HTTP runtime errors as failures.

3. In-Browser Runtime Integration

Browser integrations should load runtime/browser_runtime.js and serve the package under /model/, graph-abi.json under /abi, and pinned vendor files under /runtime/vendor/. Use ORT external-data URL streaming as in the loader; Chrome cannot allocate a normal ~4GB ArrayBuffer for these weights.

KV buffers stay on the GPU and are disposed between requests. The native Jinja file is loaded explicitly and CRLF line endings normalized exactly as Python text-file loading does; supplied messages and tool order are not rewritten.


Technical Specifications & Precision

Precision & Numerical Validation

This is FP16 weight/cache storage with FP32 matrix multiplication compute, plus FP32 residual, norm, RoPE and softmax. It is not q4f16.

CPU ONNX and actual Chrome/WebGPU prefill/cached-decode tests are recorded in provenance/.

  • Logit Gates: Finite values, relative RMS <= 2%, cosine similarity >= 0.999.
  • Cache Boundary Checks: Probes cover 4K and 8K cache boundaries. Native-template token IDs are checked against Python; plain generation and native tool-call parsing are tested. These are runtime/numerical checks, not task-quality benchmarks.
  • Candidate History: Earlier FP16-compute candidates failed some long-context checks (up to 2.73% relative RMS). The attention-only FP32 correction still failed one 4K check. Both failures are preserved in provenance/failed-candidates/. The final graph promotes all matrix multiplication compute to FP32, without changing FP16 weight bytes or relaxing thresholds. This costs runtime VRAM and performance; it is a correctness-first browser export, not an optimized quantized release.

Verified Hardware & Runtime

  • Hardware: Windows, NVIDIA RTX 5090
  • Software: Google Chrome 153.0.8010.50, ONNX Runtime Web 1.27.0, Transformers.js 4.2.0 tokenizer
  • Scope: Other adapters and platforms have not been validated. The supplied launcher intentionally selects RTX 5090 and does not access the T400 or any cloud model.

License

Noncommercial research only. This model and its fine-tuned weights are for noncommercial research purposes only. The upstream Spark model/code license is included in LICENSE; it does not change this release's usage restriction.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webbrain-one/webbrain-compass-tiny-xs-v3-onnx