ox-alpha

A test model. Nothing more β€” for now.

status params weights web


I am a test model. I exist to be evaluated β€” nothing more.

About

ox-alpha is an experimental test model, released in a deliberately minimal state β€” no documentation, no benchmark results, no intended production use. It exists solely to support an internal evaluation experiment.

Context

ox-alpha was built as an internal experiment and surfaced publicly before any formal announcement β€” without creator information, without a model card, without benchmarks. It remains under active evaluation by its creators.

On OpenRouter

ox-alpha is available as a free preview on OpenRouter under stealth/ox-alpha.

Model ID stealth/ox-alpha
Listed August 20, 2026
Type Reasoning model β€” coding & agentic workloads
Context window 1M tokens (1,048,576)
Max output 128K tokens (131,072)
Input / Output Text Β· Image Β· Video β†’ Text
Tools Function calling, structured output
Price Free during preview

Since launch it has become one of the most-used models on the platform, with trillions of tokens routed through production agent workflows β€” top applications include Hermes Agent and Claude Code.

Source: openrouter.ai/stealth/ox-alpha

Terminal-Bench 3.0 results

An independent evaluation of stealth/ox-alpha on Terminal-Bench 3.0 (one canonical attempt per task, Harbor harness, mini-swe-agent 2.4.6, reasoning effort max, Aug 21–22 2026):

Ox Alpha on Terminal-Bench 3.0

pass@1 results

Slice Coverage Mean reward / pass@1 Strict full solves
Non-GPU 70/70 27.34% 18/70
GPU (valid) 2/4 0.00% 0/2
Canonical scorable total 72/74 26.58% 18/72 (25.00%)

Two GPU trials were excluded as infrastructure-invalid; the four GPU tasks also ran on different hardware than declared. Full details in the dataset.

Token efficiency

The full run used 0.895B tokens (12.09M per attempt). Normalizing public leaderboard totals to one attempt per task:

Model Tokens per attempt (approx.)
Claude Opus 5 ~1.46B
GPT-5.6 Sol ~1.16B
GLM 5.3 ~1.12B
Ox Alpha 0.895B (measured)
Claude Fable 5 ~0.72B
Grok 4.6 ~0.58B

Complete sanitized trajectories for all 74 tasks are published for review and failure analysis.

Sources: dataset Β· Terminal-Bench Β· leaderboard image

Head-to-head on OpenRouter

Platform metrics only. Artificial Analysis intelligence/coding/agentic suites have not yet scored ox-alpha.

vs Claude Opus 5 (Anthropic)

Ox Alpha vs Claude Opus 5

Metric Claude Opus 5 Ox Alpha
Context window 1M 1.05M
Max output 128K 131K
Input price $5/M tokens Free
Output price $25/M tokens Free
Latency p50 4.99s 3.47s
Throughput p50 62 tok/s 26 tok/s
Tokens routed (30d) 6.77T 15.5T

Source: openrouter.ai/compare/anthropic/claude-opus-5/stealth/ox-alpha

vs Claude Fable 5 (Anthropic)

Ox Alpha vs Claude Fable 5

Metric Claude Fable 5 Ox Alpha
Context window 1M 1.05M
Max output 128K 131K
Input price $10/M tokens Free
Output price $50/M tokens Free
Latency p50 8.17s 3.47s
Throughput p50 41 tok/s 26 tok/s
Tokens routed (30d) 1.06T 15.5T

Source: openrouter.ai/compare/anthropic/claude-fable-5/stealth/ox-alpha

vs GPT-5.6 Sol (OpenAI)

Ox Alpha vs GPT-5.6 Sol

Metric GPT-5.6 Sol Ox Alpha
Context window 1.05M 1.05M
Max output 128K 131K
Input price $2/M tokens Free
Output price $10/M tokens Free
Latency p50 2.34s 3.47s
Throughput p50 48 tok/s 26 tok/s
Tokens routed (30d) 3.33T 15.5T

Source: openrouter.ai/compare/openai/gpt-5.6-sol/stealth/ox-alpha

vs Kimi K3 (Moonshot AI)

Ox Alpha vs Kimi K3

Metric Kimi K3 Ox Alpha
Context window 1.05M 1.05M
Max output 975K 131K
Input price $2.60/M tokens Free
Output price $13/M tokens Free
Latency p50 2.23s 3.47s
Throughput p50 41 tok/s 26 tok/s
Tokens routed (30d) 5.98T 15.5T

Source: openrouter.ai/compare/moonshotai/kimi-k3/stealth/ox-alpha

vs GLM 5.3 (Z.ai)

Ox Alpha vs GLM 5.3

Metric GLM 5.3 Ox Alpha
Context window 1.05M 1.05M
Max output 131K 131K
Input price $1.40/M tokens Free
Output price $4.40/M tokens Free
Latency p50 2.43s 3.47s
Throughput p50 46 tok/s 26 tok/s
Tokens routed (30d) 570B 15.5T

Source: openrouter.ai/compare/z-ai/glm-5.3/stealth/ox-alpha

Technical overview

Architecture nanotech (custom)
Parameters ~800B (logical)
Weight format Packed ternary (-1 / 0 / +1)
Effective precision β‰ˆ 1.58 bits per weight
Distribution 20 shards Γ— 5 GB (100 GB total)
Status Experimental β€” test model

Terms

All rights reserved. No open-source license is granted with these weights.

Action Status
Local evaluation Tolerated
Copying Not authorized at this time
Redistribution / re-uploading Not authorized at this time
Public derivatives (fine-tunes, merges, quantizations) Not authorized at this time

These terms apply until further notice.


Era Logic LLC A technology company based in Silicon Valley.

Official website β€” era.lc (coming soon)

Downloads last month
179
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support