I am a test model. I exist to be evaluated β nothing more.
About
ox-alpha is an experimental test model, released in a deliberately minimal state β no documentation, no benchmark results, no intended production use. It exists solely to support an internal evaluation experiment.
Context
ox-alpha was built as an internal experiment and surfaced publicly before any formal announcement β without creator information, without a model card, without benchmarks. It remains under active evaluation by its creators.
On OpenRouter
ox-alpha is available as a free preview on OpenRouter under stealth/ox-alpha.
| Model ID | stealth/ox-alpha |
| Listed | August 20, 2026 |
| Type | Reasoning model β coding & agentic workloads |
| Context window | 1M tokens (1,048,576) |
| Max output | 128K tokens (131,072) |
| Input / Output | Text Β· Image Β· Video β Text |
| Tools | Function calling, structured output |
| Price | Free during preview |
Since launch it has become one of the most-used models on the platform, with trillions of tokens routed through production agent workflows β top applications include Hermes Agent and Claude Code.
Source: openrouter.ai/stealth/ox-alpha
Terminal-Bench 3.0 results
An independent evaluation of stealth/ox-alpha on Terminal-Bench 3.0 (one canonical attempt per task, Harbor harness, mini-swe-agent 2.4.6, reasoning effort max, Aug 21β22 2026):
pass@1 results
| Slice | Coverage | Mean reward / pass@1 | Strict full solves |
|---|---|---|---|
| Non-GPU | 70/70 | 27.34% | 18/70 |
| GPU (valid) | 2/4 | 0.00% | 0/2 |
| Canonical scorable total | 72/74 | 26.58% | 18/72 (25.00%) |
Two GPU trials were excluded as infrastructure-invalid; the four GPU tasks also ran on different hardware than declared. Full details in the dataset.
Token efficiency
The full run used 0.895B tokens (12.09M per attempt). Normalizing public leaderboard totals to one attempt per task:
| Model | Tokens per attempt (approx.) |
|---|---|
| Claude Opus 5 | ~1.46B |
| GPT-5.6 Sol | ~1.16B |
| GLM 5.3 | ~1.12B |
| Ox Alpha | 0.895B (measured) |
| Claude Fable 5 | ~0.72B |
| Grok 4.6 | ~0.58B |
Complete sanitized trajectories for all 74 tasks are published for review and failure analysis.
Sources: dataset Β· Terminal-Bench Β· leaderboard image
Head-to-head on OpenRouter
Platform metrics only. Artificial Analysis intelligence/coding/agentic suites have not yet scored ox-alpha.
vs Claude Opus 5 (Anthropic)
| Metric | Claude Opus 5 | Ox Alpha |
|---|---|---|
| Context window | 1M | 1.05M |
| Max output | 128K | 131K |
| Input price | $5/M tokens | Free |
| Output price | $25/M tokens | Free |
| Latency p50 | 4.99s | 3.47s |
| Throughput p50 | 62 tok/s | 26 tok/s |
| Tokens routed (30d) | 6.77T | 15.5T |
Source: openrouter.ai/compare/anthropic/claude-opus-5/stealth/ox-alpha
vs Claude Fable 5 (Anthropic)
| Metric | Claude Fable 5 | Ox Alpha |
|---|---|---|
| Context window | 1M | 1.05M |
| Max output | 128K | 131K |
| Input price | $10/M tokens | Free |
| Output price | $50/M tokens | Free |
| Latency p50 | 8.17s | 3.47s |
| Throughput p50 | 41 tok/s | 26 tok/s |
| Tokens routed (30d) | 1.06T | 15.5T |
Source: openrouter.ai/compare/anthropic/claude-fable-5/stealth/ox-alpha
vs GPT-5.6 Sol (OpenAI)
| Metric | GPT-5.6 Sol | Ox Alpha |
|---|---|---|
| Context window | 1.05M | 1.05M |
| Max output | 128K | 131K |
| Input price | $2/M tokens | Free |
| Output price | $10/M tokens | Free |
| Latency p50 | 2.34s | 3.47s |
| Throughput p50 | 48 tok/s | 26 tok/s |
| Tokens routed (30d) | 3.33T | 15.5T |
Source: openrouter.ai/compare/openai/gpt-5.6-sol/stealth/ox-alpha
vs Kimi K3 (Moonshot AI)
| Metric | Kimi K3 | Ox Alpha |
|---|---|---|
| Context window | 1.05M | 1.05M |
| Max output | 975K | 131K |
| Input price | $2.60/M tokens | Free |
| Output price | $13/M tokens | Free |
| Latency p50 | 2.23s | 3.47s |
| Throughput p50 | 41 tok/s | 26 tok/s |
| Tokens routed (30d) | 5.98T | 15.5T |
Source: openrouter.ai/compare/moonshotai/kimi-k3/stealth/ox-alpha
vs GLM 5.3 (Z.ai)
| Metric | GLM 5.3 | Ox Alpha |
|---|---|---|
| Context window | 1.05M | 1.05M |
| Max output | 131K | 131K |
| Input price | $1.40/M tokens | Free |
| Output price | $4.40/M tokens | Free |
| Latency p50 | 2.43s | 3.47s |
| Throughput p50 | 46 tok/s | 26 tok/s |
| Tokens routed (30d) | 570B | 15.5T |
Source: openrouter.ai/compare/z-ai/glm-5.3/stealth/ox-alpha
Technical overview
| Architecture | nanotech (custom) |
| Parameters | ~800B (logical) |
| Weight format | Packed ternary (-1 / 0 / +1) |
| Effective precision | β 1.58 bits per weight |
| Distribution | 20 shards Γ 5 GB (100 GB total) |
| Status | Experimental β test model |
Terms
All rights reserved. No open-source license is granted with these weights.
| Action | Status |
|---|---|
| Local evaluation | Tolerated |
| Copying | Not authorized at this time |
| Redistribution / re-uploading | Not authorized at this time |
| Public derivatives (fine-tunes, merges, quantizations) | Not authorized at this time |
These terms apply until further notice.
Era Logic LLC A technology company based in Silicon Valley.
Official website β era.lc (coming soon)
- Downloads last month
- 179





