OMEGA-Reasoner (POC)
β οΈ This is a proof-of-concept β a small-scale demonstration of a self-contained LLM-driven agent framework. It is not a frontier model and is not intended to compete with production models on public leaderboards such as DeepSWE.
Honest positioning
| Property | This repo | DeepSeek-V4.1-Flash (Top-1 DeepSWE) |
|---|---|---|
| Parameters | β 100K (configurable) | 35B-class |
| Training data | none (synthetic / hand-coded) | trillions of tokens |
| Pass@1 on DeepSWE | not measured | 74.2 |
| Status | research POC | production model |
We do not submit to the official DeepSWE leaderboard. The integration
code (deep_swe.py) is provided for educational and self-evaluation
purposes only.
What this model is
OMEGA-Reasoner is a self-contained language world model + autonomous
agent framework. It implements:
- HMoE-D reasoner: hierarchical mixture-of-experts with dynamic routing
- AGI_AgentWorld: full cognitive agent with episodic / semantic / procedural memory, ReAct loop, Tree-of-Thoughts planning, Reflexion-style self-correction, and constitutional guardrails
- AgentWorldBench adapter: 5-dim scoring (format, factuality, consistency, realism, quality)
- DeepSWE adapter: Harbor-format task loader + verifier runner
Why this repo exists
To show that a complete agent framework β including memory, planning, tool-use, self-correction, and benchmark adapters β can be implemented from scratch without any external LLM API. It is meant as a research artifact, not a competitive submission.
How to use
from evo.agi_agent_world import AGIAgentWorld
from evo.agentworld_bench import load_all, WorldModelAgent, evaluate
from evo.deep_swe import load_all_tasks, run_benchmark
# 1) Run a single goal
agi = AGIAgentWorld()
result = agi.run("Buy coffee by noon")
# 2) Evaluate on AgentWorldBench (synthetic data)
records = load_all("data/agentworldbench_synth")
rep = evaluate(WorldModelAgent(), records)
print(rep.overall)
# 3) Run on DeepSWE tasks (synthetic seed data)
tasks = load_all_tasks("data/deepswe_synth")
summary = run_benchmark(tasks)
print(summary.pass_at_1)
Files
agi_agent_world.pyβ autonomous agent (memory, planning, tools)omega_arch.pyβ HMoE-D model definitionomega_nlu.pyβ NLU engine (sentences, frames, FOL)omega_math.pyβ math reasoning primitivesomega_physics.pyβ physics reasoneromega_meta.pyβ meta-cognition / auditomega_quant.pyβ quantization utilitiesagentworld_bench.pyβ AgentWorldBench loader + 5-dim scorerdeep_swe.pyβ DeepSWE loader + verifiertests/β unit tests (78+ tests, all passing)
Test results
The repository ships with 78+ unit tests covering metrics, agents, benchmarks, and the cognitive core. All tests pass on the included synthetic fixtures.
Limitations
- The underlying HMoE-D model is a demonstration with no language
pretraining. Its
predict()is template-based, not learned. Pass@1on DeepSWE is not claimed β the agent is heuristic.- Numerical / symbolic reasoning is rule-based, not learned.
Citation
If you use this code in research, please cite the framework only (no leaderboard claim):
@software{omega_reasoner_poc,
title = {OMEGA-Reasoner: a self-contained LLM-agent framework (POC)},
year = {2026},
url = {https://huggingface.co/<your-org>/omega-reasoner-poc},
}
License
Apache-2.0.
- Downloads last month
- 22