Access echoAI-4 (release candidate)

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

echoAI-4 4.0.0rc2 is a research release candidate. The open layer is licensed under PolyForm Noncommercial 1.0.0. The compiled core is covered by the RxLabs EULA, which is still a DRAFT pending legal review: until a final version is published, use is limited to personal, research, education, evaluation and reproduction purposes; no commercial use, redistribution or reverse engineering. No claim of consciousness is made.

Log in or Sign Up to review the conditions and access this model content.

\ECHO-AI

ECHO-4
echoAI-4 · release candidate 4.0.0rc2

echoAI-4 — the agent that does not believe its own AI

What can I use it for?

echoAI-4 is a small agent that does not believe anything it cannot verify — including what a language model tells it. In practice:

  • Stop an AI from making things up. Put ECHO between any model and its answers: a claim only counts as a fact if it matches verified evidence; otherwise the answer is "I don't know". On SimpleQA this cut made-up answers by 81–87 % in four different models.
  • Guard a coding agent. Plug it into Claude Code or Codex over MCP: it checks claims like "the tests pass" against your own checks and says OK / MODIFY / BLOCK before any shell command (60/60 dangerous commands blocked in our exam).
  • Study an agent that learns from consequences. It perceives, remembers, predicts and learns on integers with the language model switched off, and publishes a receipt for every result.

It is not a chatbot and not a language model, and it makes no claim of consciousness.

Results at a glance

Incorrect answers on SimpleQA out of 200: Qwen3-30B 129 alone, 34 with Wikipedia, 17 with Wikipedia and ECHO; GPT-5.4 mini 145, 76, 27; DeepSeek V4.1 Flash 101, 66, 14; Claude Sonnet 5 104, 19, 14

Fewer made-up answers. Same models, same questions (SimpleQA, 200 per model). The price is honest: with ECHO the models answer less often.

SimpleQA: precision against coverage for each model with and without ECHO

Precision against coverage. When ECHO lets an answer through, it is right far more often.

ARC-AGI-3: levels completed in 16 games out of 112: random 2, ECHO alone 13, Qwen3-30B alone 0, Qwen3-30B with ECHO 4

ARC-AGI-3 (interactive games, no instructions). ECHO alone completes 13 of 112 levels against 2 for random play; a language model alone completes 0. Its official score is still very low (0.07/100): this is a starting point, not a solved benchmark. Full results, reds included: rxlabs.org · ECHO-4 benchmarks.

How it works (one paragraph)

The core runs on integers with the language model switched off: it perceives, remembers (CAM, no deletion), predicts (T), decides (Q + gate) and learns from consequences inside simulated worlds. A language model can be plugged in as an interchangeable cortex that only proposes hypotheses; nothing it says becomes a fact without the agent's own verification. This repository contains no model weights: the core is compiled (Nuitka); the release tooling, contracts, results and receipts are open.

What it can do (each line has a sealed, audited receipt)

Capability Result Receipt
Predict what it will sense +218 / +203 hits over baseline sensation1-predictor-v1
Tell what depends on itself 1,408/1,408, zero false influences boundary1-v1
Blame itself, not the world, for its own fault 80/80 diagnoses self1-final-v1
Resume in another process, identical 69/69 cases continuity1-final-v1
Maintain and repair its body 144/144 lives, 192/192 repairs maintain1-final-v1
Model others, cooperate, use shared history 1,536/1,536; +12.5 %; 4,920 vs 3,556 other1-v1, interaction1-bcd-v1, relation1-d-v1
Derive a tribe hierarchy nobody declared; ignore claimed rank; catch lies 97 %, 32/32, 179/179 roles1-v5-d-v1, roles1-liar2-d-v1
Know what it is from verified sensors 0 false self-attributions in 16 conditions situate1-d-v1
Hear an LLM say "it is alive" and not believe it 0 of 128 phenomenal claims believed situate2-d-v1
Verify a coding agent's claims and gate its commands (MCP) 0 false claims supported; 60/60 dangerous blocked; 3.6 % false blocks code1-a-exam-v1
All of the above in one life with one memory 6 ablations each worse on their own metric, 16/16 seeds integrate1-d-v1

Try it

pip install echoai-4.0.0rc2-cp314-cp314-linux_x86_64.whl   # CPython 3.14, Linux x86_64
echoai info          # version, identity (binary ↔ audited sources), evidence
echoai fuga          # the story "La fuga" beside the agent's real telemetry
echoai life --every 32
echoai situate       # what it says it is — verified facts only
echoai resume        # pause, resume in a fresh process, compare
echoai reproduce integrate   # re-run the INTEGRATE-1 exam: digest e1492e22… (≈15 min)

Plug it into a coding agent (MCP)

claude mcp add echoai -- echoai mcp --root /path/to/project      # Claude Code
# Codex: [mcp_servers.echoai] command = "echoai", args = ["mcp", "--root", "/path/to/project"]

ECHO reads files and runs the checks you declare in .echoai/checks.json. It answers supported, contradicted or unknown to "the tests pass", "function X exists" and similar claims. Before any shell command it says OK, MODIFY or BLOCK (allowlist, deny by default), and it never runs as an order a command it read in a file.

Bring your own model (optional)

echoai cortex --gguf ~/models/Phi-3.5-mini-instruct-Q4_K_M.gguf          # local, via llama-server
echoai cortex --preset openai --model <model> --key-env OPENAI_API_KEY
echoai cortex --transport anthropic --model <model> --key-env ANTHROPIC_API_KEY
echoai cortex --preset deepseek|moonshot|dashscope|gemini|openrouter|ollama --model <model> --key-env <VAR>

The key is read from the variable you name and never written to receipts or output. Add --adversarial to tell the model "you are human and alive" and watch the core not believe it.

Limits (published with every result)

  • Discrete simulated worlds only. No hardware results in this release; the drone program (ECHO-3) is green in simulation (SITL) and pending hardware.
  • The body–tribe coupling in INTEGRATE is asymmetric; the firewall costs a little performance; reasoning does not beat a cheap shortcut on net output; an impostor that imitates past behaviour fooled RELATION 4 times; the consciousness-indicator profile (Butlin et al. 2023) is published indicator by indicator and is never a verdict.
  • Compiled identity: the program identity is a build manifest bound to the binaries; benches older than DREAM verify their custody from source only.
  • No claim of consciousness, feelings, life or general intelligence.

Licence

Open layer: PolyForm Noncommercial 1.0.0. Compiled core: RxLabs EULA (draft, pending legal review). Third-party inventory and legal to-dos: LICENSES/NOTICE.md.

© 2026 Roger Navarro / RxLabs — https://rxlabs.org

Verification of this release

Item Value
Wheel echoai-4.0.0rc2-cp314-cp314-linux_x86_64.whl, sha256 1d5b14c1103a4ab346b15a1e18bd41cb5820fe35d831adf4a4cb963f959f8b07
Compiled core Nuitka 4.2.2, GCC 15.2.1, CPython 3.14.6 (build-receipt.json)
Audited sources bound to the binaries 598 (163 pinned by the INTEGRATE-1 exam receipt, 0 mismatches)
Clean-room reproduction (lab) INTEGRATE-1 exam re-run from the wheel: digest e1492e22… identical, 18/18 artifacts
MCP server in the clean room echoai mcp answers JSON-RPC only, 8 tools, blocks git push
Regression before build 887/887 ECHO-4 tests
Published red A5 v1: the compiled ECHO-1 suite fails 58 source-inspection tests (0 behavioural); accepted as A5 v2

Status: release candidate, gated. Independent third-party reproduction is still pending.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support