YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Autonomous Innovation Engine
A 3-layer multi-agent system that generates, adversarially evaluates, and refines innovation concepts — and gets smarter with every run.
Layer 1 · Generation 3 divergent generator personas (Pragmatist · Frontier Researcher · Contrarian)
Layer 2 · Evaluation 4-critic adversarial council (Feasibility · Novelty · Market · Red Team)
Layer 3 · Memory Plane persistent outcome record → calibrates critics, detects duplicates,
makes reruns idempotent, synced to a Hub dataset across runs
How it works
- Generation — each generator persona proposes or enhances ideas as structured JSON (name, domain, pitch, mechanism, differentiator), rotating personas per seed.
- Evaluation — the council scores every idea 1–10 on four weighted dimensions; the Red Team Executioner attacks each idea and scores how well it survives. Ideas above the acceptance threshold (aggregate ≥ 7.0, red-team ≥ 6.0) are accepted; the top ideas are refined against the critics' stated weaknesses and re-scored.
- Memory Plane — every idea, score, weakness and decision is recorded to
memory_plane.jsonland synced to LoveLogicAI/innovation-engine-memory. New runs download it and calibrate against per-domain baselines, acceptance history, and the best/worst past ideas. Seeds already enhanced are skipped automatically.
Results in this repo
results/summary.json— the demo run: 3 self-generated rounds (24 ideas).results/enhanced_summary.json— 14 think-tank seed concepts enhanced and scored by the council (two runs; seeresults/enhanced_ideas.jsonlfor full records).- Dashboard: autonomous-innovation-engine-trackio
Running it
python engine_demo.py --dry-run # full loop with a mock LLM (no GPU needed)
python engine_demo.py --smoke # 1 tiny round on real hardware
python engine_demo.py # full demo: 3 rounds
ENGINE_MODE=enhance ENGINE_PUSH=1 python engine_demo.py # enhance the seed concepts
Configured for Qwen/Qwen3-1.7B (Apache-2.0). Dependencies (current at time of run):
transformers 5.19.0, torch 2.14.1, accelerate 1.15.0, trackio 0.41.0.
Findings so far
- The critic council's re-scoring of refined ideas is genuinely harder than first-pass scoring (top-3 refinements dropped from 8.0 to 7.25–7.75 in run 1).
- The 1.7B model saturates at the top of the score scale on strong ideas (clusters of 8.0); a larger critic model would sharpen discrimination.
- Batching multiple ideas into one generation call is unreliable at this scale — per-seed calls with memory-plane idempotency are the robust pattern.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support