YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Autonomous Innovation Engine

A 3-layer multi-agent system that generates, adversarially evaluates, and refines innovation concepts — and gets smarter with every run.

Layer 1 · Generation   3 divergent generator personas (Pragmatist · Frontier Researcher · Contrarian)
Layer 2 · Evaluation   4-critic adversarial council (Feasibility · Novelty · Market · Red Team)
Layer 3 · Memory Plane persistent outcome record → calibrates critics, detects duplicates,
                       makes reruns idempotent, synced to a Hub dataset across runs

How it works

  • Generation — each generator persona proposes or enhances ideas as structured JSON (name, domain, pitch, mechanism, differentiator), rotating personas per seed.
  • Evaluation — the council scores every idea 1–10 on four weighted dimensions; the Red Team Executioner attacks each idea and scores how well it survives. Ideas above the acceptance threshold (aggregate ≥ 7.0, red-team ≥ 6.0) are accepted; the top ideas are refined against the critics' stated weaknesses and re-scored.
  • Memory Plane — every idea, score, weakness and decision is recorded to memory_plane.jsonl and synced to LoveLogicAI/innovation-engine-memory. New runs download it and calibrate against per-domain baselines, acceptance history, and the best/worst past ideas. Seeds already enhanced are skipped automatically.

Results in this repo

  • results/summary.json — the demo run: 3 self-generated rounds (24 ideas).
  • results/enhanced_summary.json — 14 think-tank seed concepts enhanced and scored by the council (two runs; see results/enhanced_ideas.jsonl for full records).
  • Dashboard: autonomous-innovation-engine-trackio

Running it

python engine_demo.py --dry-run    # full loop with a mock LLM (no GPU needed)
python engine_demo.py --smoke      # 1 tiny round on real hardware
python engine_demo.py              # full demo: 3 rounds
ENGINE_MODE=enhance ENGINE_PUSH=1 python engine_demo.py   # enhance the seed concepts

Configured for Qwen/Qwen3-1.7B (Apache-2.0). Dependencies (current at time of run): transformers 5.19.0, torch 2.14.1, accelerate 1.15.0, trackio 0.41.0.

Findings so far

  • The critic council's re-scoring of refined ideas is genuinely harder than first-pass scoring (top-3 refinements dropped from 8.0 to 7.25–7.75 in run 1).
  • The 1.7B model saturates at the top of the score scale on strong ideas (clusters of 8.0); a larger critic model would sharpen discrimination.
  • Batching multiple ideas into one generation call is unreliable at this scale — per-seed calls with memory-plane idempotency are the robust pattern.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support