🧠 imagine-v11

100% on held-out. Predicate placement fixed. Analytical SQL, finally real.

Natural language in. Correct PostgreSQL out. ~1B parameters. No GPU. No API key. No metered inference.

eval gate protocol hallucinations runtime licence

Part of project imagine — Interchained


The failure that built v11

v10 was asked, live:

"Every customer with total spending on completed orders, include zeros."

It returned:

SELECT c.id, c.name, COALESCE(SUM(o.total), 0) AS total_spending
FROM customers c
LEFT JOIN orders o ON c.id = o.customer_id
WHERE o.status = 'completed'   -- ← the bug. WHERE runs AFTER the join.
GROUP BY c.id, c.name
ORDER BY total_spending DESC;

The WHERE silently converts the outer join to an inner join. Every zero-spend customer — the exact rows the question asked for — vanishes. The predicate belongs in ON:

SELECT c.id, c.name, COALESCE(SUM(o.total), 0) AS total_spending
FROM customers c
LEFT JOIN orders o ON c.id = o.customer_id AND o.status = 'completed'
GROUP BY c.id, c.name
ORDER BY total_spending DESC, c.id ASC;

v10 knew the shape of the answer but not where the filter lives. v11 does.


🎯 What v11 adds

Where v10 added complex analytical queries, v11 adds predicate placement — knowing where each clause belongs:

  • ON-clause predicates — filters on the nullable side of an outer join live in ON, never WHERE
  • Conditional aggregation — SUM(CASE WHEN ...) patterns for filtered totals
  • Anti-join patterns — NOT EXISTS / LEFT JOIN ... IS NULL for absence queries
  • Analytical templates, actually working — v10's analytical set silently generated zero candidates (a child-detection heuristic that never matched plural table names); v11 fixes the generator and ships 18 real analytical pairs

📊 The numbers

Model Held-out (telemetry) Protocol Execute rate Invented cols
imagine-v8 50.0% (20/40) 100% — —
imagine-v9 90.0% (36/40) 100% 97.5% 0
imagine-v10 100.0% (40/40) 100% 100% 0
imagine-v11 100.0% (40/40) 100% 100% 0

Perfect score, held. Zero wrong answers. Zero hallucinations. Zero refusals.


🧬 How v11 was built

Two-stage fine-tuning from Interchained/imagine-v10, on an NVIDIA B200:

Stage 1 — Full fine-tune (smoke)

  • Corpus: 3,284 admitted pairs (99.85% gate admit rate)
  • New: Predicate-placement templates (ON-clause, conditional aggregation, anti-join) + working analytical templates
  • Epochs: 3, LR 1e-5, cosine schedule
  • Final loss: 0.0007 (v10 was 0.0052 — 7x better)
  • Time: 3.3 min on B200, 33k tok/s

Stage 2 — LoRA refinement

  • Base: v11-smoke checkpoint
  • Rank: 16, LR 5e-6
  • Epochs: 3
  • Final loss: 0.0000
  • Time: 4.2 min on B200
  • Adapter folded into full checkpoint

The corpus

3,289 candidates forged, 3,284 admitted:

Schema Templates Analytical Predicate Writes Total
shop 356 3 9 396 764
clinic 545 6 36 280 867
library 481 6 0 262 749
fleet 297 3 0 231 531
audit 252 0 0 126 378

Nothing enters training that a live database hasn't agreed with.


🔬 The execution gate

Every candidate goes through L0–L4:

candidate SQL ──► L0  real PostgreSQL parser     not a regex
                 L1  read-only + bounded          no writes, no sleeps
                 L2  EXPLAIN on live schema       hallucinations die HERE
                 L3  execute, timed, capped       real rows
                 L4  SAME ANSWER as ref?          ◄── the one that matters
                                              ▼
                                       admitted to corpus

68/68 gate tests passing. The gate is the truth authority — not a bigger model, not vibes.


📦 Output protocol

SQL wrapped in sentinel blocks:

<<<SQL>>>
SELECT c.id, c.name, COALESCE(SUM(o.total), 0) AS total_spending
FROM customers c
LEFT JOIN orders o ON c.id = o.customer_id AND o.status = 'completed'
GROUP BY c.id, c.name
ORDER BY total_spending DESC, c.id ASC;
<<<END>>>

Can also refuse:

block meaning
<<<SQL>>> here is your query
<<<UNANSWERABLE>>> this schema cannot answer that
<<<CLARIFY>>> ambiguous — here's what's missing

⚠️ A truncated generation is not an answer. An unterminated block extracts to nothing.


💻 Loading v11

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "Interchained/imagine-v11"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

# Ask for a predicate-placement query — v10's failure mode
prompt = """Schema:
customers(id, name)
orders(id, customer_id, total, status)

List every customer ID, name, and total completed spending.
Include customers with no completed orders (show 0).
Sort by spending descending, then ID ascending."""

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(inputs["input_ids"], max_new_tokens=256, do_sample=False)

print(tokenizer.decode(output[0], skip_special_tokens=True))

🎯 Identity

The deployed model identity is Imagine.

Built and fine-tuned by Interchained.

DeepSeek-Coder is part of the upstream lineage (via v8 → v9 → v10 → v11), but the deployed identity is Imagine.


⚡ Local-first

Runs on your hardware. No API key. No metered inference. No cloud dependency.


📐 The rules

1 · Don't write a verifier — the engine already shipped one. Real parser. Real planner. Real rows.

2 · Assert the property, not a proxy. Execution accuracy, not string similarity.

3 · Schema goes in the prompt, not in the weights. The model learns "read the schema you were handed" — not memorise ours.


⚠️ Limitations

v11 is a research checkpoint. It may still:

  • generate incorrect SQL on novel patterns
  • misunderstand ambiguous requests
  • produce writes with wrong WHERE clauses — always review before executing

Generated SQL should be reviewed before use in production. This applies doubly to writes.


🔒 Security

Do not rely on model behavior alone for database safety. Production systems should enforce:

  • least-privilege database roles
  • statement timeouts and row limits
  • query validation and schema restrictions
  • application-level authorization
  • audit logging
  • human approval for all write statements

Built by Interchained · ownership at every layer, including the model

3 > 1 — Mark drives, oracle points, Muse builds

Downloads last month
46
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Interchained/imagine-v11