KALYPSO v1.2L

GENOMA Labs' public coding + planning model. Qwen3-Coder-30B-A3B-Instruct (30.5B total / 3.3B active MoE) with a GENOMA failure-handling LoRA (r16 on attention + expert MLPs). Successor to kalypso-v1.1l.

Why upgrade from v1.1L

Identical harness, identical budgets, all rows computed (never narrated):

Benchmark (GENOMA harness) v1.1L (14B dense) v1.2L (30B-A3B)
Coding โ€” v4_hard pass@1, hidden-test sandbox (41 tasks) 57.7% (N=3) 73.2% (N=2)
Plan-grade orchestration (30 tasks, calibrated LLM-judge) 0.832 0.848
Execution-grade orchestration (failure-handling outcome judge) 0.437 0.482
MMLU-Pro (n=300) 41.0% 50.3%
Active parameters per token 14B 3.3B (~4ร— cheaper inference)

Better on every axis at a fraction of the serving cost. 32k context (base-native), Apache-2.0.

What the fine-tune is (and honest limits)

  • Corpus (800 examples, fully public-safe): 320 failure-handling dialogues + 160 orchestration plans with explicit verification gates and rollback/compensation steps + 320 coding-hold examples. All teachers generic and local; zero proprietary content (audited).
  • Honest note: the uplift claim in the table above is vs KALYPSO v1.1L, not vs the raw Qwen base. Against its own base under identical settings, this fine-tune measures โˆ’7pp on coding pass@1 (base 82.9% vs 73.2%, N=2 each) and parity on planning/execution/knowledge. If raw one-shot coding is your only criterion, the unmodified base is the stronger pick; v1.2L remains far ahead of v1.1L on every axis. We publish all numbers as measured.
  • Execution-grade planning is an open weakness industry-wide โ€” every model we measured (ours included, plus our previous public model) scores < 0.5 on "does the plan actually handle the failure mode (rollback / compensation / gate-before-prod)". We release our benchmark methodology so the community can push on this.

Eval methodology

  • Coding: the model must emit implementation + tests as fenced files; hidden pytest tests are injected and run in a sandbox; pass = tests green. Single-shot (pass@1), temperature 0.
  • Plan-grade: 30 orchestration scenarios; a calibrated LLM-judge scores semantic plan quality (judge discrimination margin verified before use).
  • Execution-grade: an outcome judge mentally executes each plan and scores goal-achievement (0โ€“60) + failure-mode handling (0โ€“40); calibrated to penalize happy-path-only plans. All models graded by the same judge.
  • N-run where stated; single-run rows are marked by their N.

Serving

  • Fits a single 24 GB GPU in 4-bit (~19 GB); vLLM/llama.cpp/ollama supported (GGUF provided).
  • Recommended: temperature 0โ€“0.2 for coding/planning.

Provenance & license

Base: Qwen/Qwen3-Coder-30B-A3B-Instruct (Apache-2.0). Fine-tune: GENOMA Labs, Apache-2.0. No robotics, CAN-bus, or proprietary GENOMA data in the training corpus (audited pre-release).

GENOMA Labs โ€” 2026-08.

Downloads last month
-
GGUF
Model size
31B params
Architecture
qwen3moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for genomalabs/kalypso-v1.2l

Quantized
(164)
this model