WQT20M-Beta

WestQuant Transformer 20M β€” Quantum Representation Scheduler

AI schedules. Deterministic mathematics executes. Independent verification certifies.

WQT20M-Beta is a 19M-parameter Transformer that ranks quantum transformations. Given a quantum optimization state, it predicts which transformation is most promising. It does not execute transformations β€” deterministic engines do that.

WQT20M-Beta is a proof-of-concept that a 19M-parameter Transformer can learn structured quantum optimization preferences from synthetic data. Researchers building better compiler-optimization models can use this as a baseline.

This is WestQuant Open's first step toward a production-ready transformer with calibrated value predictions (wider range, objective sensitivity, real QPU training data) that could replace heuristic transpiler defaults in large complex Quantum Computing production pipelines.


What It Does β€” 3 Tested Examples

Example 1: Preference Comparison (96% accuracy)

Given two candidate transformations with their costs, the model predicts which is better:

State:    <DOMAIN:quantum_annealing> <LEVEL:ISING> <N_QUBITS:18> ...
          <RES:n_q=18 D=8 G1=12 G2=4 T=0 M=3 A=0 E=0.0100 C=1.0000>
          <OBJ_TYPE:balanced>

Candidate A: QUENCH          (cost 1278.25)
Candidate B: SET_BIAS        (cost 1245.57)

Model predicts: B>A   (SET_BIAS is better)
Correct answer: B>A   βœ“

Example 2: Value Prediction (mean error ~8)

Given a state, the model predicts the cost-to-go (remaining optimization cost):

State:    <DOMAIN:graph_optimization> <LEVEL:GRAPH> <N_NODES:21> ...
          <RES:n_q=21 D=5 G1=10 G2=3 T=0 M=2 A=0 E=0.0100 C=1.0000>
          <OBJ_TYPE:balanced>

Model predicts cost-to-go: 1275.93
Actual cost-to-go:         1264.63
Error:                     11.30 (0.9%)

Example 3: Ranked Policy (73% Top-1 via value ranking)

Given a state and all legal candidate actions, the model ranks them by predicted cost and picks the best:

State:    <DOMAIN:hardware_mapping> <LEVEL:COMPILED> <N_QUBITS:12> ...
          <RES:n_q=12 D=6 G1=8 G2=2 T=0 M=1 A=0 E=0.0050 C=1.0000>
          <OBJ_TYPE:2q_focused>

Candidates ranked by predicted cost:
  1. NOISE_AWARE        predicted=620.83   ← model picks this
  2. LAYOUT_SCORE       predicted=631.83
  3. DENSE_PLACE        predicted=639.83

Oracle (actual best):   NOISE_AWARE        βœ“ Correct!

Quick Start

pip install transformers torch
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained("WestQuantStudio/WQT20M-Beta")
tokenizer = AutoTokenizer.from_pretrained("WestQuantStudio/WQT20M-Beta")
device = "mps" if torch.backends.mps.is_available() else "cpu"
model = model.to(device).eval()

text = ("<SOLVE> <DOMAIN:quantum_annealing> <LEVEL:ISING> "
        "<N_QUBITS:18> <N_GROUND:14> <COUPLING_STRENGTH:0.8000> "
        "<RES:n_q=18 D=8 G1=12 G2=4 T=0 M=3 A=0 E=0.0100 C=1.0000> "
        "<OBJ_TYPE:balanced> "
        "<CAND_A> QUENCH <COST_A> 1278.25 "
        "<CAND_B> SET_BIAS <COST_B> 1245.57 "
        "<PREF>")

ids = tokenizer.encode(text, add_special_tokens=False, return_tensors="pt").to(device)
with torch.no_grad():
    for _ in range(5):
        out = model(input_ids=ids)
        nxt = out.logits[0, -1].argmax().unsqueeze(0).unsqueeze(0)
        ids = torch.cat([ids, nxt], dim=1)
        if nxt.item() == tokenizer.eos_token_id:
            break

print(tokenizer.decode(ids[0].tolist()).split("<PREF>")[-1].strip())
# β†’ "B>A" (SET_BIAS is better because cost 1245 < 1278)

What WQT20M-Beta Predicts

Task Input Output Accuracy
Preference State + 2 candidates with costs Which candidate is better 96.1%
Value State Predicted cost-to-go Spearman 0.98, mean error ~8
Ranked Policy State + all legal actions Best action (via value ranking) 73% Top-1
Legality State + action Is this action legal? 76.9%
Hardware State + backend Is this feasible? 91.2%

Model

Config Value
Architecture Llama-style decoder Transformer
Parameters 19.06M
Layers 8
Hidden size 384
Attention heads 6 (2 KV heads, GQA)
FFN 1536 (SwiGLU)
Context length 2048
Vocab 4561 (quantum-native structured tokens + BPE)
Model size 76 MB

Standard HuggingFace LlamaForCausalLM. Compatible with AutoModelForCausalLM, AutoTokenizer, and SafeTensors.


Validation

All 6 release gates passed:

Gate Result
Data integrity PASS
Structural learning (no cheating) PASS
Search improvement over random PASS (+35% at budget=100)
Generalization to unseen problems PASS (96%)
No regression PASS
Reproduction PASS

Anti-cheating

The model reads structure, not identifiers:

  • ID-only accuracy: 24% (domain token alone can't predict the answer)
  • Shuffled structure: 74% (scrambling the structure hurts performance)

Search improvement

Search budget Random WQT20M-Beta Improvement
10 evals baseline β€” +3.5%
50 evals baseline β€” +17.5%
100 evals baseline β€” +35.0%

QFT & Grover experiments

Independent experiments on QFT and Grover circuits across 18 configurations:

  • 0.0 regret in all 18 configs β€” the model consistently picked CANCEL, the best action
  • CANCEL ranked at position 1/15 (median model rank of true-best action)
  • Top-3 rate: 63% β€” the true-best action is in the model's top-3 63% of the time
  • CALL_PYZX is always correctly identified as the worst action

Honest Assessment

Where WQT20M-Beta has narrow real-world value

The core limitation: exhaustive search over 15 transpiler configs takes <0.01s per circuit. Any user who can afford to run Qiskit at all can afford to try all candidates and pick the best. In experiments, many configurations were ties (model and random both found the optimum), and the model did not consistently beat random on small search spaces.

Where it would add genuine value

1. Scaling to large search spaces where exhaustive is infeasible

If the candidate space grows to hundreds or thousands of configurations (e.g., combining routing methods Γ— layout methods Γ— synthesis engines Γ— basis gates Γ— opt levels Γ— seeds), exhaustive evaluation becomes expensive. The model's value-ranking approach becomes meaningful when each evaluation costs real QPU time or long simulation. A user compiling 10,000 circuits Γ— 500 candidate configs would benefit from model-guided pruning.

2. Multi-step optimization pipelines

The model predicts cost-to-go, not just immediate cost. In a sequential optimization pipeline (apply transformation A, then B, then C), the model could guide intermediate decisions where the full tree is exponentially large. This is the model's actual design intent β€” it's a scheduler, not a single-shot selector.

3. Researchers studying AI-guided quantum compilation

The model is a research artifact. Its value is as a proof-of-concept that a 19M-parameter Transformer can learn structured quantum optimization preferences from synthetic data. Researchers building better compiler-optimization models would use this as a baseline.

4. WestQuant iterating toward a production model

The Beta is explicitly a stepping stone. A future version with calibrated value predictions (wider range, objective sensitivity, real QPU training data) could replace heuristic transpiler defaults in production pipelines.

Who would NOT benefit

  • Practitioners running small circuits β€” exhaustive search is free and optimal
  • Users needing guaranteed optimality β€” the model has nonzero regret and 76.9% legality accuracy
  • Anyone with objective-dependent optimization β€” the model has 0% objective sensitivity
  • Users on real QPUs today β€” the model was trained on synthetic data, not hardware

Bottom line

The model adds real value only when the candidate search space is large enough that exhaustive evaluation is expensive, and when approximate guidance is acceptable. In its current Beta form, that threshold is well above what Qiskit's built-in transpiler exposes. The honest framing is: this is a research prototype demonstrating that learned quantum-structure preferences are feasible, not a production tool that outperforms brute force on practical workloads.


Roadmap to Production

Gap Current (Beta) Production Target
Training data Synthetic surrogate Real Qiskit/TKET transpilation outputs
Value calibration Narrow range (6.0–6.6) Actual cost magnitudes (50–25,000)
Objective sensitivity 0% Changes ranking with objective
Legality 76.9% >95%
Optimization steps Single-step Multi-step trajectory planning
Backend awareness Topology tokens only Per-edge error rates, per-qubit T1/T2
Model size 19M 100–200M
Feedback loop None Active learning from user compilations
Deployment Standalone script Qiskit transpiler plugin with budget parameter

Key changes needed

1. Real compiler outputs, not synthetic β€” Train on actual transpilation trajectories from Qiskit/TKET/PyZX across 10,000+ circuits from MQT Bench, QASMBench, and circuit libraries, with 20-50 real backend coupling maps.

2. Calibrated value prediction β€” Add a regression head that predicts cost in the actual cost range of the training data, not a narrow band.

3. Objective conditioning that works β€” Train with objective-dependent labels where the same (state, action) pair has different costs depending on the objective (depth-focused, 2q-focused, fidelity-focused, time-focused).

4. Multi-step trajectory planning β€” Train on real optimization trajectories with cumulative cost at each step, so the model learns which sequences are promising, not just which single action looks good.

5. Legality prediction β€” A binary classification head trained on real constraint data (which actions are valid for a given backend, circuit family, and optimization level). Target >95% accuracy.

6. Real backend calibration β€” Include actual per-edge error rates, per-qubit T1/T2, gate durations, and queue/wait times in the state encoding.

7. Larger model β€” 100-200M parameters, 12-16 layers, 768 hidden, 4096+ context length to encode full circuit metrics + backend calibration + action history.

8. Active learning β€” A feedback loop where the model improves from real user compilations: user compiles, model suggests, Qiskit applies, real cost is measured, (state, action, real_cost) is logged, model is fine-tuned.

9. Production deployment β€” A Qiskit transpiler plugin where the model ranks candidates, top-K are evaluated by real Qiskit, and the user specifies a budget parameter (how many real evaluations they can afford).

The model's architecture and approach (value-prediction-based ranking, quantum-native tokenization, structured state encoding) are sound. The gap is entirely in training data quality, calibration, and multi-step capability. Fix those three and it becomes a model someone would use in production.


Coverage

15 quantum optimization domains:

Domain Example Actions
Quantum chemistry UCCSD_ANSATZ, ADAPT_VQE, FROZEN_CORE
Graph optimization MAXCUT_ROUND, COLOR_GRAPH, SDP_RELAX
Error correction SYNDROME_MEASURE, DECODE_SURFACE, MAGIC_STATE_DISTILL
Quantum annealing REVERSE_ANNEAL, MINOR_EMBED, HYBRID_SOLVE
Variational QAOA_P1/P2/P3, WARM_START, ADAPTIVE_LAYER
Hamiltonian simulation TROTTER_STEP, LCU_DECOMPOSE, QUBITIZE
Circuit optimization CANCEL_GATES, ZX_SIMPLIFY, FUSE_ROTATIONS
Hardware mapping SABRE_ROUTE, PULSE_OPTIMIZE, NOISE_AWARE
Neutral atom SET_RYDBERG, PULSE_SHAPE, ADIABATIC_PASS
Photonic KLM_CNOT, CLUSTER_STATE, HERALD
Topological BRAID_ANYON, FIBONACCI_BRAID
Quantum walk COINED_WALK, GROVER_COIN, AMPLIFY
State preparation MPS_PREP, TENSOR_NETWORK, COMPRESS_STATE
Amplitude amplification GROVER_ITER, QSEARCH, ITERATIVE_QPE
Quantum ML IQP_KERNEL, DATA_REUPLOADING, FIDELITY_KERNEL

Limitations

This is a Beta release:

  • Policy Top-1 (generation): 0.8% β€” The model cannot reliably generate the best action name. However, ranked policy via value prediction achieves 73% Top-1 β€” use the value-ranking approach, not generation.
  • Legality: 76.9% β€” Below target. The model sometimes marks illegal actions as valid.
  • Objective sensitivity: 0% β€” The model does not yet change its predictions when the optimization objective changes.
  • Value calibration β€” Predicted values cluster in a narrow range (6.0–6.6), while actual costs range from 50 to 25,000. The model has good rank correlation (Spearman 0.98) but poor absolute calibration.
  • Not a circuit compiler β€” The model works on its own structured state representation, not on raw Qiskit/TKET circuits. The plugins translate between the model's representation and real frameworks.

License

Apache 2.0

Citation

@misc{westquant2026wqt20m,
  title={WQT20M-Beta: A 19M-Parameter Transformer for Quantum Representation Scheduling},
  author={Vesterlund, David},
  year={2026},
  url={https://huggingface.co/WestQuantStudio/WQT20M-Beta},
  note={WestQuant Open Source Project}
}
Downloads last month
242
Safetensors
Model size
19.1M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support