The Failed Market Maker
24 models. Promising simulations. No established live edge.
An open research archive for Polymarket BTC UP/DOWN five-minute markets: original weights, every result, and the evidence behind the failed edge.
完整中文 · Paper results · All 24 models · Every live run · Download
GitHub · Hugging Face · Full results & sources
Download complete bundles: Browse weights and companion files. Requires Python, NumPy, PyTorch, ONNX and ONNX Runtime. Keep BC’s .onnx.data, normalization and feature contracts with the weights.
| Paper · R1 deployment | Live · latest attributable report | Final platform account PnL |
|---|---|---|
| +1,791.31 USDC | -18.8755 USDC | N/A |
| June 2 run · adjusted PnL | June 22 12:06 UTC · local conservative NAV change | All 13 reports unreconciled |
| Decisions stopped June 5; not seven full days | Includes unrealized inventory | No verified cumulative account return |
Paper and live figures come from different runs and accounting methods. They are not a before/after return or a cumulative profit. Display values are rounded; linked reports retain original precision. On a narrow screen, scroll wide tables horizontally.
Every paper run, side by side.
Five independent runs. All reported gains and losses. Amounts in USDC; dates in 2026 UTC.
| Run / duration | Metric | M3 B1 | M3 B2 | M5 BC | R1 Alpha + Policy |
|---|---|---|---|---|---|
| 05.17 · 02:21 6.22 h · stopped early |
NAV change | +100.05 | -99.06 | -99.55 | N/A · not included |
| 05.17 · 22:34 39.30 h · stopped early |
Adjusted PnL | -55.79 | +78.42 | -47.72 | N/A · not included |
| 05.19 → 05.22 72.07 h |
Adjusted PnL | -104.69 | +43.09 | -0.45 | N/A · not included |
| 05.27 → 05.30 71.98 h |
Adjusted PnL | -127.80 | -100.67 | +20.20 | +1,114.97 |
| 06.02 → 06.05 Decisions stopped June 5 |
Adjusted PnL | -298.04 | -224.99 | +84.78 | +1,791.31 |
Read row by row: the first run uses NAV change; the other four use adjusted PnL, which includes closed risk segments and the current segment. Capital top-ups are not profit. Never sum these runs into a lifetime return.
June 2 was not a completed seven-day test. B1/B2 last decisions: June 5 11:37:05 / 11:44:01 UTC; BC and R1: 18:09:59. Process termination: June 7. B1/B2/BC used signal outputs plus external execution and risk rules.
Simulated fills, queueing and rebates depend on the original assumptions; inspect the full results and sensitivity evidence.
One scorecard for every model.
Every retained training candidate, grouped by family. Short names link to exact model IDs, files and evidence. Different datasets, targets and metrics are not one global ranking.
M3 · 3 · M5 BC · 1 · IQL · 3 · Alpha · 8 · Policy · 7 + 2
M3 · supervised signals
| Model | Best epoch | Train loss | Val AUC | Test AUC | Test top-decile EV |
|---|---|---|---|---|---|
| M3 run3 | 1 | 1.941012 | 0.773355 | 0.743551 | +0.245144 |
| M3 B1 | 13 | 0.277852 | 0.735678 | 0.713834 | -0.105901 |
| M3 B2 | 12 | 0.299922 | 0.735113 | 0.713882 | -0.086622 |
The original supervised checkpoint uses 65 inputs; B1/B2 collection adaptations use 76. Fixed-execution test PnL: B1 +92.2782; B2 +137.8349 USDC. These are not paper results, nor scores for newly added adaptation heads. Run3 paper/live: N/A — no corresponding report; B1/B2 live: N/A — no corresponding run evidence.
M5 BC · behavior cloning
| Model | Best epoch | Train accuracy | Val accuracy | Test accuracy |
|---|---|---|---|---|
| M5 BC | 1 | 0.848722 | 0.710701 | 0.682590 |
Final epoch 4: validation accuracy 0.703286; test N/A — not independently evaluated. Paper alias: m5_baseline. Live N/A — no corresponding run evidence.
M5 IQL · fixed versus native execution
| Model | Best epoch | Policy loss | Val Q loss | Hold best/final | Fixed return | Native return |
|---|---|---|---|---|---|---|
| IQL-0 | 46 | 4.649556 | 0.029637 | 92.60% / 93.85% | +4.9119% | +0.0221% |
| IQL-1 | 35 | 4.590885 | 0.028484 | 85.25% / 89.20% | +2.9932% | -2.1677% |
| IQL-2 | 43 | 4.635094 | 0.035254 | 94.25% / 95.30% | +5.2949% | -2.1522% |
Returns are mean window returns over 24 offline validation windows. ONNX uses the best checkpoint; final checkpoints are epoch 50. Independent test, paper and live: N/A — no corresponding independent report.
Alpha V2 · two AUC definitions, both visible
| Model | Best epoch | Train loss | val.auc | true_best_side_auc | Val selected-side top-decile EV |
|---|---|---|---|---|---|
| R1 A3 · s2 | 1 | 1.416982 | 0.868495 | 0.595098 | +0.065216 |
| R1 A3 · s3 | 2 | 1.352764 | 0.861482 | 0.571024 | +0.138491 |
| R2 A1 · s2 | 4 | 1.195743 | 0.872110 | 0.667253 | +0.042430 |
| R3 A6 · s2 | 3 | 1.185670 | 0.866127 | 0.665461 | +0.013070 |
| R4 LB2 · s1 | 2 | 1.138801 | 0.853418 | 0.649595 | +0.093210 |
| R4 LB2 · s2 | 1 | 1.236778 | 0.871611 | 0.674253 | +0.047972 |
| R5 A4 · s3 | 1 | 2.359265 | 0.867386 | 0.590465 | +0.026660 |
| R6 C2 · s2 | 3 | 0.714006 | 0.864232 | 0.776629 | +0.068376 |
Independent test N/A — not executed. Standalone paper/live N/A — no separate strategy report. R1 s2 participates in the deployment; deployment PnL is not copied to Alpha.
8 more Alpha experiments: metrics retained, weights missing
| Experiment | val.auc | true_best_side_auc | Val selected-side top-decile EV | Weights |
|---|---|---|---|---|
| R2 A1 · s1 | 0.845420 | 0.632371 | -0.071830 | N/A · not retained |
| R3 A6 · s1 | 0.835604 | 0.631378 | -0.000114 | N/A · not retained |
| R2 A1 · s3 | 0.843929 | 0.620569 | -0.055705 | N/A · not retained |
| R5 A4 · s2 | 0.805230 | 0.561291 | -0.089587 | N/A · not retained |
| R5 A4 · s1 | 0.795446 | 0.554888 | -0.229608 | N/A · not retained |
| R4 LB2 · s3 | 0.782713 | 0.587704 | -0.121413 | N/A · not retained |
| R1 A3 · s1 | 0.806882 | 0.554470 | -0.142231 | N/A · not retained |
| R3 A6 · s3 | 0.831988 | 0.626608 | -0.083801 | N/A · not retained |
MakerPolicy · 7 candidates + 2 repair/control runs
| Model | Phase | Last BC loss | Last RL loss | 263-window sum | 30-window sum |
|---|---|---|---|---|---|
| R1 A3 · s2 | 6.5 | 0.000110815 | 0.000076921 | +46,401.53 | +5,336.60 |
| R1 A3 · s3 | 6.5 | 0.000117271 | 0.000079444 | +46,360.58 | N/A · not independently tested |
| R2 A1 · s2 | 6.5 | 0.000113844 | 0.000065989 | +46,367.90 | N/A · not independently tested |
| R3 A6 · s2 | 6.5 | 0.000113422 | 0.000071329 | +46,373.30 | +5,334.89 |
| R4 LB2 · s1 | 6.5 | 0.000149778 | 0.000161845 | +46,356.43 | N/A · not independently tested |
| R4 LB2 · s2 | 6.5 | 0.000143163 | 0.000158601 | +46,366.36 | N/A · not independently tested |
| R6 C2 · s2 | 6.5 | 0.000122122 | 0.000061272 | +46,395.70 | +5,332.34 |
| 6.7 C0 | 6.7 | 0.156779423 | 0.001588044 | +45,844.68 | N/A · not independently tested |
| 6.7 P1 | 6.7 | 0.576368110 | 0.303230513 | +6,867.81 | N/A · not independently tested |
Positive replay PnL did not establish maker edge. Phase 6.5/6.6: FAIL_NO_MAKER_EDGE. Phase 6.7: FAIL_CONTROL_AMPLIFICATION. The sums are over independent windows, not a continuous account. Independent val/test loss: N/A — no report. Phase 6.7 independent test: N/A — not started after failure.
Policy standalone paper/live N/A — no corresponding standalone report. Only R1 s2 has a retained complete historical Alpha + Policy ONNX pair; six other Phase 6.5 policies have weights but lack the paired historical Alpha ONNX.
One deployment: Alpha R1 A3 s2 + Policy R1 A3 s2. It owns the paper/live records on this page; it is not a 25th trained candidate.
What the live ledger actually recorded.
13 live reports for the R1 deployment: 6 observable local NAV changes, all 13 final platform results unreconciled. Original local-ledger figures; no new account reconciliation. All times are UTC in 2026.
Every live run · USDC
| UTC run | Local NAV change | Realized only | Final platform PnL |
|---|---|---|---|
| 06.18 19:42:41 | +0.290000 | +0.050000 | N/A · not reconciled |
| 06.18 21:12:30 | -5.500000 | +0.000000 | N/A · not reconciled |
| 06.18 21:57:47 | +0.250000 | +0.000000 | N/A · not reconciled |
| 06.18 23:02:37 | N/A · unobservable | +0.000000 | N/A · not reconciled |
| 06.19 00:22:00 | N/A · unobservable | +0.000000 | N/A · not reconciled |
| 06.19 02:00:53 | N/A · unobservable | +0.000000 | N/A · not reconciled |
| 06.19 09:28:50 | N/A · unobservable | -0.316667 | N/A · not reconciled |
| 06.21 02:24:02 | +0.000000 | +0.000000 | N/A · not reconciled |
| 06.21 06:16:51 | N/A · unobservable | +0.000000 | N/A · not reconciled |
| 06.21 13:18:04 | N/A · unobservable | -2.929107 | N/A · not reconciled |
| 06.21 15:08:00 | N/A · unobservable | +0.000000 | N/A · not reconciled |
| 06.22 08:26:24 | -8.505000 | -0.329808 | N/A · not reconciled |
| 06.22 12:06:33 | -18.875455 | -4.150730 | N/A · not reconciled |
June 21 02:24:02 had no orders or fills. A zero realized result does not imply zero total PnL when NAV is unobservable. Never sum these runs or infer ROI, maximum drawdown or final equity.
The experiments behind the numbers.
M1 / M2 → M3 → M5 BC / IQL → Alpha + Policy → 6.7
M1/M2 are replay and feature infrastructure, not trained models. M3 A6 and M5 BC-6 weights were not retained; BC-6 test PnL −11.0374 USDC remains part of the failure record. The newer architecture has no formal trained release here.
Earlier live runtime, SDK unit errors, startup failures and platform probes remain in the full results. They are not omitted or relabeled as strategy profits.
Read every stage and limitation →
Explore the models. Revisit the evidence.
Original weights, normalization, feature contracts and a minimal inference example. Keep companion files together: M5 BC requires its .onnx.data file.
Choose a bundle in the model index, download its complete directory from Hugging Face, and follow the inference guide. Use Python with NumPy, PyTorch, ONNX and ONNX Runtime; see requirements.txt.
The v1.0.0 GitHub Release is a historical full archive, including older documentation. To add its model files to a current checkout, extract models/ while excluding its historical model cards (archive-relative paths):
tar -xzf failed-market-maker-v1.0.0.tar.gz --exclude='models/*/README.md' models/
python -m inference.run models/m3_run3/best_model.pt --family m3 --smoke
python -m inference.run models/m5_v7v72_bc/agent_model_m5.onnx --smoke
python -m inference.run models/mpv2_r1_a3_seq120_seed2 --smoke
Alternatively unpack the historical full archive into a separate directory. Synthetic forward checks demonstrate executability only; they do not reproduce historical training or trading performance. Keep normalization, feature contracts and configurations with every model; BC also requires agent_model_m5.onnx.data.
Development story · Model cards · Provenance & redaction · Unified file manifest
Original code and weights: Apache-2.0. Curated documents and reports: CC-BY-4.0. See NOTICE for attribution. Raw market datasets are not distributed. Historical evidence is a source record, not current deployment instructions.