The Failed Market Maker

24 models. Promising simulations. No established live edge.

24 models. Promising simulations. No established live edge.

An open research archive for Polymarket BTC UP/DOWN five-minute markets: original weights, every result, and the evidence behind the failed edge.

完整中文 · Paper results · All 24 models · Every live run · Download

GitHub · Hugging Face · Full results & sources

Download complete bundles: Browse weights and companion files. Requires Python, NumPy, PyTorch, ONNX and ONNX Runtime. Keep BC’s .onnx.data, normalization and feature contracts with the weights.

Paper · R1 deployment Live · latest attributable report Final platform account PnL
+1,791.31 USDC -18.8755 USDC N/A
June 2 run · adjusted PnL June 22 12:06 UTC · local conservative NAV change All 13 reports unreconciled
Decisions stopped June 5; not seven full days Includes unrealized inventory No verified cumulative account return

Paper and live figures come from different runs and accounting methods. They are not a before/after return or a cumulative profit. Display values are rounded; linked reports retain original precision. On a narrow screen, scroll wide tables horizontally.

Every paper run, side by side.

Five independent runs. All reported gains and losses. Amounts in USDC; dates in 2026 UTC.

June 2 paper run, adjusted PnL for all four profiles; decisions stopped June 5

Run / duration Metric M3 B1 M3 B2 M5 BC R1 Alpha + Policy
05.17 · 02:21
6.22 h · stopped early
NAV change +100.05 -99.06 -99.55 N/A · not included
05.17 · 22:34
39.30 h · stopped early
Adjusted PnL -55.79 +78.42 -47.72 N/A · not included
05.19 → 05.22
72.07 h
Adjusted PnL -104.69 +43.09 -0.45 N/A · not included
05.27 → 05.30
71.98 h
Adjusted PnL -127.80 -100.67 +20.20 +1,114.97
06.02 → 06.05
Decisions stopped June 5
Adjusted PnL -298.04 -224.99 +84.78 +1,791.31

Read row by row: the first run uses NAV change; the other four use adjusted PnL, which includes closed risk segments and the current segment. Capital top-ups are not profit. Never sum these runs into a lifetime return.

June 2 was not a completed seven-day test. B1/B2 last decisions: June 5 11:37:05 / 11:44:01 UTC; BC and R1: 18:09:59. Process termination: June 7. B1/B2/BC used signal outputs plus external execution and risk rules.
Simulated fills, queueing and rebates depend on the original assumptions; inspect the full results and sensitivity evidence.

One scorecard for every model.

Every retained training candidate, grouped by family. Short names link to exact model IDs, files and evidence. Different datasets, targets and metrics are not one global ranking.

M3 · 3 · M5 BC · 1 · IQL · 3 · Alpha · 8 · Policy · 7 + 2

M3 · supervised signals

Model Best epoch Train loss Val AUC Test AUC Test top-decile EV
M3 run3 1 1.941012 0.773355 0.743551 +0.245144
M3 B1 13 0.277852 0.735678 0.713834 -0.105901
M3 B2 12 0.299922 0.735113 0.713882 -0.086622

The original supervised checkpoint uses 65 inputs; B1/B2 collection adaptations use 76. Fixed-execution test PnL: B1 +92.2782; B2 +137.8349 USDC. These are not paper results, nor scores for newly added adaptation heads. Run3 paper/live: N/A — no corresponding report; B1/B2 live: N/A — no corresponding run evidence.

M5 BC · behavior cloning

Model Best epoch Train accuracy Val accuracy Test accuracy
M5 BC 1 0.848722 0.710701 0.682590

Final epoch 4: validation accuracy 0.703286; test N/A — not independently evaluated. Paper alias: m5_baseline. Live N/A — no corresponding run evidence.

M5 IQL · fixed versus native execution

Model Best epoch Policy loss Val Q loss Hold best/final Fixed return Native return
IQL-0 46 4.649556 0.029637 92.60% / 93.85% +4.9119% +0.0221%
IQL-1 35 4.590885 0.028484 85.25% / 89.20% +2.9932% -2.1677%
IQL-2 43 4.635094 0.035254 94.25% / 95.30% +5.2949% -2.1522%

IQL: fixed versus native execution / 固定执行与原生执行

Returns are mean window returns over 24 offline validation windows. ONNX uses the best checkpoint; final checkpoints are epoch 50. Independent test, paper and live: N/A — no corresponding independent report.

Alpha V2 · two AUC definitions, both visible

Model Best epoch Train loss val.auc true_best_side_auc Val selected-side top-decile EV
R1 A3 · s2 1 1.416982 0.868495 0.595098 +0.065216
R1 A3 · s3 2 1.352764 0.861482 0.571024 +0.138491
R2 A1 · s2 4 1.195743 0.872110 0.667253 +0.042430
R3 A6 · s2 3 1.185670 0.866127 0.665461 +0.013070
R4 LB2 · s1 2 1.138801 0.853418 0.649595 +0.093210
R4 LB2 · s2 1 1.236778 0.871611 0.674253 +0.047972
R5 A4 · s3 1 2.359265 0.867386 0.590465 +0.026660
R6 C2 · s2 3 0.714006 0.864232 0.776629 +0.068376

Independent test N/A — not executed. Standalone paper/live N/A — no separate strategy report. R1 s2 participates in the deployment; deployment PnL is not copied to Alpha.

8 more Alpha experiments: metrics retained, weights missing
Experiment val.auc true_best_side_auc Val selected-side top-decile EV Weights
R2 A1 · s1 0.845420 0.632371 -0.071830 N/A · not retained
R3 A6 · s1 0.835604 0.631378 -0.000114 N/A · not retained
R2 A1 · s3 0.843929 0.620569 -0.055705 N/A · not retained
R5 A4 · s2 0.805230 0.561291 -0.089587 N/A · not retained
R5 A4 · s1 0.795446 0.554888 -0.229608 N/A · not retained
R4 LB2 · s3 0.782713 0.587704 -0.121413 N/A · not retained
R1 A3 · s1 0.806882 0.554470 -0.142231 N/A · not retained
R3 A6 · s3 0.831988 0.626608 -0.083801 N/A · not retained

All 16 original experiment records

MakerPolicy · 7 candidates + 2 repair/control runs

Model Phase Last BC loss Last RL loss 263-window sum 30-window sum
R1 A3 · s2 6.5 0.000110815 0.000076921 +46,401.53 +5,336.60
R1 A3 · s3 6.5 0.000117271 0.000079444 +46,360.58 N/A · not independently tested
R2 A1 · s2 6.5 0.000113844 0.000065989 +46,367.90 N/A · not independently tested
R3 A6 · s2 6.5 0.000113422 0.000071329 +46,373.30 +5,334.89
R4 LB2 · s1 6.5 0.000149778 0.000161845 +46,356.43 N/A · not independently tested
R4 LB2 · s2 6.5 0.000143163 0.000158601 +46,366.36 N/A · not independently tested
R6 C2 · s2 6.5 0.000122122 0.000061272 +46,395.70 +5,332.34
6.7 C0 6.7 0.156779423 0.001588044 +45,844.68 N/A · not independently tested
6.7 P1 6.7 0.576368110 0.303230513 +6,867.81 N/A · not independently tested

Positive replay PnL did not establish maker edge. Phase 6.5/6.6: FAIL_NO_MAKER_EDGE. Phase 6.7: FAIL_CONTROL_AMPLIFICATION. The sums are over independent windows, not a continuous account. Independent val/test loss: N/A — no report. Phase 6.7 independent test: N/A — not started after failure.

Policy standalone paper/live N/A — no corresponding standalone report. Only R1 s2 has a retained complete historical Alpha + Policy ONNX pair; six other Phase 6.5 policies have weights but lack the paired historical Alpha ONNX.

mpv2_r1_a3_seq120_seed2

One deployment: Alpha R1 A3 s2 + Policy R1 A3 s2. It owns the paper/live records on this page; it is not a 25th trained candidate.

What the live ledger actually recorded.

13 live reports for the R1 deployment: 6 observable local NAV changes, all 13 final platform results unreconciled. Original local-ledger figures; no new account reconciliation. All times are UTC in 2026.

Every live run · USDC
UTC run Local NAV change Realized only Final platform PnL
06.18 19:42:41 +0.290000 +0.050000 N/A · not reconciled
06.18 21:12:30 -5.500000 +0.000000 N/A · not reconciled
06.18 21:57:47 +0.250000 +0.000000 N/A · not reconciled
06.18 23:02:37 N/A · unobservable +0.000000 N/A · not reconciled
06.19 00:22:00 N/A · unobservable +0.000000 N/A · not reconciled
06.19 02:00:53 N/A · unobservable +0.000000 N/A · not reconciled
06.19 09:28:50 N/A · unobservable -0.316667 N/A · not reconciled
06.21 02:24:02 +0.000000 +0.000000 N/A · not reconciled
06.21 06:16:51 N/A · unobservable +0.000000 N/A · not reconciled
06.21 13:18:04 N/A · unobservable -2.929107 N/A · not reconciled
06.21 15:08:00 N/A · unobservable +0.000000 N/A · not reconciled
06.22 08:26:24 -8.505000 -0.329808 N/A · not reconciled
06.22 12:06:33 -18.875455 -4.150730 N/A · not reconciled

June 21 02:24:02 had no orders or fills. A zero realized result does not imply zero total PnL when NAV is unobservable. Never sum these runs or infer ROI, maximum drawdown or final equity.

The experiments behind the numbers.

M1 / M2 → M3 → M5 BC / IQL → Alpha + Policy → 6.7

M1/M2 are replay and feature infrastructure, not trained models. M3 A6 and M5 BC-6 weights were not retained; BC-6 test PnL −11.0374 USDC remains part of the failure record. The newer architecture has no formal trained release here.

Earlier live runtime, SDK unit errors, startup failures and platform probes remain in the full results. They are not omitted or relabeled as strategy profits.

Read every stage and limitation →

Explore the models. Revisit the evidence.

Original weights, normalization, feature contracts and a minimal inference example. Keep companion files together: M5 BC requires its .onnx.data file.

Choose a bundle in the model index, download its complete directory from Hugging Face, and follow the inference guide. Use Python with NumPy, PyTorch, ONNX and ONNX Runtime; see requirements.txt.

The v1.0.0 GitHub Release is a historical full archive, including older documentation. To add its model files to a current checkout, extract models/ while excluding its historical model cards (archive-relative paths):

tar -xzf failed-market-maker-v1.0.0.tar.gz --exclude='models/*/README.md' models/
python -m inference.run models/m3_run3/best_model.pt --family m3 --smoke
python -m inference.run models/m5_v7v72_bc/agent_model_m5.onnx --smoke
python -m inference.run models/mpv2_r1_a3_seq120_seed2 --smoke

Alternatively unpack the historical full archive into a separate directory. Synthetic forward checks demonstrate executability only; they do not reproduce historical training or trading performance. Keep normalization, feature contracts and configurations with every model; BC also requires agent_model_m5.onnx.data.

Development story · Model cards · Provenance & redaction · Unified file manifest

Original code and weights: Apache-2.0. Curated documents and reports: CC-BY-4.0. See NOTICE for attribution. Raw market datasets are not distributed. Historical evidence is a source record, not current deployment instructions.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading