BarScript C1.3 (sc2_c1_v13_s1234_best) release package: bf16 weights, frozen serving contract, model card
09e2da6 verified | license: mit | |
| tags: | |
| - taiko | |
| - rhythm-game | |
| - chart-generation | |
| - music | |
| - audio-to-symbolic | |
| # BarScript (experimental preview) β C1.3, denominator augmentation, seed 1234 | |
| Audio log-mel + a bar grid β a Taiko no Tatsujin chart, decoded under a | |
| finite-state grammar over bar-scoped plan, skeleton and realization stages. | |
| **BarScript** is the release name of this charting model; it is also the name of | |
| the bar-scoped multi-stage token encoding it decodes into. It is a **preview | |
| under evaluation**, published so it can be tried, not as a finished model: the | |
| champion comparison below is `inconclusive`, and Β§2 lists measured defects β one | |
| of which fires a pre-registered kill condition. The seven published SoftChart | |
| 1.x models are a separate, unaffected line. | |
| | | | | |
| |---|---| | |
| | source checkpoint | `runs/sc2_c1_v13_s1234/best.pt` | | |
| | sha256 | `44980bf64b1611ea73c1433c41adceff7596cfeece30dfaad5d228eae4a14e6e` | | |
| | seed / step | 1234 / 96 000 of 100 000 | | |
| | best validation CE | 0.5769365892960475 | | |
| | parameters | 9 107 291 | | |
| | shipped weights | bfloat16 safetensors, 18.2 MB (fp32 checkpoint preserved in the run dir) | | |
| | training data | `JacobLinCool/taiko-1000-parsed-clean`, revision `b72da4616d643018e81f372cea06ce51349285e0` | | |
| | label spec | `barscript_labels_v1`, frozen, shipped as `spec.json` (sha1 `cd876768β¦`) | | |
| | training flags vs C1 | `--cal-axes big span density` Β· `--balance-exclude-big` Β· `--denom-augment 0.5 --denom-augment-mode lattice` | | |
| | external pretraining | none | | |
| **This card is written for THIS checkpoint.** Every number below was measured on | |
| C1.3's own generations under the serving code this package ships | |
| (`generate.py` md5 `73b6e49bβ¦`, `min_onset_gap_sec = 0.020`). Nothing is | |
| inherited from the C1 card. Where an item on the C1 card has no C1.3 | |
| measurement, it is marked **not measured on this arm** rather than carried over. | |
| **Read the KNOWN LIMITATIONS section before using any number above it.** This | |
| model has real, measured defects, one of which fires a pre-registered kill | |
| condition, and its difficulty ladder is roughly 25 coverage points less nested | |
| than an authored one. | |
| --- | |
| ## 0. Status: the champion comparison is `inconclusive` | |
| `scripts/sc2_gates.py champion`, three pairings on one serving code base | |
| (`experiments/sc2_eval/FINAL_CHAMPION.md` Β§1): | |
| | pairing | `champion_verdict` | reason code | | |
| |---|---|---| | |
| | C1.2 vs C1 | `inconclusive` | `no_seed_replicate` | | |
| | **C1.3 vs C1** | **`inconclusive`** | `no_seed_replicate` | | |
| | C1.3 vs C1.2 | `inconclusive` | `no_seed_replicate` | | |
| `MIN_SEEDS_FOR_VERDICT` is 2 in **both** arms and C1.3 has one seed. The rule | |
| fires before any metric is read. **No arm has been declared champion.** | |
| What the clause ledger says, which is one-directional: C1.3 **passes** the | |
| required Β§4 clause `plan_token_calibration`, which **C1 fails on both its | |
| seeds**. No clause is passed by C1 and failed by C1.3. The two `a2` span clauses | |
| fail on **all four arms** and belong to the campaign, not to this checkpoint. | |
| The shipping rationale, the two named compromises and the decision rule for | |
| choosing a different checkpoint are in | |
| `experiments/sc2_eval/SHIP_DECISION.md`. | |
| Gate items on the 71-chart hard+oni population, this arm against the C1 seed-1234 | |
| arm under identical code: | |
| | | C1 | **C1.3** | | |
| |---|---:|---:| | |
| | fail / pass / na | 9 / 10 / 6 | **8 / 11 / 6** | | |
| | newly passing | β | `control_own_axis_density`, `control_named_regression_density_big_count` | | |
| | newly failing | β | `a2_span_placement_fit` (knife-edge, Β§2.6) | | |
| --- | |
| ## 1. What it does | |
| * **Input** β 128-bin log-mel at 22 050 Hz (`preprocessor_config.json` freezes | |
| the exact decode/STFT/mel contract), plus a bar grid with explicit measure | |
| edges, exact rational meters and per-bar lattice denominators. | |
| * **Conditions** β course, authored level, density bucket, and the split-axis | |
| knobs `big_rate` / `span_rate` / `stream` / `sync`. | |
| * **Output** β a BarScript token stream decoded under `BarscriptFSM`, exported | |
| as TJA on the 96-slot lattice. | |
| * **Family decoding** β `generate_chart_family` decodes a song's courses as one | |
| nested ladder, hardest first, each easier course biased toward its harder | |
| sibling's onsets. **Off by default**; the campaign's headline population was | |
| decoded independently (Β§2.1). | |
| Capabilities this checkpoint carries (`config.json:capabilities`): `aux`, | |
| `beat_head` (hi-res), `hierarchical_ctx`, `axis_knobs`, `sync_token`, | |
| `stage_emb`, `span_duration_head`, `tempo_head`, `rich_section_stats`. It | |
| carries **no** `style`, `sibling`, `ctx`, `plan`-prefix, `slot`, `dual`, | |
| `align`, `mask_infill`, `func_time`, `global_ctx` or `complexity` conditioning; | |
| requests on those axes reach nothing. | |
| ### What works | |
| Evidence marks: **β** n β₯ 71 charts with a measured C1 seed band; **β** n = 19β71 | |
| charts or 8β12 sweep cells, one seed; **β** n < 19 or a CI covering zero. | |
| | axis | this checkpoint | reference | mark | | |
| |---|---|---|---| | |
| | **Density knob** β median within-song Ο **0.850**, sign p **0.0078**, 8/8 cells positive, **+1.413 nps** firstβlast, 12 % monotone adjacency | C1 moves the chart +0.024 / +0.151 nps over the same range and **fails** the gate; C1 seed-to-seed ΞΟ 0.003 (p 1.000) | 8 songs Γ 5 buckets, oni | β | | |
| | **`big_rate` knob** β Ο **1.000**, p 0.00049, 91.7 % monotone, and **`calibrated: true`** β mean absolute bucket error **0.533**, within Β±1 bucket **93.3 %** | first axis in the whole campaign to clear `calibrated`; C1 0.533 vs 1.200 / 0.967 | 12 songs Γ 5 buckets, oni | β | | |
| | **`span_rate` knob** β Ο 0.810, p 0.0117, `calibrated: true`, mean abs err 1.483; and unlike C1 it **no longer drags hit density with it** | C1's span knob flags `onset_nps` / `hit_nps` / `n_onsets` as interference; C1.3 flags only the definitional `span_count` | 12 songs Γ 5 buckets, oni | β | | |
| | **Span budget** β 739 spans against the authored 762 (**0.970Γ**); `span_rate_gen_per_min` gap to authored **0.049** against a C1 seed band of 0.198 | C1 1.160β1.244Γ, C1.2 1.259Γ | 97 charts | β | | |
| | **FSM guarantees** β **0 unclosed, 0 orphan ends, 0 swallowed hits** over 739 generated spans | holds on all four arms | 97 charts | β | | |
| | **Placement does not degrade** β precision 0.6477 (C1 band 0.6480β0.6502), median \|offset\| **0 ms**, `exact_slot_lift_over_null` **1.1755** (above the C1 band), long-song drift tests significant **2** vs C1's 6 / 5 | see Β§2.3 for why apparent recall drops | 71 charts | β | | |
| | **Accent over-emission reduced furthest** β pooled hard+oni big share **0.0873** against C1's 0.1082 / 0.1027 (3.3 seed bands) and an authored 0.0589 | still outside the pre-registered stop window 0.045β0.075 | 71 charts | β | | |
| | **Cheaper** β 1 618.6 tokens/chart, 19.22 tokens/bar, `max_window_tokens_p99` **449** | C1 1 731.2 / 20.51 / 515 | 71 charts | β | | |
| | **Deployment grid repaired** β on a forced `/16` BPM grid, notes Γ· authored **1.130 / 1.112 / 1.067 / 1.027** and onset F1 **0.333 / 0.374 / 0.466 / 0.580** | C1 1.688 / 1.461 / 1.297 / 1.210 and F1 0.318 / 0.376 / 0.477 / 0.599; C1.3 beats C1.2 on **12 of 12** deployable cells | 13 songs Γ 4 courses, one seed | β | | |
| | **Empty-bar declaration restored** β `/16` empty bars 12.82 / 8.60 / 4.55 / 3.14 % against C1.2's 5.96 / 3.23 / 1.08 / 0.33 %, sign-significant on 3 of 4 courses | authored 15.69 / 12.30 / 9.08 / 6.27 % | 13 songs | β | | |
| --- | |
| ## 2. KNOWN LIMITATIONS | |
| Nothing in this section is softened. Where a limitation is invisible to the gate | |
| suite, that is said. | |
| ### 2.1 The difficulty ladder is 25β32 coverage points less nested than authored | |
| Coverage = the fraction of the easier chart's notes that have a note in the | |
| harder chart within tolerance. **Two decode modes, two different numbers, and | |
| both belong on the record.** | |
| | adjacent pair | authored charts | **independent decode** (package default) | **family decode Ξ² = 2** (authored grid) | **family decode Ξ² = 2** (`/16` deploy grid) | | |
| |---|---:|---:|---:|---:| | |
| | easy β normal | 0.970 | **0.658** | 0.730 | 0.726 | | |
| | normal β hard | 0.979 | **0.667** | 0.797 | 0.811 | | |
| | hard β oni | 0.985 | **0.660** | 0.822 | 0.834 | | |
| Independent-decode figures: `FINAL_CHAMPION.md` Β§8.1, n = 13 / 13 / 35 songs, one | |
| seed, same serving code as this package (β/β). Family-decode figures: | |
| `C13_PREDICTION.md` Β§3.3, n = 13 songs, one seed (β). | |
| Five things make this worse than the headline: | |
| * **`nesting_coverage` and `hand_agreement` FAIL the family gate on every arm | |
| measured**, C1 included. `difficulty_monotonicity` passes on all of them. | |
| * **C1.3 is slightly worse than C1 here.** C1's independent-decode coverage on | |
| the same population is 0.690 / 0.704 / 0.737; C1.3 is β0.032 / β0.037 / β0.077 | |
| against C1 seed bands of 0.010 / 0.021 / 0.025. The move is small against the | |
| ~0.30 gap to authored that every arm shares, but it is in the wrong direction. | |
| * **Hand agreement at coinciding hits is near chance on hardβoni**: C1.3 0.539 | |
| against a marginal-chance null of 0.500 and an authored 0.753. When two | |
| generated courses agree that a note belongs somewhere, which drum they pick is | |
| near-independent across courses. | |
| * **The `/16` figure overstates the model's nesting.** The same checkpoint on a | |
| `/96` grid drops to 0.649 / 0.730 / 0.752, worse on 11β12 of 13 songs | |
| (p = .003 / .022 / .022). A coarse lattice manufactures agreement by leaving | |
| few places to disagree (`C13_PREDICTION.md` Β§5.2). | |
| * **`family_bias = 2.0` / `family_hand_bias = 1.5` are PROVISIONAL** β described | |
| in `generate.py` as logit offsets on a decoder never trained at that setting, | |
| and never calibrated against the authored target. | |
| ### 2.2 The note-type channel carries almost no information | |
| 71 charts, **22 260 matched slots** (`FINAL_CHAMPION.md` Β§5.2, β). | |
| | quantity | this model | its own floor | corpus Β§2.1 floor | | |
| |---|---:|---:|---:| | |
| | 4-way accuracy | **0.4776** | majority **0.5559** | majority 4-way **0.563** | | |
| | β margin over majority | **β0.0783**, 95 % CI [β0.1022, β0.0530] | | | | |
| | hand accuracy | **0.5493** | always-don **0.5975** | always-don **0.620** | | |
| | β margin | **β0.0482**, CI [β0.0721, β0.0229] | | | | |
| | **MI(generated; authored), 4-way** | **0.01028 bits** | β | **0.79 % of H(authored) = 1.300 bits** | | |
| | MI, hand | 0.00534 bits | β | 0.55 % of 0.972 bits | | |
| * **The model sits below its own majority floor on both endpoints**, and below | |
| the campaign's corpus baselines (56.3 % / 62.0 %), which are a fixed reference | |
| line from a separate 59 961-hit census, not this population. | |
| * **The honest caveat, which cuts the other way.** The published floors are | |
| *argmax* predictors scored against a temperature-1.0 top-p-0.95 *sample*. | |
| Against the sampler-appropriate i.i.d. floor (Ξ£ pΒ² = 0.4582 4-way, 0.5190 | |
| hand) this model is **above** baseline by **+0.0194** and **+0.0303** β the | |
| largest 4-way excess of any arm in the campaign. Both readings belong on the | |
| record. | |
| * **MI is flat at ~0.010β0.012 bits on all four arms** (seed band 0.0008). The | |
| accuracy differences between arms are marginal-matching, not information. | |
| * Accuracies are conditional on coverage **0.6442** (precision side) / | |
| **0.5842** (recall side); more than a third of generated hits have no authored | |
| partner, and coverage is 9 % lower than C1's, so this population is smaller | |
| and differently selected than C1's. | |
| * **No inter-charter agreement ceiling exists** for this split β no | |
| (song, course) carries two independent authored charts β so 100 % is not a | |
| legitimate target and is not used as one. | |
| * **Greedy re-decoding was not repeated on this arm.** The C1-era finding that | |
| greedy closes part of the gap at an unacceptable `motif_reuse` cost is **not | |
| measured on this arm** and must not be quoted for it. | |
| ### 2.3 The note budget is 9 % short, and it is the whole of the apparent timing regression | |
| | | C1 (s1234 / s4321) | **C1.3** | | |
| |---|---:|---:| | |
| | `onset_precision` | 0.6480 / 0.6502 | **0.6477** | | |
| | note budget (gen Γ· authored notes) | 1.0046 / 0.9912 | **0.9069** | | |
| | `onset_recall` | 0.6510 / 0.6445 | **0.5874** | | |
| | `frac_of_authored_within_jnd` | 0.6463 / 0.6397 | 0.5840 | | |
| | `type_coverage_recall` | 0.6465 / 0.6401 | 0.5842 | | |
| `precision Γ budget` equals `onset_recall` to four decimals on **every** arm, by | |
| construction. Those three "timing" rows are one quantity, and it is the note | |
| budget, not placement: precision is inside the C1 seed band, median offset is | |
| 0 ms, `exact_slot_lift_over_null` is *above* the band, and significant drift | |
| tests fall from 6 / 5 to 2. | |
| **The shortfall itself is a real regression and its cause is unknown.** It | |
| appears on both C1.2 and C1.3, so it belongs to the `--cal-axes` / | |
| `--balance-exclude-big` pair rather than to the denominator augmentation. It is | |
| the largest practical regression in this release. | |
| ### 2.4 Raising density squeezes drumrolls out β this is the compromise a user can hit | |
| Median spans/min over eight oni songs, per requested density bucket | |
| (authored **3.885**; `FINAL_CHAMPION.md` Β§4.3, 8 songs Γ 5 buckets, β): | |
| | requested density | 0 | 3 | 7 | 11 | **15** | | |
| |---|---:|---:|---:|---:|---:| | |
| | C1 (s1234) | 3.629 | 3.055 | 4.153 | 2.818 | **2.971** | | |
| | **C1.3** | 3.550 | 4.413 | 2.172 | 2.727 | **1.279** | | |
| Ο(density β `span_per_min`) = **β0.759**, sign p 0.0078, firstβlast β2.443 /min. | |
| At the top of the density request C1.3 delivers **33 %** of the authored span | |
| rate; C1 delivers 76β79 %, C1.2 44 %. | |
| * **Do not advertise the density knob at its top setting** until this is fixed. | |
| * There is a named suspect: `cal_span` never left **0.1667** on either | |
| challenger, so this is plausibly a calibration term that never trained rather | |
| than an intrinsic cost of the knob. | |
| * The mirror image is a *gain*: C1.3's `span_rate` knob no longer drags hit | |
| density, which C1's does. | |
| ### 2.5 Absolute bucket calibration is still wrong on the density axis | |
| | axis | exact bucket | within Β±1 | mean abs error | mean signed | `calibrated` | | |
| |---|---:|---:|---:|---:|:--:| | |
| | density | 5.0 % | 25.0 % | 4.125 | β1.075 | **false** | | |
| | `big_rate` | **53.3 %** | **93.3 %** | **0.533** | +0.267 | **true** | | |
| | `span_rate` | β | β | 1.483 | β | true | | |
| The density knob orders the request correctly and lands in the wrong bucket, | |
| under-shooting by ~1 bucket on average β more than C1.2 does. Rank control is | |
| real; absolute density targeting is not delivered. | |
| Note the cost that comes with the `big_rate` calibration win: C1.3 has the | |
| **shortest** `big_rate` ladder of any arm (firstβlast +0.159 against C1's | |
| 0.183 / 0.224), so the top of that request now delivers less accent than C1's | |
| did. | |
| ### 2.6 It fails its own span-quality gate and triggers a pre-registered kill condition | |
| 71 charts hard+oni (467 closed spans) and 26 charts easy+normal (272): | |
| | axis | hard+oni | easy+normal | | |
| |---|---|---| | |
| | `duration_low_tail` | **fail** (rate verdict passes; balloon 0.250 beats = 0.75Γ the shortest authored balloon, roll 0.250 = 0.60Γ) | **fail** (balloon 0.250 beats = 0.20Γ the shortest authored balloon; 45/159 balloons below authored support) | | |
| | `forced_close` | **fail** β **7 / 467 = 1.499 %** (0 unclosed, 7 clamp-forced) | **fail** β **7 / 272 = 2.574 %** | | |
| | `over_span_max` | **fail** β one balloon 7.72 s = 1.19Γ the 6.5 s ceiling | **fail** β one balloon 7.58 s = 1.17Γ | | |
| | `placement_fit` | **fail** β AUC **0.560** (CI [0.529, 0.593]) against a 0.56 floor; authored reference 0.668 | pass β AUC 0.604 | | |
| | `orphans` / `swallowed` / `rate` / `duration` | pass | pass | | |
| | **`a2_kill_condition_not_triggered`** | **fail β the kill condition FIRED** | fail | | |
| The pre-registered kill condition (`SOFTCHART2_DESIGN.md` Β§4-A2) reads: *rate is | |
| calibrated but placement/length quality is bad β concede that the bar-scoped | |
| commitment mechanism is insufficient and escalate to an independent span-state | |
| objective.* It fires on **all four arms**, including C1. Shipping does not | |
| discharge it. | |
| **`placement_fit` is a knife-edge, not a finding.** C1.2 passes at 0.561 and | |
| C1.3 fails at 0.560 with CIs overlapping every other arm across their whole | |
| width. One thousandth of AUC decides the verdict. It is reported, not used. | |
| ### 2.7 One span in seven is a micro-span | |
| All generated spans, 97 charts, `FINAL_CHAMPION.md` Β§2 (β): | |
| | | n spans | p1 | p5 | p50 | p90 | min | **share < 0.25 s** | | |
| |---|---:|---:|---:|---:|---:|---:|---:| | |
| | authored | 762 | 0.177 | 0.273 | 0.963 | 2.419 | 0.083 | **3.8 %** | | |
| | **this model** | 739 | 0.089 | 0.150 | 0.818 | 2.258 | **0.072** | **14.7 %** | | |
| Per course: easy 6.6 %, normal 8.1 %, hard 13.7 %, oni **25.2 %** β authored | |
| 3.0 / 3.3 / 2.1 / 7.1 %. | |
| * C1.3 does **not** inherit C1.2's worsening (16.4 %); 14.7 % is exactly at the | |
| top of the C1 seed band (13.6β14.7 %). It is not an improvement either. | |
| * **Mechanism is known and unfixed on every arm.** Duration bucket 0 is open at | |
| the bottom; 123 of C1.3's 739 spans sit in it, at a median 0.5 beats against | |
| the authored 0.917. | |
| * **The shipped floor does not fix it.** `span_min_sec` / `span_min_beats` fired | |
| 263 times (`span_close_floor`) plus 33 carries across 97 charts and the | |
| < 0.25 s share is still 14.7 %. | |
| * **The mitigation is unbuilt.** `span_bucket0_profile` is `null` in this | |
| package because `scripts/build_span_bucket0_profile.py` has never been run on | |
| the train split. This is the shortest available fix and it is not in the box. | |
| * **The gate cannot see it directly.** `a2_span_duration` compares p50/p90/p99 | |
| only and passes; `duration_low_tail` catches the support violation but not the | |
| rate (its rate verdict passes on this arm). | |
| * On a `/16` deployment grid the same quantity is 14.7 / 12.4 / **30.6** / 21.8 % | |
| β hard is untouched by the augmentation and is the worst cell in the release. | |
| ### 2.8 The deployment grid has a two-role conflict that no checkpoint dissolves | |
| `BAR_DENOM` is simultaneously an input the grid supplies and a **trained output | |
| token** the model reads as a density announcement. `bar_denoms="supplied"`, the | |
| shipped default, teacher-forces it. | |
| 13 songs Γ 4 courses, one seed, `--cond authored`, `C13_PREDICTION.md`: | |
| | easy / normal / hard / oni | authored grid | **forced `/16`** | `/96` | `/24` | | |
| |---|---|---|---|---| | |
| | notes Γ· authored | 0.917 / 0.919 / 0.875 / 0.899 | **1.130 / 1.112 / 1.067 / 1.027** | 0.944 / 0.950 / 0.943 / 1.027 | 1.102 / 1.070 / 0.950 / 0.921 | | |
| | onset F1 | 0.613 / 0.565 / 0.595 / 0.647 | **0.333 / 0.374 / 0.466 / 0.580** | 0.303 / 0.325 / 0.393 / 0.535 | 0.342 / 0.370 / 0.442 / 0.542 | | |
| | genuine triplets (den 3/6/12/24) % | 1.13 / 2.48 / 6.34 / 9.43 | **0.00** (arithmetically impossible) | 3.45 / 5.78 / 10.57 / 11.46 | 14.13 / 20.31 / 28.57 / 38.70 | | |
| | ultra-fine (den 48/96) % | ~0 | 0.00 | **11.53 / 13.22 / 15.62 / 6.60** | 0.00 | | |
| | nesting (family Ξ² = 2) | 0.730 / 0.797 / 0.822 | 0.726 / 0.811 / 0.834 | 0.649 / 0.730 / 0.752 | 0.714 / 0.769 / 0.779 | | |
| Authored triplet rates for reference: 4.20 / 5.72 / 7.43 / 9.95 %; ultra-fine | |
| ~0.00 %. | |
| **The recommended deployment configuration is `--grid bpm --grid-bpm-denom 16` | |
| with the package default `bar_denoms="supplied"`.** Its two costs, stated | |
| plainly: **triplets are exactly 0 and cannot be otherwise**, and hard micro-spans | |
| stay at 30.6 % against an authored 2.1 %. | |
| Four consequences a deployer must carry: | |
| * **Deployment costs 10β46 % of onset F1** relative to the authored grid: | |
| easy β46 %, normal β34 %, hard β22 %, oni β10 %. Every `--grid authored` | |
| number in this card is therefore optimistic for a user who supplies only BPM. | |
| * **`p90 |offset|` is 0 ms on the authored grid and 14.6β36.7 ms on the bpm | |
| grids.** That residual belongs to the bar edges; no denominator choice touches | |
| it. | |
| * **The model's off-dyadic rate is a property of the model, not the lattice.** | |
| Given any 3-divisible grid it puts 2β3Γ the authored share off the dyadic | |
| grid; the grid only decides whether that mass gets called "triplets" (`/24`) | |
| or "ultra-fine jitter" (`/96`). This is a training-side defect and no | |
| deployment grid fixes it. | |
| * **The gate suite cannot score the deployment path.** `robustness.py:629-642` | |
| raises `ContractError` on uniform-BPM arms, so every gate timing metric is | |
| `null` for the path the model ships on; the `/16` figures above come from | |
| `timing.py`, a second implementation that reproduces the gate on the authored | |
| grid. | |
| **The `deploy_bpm_grid` profile in this package is NOT the recommendation.** It | |
| sets `bar_denoms="model"`, `bar_denom_mask="div96"`, per `DEPLOY_GRID_FIX.md` | |
| Β§7 β a recommendation made explicitly conditional on C1.3 failing, which it did | |
| not. That arm was measured on **C1 only**: density 1.06 / 1.08 / 0.97 / 1.01 and | |
| nesting 0.635 / 0.706 / 0.778. **It has never been run on this checkpoint.** The | |
| profile is kept because it is implemented and because the arm is the obvious | |
| next experiment, not because it is advised here. | |
| **One contradiction inside this package, stated so nobody has to find it.** | |
| `config.json:serving_profiles.deploy_bpm_grid.basis` still quotes | |
| `DEPLOY_GRID_FIX.md` Β§7 verbatim, including the words *"Recommended now"*, and | |
| still carries C1's nesting figures. That string is generated by | |
| `scripts/sc2_package.py`, which is read-only for this release. **This section | |
| supersedes it.** The generated string is left intact rather than silently | |
| diverging from the shipped code. | |
| ### 2.9 Playability floor: gaps no human hand can play | |
| Recomputed on this arm's own events, 71 charts hard+oni, generated vs authored | |
| on the same songs (producer `experiments/sc2_eval/ship_c13/playability.py`, | |
| records `.../records/playability_all_arms.json`; β): | |
| | | this model | authored | | |
| |---|---:|---:| | |
| | inter-onset gaps < 40 ms | **102** across **29** charts | 2 across 1 chart | | |
| | β as a share of adjacent pairs | **0.296 %** (102 / 34 484) | **0.0053 %** (2 / 38 032) | | |
| | gaps below the series record 29.4118 ms | **19** | 0 | | |
| | shortest gap | **20.8 ms** (the serving floor) | 39.7 ms | | |
| | peak burst in any 1 s window | **17 notes** | 15 notes | | |
| C1 on the same population: 151 gaps (0.395 %), shortest 16.7 ms, peak 22 | |
| notes/s. C1.3 is better on every row and still **56Γ the authored rate**. | |
| * `min_onset_gap_sec = 0.020` is a **degeneracy guard, not a fix**: it is set | |
| below the fastest thing the series has shipped (29.4118 ms, | |
| TAIKO-TONGUE-TWISTER oni, BPM 170, 48th notes) precisely so it cannot refuse a | |
| chart the domain writes. It fired **125 times** across 97 charts with **0** | |
| fail-opens. *A mask cannot fix a distribution.* | |
| * The `generate.py` constant comment quotes the generated sub-40 ms rate as | |
| ~0.35 % against ~0.045 % authored. **That authored figure does not reconcile | |
| with the 0.0053 % measured here**, and the two have never been put on the same | |
| denominator. The excess is 8Γ on the comment's accounting and 56Γ on this one. | |
| * **Good news that replaces a C1-card claim.** The C1 card reported the model | |
| placing hits in 19.6 % of sub-0.25 s scaffolding bars. Under today's serving | |
| floors this arm places **0 hits in all 112 such bars** across 97 charts, exactly | |
| as the authored side does; `degenerate_bar_hits_blocked` fired 12 times with 0 | |
| fail-opens. | |
| ### 2.10 Pattern proxies regress against C1 on hard+oni | |
| 71 charts, one code version, C1 band = |s1234 β s4321| (`FINAL_CHAMPION.md` Β§3): | |
| | metric (lower is better) | C1 mean | C1 band | **C1.3** | Γband | relative | | |
| |---|---:|---:|---:|---:|---:| | |
| | **`motif_reuse_gap`** | 0.3630 | 0.0042 | **0.4222** | +14.1 | **+16.3 %** | | |
| | `compression_gap` | 0.0976 | 0.0006 | 0.1116 | +23.3 | +14.3 % | | |
| | `motif_ref_marginal_js` | 0.0331 | 0.0046 | 0.0452 | +2.6 | +36.6 % | | |
| | `ioi_js_per_chart` | 0.0295 | 0.0006 | **0.0352** | **+9.7** | +19.3 % | | |
| | `ul_4gram_js_per_chart` | 0.2402 | 0.0055 | 0.2530 | +2.3 | +5.3 % | | |
| | `class_4gram_js_per_chart` | 0.1337 | 0.0080 | 0.1357 | +0.2 | +1.5 % | | |
| The `Γband` column overstates the case β several bands are under 1 % of their own | |
| level β so the relative column is the honest one. On `motif_reuse` the bootstrap | |
| CIs are nonetheless **disjoint**: C1 [β0.392, β0.334] / [β0.388, β0.334] against | |
| C1.3 [β0.447, β0.396]. **The regression is real and not seed noise.** | |
| Four bounds, all measured, none of them a dismissal: | |
| * **It does not happen on easy+normal.** Every pattern proxy there is inside the | |
| C1 seed band, `motif_reuse_gap` is marginally *better* than C1's mean, and | |
| `class_4gram_js` is better on both challengers. | |
| * **The decomposition puts the loss on rhythm and hand, while the accent layer | |
| improves**: rhythm 0.1251 β 0.1640, hand 0.1774 β 0.2102, accent 0.0606 β | |
| **0.0480**. | |
| * **`ioi_js_per_chart` is where the augmentation shows up** β it reads the exact | |
| rational IOI lattice, and it is C1.3's worst pattern metric relative to C1.2 | |
| (+9.7 band against +1.5). | |
| * **The proxy reads mostly a channel carrying ~0.010 bits** about the authored | |
| type (Β§2.2), and it has **never been validated against a listener**. It is | |
| quoted because Β§4 names it, not because it is strong. | |
| Realized `motif_reuse` is **0.3165** against an authored **0.7387** β under half | |
| the authored repetition. | |
| ### 2.11 The note-type marginal, and one C1 finding that does NOT reproduce | |
| Big notes are **1.48Γ the authored rate** pooled hard+oni (0.0873 vs 0.0589), | |
| worst on the easiest course: | |
| | course | authored | **C1.3** | C1 (s1234 / s4321) | | |
| |---|---:|---:|---:| | |
| | easy | 18.80 % | **26.22 % (1.40Γ)** | 32.68 % / 29.65 % | | |
| | normal | 11.75 % | **17.41 % (1.48Γ)** | 19.21 % / 20.77 % | | |
| | hard | 7.38 % | **10.61 % (1.44Γ)** | 13.11 % / 12.73 % | | |
| | oni | 4.88 % | **7.40 % (1.52Γ)** | 9.29 % / 8.58 % | | |
| The pre-registered stop window (0.045β0.075 pooled hard+oni) is **not met**. | |
| **The C1 card's Β§2.8 claim does not reproduce here and is corrected.** On C1 the | |
| realized class marginal was 13Γ closer to the training loss weights than to the | |
| corpus. On C1.3 it is the other way round: JS(gen β authored) = **0.00588** | |
| against JS(gen β class-weight prediction) = **0.00746** | |
| (JS(authored β weights) = 0.01452). `--balance-exclude-big` moved the marginal | |
| off the loss weights and toward the corpus. The log-log fit of realized share | |
| against class weight collapses from slope 0.378 (Ο 0.436) on C1 to slope 0.078 | |
| (Ο 0.156) on C1.3. | |
| ### 2.12 Fine-lattice, tuplet and syncopation accuracy are much worse than average | |
| `robustness.py` timing strata, 71 charts, exact-slot rate (overall **0.5840**): | |
| | stratum | n authored | exact-slot | C1 (s1234) | | |
| |---|---:|---:|---:| | |
| | lattice step 11.6β23.2 ms | 114 | **0.342** | 0.456 | | |
| | lattice step 23.2β46.4 ms | 790 | **0.443** | 0.465 | | |
| | lattice step β₯ 92.9 ms | 24 228 | 0.594 | 0.651 | | |
| | positions with a denominator divisible by 3 | 3 480 | **0.345** | 0.496 | | |
| | syncopation band 0 β band 5 | 20 846 β 1 669 | **0.615 β 0.434** | 0.672 β 0.514 | | |
| | local IOI 25β50 ms | 189 | **0.402** | β | | |
| | BPM > 250 | 2 570 | **0.499** | β | | |
| Every stratum is lower than C1's, which is the note-budget effect of Β§2.3 acting | |
| on a per-stratum recall, but the *shape* is the finding: tuplet positions and | |
| fast lattice steps are 0.24 below the overall rate, and accuracy decays | |
| monotonically with syncopation. | |
| ### 2.13 Smaller, but on the record | |
| | finding | number | n | source | | |
| |---|---|---|---| | |
| | `long_song_no_drift` fails | **2** of 21 metrics significant at BH q β€ 0.05 (`dens_signed_err`, `plan_dens_bucket_tvd`) β C1 fails 6 / 5 | 71 charts | `robustness.json` β | | |
| | Harness-level generation failures | **6** of 71 song-courses refused (`#BRANCHSTART` and non-representable meters); **1** chart exported `gen_tja = null` β identical to C1 on the same population | 71 charts | `index.json` β | | |
| | Slot-export off-lattice events | 8 across 97 charts (C1: 5) β the documented contract path, not a crash | 97 charts | decoder counters β | | |
| | `escape_hatch` fires | 91 times across 97 charts (C1: 84) | 97 charts | decoder counters β | | |
| | Generation wall time | 9.29 s/chart, 4.33 s per audio-minute β **contended**, the GPU was shared, quoted as a declaration only | 71 charts | β | | |
| | Complete-bar rest placement | **not measured on this arm** (C1: Jaccard 0.474, recall 0.623) | β | β | | |
| | Greedy-decode contrast | **not measured on this arm** | β | β | | |
| | Micro-span seed sensitivity | **not measured on this arm** (C1: 24.0 % vs 32.4 % across two sampling seeds on 19 oni charts) | β | β | | |
| ### 2.14 The evidence base is thinner than it looks | |
| * **One seed.** Every number in this card is n = 1 in the training seed. The C1 | |
| seed band quoted throughout is C1's, used as a reproducibility scale; it is a | |
| range over n = 2 and carries no confidence statement. | |
| * **easy and normal are evaluated on 13 songs.** The 71-chart population is | |
| `hard` + `oni` only; the easy/normal population is 26 charts from 13 songs, and | |
| its seed band is 4β5Γ wider than hard+oni's. | |
| * Every difficulty-family and deployment-grid number is **13 songs, one seed** | |
| (β). Every knob sweep is **8β12 songs Γ 5 buckets** (β). | |
| * At these sizes a gate `pass` carries little information. Worked example from | |
| this campaign: the `big_share_of_hits` gate **flips between two C1 training | |
| seeds** while the failing seed's point estimate is *smaller*. At n = 36 that | |
| gate cannot rank checkpoints, and no pass/fail on it should be quoted as | |
| evidence for any arm. | |
| * **Nothing here is audio-referenced** except `a2_span_placement_fit`. Every | |
| other quantity is chart-vs-chart on the exact rational lattice or on hit order. | |
| No timing window and no game judgement parameter appears anywhere in this path. | |
| * **`best_val` is comparable across C1 / C1.2 / C1.3** (one training cache) and | |
| **not** comparable to C1.1 or to any arm on a different cache. C1.3's 0.5769 is | |
| **2.9 % worse** than C1.2's 0.5605; that was the registered trade and the | |
| deployment side won it. | |
| * **No plan-neutral arm was generated**, so the `plan_neutral_fallback` Β§4 clause | |
| is `na` and the ledger is incomplete by one required clause. | |
| ### 2.15 Serving regime: what `motif_constraint: auto` resolves to, and why it matters less here | |
| `motif_constraint: "auto"` resolves to **OFF** for this checkpoint: the training | |
| cache's `barscript_md5` (`634e3dc3β¦`) differs from the serving `barscript.py` | |
| (`9afc715aβ¦`), so the gold sequences it trained on never satisfied the MOTIF hard | |
| constraint. `train_serve_matched = true` β OFF is the train/serve-matched choice | |
| and it is the right default. | |
| **Unlike the C1 package, this card needs no correction for it.** Every number in | |
| this card was measured with the constraint **OFF**, i.e. under exactly this | |
| package's resolved default; the 20 evaluation runs all record | |
| `serving_fsm.motif_constraint: off, verified: true`. In particular | |
| `motif_ref_marginal_js` here is **0.0452** (hard+oni) and **0.0114** | |
| (easy+normal), both measured OFF. | |
| Two things that still belong on the record: | |
| * **The ON/OFF blast radius was measured on C1 only**, where switching the | |
| constraint on moved `motif_ref_marginal_js` 0.0311 β 0.0189 (β39 %). On the C1 | |
| package the shipped default therefore makes the correct value **0.0311, not the | |
| 0.0189 in the older tables**. **That contrast has never been measured on | |
| C1.3**, so no ON-constraint number should be quoted for this checkpoint. | |
| * Serving under the constraint ON would be train/serve **mismatched** for this | |
| checkpoint and is not a supported configuration. | |
| --- | |
| ## 3. Serving contract | |
| Every value below is in `config.json:serving`, with its justification in | |
| `config.json:serving_basis`. Pass them explicitly β the package's | |
| `serving_kwargs()` does β so the recorded contract is the one that reaches the | |
| decoder. | |
| | parameter | value | one-line basis | | |
| |---|---|---| | |
| | `min_onset_gap_sec` | **0.020** | Degeneracy guard, **not** a corpus percentile. Must stay strictly below the series record of 29.4118 ms (BPM 170, 48ths). Two earlier corpus-derived values (0.0395, 0.0300) were both wrong. See Β§2.9. | | |
| | `degenerate_bar_sec` | **0.25** | Authored charts place zero hits in any bar under 0.375 s (3 010 bars). Largest round threshold with zero counterexamples and 1.5Γ margin; identical to the gate's threshold. Blocks hits only, never span geometry. | | |
| | `span_min_sec` / `span_min_beats` | 0.0833 / 1β3 | Authored population minima. Fail-open, counted. Does not fix Β§2.7. | | |
| | `span_bucket0_profile` | **null** | Not built β see Β§2.7. | | |
| | `family_bias` / `family_hand_bias` / `family_mode` | 2.0 / 1.5 / `bias` | **Provisional**, never calibrated β see Β§2.1. Family decode is opt-in. | | |
| | `motif_constraint` | `auto` β resolves **off** here | Train/serve matching by `barscript.py` md5 β see Β§2.15. | | |
| | `bar_denoms` / `bar_denom_mask` | `supplied` / `lattice` | The evaluated regime, and the recommended deployment regime on a uniform `/16` grid. The `deploy_bpm_grid` profile exists but is **not** recommended for this checkpoint β see Β§2.8. | | |
| | `greedy` / `temperature` / `top_p` | false / 1.0 / 0.95 | As evaluated. | | |
| | `plan_temperature` | `null` (follows `temperature`) | Never exercised in evaluation; shipped unset rather than tuned. | | |
| `_has_sync` **must** be forced on at load. `train.py` records | |
| `sync_token=False` while the training prefix carries the SYNC slot, so a loader | |
| that trusts the recorded flag drops the slot and every sync bucket reaches the | |
| decoder as an identical prefix. The package loader does this and records why in | |
| `config.json:serving_prefix_fix`. This is a workaround; the fix belongs in | |
| `train.py`. | |
| ### Recommended deployment configuration | |
| ``` | |
| --grid bpm --grid-bpm-denom 16 # uniform /16 grid | |
| bar_denoms = "supplied" # the package default | |
| motif_constraint = auto -> off # resolved at package time | |
| family decode: optional; Ξ² = 2.0 / 1.5 is provisional | |
| ``` | |
| Expected behaviour under it, 13 songs Γ 4 courses, one seed: notes within | |
| 13 / 11 / 7 / 3 % of authored; onset F1 0.333 / 0.374 / 0.466 / 0.580; | |
| **zero triplets**; hard micro-spans ~30 %. | |
| --- | |
| ## 4. Weights are bfloat16 | |
| Training ran bf16 autocast and CUDA inference runs bf16 autocast, so fp32 | |
| storage carried no information the forward pass could use β the argument the | |
| v1.5 release made and verified. Measured cast cost: max absolute delta | |
| **0.00711** on `frontend.2.weight` (0.349 % of that tensor's max) over 9 107 291 | |
| float elements. | |
| This is an argument about the **compute** dtype, not a proof that the two | |
| checkpoints decode identically. Sampling is chaotic in the logits, so individual | |
| charts can differ. The paired fp32-vs-bf16 evaluation has **not** been repeated | |
| for BarScript. The fp32 checkpoint is preserved in the source run directory and | |
| is the reference for any bit-level comparison. | |
| --- | |
| ## 5. Provenance | |
| | | | | |
| |---|---| | |
| | serving code | `generate.py` md5 `73b6e49b00fb60f3229c8539bc7a1029`; all 18 `src/softchart/*.py` md5s in `config.json:code.module_md5` | | |
| | evaluation harness | `sc2_generate_eval_v2_2_serving_floors`, contract `sc2_eval_v2` | | |
| | evaluation runs | 20 run indices, 1 028 charts, one `serving_floors` signature, `comparable: true`, 6/6 floor counters present | | |
| | headline population | 71 charts hard+oni (36 songs) + 26 charts easy+normal (13 songs), `--cond authored`, `--grid authored`, seed 1 | | |
| | deployment population | 13 songs Γ 4 courses Γ 10 arms, `--cond authored`, one serving tree, one signature | | |
| | label spec | `spec.json`, field-for-field equal to the cache manifest's copy (`spec_parameters_match_cache_manifest: true`) | | |
| | uploaded | **no** (`release_manifest.json: "uploaded": false`) | | |
| Reports: `experiments/sc2_eval/SHIP_DECISION.md`, | |
| `experiments/sc2_eval/FINAL_CHAMPION.md`, | |
| `experiments/sc2_eval/C13_PREDICTION.md`, | |
| `experiments/sc2_eval/C12_EVAL.md`, `experiments/sc2_eval/DEPLOY_GRID_FIX.md`. | |