softchart-barscript / MODEL_CARD.md
JacobLinCool's picture
BarScript C1.3 (sc2_c1_v13_s1234_best) release package: bf16 weights, frozen serving contract, model card
09e2da6 verified
|
Raw
History Blame Contribute Delete
36 kB
---
license: mit
tags:
- taiko
- rhythm-game
- chart-generation
- music
- audio-to-symbolic
---
# BarScript (experimental preview) β€” C1.3, denominator augmentation, seed 1234
Audio log-mel + a bar grid β†’ a Taiko no Tatsujin chart, decoded under a
finite-state grammar over bar-scoped plan, skeleton and realization stages.
**BarScript** is the release name of this charting model; it is also the name of
the bar-scoped multi-stage token encoding it decodes into. It is a **preview
under evaluation**, published so it can be tried, not as a finished model: the
champion comparison below is `inconclusive`, and Β§2 lists measured defects β€” one
of which fires a pre-registered kill condition. The seven published SoftChart
1.x models are a separate, unaffected line.
| | |
|---|---|
| source checkpoint | `runs/sc2_c1_v13_s1234/best.pt` |
| sha256 | `44980bf64b1611ea73c1433c41adceff7596cfeece30dfaad5d228eae4a14e6e` |
| seed / step | 1234 / 96 000 of 100 000 |
| best validation CE | 0.5769365892960475 |
| parameters | 9 107 291 |
| shipped weights | bfloat16 safetensors, 18.2 MB (fp32 checkpoint preserved in the run dir) |
| training data | `JacobLinCool/taiko-1000-parsed-clean`, revision `b72da4616d643018e81f372cea06ce51349285e0` |
| label spec | `barscript_labels_v1`, frozen, shipped as `spec.json` (sha1 `cd876768…`) |
| training flags vs C1 | `--cal-axes big span density` Β· `--balance-exclude-big` Β· `--denom-augment 0.5 --denom-augment-mode lattice` |
| external pretraining | none |
**This card is written for THIS checkpoint.** Every number below was measured on
C1.3's own generations under the serving code this package ships
(`generate.py` md5 `73b6e49b…`, `min_onset_gap_sec = 0.020`). Nothing is
inherited from the C1 card. Where an item on the C1 card has no C1.3
measurement, it is marked **not measured on this arm** rather than carried over.
**Read the KNOWN LIMITATIONS section before using any number above it.** This
model has real, measured defects, one of which fires a pre-registered kill
condition, and its difficulty ladder is roughly 25 coverage points less nested
than an authored one.
---
## 0. Status: the champion comparison is `inconclusive`
`scripts/sc2_gates.py champion`, three pairings on one serving code base
(`experiments/sc2_eval/FINAL_CHAMPION.md` Β§1):
| pairing | `champion_verdict` | reason code |
|---|---|---|
| C1.2 vs C1 | `inconclusive` | `no_seed_replicate` |
| **C1.3 vs C1** | **`inconclusive`** | `no_seed_replicate` |
| C1.3 vs C1.2 | `inconclusive` | `no_seed_replicate` |
`MIN_SEEDS_FOR_VERDICT` is 2 in **both** arms and C1.3 has one seed. The rule
fires before any metric is read. **No arm has been declared champion.**
What the clause ledger says, which is one-directional: C1.3 **passes** the
required Β§4 clause `plan_token_calibration`, which **C1 fails on both its
seeds**. No clause is passed by C1 and failed by C1.3. The two `a2` span clauses
fail on **all four arms** and belong to the campaign, not to this checkpoint.
The shipping rationale, the two named compromises and the decision rule for
choosing a different checkpoint are in
`experiments/sc2_eval/SHIP_DECISION.md`.
Gate items on the 71-chart hard+oni population, this arm against the C1 seed-1234
arm under identical code:
| | C1 | **C1.3** |
|---|---:|---:|
| fail / pass / na | 9 / 10 / 6 | **8 / 11 / 6** |
| newly passing | β€” | `control_own_axis_density`, `control_named_regression_density_big_count` |
| newly failing | β€” | `a2_span_placement_fit` (knife-edge, Β§2.6) |
---
## 1. What it does
* **Input** β€” 128-bin log-mel at 22 050 Hz (`preprocessor_config.json` freezes
the exact decode/STFT/mel contract), plus a bar grid with explicit measure
edges, exact rational meters and per-bar lattice denominators.
* **Conditions** β€” course, authored level, density bucket, and the split-axis
knobs `big_rate` / `span_rate` / `stream` / `sync`.
* **Output** β€” a BarScript token stream decoded under `BarscriptFSM`, exported
as TJA on the 96-slot lattice.
* **Family decoding** β€” `generate_chart_family` decodes a song's courses as one
nested ladder, hardest first, each easier course biased toward its harder
sibling's onsets. **Off by default**; the campaign's headline population was
decoded independently (Β§2.1).
Capabilities this checkpoint carries (`config.json:capabilities`): `aux`,
`beat_head` (hi-res), `hierarchical_ctx`, `axis_knobs`, `sync_token`,
`stage_emb`, `span_duration_head`, `tempo_head`, `rich_section_stats`. It
carries **no** `style`, `sibling`, `ctx`, `plan`-prefix, `slot`, `dual`,
`align`, `mask_infill`, `func_time`, `global_ctx` or `complexity` conditioning;
requests on those axes reach nothing.
### What works
Evidence marks: **●** n β‰₯ 71 charts with a measured C1 seed band; **◐** n = 19–71
charts or 8–12 sweep cells, one seed; **β—‹** n < 19 or a CI covering zero.
| axis | this checkpoint | reference | mark |
|---|---|---|---|
| **Density knob** β€” median within-song ρ **0.850**, sign p **0.0078**, 8/8 cells positive, **+1.413 nps** firstβ†’last, 12 % monotone adjacency | C1 moves the chart +0.024 / +0.151 nps over the same range and **fails** the gate; C1 seed-to-seed Δρ 0.003 (p 1.000) | 8 songs Γ— 5 buckets, oni | ◐ |
| **`big_rate` knob** β€” ρ **1.000**, p 0.00049, 91.7 % monotone, and **`calibrated: true`** β€” mean absolute bucket error **0.533**, within Β±1 bucket **93.3 %** | first axis in the whole campaign to clear `calibrated`; C1 0.533 vs 1.200 / 0.967 | 12 songs Γ— 5 buckets, oni | ◐ |
| **`span_rate` knob** β€” ρ 0.810, p 0.0117, `calibrated: true`, mean abs err 1.483; and unlike C1 it **no longer drags hit density with it** | C1's span knob flags `onset_nps` / `hit_nps` / `n_onsets` as interference; C1.3 flags only the definitional `span_count` | 12 songs Γ— 5 buckets, oni | ◐ |
| **Span budget** β€” 739 spans against the authored 762 (**0.970Γ—**); `span_rate_gen_per_min` gap to authored **0.049** against a C1 seed band of 0.198 | C1 1.160–1.244Γ—, C1.2 1.259Γ— | 97 charts | ● |
| **FSM guarantees** β€” **0 unclosed, 0 orphan ends, 0 swallowed hits** over 739 generated spans | holds on all four arms | 97 charts | ● |
| **Placement does not degrade** β€” precision 0.6477 (C1 band 0.6480–0.6502), median \|offset\| **0 ms**, `exact_slot_lift_over_null` **1.1755** (above the C1 band), long-song drift tests significant **2** vs C1's 6 / 5 | see Β§2.3 for why apparent recall drops | 71 charts | ● |
| **Accent over-emission reduced furthest** β€” pooled hard+oni big share **0.0873** against C1's 0.1082 / 0.1027 (3.3 seed bands) and an authored 0.0589 | still outside the pre-registered stop window 0.045–0.075 | 71 charts | ● |
| **Cheaper** β€” 1 618.6 tokens/chart, 19.22 tokens/bar, `max_window_tokens_p99` **449** | C1 1 731.2 / 20.51 / 515 | 71 charts | ● |
| **Deployment grid repaired** β€” on a forced `/16` BPM grid, notes Γ· authored **1.130 / 1.112 / 1.067 / 1.027** and onset F1 **0.333 / 0.374 / 0.466 / 0.580** | C1 1.688 / 1.461 / 1.297 / 1.210 and F1 0.318 / 0.376 / 0.477 / 0.599; C1.3 beats C1.2 on **12 of 12** deployable cells | 13 songs Γ— 4 courses, one seed | ◐ |
| **Empty-bar declaration restored** β€” `/16` empty bars 12.82 / 8.60 / 4.55 / 3.14 % against C1.2's 5.96 / 3.23 / 1.08 / 0.33 %, sign-significant on 3 of 4 courses | authored 15.69 / 12.30 / 9.08 / 6.27 % | 13 songs | ◐ |
---
## 2. KNOWN LIMITATIONS
Nothing in this section is softened. Where a limitation is invisible to the gate
suite, that is said.
### 2.1 The difficulty ladder is 25–32 coverage points less nested than authored
Coverage = the fraction of the easier chart's notes that have a note in the
harder chart within tolerance. **Two decode modes, two different numbers, and
both belong on the record.**
| adjacent pair | authored charts | **independent decode** (package default) | **family decode Ξ² = 2** (authored grid) | **family decode Ξ² = 2** (`/16` deploy grid) |
|---|---:|---:|---:|---:|
| easy β†’ normal | 0.970 | **0.658** | 0.730 | 0.726 |
| normal β†’ hard | 0.979 | **0.667** | 0.797 | 0.811 |
| hard β†’ oni | 0.985 | **0.660** | 0.822 | 0.834 |
Independent-decode figures: `FINAL_CHAMPION.md` Β§8.1, n = 13 / 13 / 35 songs, one
seed, same serving code as this package (●/◐). Family-decode figures:
`C13_PREDICTION.md` Β§3.3, n = 13 songs, one seed (◐).
Five things make this worse than the headline:
* **`nesting_coverage` and `hand_agreement` FAIL the family gate on every arm
measured**, C1 included. `difficulty_monotonicity` passes on all of them.
* **C1.3 is slightly worse than C1 here.** C1's independent-decode coverage on
the same population is 0.690 / 0.704 / 0.737; C1.3 is βˆ’0.032 / βˆ’0.037 / βˆ’0.077
against C1 seed bands of 0.010 / 0.021 / 0.025. The move is small against the
~0.30 gap to authored that every arm shares, but it is in the wrong direction.
* **Hand agreement at coinciding hits is near chance on hard→oni**: C1.3 0.539
against a marginal-chance null of 0.500 and an authored 0.753. When two
generated courses agree that a note belongs somewhere, which drum they pick is
near-independent across courses.
* **The `/16` figure overstates the model's nesting.** The same checkpoint on a
`/96` grid drops to 0.649 / 0.730 / 0.752, worse on 11–12 of 13 songs
(p = .003 / .022 / .022). A coarse lattice manufactures agreement by leaving
few places to disagree (`C13_PREDICTION.md` Β§5.2).
* **`family_bias = 2.0` / `family_hand_bias = 1.5` are PROVISIONAL** β€” described
in `generate.py` as logit offsets on a decoder never trained at that setting,
and never calibrated against the authored target.
### 2.2 The note-type channel carries almost no information
71 charts, **22 260 matched slots** (`FINAL_CHAMPION.md` Β§5.2, ●).
| quantity | this model | its own floor | corpus Β§2.1 floor |
|---|---:|---:|---:|
| 4-way accuracy | **0.4776** | majority **0.5559** | majority 4-way **0.563** |
| β€” margin over majority | **βˆ’0.0783**, 95 % CI [βˆ’0.1022, βˆ’0.0530] | | |
| hand accuracy | **0.5493** | always-don **0.5975** | always-don **0.620** |
| β€” margin | **βˆ’0.0482**, CI [βˆ’0.0721, βˆ’0.0229] | | |
| **MI(generated; authored), 4-way** | **0.01028 bits** | β€” | **0.79 % of H(authored) = 1.300 bits** |
| MI, hand | 0.00534 bits | β€” | 0.55 % of 0.972 bits |
* **The model sits below its own majority floor on both endpoints**, and below
the campaign's corpus baselines (56.3 % / 62.0 %), which are a fixed reference
line from a separate 59 961-hit census, not this population.
* **The honest caveat, which cuts the other way.** The published floors are
*argmax* predictors scored against a temperature-1.0 top-p-0.95 *sample*.
Against the sampler-appropriate i.i.d. floor (Ξ£ pΒ² = 0.4582 4-way, 0.5190
hand) this model is **above** baseline by **+0.0194** and **+0.0303** β€” the
largest 4-way excess of any arm in the campaign. Both readings belong on the
record.
* **MI is flat at ~0.010–0.012 bits on all four arms** (seed band 0.0008). The
accuracy differences between arms are marginal-matching, not information.
* Accuracies are conditional on coverage **0.6442** (precision side) /
**0.5842** (recall side); more than a third of generated hits have no authored
partner, and coverage is 9 % lower than C1's, so this population is smaller
and differently selected than C1's.
* **No inter-charter agreement ceiling exists** for this split β€” no
(song, course) carries two independent authored charts β€” so 100 % is not a
legitimate target and is not used as one.
* **Greedy re-decoding was not repeated on this arm.** The C1-era finding that
greedy closes part of the gap at an unacceptable `motif_reuse` cost is **not
measured on this arm** and must not be quoted for it.
### 2.3 The note budget is 9 % short, and it is the whole of the apparent timing regression
| | C1 (s1234 / s4321) | **C1.3** |
|---|---:|---:|
| `onset_precision` | 0.6480 / 0.6502 | **0.6477** |
| note budget (gen Γ· authored notes) | 1.0046 / 0.9912 | **0.9069** |
| `onset_recall` | 0.6510 / 0.6445 | **0.5874** |
| `frac_of_authored_within_jnd` | 0.6463 / 0.6397 | 0.5840 |
| `type_coverage_recall` | 0.6465 / 0.6401 | 0.5842 |
`precision Γ— budget` equals `onset_recall` to four decimals on **every** arm, by
construction. Those three "timing" rows are one quantity, and it is the note
budget, not placement: precision is inside the C1 seed band, median offset is
0 ms, `exact_slot_lift_over_null` is *above* the band, and significant drift
tests fall from 6 / 5 to 2.
**The shortfall itself is a real regression and its cause is unknown.** It
appears on both C1.2 and C1.3, so it belongs to the `--cal-axes` /
`--balance-exclude-big` pair rather than to the denominator augmentation. It is
the largest practical regression in this release.
### 2.4 Raising density squeezes drumrolls out β€” this is the compromise a user can hit
Median spans/min over eight oni songs, per requested density bucket
(authored **3.885**; `FINAL_CHAMPION.md` Β§4.3, 8 songs Γ— 5 buckets, ◐):
| requested density | 0 | 3 | 7 | 11 | **15** |
|---|---:|---:|---:|---:|---:|
| C1 (s1234) | 3.629 | 3.055 | 4.153 | 2.818 | **2.971** |
| **C1.3** | 3.550 | 4.413 | 2.172 | 2.727 | **1.279** |
ρ(density β†’ `span_per_min`) = **βˆ’0.759**, sign p 0.0078, firstβ†’last βˆ’2.443 /min.
At the top of the density request C1.3 delivers **33 %** of the authored span
rate; C1 delivers 76–79 %, C1.2 44 %.
* **Do not advertise the density knob at its top setting** until this is fixed.
* There is a named suspect: `cal_span` never left **0.1667** on either
challenger, so this is plausibly a calibration term that never trained rather
than an intrinsic cost of the knob.
* The mirror image is a *gain*: C1.3's `span_rate` knob no longer drags hit
density, which C1's does.
### 2.5 Absolute bucket calibration is still wrong on the density axis
| axis | exact bucket | within Β±1 | mean abs error | mean signed | `calibrated` |
|---|---:|---:|---:|---:|:--:|
| density | 5.0 % | 25.0 % | 4.125 | βˆ’1.075 | **false** |
| `big_rate` | **53.3 %** | **93.3 %** | **0.533** | +0.267 | **true** |
| `span_rate` | β€” | β€” | 1.483 | β€” | true |
The density knob orders the request correctly and lands in the wrong bucket,
under-shooting by ~1 bucket on average β€” more than C1.2 does. Rank control is
real; absolute density targeting is not delivered.
Note the cost that comes with the `big_rate` calibration win: C1.3 has the
**shortest** `big_rate` ladder of any arm (first→last +0.159 against C1's
0.183 / 0.224), so the top of that request now delivers less accent than C1's
did.
### 2.6 It fails its own span-quality gate and triggers a pre-registered kill condition
71 charts hard+oni (467 closed spans) and 26 charts easy+normal (272):
| axis | hard+oni | easy+normal |
|---|---|---|
| `duration_low_tail` | **fail** (rate verdict passes; balloon 0.250 beats = 0.75Γ— the shortest authored balloon, roll 0.250 = 0.60Γ—) | **fail** (balloon 0.250 beats = 0.20Γ— the shortest authored balloon; 45/159 balloons below authored support) |
| `forced_close` | **fail** β€” **7 / 467 = 1.499 %** (0 unclosed, 7 clamp-forced) | **fail** β€” **7 / 272 = 2.574 %** |
| `over_span_max` | **fail** β€” one balloon 7.72 s = 1.19Γ— the 6.5 s ceiling | **fail** β€” one balloon 7.58 s = 1.17Γ— |
| `placement_fit` | **fail** β€” AUC **0.560** (CI [0.529, 0.593]) against a 0.56 floor; authored reference 0.668 | pass β€” AUC 0.604 |
| `orphans` / `swallowed` / `rate` / `duration` | pass | pass |
| **`a2_kill_condition_not_triggered`** | **fail β€” the kill condition FIRED** | fail |
The pre-registered kill condition (`SOFTCHART2_DESIGN.md` Β§4-A2) reads: *rate is
calibrated but placement/length quality is bad β†’ concede that the bar-scoped
commitment mechanism is insufficient and escalate to an independent span-state
objective.* It fires on **all four arms**, including C1. Shipping does not
discharge it.
**`placement_fit` is a knife-edge, not a finding.** C1.2 passes at 0.561 and
C1.3 fails at 0.560 with CIs overlapping every other arm across their whole
width. One thousandth of AUC decides the verdict. It is reported, not used.
### 2.7 One span in seven is a micro-span
All generated spans, 97 charts, `FINAL_CHAMPION.md` Β§2 (●):
| | n spans | p1 | p5 | p50 | p90 | min | **share < 0.25 s** |
|---|---:|---:|---:|---:|---:|---:|---:|
| authored | 762 | 0.177 | 0.273 | 0.963 | 2.419 | 0.083 | **3.8 %** |
| **this model** | 739 | 0.089 | 0.150 | 0.818 | 2.258 | **0.072** | **14.7 %** |
Per course: easy 6.6 %, normal 8.1 %, hard 13.7 %, oni **25.2 %** β€” authored
3.0 / 3.3 / 2.1 / 7.1 %.
* C1.3 does **not** inherit C1.2's worsening (16.4 %); 14.7 % is exactly at the
top of the C1 seed band (13.6–14.7 %). It is not an improvement either.
* **Mechanism is known and unfixed on every arm.** Duration bucket 0 is open at
the bottom; 123 of C1.3's 739 spans sit in it, at a median 0.5 beats against
the authored 0.917.
* **The shipped floor does not fix it.** `span_min_sec` / `span_min_beats` fired
263 times (`span_close_floor`) plus 33 carries across 97 charts and the
< 0.25 s share is still 14.7 %.
* **The mitigation is unbuilt.** `span_bucket0_profile` is `null` in this
package because `scripts/build_span_bucket0_profile.py` has never been run on
the train split. This is the shortest available fix and it is not in the box.
* **The gate cannot see it directly.** `a2_span_duration` compares p50/p90/p99
only and passes; `duration_low_tail` catches the support violation but not the
rate (its rate verdict passes on this arm).
* On a `/16` deployment grid the same quantity is 14.7 / 12.4 / **30.6** / 21.8 %
β€” hard is untouched by the augmentation and is the worst cell in the release.
### 2.8 The deployment grid has a two-role conflict that no checkpoint dissolves
`BAR_DENOM` is simultaneously an input the grid supplies and a **trained output
token** the model reads as a density announcement. `bar_denoms="supplied"`, the
shipped default, teacher-forces it.
13 songs Γ— 4 courses, one seed, `--cond authored`, `C13_PREDICTION.md`:
| easy / normal / hard / oni | authored grid | **forced `/16`** | `/96` | `/24` |
|---|---|---|---|---|
| notes Γ· authored | 0.917 / 0.919 / 0.875 / 0.899 | **1.130 / 1.112 / 1.067 / 1.027** | 0.944 / 0.950 / 0.943 / 1.027 | 1.102 / 1.070 / 0.950 / 0.921 |
| onset F1 | 0.613 / 0.565 / 0.595 / 0.647 | **0.333 / 0.374 / 0.466 / 0.580** | 0.303 / 0.325 / 0.393 / 0.535 | 0.342 / 0.370 / 0.442 / 0.542 |
| genuine triplets (den 3/6/12/24) % | 1.13 / 2.48 / 6.34 / 9.43 | **0.00** (arithmetically impossible) | 3.45 / 5.78 / 10.57 / 11.46 | 14.13 / 20.31 / 28.57 / 38.70 |
| ultra-fine (den 48/96) % | ~0 | 0.00 | **11.53 / 13.22 / 15.62 / 6.60** | 0.00 |
| nesting (family Ξ² = 2) | 0.730 / 0.797 / 0.822 | 0.726 / 0.811 / 0.834 | 0.649 / 0.730 / 0.752 | 0.714 / 0.769 / 0.779 |
Authored triplet rates for reference: 4.20 / 5.72 / 7.43 / 9.95 %; ultra-fine
~0.00 %.
**The recommended deployment configuration is `--grid bpm --grid-bpm-denom 16`
with the package default `bar_denoms="supplied"`.** Its two costs, stated
plainly: **triplets are exactly 0 and cannot be otherwise**, and hard micro-spans
stay at 30.6 % against an authored 2.1 %.
Four consequences a deployer must carry:
* **Deployment costs 10–46 % of onset F1** relative to the authored grid:
easy βˆ’46 %, normal βˆ’34 %, hard βˆ’22 %, oni βˆ’10 %. Every `--grid authored`
number in this card is therefore optimistic for a user who supplies only BPM.
* **`p90 |offset|` is 0 ms on the authored grid and 14.6–36.7 ms on the bpm
grids.** That residual belongs to the bar edges; no denominator choice touches
it.
* **The model's off-dyadic rate is a property of the model, not the lattice.**
Given any 3-divisible grid it puts 2–3Γ— the authored share off the dyadic
grid; the grid only decides whether that mass gets called "triplets" (`/24`)
or "ultra-fine jitter" (`/96`). This is a training-side defect and no
deployment grid fixes it.
* **The gate suite cannot score the deployment path.** `robustness.py:629-642`
raises `ContractError` on uniform-BPM arms, so every gate timing metric is
`null` for the path the model ships on; the `/16` figures above come from
`timing.py`, a second implementation that reproduces the gate on the authored
grid.
**The `deploy_bpm_grid` profile in this package is NOT the recommendation.** It
sets `bar_denoms="model"`, `bar_denom_mask="div96"`, per `DEPLOY_GRID_FIX.md`
Β§7 β€” a recommendation made explicitly conditional on C1.3 failing, which it did
not. That arm was measured on **C1 only**: density 1.06 / 1.08 / 0.97 / 1.01 and
nesting 0.635 / 0.706 / 0.778. **It has never been run on this checkpoint.** The
profile is kept because it is implemented and because the arm is the obvious
next experiment, not because it is advised here.
**One contradiction inside this package, stated so nobody has to find it.**
`config.json:serving_profiles.deploy_bpm_grid.basis` still quotes
`DEPLOY_GRID_FIX.md` Β§7 verbatim, including the words *"Recommended now"*, and
still carries C1's nesting figures. That string is generated by
`scripts/sc2_package.py`, which is read-only for this release. **This section
supersedes it.** The generated string is left intact rather than silently
diverging from the shipped code.
### 2.9 Playability floor: gaps no human hand can play
Recomputed on this arm's own events, 71 charts hard+oni, generated vs authored
on the same songs (producer `experiments/sc2_eval/ship_c13/playability.py`,
records `.../records/playability_all_arms.json`; ●):
| | this model | authored |
|---|---:|---:|
| inter-onset gaps < 40 ms | **102** across **29** charts | 2 across 1 chart |
| β€” as a share of adjacent pairs | **0.296 %** (102 / 34 484) | **0.0053 %** (2 / 38 032) |
| gaps below the series record 29.4118 ms | **19** | 0 |
| shortest gap | **20.8 ms** (the serving floor) | 39.7 ms |
| peak burst in any 1 s window | **17 notes** | 15 notes |
C1 on the same population: 151 gaps (0.395 %), shortest 16.7 ms, peak 22
notes/s. C1.3 is better on every row and still **56Γ— the authored rate**.
* `min_onset_gap_sec = 0.020` is a **degeneracy guard, not a fix**: it is set
below the fastest thing the series has shipped (29.4118 ms,
TAIKO-TONGUE-TWISTER oni, BPM 170, 48th notes) precisely so it cannot refuse a
chart the domain writes. It fired **125 times** across 97 charts with **0**
fail-opens. *A mask cannot fix a distribution.*
* The `generate.py` constant comment quotes the generated sub-40 ms rate as
~0.35 % against ~0.045 % authored. **That authored figure does not reconcile
with the 0.0053 % measured here**, and the two have never been put on the same
denominator. The excess is 8Γ— on the comment's accounting and 56Γ— on this one.
* **Good news that replaces a C1-card claim.** The C1 card reported the model
placing hits in 19.6 % of sub-0.25 s scaffolding bars. Under today's serving
floors this arm places **0 hits in all 112 such bars** across 97 charts, exactly
as the authored side does; `degenerate_bar_hits_blocked` fired 12 times with 0
fail-opens.
### 2.10 Pattern proxies regress against C1 on hard+oni
71 charts, one code version, C1 band = |s1234 βˆ’ s4321| (`FINAL_CHAMPION.md` Β§3):
| metric (lower is better) | C1 mean | C1 band | **C1.3** | Γ—band | relative |
|---|---:|---:|---:|---:|---:|
| **`motif_reuse_gap`** | 0.3630 | 0.0042 | **0.4222** | +14.1 | **+16.3 %** |
| `compression_gap` | 0.0976 | 0.0006 | 0.1116 | +23.3 | +14.3 % |
| `motif_ref_marginal_js` | 0.0331 | 0.0046 | 0.0452 | +2.6 | +36.6 % |
| `ioi_js_per_chart` | 0.0295 | 0.0006 | **0.0352** | **+9.7** | +19.3 % |
| `ul_4gram_js_per_chart` | 0.2402 | 0.0055 | 0.2530 | +2.3 | +5.3 % |
| `class_4gram_js_per_chart` | 0.1337 | 0.0080 | 0.1357 | +0.2 | +1.5 % |
The `Γ—band` column overstates the case β€” several bands are under 1 % of their own
level β€” so the relative column is the honest one. On `motif_reuse` the bootstrap
CIs are nonetheless **disjoint**: C1 [βˆ’0.392, βˆ’0.334] / [βˆ’0.388, βˆ’0.334] against
C1.3 [βˆ’0.447, βˆ’0.396]. **The regression is real and not seed noise.**
Four bounds, all measured, none of them a dismissal:
* **It does not happen on easy+normal.** Every pattern proxy there is inside the
C1 seed band, `motif_reuse_gap` is marginally *better* than C1's mean, and
`class_4gram_js` is better on both challengers.
* **The decomposition puts the loss on rhythm and hand, while the accent layer
improves**: rhythm 0.1251 β†’ 0.1640, hand 0.1774 β†’ 0.2102, accent 0.0606 β†’
**0.0480**.
* **`ioi_js_per_chart` is where the augmentation shows up** β€” it reads the exact
rational IOI lattice, and it is C1.3's worst pattern metric relative to C1.2
(+9.7 band against +1.5).
* **The proxy reads mostly a channel carrying ~0.010 bits** about the authored
type (Β§2.2), and it has **never been validated against a listener**. It is
quoted because Β§4 names it, not because it is strong.
Realized `motif_reuse` is **0.3165** against an authored **0.7387** β€” under half
the authored repetition.
### 2.11 The note-type marginal, and one C1 finding that does NOT reproduce
Big notes are **1.48Γ— the authored rate** pooled hard+oni (0.0873 vs 0.0589),
worst on the easiest course:
| course | authored | **C1.3** | C1 (s1234 / s4321) |
|---|---:|---:|---:|
| easy | 18.80 % | **26.22 % (1.40Γ—)** | 32.68 % / 29.65 % |
| normal | 11.75 % | **17.41 % (1.48Γ—)** | 19.21 % / 20.77 % |
| hard | 7.38 % | **10.61 % (1.44Γ—)** | 13.11 % / 12.73 % |
| oni | 4.88 % | **7.40 % (1.52Γ—)** | 9.29 % / 8.58 % |
The pre-registered stop window (0.045–0.075 pooled hard+oni) is **not met**.
**The C1 card's Β§2.8 claim does not reproduce here and is corrected.** On C1 the
realized class marginal was 13Γ— closer to the training loss weights than to the
corpus. On C1.3 it is the other way round: JS(gen β€– authored) = **0.00588**
against JS(gen β€– class-weight prediction) = **0.00746**
(JS(authored β€– weights) = 0.01452). `--balance-exclude-big` moved the marginal
off the loss weights and toward the corpus. The log-log fit of realized share
against class weight collapses from slope 0.378 (ρ 0.436) on C1 to slope 0.078
(ρ 0.156) on C1.3.
### 2.12 Fine-lattice, tuplet and syncopation accuracy are much worse than average
`robustness.py` timing strata, 71 charts, exact-slot rate (overall **0.5840**):
| stratum | n authored | exact-slot | C1 (s1234) |
|---|---:|---:|---:|
| lattice step 11.6–23.2 ms | 114 | **0.342** | 0.456 |
| lattice step 23.2–46.4 ms | 790 | **0.443** | 0.465 |
| lattice step β‰₯ 92.9 ms | 24 228 | 0.594 | 0.651 |
| positions with a denominator divisible by 3 | 3 480 | **0.345** | 0.496 |
| syncopation band 0 β†’ band 5 | 20 846 β†’ 1 669 | **0.615 β†’ 0.434** | 0.672 β†’ 0.514 |
| local IOI 25–50 ms | 189 | **0.402** | β€” |
| BPM > 250 | 2 570 | **0.499** | β€” |
Every stratum is lower than C1's, which is the note-budget effect of Β§2.3 acting
on a per-stratum recall, but the *shape* is the finding: tuplet positions and
fast lattice steps are 0.24 below the overall rate, and accuracy decays
monotonically with syncopation.
### 2.13 Smaller, but on the record
| finding | number | n | source |
|---|---|---|---|
| `long_song_no_drift` fails | **2** of 21 metrics significant at BH q ≀ 0.05 (`dens_signed_err`, `plan_dens_bucket_tvd`) β€” C1 fails 6 / 5 | 71 charts | `robustness.json` ● |
| Harness-level generation failures | **6** of 71 song-courses refused (`#BRANCHSTART` and non-representable meters); **1** chart exported `gen_tja = null` β€” identical to C1 on the same population | 71 charts | `index.json` ● |
| Slot-export off-lattice events | 8 across 97 charts (C1: 5) β€” the documented contract path, not a crash | 97 charts | decoder counters ● |
| `escape_hatch` fires | 91 times across 97 charts (C1: 84) | 97 charts | decoder counters ● |
| Generation wall time | 9.29 s/chart, 4.33 s per audio-minute β€” **contended**, the GPU was shared, quoted as a declaration only | 71 charts | ● |
| Complete-bar rest placement | **not measured on this arm** (C1: Jaccard 0.474, recall 0.623) | β€” | β€” |
| Greedy-decode contrast | **not measured on this arm** | β€” | β€” |
| Micro-span seed sensitivity | **not measured on this arm** (C1: 24.0 % vs 32.4 % across two sampling seeds on 19 oni charts) | β€” | β€” |
### 2.14 The evidence base is thinner than it looks
* **One seed.** Every number in this card is n = 1 in the training seed. The C1
seed band quoted throughout is C1's, used as a reproducibility scale; it is a
range over n = 2 and carries no confidence statement.
* **easy and normal are evaluated on 13 songs.** The 71-chart population is
`hard` + `oni` only; the easy/normal population is 26 charts from 13 songs, and
its seed band is 4–5Γ— wider than hard+oni's.
* Every difficulty-family and deployment-grid number is **13 songs, one seed**
(◐). Every knob sweep is **8–12 songs Γ— 5 buckets** (◐).
* At these sizes a gate `pass` carries little information. Worked example from
this campaign: the `big_share_of_hits` gate **flips between two C1 training
seeds** while the failing seed's point estimate is *smaller*. At n = 36 that
gate cannot rank checkpoints, and no pass/fail on it should be quoted as
evidence for any arm.
* **Nothing here is audio-referenced** except `a2_span_placement_fit`. Every
other quantity is chart-vs-chart on the exact rational lattice or on hit order.
No timing window and no game judgement parameter appears anywhere in this path.
* **`best_val` is comparable across C1 / C1.2 / C1.3** (one training cache) and
**not** comparable to C1.1 or to any arm on a different cache. C1.3's 0.5769 is
**2.9 % worse** than C1.2's 0.5605; that was the registered trade and the
deployment side won it.
* **No plan-neutral arm was generated**, so the `plan_neutral_fallback` Β§4 clause
is `na` and the ledger is incomplete by one required clause.
### 2.15 Serving regime: what `motif_constraint: auto` resolves to, and why it matters less here
`motif_constraint: "auto"` resolves to **OFF** for this checkpoint: the training
cache's `barscript_md5` (`634e3dc3…`) differs from the serving `barscript.py`
(`9afc715a…`), so the gold sequences it trained on never satisfied the MOTIF hard
constraint. `train_serve_matched = true` β€” OFF is the train/serve-matched choice
and it is the right default.
**Unlike the C1 package, this card needs no correction for it.** Every number in
this card was measured with the constraint **OFF**, i.e. under exactly this
package's resolved default; the 20 evaluation runs all record
`serving_fsm.motif_constraint: off, verified: true`. In particular
`motif_ref_marginal_js` here is **0.0452** (hard+oni) and **0.0114**
(easy+normal), both measured OFF.
Two things that still belong on the record:
* **The ON/OFF blast radius was measured on C1 only**, where switching the
constraint on moved `motif_ref_marginal_js` 0.0311 β†’ 0.0189 (βˆ’39 %). On the C1
package the shipped default therefore makes the correct value **0.0311, not the
0.0189 in the older tables**. **That contrast has never been measured on
C1.3**, so no ON-constraint number should be quoted for this checkpoint.
* Serving under the constraint ON would be train/serve **mismatched** for this
checkpoint and is not a supported configuration.
---
## 3. Serving contract
Every value below is in `config.json:serving`, with its justification in
`config.json:serving_basis`. Pass them explicitly β€” the package's
`serving_kwargs()` does β€” so the recorded contract is the one that reaches the
decoder.
| parameter | value | one-line basis |
|---|---|---|
| `min_onset_gap_sec` | **0.020** | Degeneracy guard, **not** a corpus percentile. Must stay strictly below the series record of 29.4118 ms (BPM 170, 48ths). Two earlier corpus-derived values (0.0395, 0.0300) were both wrong. See Β§2.9. |
| `degenerate_bar_sec` | **0.25** | Authored charts place zero hits in any bar under 0.375 s (3 010 bars). Largest round threshold with zero counterexamples and 1.5Γ— margin; identical to the gate's threshold. Blocks hits only, never span geometry. |
| `span_min_sec` / `span_min_beats` | 0.0833 / 1⁄3 | Authored population minima. Fail-open, counted. Does not fix Β§2.7. |
| `span_bucket0_profile` | **null** | Not built β€” see Β§2.7. |
| `family_bias` / `family_hand_bias` / `family_mode` | 2.0 / 1.5 / `bias` | **Provisional**, never calibrated β€” see Β§2.1. Family decode is opt-in. |
| `motif_constraint` | `auto` β†’ resolves **off** here | Train/serve matching by `barscript.py` md5 β€” see Β§2.15. |
| `bar_denoms` / `bar_denom_mask` | `supplied` / `lattice` | The evaluated regime, and the recommended deployment regime on a uniform `/16` grid. The `deploy_bpm_grid` profile exists but is **not** recommended for this checkpoint β€” see Β§2.8. |
| `greedy` / `temperature` / `top_p` | false / 1.0 / 0.95 | As evaluated. |
| `plan_temperature` | `null` (follows `temperature`) | Never exercised in evaluation; shipped unset rather than tuned. |
`_has_sync` **must** be forced on at load. `train.py` records
`sync_token=False` while the training prefix carries the SYNC slot, so a loader
that trusts the recorded flag drops the slot and every sync bucket reaches the
decoder as an identical prefix. The package loader does this and records why in
`config.json:serving_prefix_fix`. This is a workaround; the fix belongs in
`train.py`.
### Recommended deployment configuration
```
--grid bpm --grid-bpm-denom 16 # uniform /16 grid
bar_denoms = "supplied" # the package default
motif_constraint = auto -> off # resolved at package time
family decode: optional; Ξ² = 2.0 / 1.5 is provisional
```
Expected behaviour under it, 13 songs Γ— 4 courses, one seed: notes within
13 / 11 / 7 / 3 % of authored; onset F1 0.333 / 0.374 / 0.466 / 0.580;
**zero triplets**; hard micro-spans ~30 %.
---
## 4. Weights are bfloat16
Training ran bf16 autocast and CUDA inference runs bf16 autocast, so fp32
storage carried no information the forward pass could use β€” the argument the
v1.5 release made and verified. Measured cast cost: max absolute delta
**0.00711** on `frontend.2.weight` (0.349 % of that tensor's max) over 9 107 291
float elements.
This is an argument about the **compute** dtype, not a proof that the two
checkpoints decode identically. Sampling is chaotic in the logits, so individual
charts can differ. The paired fp32-vs-bf16 evaluation has **not** been repeated
for BarScript. The fp32 checkpoint is preserved in the source run directory and
is the reference for any bit-level comparison.
---
## 5. Provenance
| | |
|---|---|
| serving code | `generate.py` md5 `73b6e49b00fb60f3229c8539bc7a1029`; all 18 `src/softchart/*.py` md5s in `config.json:code.module_md5` |
| evaluation harness | `sc2_generate_eval_v2_2_serving_floors`, contract `sc2_eval_v2` |
| evaluation runs | 20 run indices, 1 028 charts, one `serving_floors` signature, `comparable: true`, 6/6 floor counters present |
| headline population | 71 charts hard+oni (36 songs) + 26 charts easy+normal (13 songs), `--cond authored`, `--grid authored`, seed 1 |
| deployment population | 13 songs Γ— 4 courses Γ— 10 arms, `--cond authored`, one serving tree, one signature |
| label spec | `spec.json`, field-for-field equal to the cache manifest's copy (`spec_parameters_match_cache_manifest: true`) |
| uploaded | **no** (`release_manifest.json: "uploaded": false`) |
Reports: `experiments/sc2_eval/SHIP_DECISION.md`,
`experiments/sc2_eval/FINAL_CHAMPION.md`,
`experiments/sc2_eval/C13_PREDICTION.md`,
`experiments/sc2_eval/C12_EVAL.md`, `experiments/sc2_eval/DEPLOY_GRID_FIX.md`.