YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

No fMRI claim may be read off this model, and this time we know why

The measured BOLD path in this run does integrate the Balloon-Windkessel ODE -- that is what run 4 fixed (ISSUE-008). The haemodynamic likelihood is real here in a way it was not in run 3.

It diverges during training, and the full run settled how far. real_bold_nll ran 1.99 to 36,472 over 14,600 steps -- a factor of about 18,000 -- while eeg_nll IMPROVED (1.74 to 1.50), the total loss stayed flat near 1.0, and bold_log_scale held at 5.3-5.9. So this is not a variance explosion and the fMRI term is not dominating the mixture: it is the one term getting worse while everything around it gets better. It never plateaus: T4 alone spans 1,530 to 650,815, four orders of magnitude inside one stage. And T5's measured return does not repair it -- that stage grants bold.* again and ends at 36,472. bold_parcels_covered held at full value throughout, so the ODE ran on every parcel for 46 hours and the likelihood it produced is worthless.

Four diagnostic arms located the cause: the shared trunk moves out from under the BOLD head. ds002336_real is 5.39% of the source mixture and is outvoted 17.6 : 1 by the EEG-like sources, so the latent state converges on what they want. Freezing the five Balloon parameters changes nothing; freezing the trunk makes the BOLD likelihood improve.

ds002336_real appears in contributed_sources and that is accurate: its BOLD channel contributed a gradient. Whether that gradient carried information is a separate question, and this run's leave-one-source-out answers it: removing the corpus made measured EEG prediction worse (+0.0010), so the information went into the shared state and does not come back out through the BOLD head. Both things are true at once and neither licenses an fMRI claim.

No fMRI, haemodynamic or neurovascular claim about this artifact is supported. ISSUE-016 is open; the remedy is unshared capacity for the slow modality, which is a later run's design, not a caveat on this one.

No inference or parameter-recovery claim may be read off this model. ISSUE-012 is open and this run MEASURED it rather than inheriting it. The learning-rate repair worked -- best log_G R^2 0.284, against ~0 on every parameter in run 3, so the flow now reads its conditioning -- and it overshot: worst posterior_z_sd 59.3 where a calibrated posterior sits near 1.0, SBC KS p_min 1.03e-147, coverage MAE 0.203. The posterior narrowed far more than its accuracy earned, so it is confidently wrong rather than uninformative. 4 of 6 parameters still explain no variance (log_velocity, ei_global, ei_gradient, drive). Run 3's posterior was uninformative and honest; this one is partly informative and overconfident. Neither supports inference.

No individualisation or personalisation claim may be read off this model, and this run MEASURED that rather than inheriting it.

session_individualisation scored 75 participants over 1500 held-out second-night windows -- the same people on both sides of a SESSION split, which is the only arrangement on which a person effect is measurable at all. Held-out session NLL 2.0436 [2.0048, 2.0897], bootstrapped over participants rather than windows.

That score is not the finding. The between-participant spread of the applied theta shift is 0.000706 -- 0.67% of the scale the model allocated for that effect. 30 of the scored person-effect rows are exactly zero. The individualiser applied essentially nothing on a split built specifically to let it apply something, which is the falsifier this evaluation declares for the capability. Earlier runs reported individualisation as unmeasurable on a participant-disjoint split; this one built the split, trained the effect and measured it, and the effect is a fraction of a percent of its own scale.

The held-out NLL is reported because withholding a measured number is its own distortion. It answers a different question than it appears to: what separates an individualised model from the population model here is the shift, not the score.

Which sources earn their place: leave-one-source-out on the MEASURED holdout

Each arm drops one source family, retrains 200 steps, and is scored on the same held-out participants the headline rests on. Positive delta means removing the family made measured prediction WORSE, i.e. the family contributed.

2 of 10 families contributed on the measured holdout; 8 showed negative transfer (anatomical_prior, ds000117_behaviour, ds000117_real, ds004024_perturb, ds004024_rest_real, montage_calibration, sim_wholebrain, sleepedf_real). Largest positive deltas: eegmmidb_real +0.0144, ds002336_real +0.0010.

The simulated-holdout arm of the same run is retained for comparability with earlier runs and is NOT the result: 5 of 10 families come back as negative transfer there. Scoring a measured source against the simulator asks whether dropping it helps the model fit the simulator, which is a different question from the one above and is not evidence about the source.

The EEG lead field remains an analytic sphere, not a head model, so no source-localisation claim is available.


license: other license_name: see-licence-section tags: - neuroscience - eeg - brain-dynamics pretty_name: scwbd-004 (treatment arm)

scwbd-004

Read this first: this checkpoint loses to copying the last observed sample forward. It is published as a negative result and as a reference artifact for others, not as a working model. If you are looking for a brain-dynamics model that works, this is not it.

This model never saw measured data during training, and four other training mechanisms were silently off. This run renamed its training stages; six gates in the trainer match on the previous run's stage names, and five of the six therefore gave the wrong answer:

gate decides result for this run
admit measured sources refused -- no gradient was ever taken on real EEG
per-stage gradient allowlist wildcard -- no restriction applied
boundary randomisation of sim inputs off
haemodynamic state in the rollout off
build the individualizer off -- the stage named for it ran ordinary training
admit simulated sources admitted, and correct only by accident

Nothing raised, because every gate fails toward permissive: an unmatched stage name means "no restriction" rather than "unknown stage". The scores below are therefore a simulation-to-measurement transfer result from a partially configured trainer -- not a held-out-performance result -- and a stage named for measurement is not evidence that measurement occurred. See reports/RUN2.md section 2b.

The comparison below flatters this model, and the correction is not applied to the numbers. evaluate.py scores SC-WBD on y = target / s, where s is each window's own standard deviation, with the Jacobian folded into the log-variance. Every baseline is scored on the raw target. The algebra is exact and model-independent: NLL_scaled = NLL_raw - log s, so the two sides are different random variables and SC-WBD's figure is the smaller one.

Measured on this test fold, mean(log s) = 0.5694 nats -- against a spread of 0.035 nats across the three non-trivial baselines, so the offset is roughly 17x the entire spread it is being compared against. In the baselines' units SC-WBD's NLL is approximately 3.75 rather than the 3.179 tabulated, and the gap to the best baseline is about 1.70 nats rather than 1.13. MSE is off by 1/s^2 and does not cancel.

The rescale is harmless during training -- s does not depend on the parameters -- and is a pure unearned advantage at evaluation. No verdict changes: every paired interval already excluded zero and the correction moves all of them further from SC-WBD. The numbers are left as measured rather than silently adjusted, because re-scoring both sides on the raw target is the fix and arithmetic on published figures is not.

The headline

No baseline beats scwbd-004, and scwbd-004 is not shown to beat ar16, var4 (3 of 5 comparators separated) on the paired participant-clustered 95% interval of the per-window NLL difference

Metric: gaussian NLL, nats per channel per sample, sensor space, participant-clustered 95% CI

arm NLL 95% CI MSE params
scwbd-004 ← 2.0244 [1.9750, 2.0903] 3.9822 26,304,657
ar16 2.0345 [1.9696, 2.1349] 3.9856 5,184
var4 2.0371 [1.9689, 2.1439] 3.9622 20,544
population_gaussian 2.0589 [1.9975, 2.1526] 4.1505 2,208
persistence 2.3274 [2.2546, 2.4305] 7.3341 4,096
dense_neural 3.4062 [3.2229, 3.6615] 14.3610 26,296,869

Paired participant-clustered differences (positive = SC-WBD worse):

vs Ξ” NLL 95% CI excludes zero
ar16 -0.0100 [-0.0480, 0.0144] False
var4 -0.0127 [-0.0561, 0.0146] False
population_gaussian -0.0345 [-0.0654, -0.0133] True
persistence -0.3030 [-0.3427, -0.2691] True
dense_neural -1.3818 [-1.5750, -1.2380] True

Why it lost β€” the diagnosis

Two things, and neither is 'the architecture does not work'. Both are documented in the repository, not inferred here.

1. It is the treatment arm, and that is still not the ablation. Regional state here is family-indexed and heterogeneous, so unlike run 1 these numbers describe the architecture the thesis argues for rather than its control. But the thesis claim is comparative -- structured regional state against pooled state at matched capacity -- and the pre-registered ablation needs five further arms that do not exist: two capacity-matched pooled controls, a scalar floor, a theta-conditioned control, and a permuted-family attribution control. Against generic forecasting baselines this artifact can lose or win without either outcome attributing anything to the structure. Read it as the candidate arm measured, not as the hypothesis tested.

2. The whole loss is in the variance channel. On the conditional mean this model beats every baseline including persistence β€” its MSE is the best in the table above. It loses on NLL because a single per-channel scalar (eeg.log_noise) sets the predictive variance, was left to SGD instead of its closed-form optimum, and ended up uniformly overconfident. The scalar cannot represent horizon-dependence at all; the baselines' variance can.

This is a useful shape to know about: a model can win on point prediction and still lose decisively on likelihood because one uncalibrated scalar dominates the score.

What it was trained on

  • Anatomy: 414 regions β€” atlas Schaefer400x7 in fsLR/32k, subcortex Aseg14T, 10 declared sources (enigma_hcp_sc, hansen_receptors, hcps1200_maps, hill2010, margulies2016, neuromaps, raichle_metabolism, schaefer2018, sydnor2021, tian2020); is_biological = True. Full provenance, including every licence and citation, is carried inside the checkpoint under extra.anatomy and in reports/anatomy_prior.md.
  • Lead field: 64 channels, provenance analytic_sphere_fallback, individual head model: False.
  • Evaluation split (this model was not fit to any of it β€” see the disclosure above; the baselines are fit to the first half): 2010 fitting windows / 67 participants; 1000 test windows / 25 participants, participant-disjoint.

Known defects

  • Five of six training-stage gates gave the wrong answer (see the disclosure at the top). No measured-data gradient, no per-stage gradient restriction, no boundary randomisation, no haemodynamic state in the rollout, and no individualizer. A complete fix existed in the repository before this run started and was not applied; six tests naming the defect were failing on the main branch throughout. This is the defect that most changes how the scores should be read.
  • The modules ['tau_prior'] are named by no card's gradient_permission. They carry no parameters in this checkpoint's parameter report, so the share of the model that could not receive a gradient is 0.0% -- this is a completeness note about the cards, not a finding about the weights. Computed from the source cards at b20f368, the commit this checkpoint records. That commit is recorded with a -dirty suffix, so the tree that trained carried uncommitted changes and the cards it used may differ from the cards at the commit.
  • Individualisation did not happen: 25 participants scored, 0 individualised, 25 still at initialisation.
  • subject_specific_ar is not in the table. DROPPED from this table under baseline protocol v2 (ISSUE-013). Refusal R10 makes the fit and score participant sets disjoint, so 100% of scored windows routed to the pooled fallback and the row was bit-for-bit ar16 -- a duplicate carrying the name of the hardest baseline the thesis names. Protocol v1 (runs 1-3) reported it; those numbers stand as ar16's. The quantity it was supposed to measure is measured instead by within_participant_holdout, on a within-participant temporal split, and is NOT comparable with the rows here: it has seen the scored participant and every row here has not.

What it is legitimately good for

  • A treatment-arm artifact: family-indexed heterogeneous regional state with published weights and a published loss, for anyone running the same ablation against their own control β€” provided that control is trained the same way. These weights saw simulation only (see above), so a control fitted to recordings is not a comparison of state structure; it is a comparison of what each model was shown.
  • A worked example of a variance-channel failure, with the mean/variance decomposition available in the repository.
  • It is not evidence for or against the SC-WBD thesis.

Licence and attribution

Computed union: non-commercial: yes; share-alike: yes; attribution: required; redistribution: unknown; SHARE-ALIKE IN FORCE: derivative works must be released under the same licence; 1 source(s) with UNKNOWN licence (montage_calibration) β€” unknown is not permissive

  • non-commercial: True
    • forced by: anatomical_prior
  • share-alike: True
    • forced by: anatomical_prior
  • sources stating no terms: montage_calibration

These are derived from each source's own licence text by scwbd.release.licence.union_of, not asserted here.

Citations (a licence condition, not a courtesy)

The Melbourne Subcortex Atlas grants unrestricted use subject to citation; several other inputs carry attribution as their only obligation. Using this artifact requires reproducing these:

  • Wakeman DG, Henson RN (2015). A multi-subject, multi-modal human neuroimaging dataset. Scientific Data 2:150001, doi:10.1038/sdata.2015.1. OpenNeuro dataset ds000117 v1.1.0, doi:10.18112/openneuro.ds000117.v1.1.0.
  • Lioi G, Cury C, Perronnet L, Mano M, Bannier E, Lecuyer A, Barillot C (2020). Simultaneous MRI-EEG during a motor imagery neurofeedback task: an open access brain imaging dataset for multi-modal data integration. Scientific Data 7:173, doi:10.1038/s41597-020-0498-3. OpenNeuro dataset ds002336 v2.0.2, doi:10.18112/openneuro.ds002336.v2.0.2. Paradigm: Perronnet L et al. (2017), Front Hum Neurosci 11:193.
  • Hernandez Pavon JC, Schneider Garces N, Begnoche JP, Miller LE, Raij T (2022). OpenNeuro dataset ds004024, doi:10.18112/openneuro.ds004024.v1.0.0. Cortico-cortical paired associative stimulation (ccPAS) with bi-focal MRI-navigated TMS-EEG of left and right M1.
  • Schalk G, McFarland DJ, Hinterberger T, Birbaumer N, Wolpaw JR (2004). BCI2000: A General-Purpose Brain-Computer Interface (BCI) System. IEEE Trans Biomed Eng 51(6):1034-1043. Dataset: Schalk G (2009), EEG Motor Movement/Imagery Dataset (version 1.0.0), PhysioNet, RRID:SCR_007345, https://doi.org/10.13026/C28G6P
  • Kemp B, Zwinderman AH, Tuk B, Kamphuisen HAC, Oberye JJL (2000). Analysis of a sleep-dependent neuronal feedback loop: the slow-wave microcontinuity of the EEG. IEEE Trans Biomed Eng 47(9):1185-1194. Dataset: Kemp B, Zwinderman AH, Tuk B, Kamphuisen HAC, Oberye JJL. Sleep-EDF Database Expanded (version 1.0.0), PhysioNet, https://doi.org/10.13026/C2X676
Full attribution block (generated)
ATTRIBUTION
checkpoint: scwbd-004
============================================================

DATASET INPUTS (7)
  ds000117 (dataset)
    cite:    Wakeman DG, Henson RN (2015). A multi-subject, multi-modal human neuroimaging dataset. Scientific Data 2:150001, doi:10.1038/sdata.2015.1. OpenNeuro dataset ds000117 v1.1.0, doi:10.18112/openneuro.ds000117.v1.1.0.
    licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
    doi:     10.18112/openneuro.ds000117.v1.1.0
    from:    scwbd/sources/cards/ds000117.yaml
  ds000117 (dataset)
    cite:    Wakeman DG, Henson RN (2015). A multi-subject, multi-modal human neuroimaging dataset. Scientific Data 2:150001, doi:10.1038/sdata.2015.1. OpenNeuro dataset ds000117 v1.1.0, doi:10.18112/openneuro.ds000117.v1.1.0.
    licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
    doi:     10.18112/openneuro.ds000117.v1.1.0
    from:    scwbd/sources/cards/ds000117.yaml
  ds002336 (dataset)
    cite:    Lioi G, Cury C, Perronnet L, Mano M, Bannier E, Lecuyer A, Barillot C (2020). Simultaneous MRI-EEG during a motor imagery neurofeedback task: an open access brain imaging dataset for multi-modal data integration. Scientific Data 7:173, doi:10.1038/s41597-020-0498-3. OpenNeuro dataset ds002336 v2.0.2, doi:10.18112/openneuro.ds002336.v2.0.2. Paradigm: Perronnet L et al. (2017), Front Hum Neurosci 11:193.
    licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
    doi:     10.18112/openneuro.ds002336.v2.0.2
    from:    scwbd/sources/cards/ds002336.yaml
  ds004024 (dataset)
    cite:    Hernandez Pavon JC, Schneider Garces N, Begnoche JP, Miller LE, Raij T (2022). OpenNeuro dataset ds004024, doi:10.18112/openneuro.ds004024.v1.0.0. Cortico-cortical paired associative stimulation (ccPAS) with bi-focal MRI-navigated TMS-EEG of left and right M1.
    licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
    doi:     10.18112/openneuro.ds004024.v1.0.0
    from:    scwbd/sources/cards/ds004024.yaml
  ds004024 (dataset)
    cite:    Hernandez Pavon JC, Schneider Garces N, Begnoche JP, Miller LE, Raij T (2022). OpenNeuro dataset ds004024, doi:10.18112/openneuro.ds004024.v1.0.0. Cortico-cortical paired associative stimulation (ccPAS) with bi-focal MRI-navigated TMS-EEG of left and right M1.
    licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
    doi:     10.18112/openneuro.ds004024.v1.0.0
    from:    scwbd/sources/cards/ds004024.yaml
  eegmmidb (dataset)
    cite:    Schalk G, McFarland DJ, Hinterberger T, Birbaumer N, Wolpaw JR (2004). BCI2000: A General-Purpose Brain-Computer Interface (BCI) System. IEEE Trans Biomed Eng 51(6):1034-1043. Dataset: Schalk G (2009), EEG Motor Movement/Imagery Dataset (version 1.0.0), PhysioNet, RRID:SCR_007345, https://doi.org/10.13026/C28G6P
    licence: [ODC-By-1.0] Open Data Commons Attribution License v1.0 (ODC-By 1.0)
    doi:     10.13026/C28G6P
    from:    scwbd/sources/cards/eegmmidb.yaml
  sleep-edfx (dataset)
    cite:    Kemp B, Zwinderman AH, Tuk B, Kamphuisen HAC, Oberye JJL (2000). Analysis of a sleep-dependent neuronal feedback loop: the slow-wave microcontinuity of the EEG. IEEE Trans Biomed Eng 47(9):1185-1194. Dataset: Kemp B, Zwinderman AH, Tuk B, Kamphuisen HAC, Oberye JJL. Sleep-EDF Database Expanded (version 1.0.0), PhysioNet, https://doi.org/10.13026/C2X676
    licence: [ODC-By-1.0] Open Data Commons Attribution License v1.0 (ODC-By 1.0)
    doi:     10.13026/C2X676
    from:    scwbd/sources/cards/sleep-edfx.yaml

The project

Source code: https://github.com/JacobFV/sc-wbd

The SC-WBD repository itself is licensed CC-BY-NC-SA-4.0, the most restrictive term any of its inputs imposes. Treat that as a floor: the licence section above is computed from this artifact's own inputs and an artifact can inherit more than the floor, never less.

This is research code. It is not a medical device, not a clinical tool, and nothing here should be used to make a decision about a person.

How this card was produced

Every figure above was read at build time from a file in the SC-WBD repository by scwbd/release/publish.py. None of them is typed into the card generator. The sources:

  • reports/training/evaluation_run4.json β€” every score, CI, parameter count and split size
  • configs/scwbd_001_beta.yaml β€” the training mixture
  • scwbd/sources/cards/*.yaml β€” dataset citations and licences
  • reports/scope_gap.md, reports/training/p0_variance_channel.md β€” the two diagnoses, stated in prose above
Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support