MuonShower models

Flow-matching generators and muon-count heads for simulating cosmic-ray muon showers (CORSIKA I3MCTree_preMuonProp, primaries with zenith in [55, 65] degrees).

Code lives in the MuonShower repo. Clone it, install environment.yml, then either run the end-to-end script or load models directly -- don't load .safetensors/.pth files by hand, hf_pipeline.py reconstructs the right model class from the config.json next to each weight file and loads the state dict for you.

Quick start: end-to-end reproduction

CUDA_VISIBLE_DEVICES=0 ./venv/bin/python reproduce.py \
    --generator block_diffusion_bs200 --counter block_count_200 \
    --data_dir /path/to/your/corsika/test/split

Downloads every checkpoint in this repo (~3.1 GB, cached after the first run), predicts each event's muon count, samples the shower, decodes to physical units, and writes events.csv / particles.csv to --out_dir (default ./reproduce_out). --data_dir must point at your own CORSIKA HDF5 split -- that dataset is not hosted here, only the weights are. For a quick check instead of the full split: add --max_files 1 --n_events 5.

Loading a model directly

from hf_pipeline import load_generator, load_counter

gen = load_generator("block_diffusion_bs200")   # {"model", "matcher", "config"}
cnt = load_counter("block_count_200")           # {"nblocks", "count", "lastblock", "config"}

Generators (generators/)

name block layout objective
original_2000 1 block x 2000 slots plain flow matching, no auxiliary terms
block_diffusion_bs200 10 blocks x 200 slots two-stream block diffusion (BD3-LM style), flow-matching loss on real slots only
block_diffusion_bs500 4 blocks x 500 slots same recipe, different block size

All are MuonDiffusionTransformer (dim=1024, depth=16), trained with DELTA_ANGLES=True (muon zenith/azimuth encoded relative to the primary) -- decoding requires that setting in muon_dataset.py/flow_matching.py to match.

Counters (counters/)

name what it predicts input
poisson_2000 total muon count N, direct Poisson regression primary only
block_count_10 n_blocks (Poisson) + per-block count, block_size=10 primary [+ muons for the count head]
block_count_200 n_blocks + per-block count + a dedicated last-block head, block_size=200 primary [+ muons]
block_count_500 n_blocks + per-block count, block_size=500 primary [+ muons]

block_count_200/lastblock_classifier.pth is generator-specific. It was trained on muon prefixes generated by generators/block_diffusion_bs200 (checkpoint_044), not on real data, so that training matches what it sees at inference. Using it with a different generator checkpoint is untested and likely to underperform its measured numbers:

stage-3 head full-N MAE (400-event stratified test) predictions landing on a multiple of 200
naively trained (target-alignment bug) 107.2 0.000
trained on true prefixes (ceiling, not deployable) 76.8 0.177
trained on generated prefixes (shipped) 65.5 0.000
poisson_2000 baseline (no generator involved) 57.2 0.003

No count read off a generated grid can beat the primary-only baseline: the generated muons carry no information about the true count beyond what the primary already gives (x_gen = f(primary, noise), noise independent of the true shower). The value of the two-stage pipeline is removing degenerate behavior (the 200/400/600/800 quantization above), not beating the baseline on accuracy.

classifiers_nblocks (the nblocks head in block_count_200) was not trained with tail-balancing; its exact-match accuracy degrades from ~0.99 (single-block events, 96.9% of the data) to 0.14-0.55 on events needing 2-6 blocks.

Caveats carried over from the source repo

  • block_diffusion_* checkpoints should not be re-selected by val/flow_loss alone -- that metric improved monotonically through training while across-event generation diversity (what any downstream count read-off depends on) got worse. These checkpoints are the ones actually evaluated above, not the training-loss optimum.
  • original_2000 and the block_diffusion_* runs are trained on different data splits / epoch counts and are not loss-comparable to each other; compare downstream task metrics, not raw training loss, across families.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support