VanillaSort pretrained models

These are the original VanillaDet and HuiduRep state dictionaries for VanillaSort, model version hybrid-janelia-2026.09. The files are unchanged from the published inference repository.

VanillaSort combines learned spike detection and denoised waveform embeddings with relative-amplitude features, Gaussian-mixture clustering and template-residual refinement.

Usage

Install the Python package from the VanillaSort repository (the inference package API requires version 0.1.0 or newer):

git clone https://github.com/IgarashiAkatuki/VanillaSort.git
cd VanillaSort
pip install -e .
import vanillasort

# recording: a SpikeInterface BaseRecording with 2D channel locations
sorting = vanillasort.sort(recording, components=22, device="auto", seed=0)

The package downloads the two checkpoints at a pinned revision on first model load, validates their SHA-256 hashes and reuses the local Hugging Face cache. Ordinary import does not download weights. components is the chosen GMM K; the default 22 is a historical preset, not a neuron-count estimate.

Files and integrity

File SHA-256
detector_mask_r4_best_ap.pt c112499b0077d3613ef5b791f61130eacb75d090bea4f6fa9d59aabc0c2aa1a6
HuiduRep.pt 048203dbb326e159f4320bc4a7204c93a9951477b8c9995a75728bb87378da68

config.json specifies the architectures and published inference profiles. Both .pt files are loaded as state dictionaries using torch.load(..., weights_only=True); use the model classes from vanillasort.models. The checkpoint pair totals about 36 MiB.

Intended use and limitations

The published checkpoint workflow uses four-channel extracellular recordings at approximately 30 kHz, with a common voltage scale and 2D coordinates in micrometres. VanillaDet requires four real channels and has no channel mask. HuiduRep uses the original 60-sample waveform preprocessing, interpolation to 90 samples and repetition/cropping to 11 input channels.

The package's larger-probe adapter uses experimental nearest-four neighborhoods; it needs validation on broader probes and datasets. There is no cross-neighborhood unit merging or drift correction. Segments are clustered independently. Exact whole-segment filtering and normalization require host RAM proportional to duration. Synthetic smoke tests verify implementation consistency, not biological accuracy. For training, benchmark design, datasets and evaluation, consult the paper and repository.

Citation

Zishuo Feng and Feng Cao. Spike Sorting with VanillaSort. bioRxiv, 2026. doi:10.64898/2026.09.18.752552.

@article{feng2026vanillasort,
  title = {Spike Sorting with {VanillaSort}},
  author = {Feng, Zishuo and Cao, Feng},
  journal = {bioRxiv},
  year = {2026},
  doi = {10.64898/2026.09.18.752552}
}

Use of VanillaSort requires citation of this paper. See LICENSE for the GNU Affero General Public License v3.0.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support