dIon 0.1

dIon is a foundation encoder for tandem mass spectra. This release contains the foundation checkpoint and two complete de novo peptide-sequencing checkpoints. The implementation and loading utilities are provided by the dIon source repository.

Checkpoints

File Intended use Peak cap Loading argument
checkpoints/dion-v0.1-foundation.ckpt Foundation encoder and downstream initialization 200 --encoder_weights
checkpoints/dion-v0.1-denovo-200peaks.ckpt Recommended de novo model 200 --downstream_weights
checkpoints/dion-v0.1-denovo-1000peaks.ckpt Slower de novo model with slightly stronger predictions 1,000 --downstream_weights

The foundation file is a PyTorch Lightning pretraining checkpoint containing student and EMA-teacher state. dIon's downstream loaders use the EMA teacher backbone as the released representation. The de novo files contain the complete fine-tuned encoder and Casanovo-style decoder.

Do not load a de novo checkpoint through --encoder_weights. Use --downstream_weights so that both encoder and decoder state are restored.

Download

Install the Hugging Face CLI, then download the complete release:

hf download alfred-n/dIon --local-dir dion-v0.1

Or download one checkpoint:

hf download alfred-n/dIon \
  checkpoints/dion-v0.1-denovo-200peaks.ckpt \
  --local-dir dion-v0.1

Use the Foundation Encoder

From a dIon source checkout, initialize de novo fine-tuning with a new decoder:

python -m src.main \
  --config configs/master_denovo_dion_finetune.yaml \
  --encoder_weights /path/to/dion-v0.1-foundation.ckpt \
  --downstream_root_dir /path/to/lance_dataset \
  --output_dir /path/to/output/checkpoints \
  --log_dir /path/to/output/logs \
  --log_wandb 0

The input and adaptation contract is documented in the source repository's docs/finetune_denovo.md.

Use a De Novo Checkpoint

Evaluate the 200-peak model on a compatible labeled Lance dataset:

python -m src.main \
  --config configs/master_denovo_dion_finetune.yaml \
  --downstream_weights /path/to/dion-v0.1-denovo-200peaks.ckpt \
  --downstream_root_dir /path/to/lance_dataset \
  --eval_only 1 \
  --validate_on_end 0 \
  --test_on_end 1 \
  --save_top_k 0 \
  --save_last 0 \
  --log_wandb 0

For the 1,000-peak model, also pass:

--max_peaks 1000 --disable_cudnn_sdp 1

The 1,000-peak model is substantially slower and requires more GPU memory.

Input Contract

The released models consume centroided tandem mass spectra with:

  • fragment m/z values in [0, 2500];
  • base-peak-scaled intensities;
  • precursor m/z and charge conditioning;
  • at most 200 or 1,000 retained peaks, according to the checkpoint;
  • peptide labels represented by the bundled PA1.1 numeric mass-delta tokenizer.

For labeled Lance datasets, expected columns are mz_array, intensity_array, precursor_mz, precursor_charge, and seq.

Training Provenance

The foundation model adapts DINO with a dual objective comprising two latent prediction tasks, both recovering a clean teacher representation: one from a spectrum mixture using the precursor as a selection query, and one from a partial spectrum with the precursor withheld. iBOT and KoLeo are disabled in this release. The checkpoint is from epoch 279 at global step 315,560 of the nominal 300-epoch training schedule.

The de novo models were initialized from this foundation checkpoint and fine-tuned end to end on dIon-de-novo-labeled-v1 (DNLv1). Checkpoints were selected by native DNLv1 validation peptide precision, without test-set model selection. The 200-peak checkpoint is epoch 37; the 1,000-peak checkpoint is epoch 27.

Portable release configurations are included under configs/. Per-checkpoint machine-readable provenance, sizes, and hashes are under metadata/.

Limitations

  • These models are research software and are not intended for clinical use.
  • Performance depends on instrument, fragmentation regime, precursor metadata, modification vocabulary, charge distribution, and preprocessing.
  • Unsupported modifications must not be silently coerced into supported tokens.
  • The 1,000-peak variant trades substantially greater compute and memory use for a modest prediction improvement.
  • Confidence values require task- and dataset-specific calibration before use as error probabilities.

Method Attribution

The self-distillation components are adapted from DINO and DINOv2. The de novo decoder and evaluation conventions are Casanovo-style. See the dIon source repository for exact upstream references and implementation notes.

Integrity

Verify downloads with SHA256SUMS. The three checkpoint hashes are also stored in their corresponding JSON files under metadata/.

Citation

The dIon publication citation will be added when it is publicly available. Until then, cite this model repository and release (dIon 0.1, v0.1.0).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support