DeMiR: Density-Matrix-Informed Representations for Molecular Property Prediction

Pretrained weights for DeMiR, a molecular representation model pretrained by predicting DFT electron density matrices. Code: https://github.com/dhyeonk-0008/DeMiR

Files

File Description SHA-256
demir.pth PyTorch state_dict (6.4M parameters, 25.8 MB) dbaefe607b31c4cccd6e0ea0c63eef40191c8903f055ee12eb89c614a88ed815

This is the checkpoint used for all results in the paper (validation-loss minimum of the joint QM9 + QMugs training run).

Usage

git clone https://github.com/dhyeonk-0008/DeMiR && cd DeMiR
# install as in the repository README, then:
hf download DeonX/DeMiR demir.pth --local-dir checkpoints

python scripts/inference/predict_density.py \
    --smiles "CCO" \
    --checkpoint checkpoints/demir.pth \
    --output-npz ethanol.npz \
    --device cpu

The architecture is inferred from the weights. See the repository README for embedding extraction and property-prediction probes.

Training data

B3LYP/cc-pVDZ density matrices computed with PySCF for QM9 molecules and a representative subset of QMugs (adding S, P, Cl, Br and larger drug-like molecules). Median real-space density NMAE on held-out molecules: 1.00% (QM9), 2.16% (QMugs).

Limitations

  • Neutral, closed-shell molecules in the cc-pVDZ basis only (H, C, N, O, F, P, S, Cl, Br).
  • Bromine-containing molecules reconstruct poorly (NMAE > 8%).
  • The eSCN backbone samples per-edge frames, so repeated forward passes differ slightly (up to ~1e-2 in individual density-matrix elements) unless the random seed is fixed.

License

MIT. The QM9 and QMugs source datasets keep their own licenses.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support