DeMiR: Density-Matrix-Informed Representations for Molecular Property Prediction
Pretrained weights for DeMiR, a molecular representation model pretrained by predicting DFT electron density matrices. Code: https://github.com/dhyeonk-0008/DeMiR
Files
| File | Description | SHA-256 |
|---|---|---|
demir.pth |
PyTorch state_dict (6.4M parameters, 25.8 MB) |
dbaefe607b31c4cccd6e0ea0c63eef40191c8903f055ee12eb89c614a88ed815 |
This is the checkpoint used for all results in the paper (validation-loss minimum of the joint QM9 + QMugs training run).
Usage
git clone https://github.com/dhyeonk-0008/DeMiR && cd DeMiR
# install as in the repository README, then:
hf download DeonX/DeMiR demir.pth --local-dir checkpoints
python scripts/inference/predict_density.py \
--smiles "CCO" \
--checkpoint checkpoints/demir.pth \
--output-npz ethanol.npz \
--device cpu
The architecture is inferred from the weights. See the repository README for embedding extraction and property-prediction probes.
Training data
B3LYP/cc-pVDZ density matrices computed with PySCF for QM9 molecules and a representative subset of QMugs (adding S, P, Cl, Br and larger drug-like molecules). Median real-space density NMAE on held-out molecules: 1.00% (QM9), 2.16% (QMugs).
Limitations
- Neutral, closed-shell molecules in the cc-pVDZ basis only (H, C, N, O, F, P, S, Cl, Br).
- Bromine-containing molecules reconstruct poorly (NMAE > 8%).
- The eSCN backbone samples per-edge frames, so repeated forward passes differ slightly (up to ~1e-2 in individual density-matrix elements) unless the random seed is fixed.
License
MIT. The QM9 and QMugs source datasets keep their own licenses.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support