QuantaMind LoRA Adapters

QuantaMind is a 32-billion-parameter chemistry language model for structured molecular-property prediction and multi-objective molecular screening. Given a chemistry instruction and a molecular SMILES string, it generates a fixed-order property profile covering photophysical, excited-state, safety, accessibility, and physicochemical targets.

The model was created by parameter-efficient instruction tuning of Qwen-2.5-32B on QuantumChem-200K. The ICLR 2027 manuscript evaluates QuantaMind as a high-throughput prioritization layer for organic molecular screening, with photoinitiator discovery as a demanding case study. It is not a replacement for electronic-structure calculations, chemical-safety review, or experiments.

Research status: The associated manuscript is under review at ICLR 2027. Reported results are computational benchmark results, not wet-lab validation.

Resources

Repository contents

models_LoRA/
├── forward_model_LoRA/   # SMILES → molecular-property profile
└── reverse_model_LoRA/   # bundled experimental adapter

The ICLR manuscript, evaluation tables, and usage example in this card describe forward_model_LoRA. The bundled reverse_model_LoRA is not documented or evaluated in the manuscript, so this card makes no validated performance claim for it.

Model details

Item Value
Backbone Qwen-2.5-32B decoder-only language model
Released artifact LoRA adapters plus tokenizer files
Adaptation LoRA on a 4-bit quantized backbone
LoRA configuration Rank 16, alpha 16, dropout 0, no bias
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Training objective Causal language-model loss over ordered response tokens
Supervision More than 210K instruction-SMILES-response examples
Selected schedule Six epochs; learning rate 2e-4
Training compute Four NVIDIA A100 80 GB GPUs; approximately 4–5 days
License Apache-2.0 for the released model artifacts

Inputs and outputs

The forward adapter accepts an instruction and one SMILES string in the Alpaca prompt format used during training. It emits a parseable molecular-property profile.

The seven primary benchmark targets are:

Group Property Unit or interpretation Screening direction used in the paper
Photophysical TPA cross section at 780 nm, sigma_780 GM Higher
Photophysical Maximum TPA cross section, sigma_max GM Higher
Excited state Singlet-triplet ISC energy gap eV Lower
Practical Toxicity score Surrogate score in [0, 1] Lower
Practical Synthetic-accessibility score Surrogate score in [0, 1] Higher
Physicochemical Boiling point °C Objective-dependent
Physicochemical Solubility Dataset-reported scale Objective-dependent

The training response schema also contains an absorption wavelength range, logP, aromaticity, and molecular weight. These auxiliary fields were excluded from the primary seven-property benchmark.

Intended use

QuantaMind is intended for:

  • rapid, preliminary property profiling of supplied organic molecules;
  • prioritizing an existing candidate library under user-defined objectives;
  • photoinitiator screening and related research workflows;
  • generating structured estimates for later parsing, ranking, and expert review.

Shortlisted molecules should be advanced to independent quantum-chemical calculations and, where appropriate, experimental validation.

Out-of-scope use

Do not use this model as:

  • a biological-safety or toxicity determination;
  • a substitute for electronic-structure calculations or wet-lab measurements;
  • a synthesis-route planner or proof of synthetic feasibility;
  • a molecule generator—the paper evaluates ranking of supplied candidates, not molecular generation;
  • the sole basis for decisions involving hazardous substances, human health, environmental release, or regulatory compliance.

Getting started

Requirements

The authors' reference workflow uses Python 3.10/3.11, CUDA, PyTorch, Unsloth, Transformers, PEFT, Accelerate, bitsandbytes, and huggingface_hub. A 32B model in 4-bit mode still requires substantial GPU memory.

pip install torch transformers peft accelerate bitsandbytes unsloth huggingface_hub

Forward-model inference

The adapters are stored in a subdirectory, so this example downloads only the forward adapter and loads it from its local snapshot path.

from pathlib import Path

from huggingface_hub import snapshot_download
from unsloth import FastLanguageModel

REPO_ID = "QuantumChem/QuantaMind"
ADAPTER_SUBDIR = "models_LoRA/forward_model_LoRA"

snapshot_path = snapshot_download(
    repo_id=REPO_ID,
    allow_patterns=[f"{ADAPTER_SUBDIR}/*"],
)
adapter_path = str(Path(snapshot_path) / ADAPTER_SUBDIR)

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name=adapter_path,
    max_seq_length=2048,
    dtype=None,
    load_in_4bit=True,
)
FastLanguageModel.for_inference(model)

alpaca_prompt = """Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.

### Instruction:
{}

### Input:
{}

### Response:
{}"""

instruction = (
    "Based on the given SMILES string, predict the monomer's relevant "
    "properties. Given toxicity and SA score ranging from 0 to 1, "
    "aromaticity can only take 0 or 1. Answer the task in this format: "
    "The monomer compound has sigma of ? GM at 780 nm, maximum sigma "
    "of ? GM, ISC of ? eV, toxicity score of ?, SA score of ?, boiling "
    "point of ? °C, logP of ?, aromaticity of ?, solubility of ? ug/mol, "
    "molecular weight of ? g/mol."
)
smiles = "C=C(C)OC"
prompt = alpaca_prompt.format(instruction, smiles, "")

inputs = tokenizer([prompt], return_tensors="pt").to("cuda")
outputs = model.generate(
    **inputs,
    max_new_tokens=256,
    do_sample=False,
    use_cache=True,
)

generated_tokens = outputs[0, inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(generated_tokens, skip_special_tokens=True))

For reproducible batch screening, retain the exact property order and parse named fields rather than relying on free-form text positions. Treat blank, malformed, or non-finite fields as missing; do not silently replace them with chemically meaningful values.

Training data

QuantaMind was adapted on QuantumChem-200K, which uses QM9 and the Open Macromolecular Genome as molecular-structure pools. After validity filtering, canonicalization, deduplication, and compatibility checks, the structures were annotated through a hybrid workflow:

  • Two-photon absorption: MLatom-based spectral calculations over 600–850 nm, retaining the response at 780 nm and the spectral maximum.
  • Intersystem crossing: B3LYP/6-31G(2df,p) geometry optimization followed by AIQM1/MNDO-CIS singlet and triplet calculations.
  • Toxicity and synthetic accessibility: eToxPred-derived surrogate scores.
  • Boiling point and solubility: JRgui/RDKit-centered property workflows.
  • Additional descriptors: molecular weight, logP, aromaticity, and wavelength range.

Training examples serialize instruction, input (SMILES), and output using the Alpaca wrapper shown above. Some released records contain missing property values represented by blank fields.

The external 3,000-molecule testbank contains broader scaffolds from VQM24 and ZINC20 and was held out from adaptation.

Evaluation

The paper evaluates absolute prediction error, structure-property correlation, and decision-oriented multi-objective screening. All language models receive the same fixed property order, units, response schema, and field-wise numerical parser.

Molecular-property prediction

On the 3,000-molecule external testbank, QuantaMind achieves an overall weighted mean absolute error (wMAE) of 0.1980.

Model Overall wMAE ↓
Base Qwen-2.5-32B 3.3040
Base Gemma-3-27B 3.1480
EdgeCNN 0.5717
Fine-tuned Gemma-3-27B 0.5297
QuantaMind 0.1980

QuantaMind's property-level wMAE values are:

sigma_780 sigma_max ISC Toxicity SA Boiling point Solubility
0.0106 0.0104 0.0273 0.0446 0.0227 0.0074 0.0057

Reported Pearson correlations are 0.433 for sigma_780, 0.416 for sigma_max, 0.551 for ISC, 0.966 for toxicity, and 0.911 for synthetic accessibility.

Multi-objective screening

Under the paper's seven-property ranking objective, QuantaMind recovers 35 of the ground-truth top 100 candidates.

Model Precision@100 ↑ Recall@100 ↑ Enrichment@100 ↑ ε-Pareto precision@20 ↑ Normalized HV regret@20 ↓
Fine-tuned Gemma-3-27B 0.31 0.31 9.30 0.85 42.4%
QuantaMind 0.35 0.35 10.50 0.80 3.26%

These results measure recovery against computational reference labels. They do not establish experimental success or material performance.

Bias, risks, and limitations

  • SMILES-only input: SMILES does not explicitly encode three-dimensional conformation, solvent, concentration, formulation, irradiation conditions, or competing relaxation pathways.
  • No calibrated uncertainty: The released model does not provide a validated uncertainty threshold. Numerical-looking outputs should not be interpreted as calibrated confidence.
  • Supervision quality: Predictions inherit the approximations, bias, missingness, and domain limits of computed and model-derived labels.
  • Excited-state limits: The paper reports only moderate agreement between the inexpensive ISC/TPA supervision workflows and higher-level or experimental references. These targets are useful for preliminary ranking, not definitive calculation.
  • Surrogate safety scores: Toxicity and synthetic-accessibility values are screening surrogates, not biological-safety findings or guaranteed synthesis outcomes.
  • Chemical-domain shift: Accuracy may degrade for unfamiliar scaffolds, heavy elements, unusual charge states, stereochemical regimes, very large molecules, and chemistry outside the training distribution.
  • Failure modes: The paper observes TPA underprediction for extended conjugated scaffolds and larger ISC errors for some nitrogen-rich structures.
  • Unit caution: The manuscript prompt writes solubility as ug/mol, while released training records serialize solubility as ug/ml. Preserve the convention used by the workflow being reproduced and verify units before scientific use.
  • Dual use: Molecular-screening systems may be misused to prioritize harmful substances. Use domain-expert oversight, appropriate chemical-safety controls, and applicable legal and institutional safeguards.

Recommendations

  1. Validate SMILES strings and canonicalize them consistently before inference.
  2. Use deterministic decoding and the training-time prompt wrapper for benchmark reproduction.
  3. Parse each named field explicitly and preserve missing outputs.
  4. Treat the model as a prioritization layer, not an oracle.
  5. Recompute shortlisted candidates with independent quantum-chemical methods and appropriate experimental controls.
  6. Obtain expert chemical-safety review before acting on toxicity-related or potentially hazardous candidates.

Environmental impact

The paper reports training on four A100 80 GB GPUs for approximately four to five days. Energy use and carbon emissions were not measured, so no emissions estimate is provided.

Citation

If you use QuantaMind, QuantumChem-200K, or the evaluation protocol, cite the ICLR 2027 submission:

@article{quantamind2026,
  title   = {QuantaMind: A Large Chemistry Language Model for Structured Molecular Screening and Discovery},
  author  = {Anonymous Authors},
  journal = {Under review at ICLR 2027},
  year    = {2026},
  url     = {https://github.com/AnonymousUser-3/QuantaMind}
}

License and attribution

The released model artifacts are provided under Apache-2.0. The companion code repository uses the MIT License. Dataset reuse is governed by the dataset card and by the licenses and attribution requirements of the underlying source resources.

Contact

For questions or corrections, open a discussion in this model repository or an issue in the companion code repository.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train QuantumChem/QuantaMind