Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
Abstract
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.
Community
3D-Fit is a novel benchmark for evaluating the ability of large language models to generate 3D molecules under realistic spatial constraints for structure-based drug design. It tests whether general-purpose LLMs can produce ligands conditioned on a protein pocket while also satisfying additional requirements such as anchor fragments, pharmacophore points, and mandatory protein–ligand interactions. The benchmark uses token-efficient textual descriptions of 3D conditions and a structured output format for generated molecules.
The results show emerging 3D instruction-following: general-purpose LLMs can often satisfy explicit local constraints—especially coordinate-like anchors and pharmacophores—and can attempt multiple heterogeneous conditions in a single prompt. At the same time, satisfying those constraints is not enough for physically reliable binders. LLM poses frequently show steric clashes, weaker intramolecular validity, and worse docking scores than specialized diffusion models, even after local optimization.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Molexar: A Unified Multimodal Molecular Foundation Model for Drug Design (2026)
- HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration (2026)
- Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization (2026)
- ShallowBench: Benchmarking Generative Drug Design Models on Shallow-Pocket Targets (2026)
- MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents (2026)
- Sesame: Structure-Aware Molecular Generation via Spatial Density-Map Conditioning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper