ArXiv

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

Authors
Thomas MacDougall, Maksim Kuznetsov, Roman Schutski...
Categories
cs.LG, cs.AI, cs.CL
arXiv
https://arxiv.org/abs/2607.18144v1
PDF
https://arxiv.org/pdf/2607.18144v1

Brief

The paper evaluates whether general-purpose LLMs can reason about 3D spatial constraints in structure-based drug design by comparing them to established diffusion-model baselines. Using an introduced benchmark, 3D-Fit, the authors test pocket-conditioned ligand generation under ligand- and interaction-derived constraints (anchor fragments, pharmacophore points, mandatory pocket–ligand interactions). Based on the abstract, LLMs are promising and handle multiple constraints but remain behind diffusion approaches; full text was not available for deeper quantitative details.

Why it matters

Authors (MacDougall et al., published 2026-07-20) find that general-purpose LLMs can satisfy multiple 3D spatial constraints simultaneously — including anchor fragments, pharmacophore points, and mandatory pocket–ligand interactions — but still trail state-of-the-art diffusion-based 3D molecule generation models in overall performance.

Key details

  • The paper introduces 3D-Fit, a token-efficient benchmarking strategy for multi-conditioned spatial molecule generation that systematically evaluates pocket-conditioned ligand design under heterogeneous spatial constraints, aiming to compare LLM-based methods against diffusion-model baselines.
Source evidence

Abstract

Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.