>News Center>News Detail

JACS | From Generation to Interaction Awareness: Momed Biotech & Prof. Hou Tingjun’s Group (Zhejiang University) Present EIP-Diff to Enable High-Fidelity, Controllable 3D Molecular Generation

2026.08.11
图片1.png

In Structure-Based Drug Design (SBDD), generative AI is already capable of rapidly generating a large number of novel molecules targeting protein pockets.

Nevertheless, an unavoidable open question remains: does the model truly capture authentic protein–ligand interactions, or merely identify shortcuts that yield high scores within the training data and scoring functions?


To address this question, a joint research team comprising Momed Biotech, Zhejiang University, Zhejiang Provincial People’s Hospital, Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, and the University of Southern California has proposed the EIP-Diff (Explicit Interaction-Prompted Diffusion) framework.

This work integrates high-fidelity crystallographic data and residue-level interaction prompts into 3D molecular generation, aiming to advance models from merely “generating molecules” toward “understanding and adhering to critical biological interactions”.



01丨Why Can High-Scoring Models Still Produce “Wrong Answers”?

At present, most target-specific 3D generative models adopt implicit learning paradigms: neural networks autonomously summarize from training data how ligands should occupy pockets, form hydrogen bonds, and maintain three-dimensional conformations.

The crux is that systematic biases within training data risk being learned by the model as inherent rules, which are then further amplified in generated outputs. The paper defines this phenomenon as:

Algorithmic Bias Amplification

The research team analyzed the widely adopted Crossdocked2020 dataset. Compared with experimentally resolved native crystal complexes, the Crossdocked dataset exhibits multiple prominent biases:

• The average molecule duplication count reaches 11.52, versus only 1.67 in crystal data;

• Artificial cross-docking pairing disrupts the native volume complementarity between pockets and ligands;

• Hydrogen-bond probability of ILE residues is overestimated by approximately 12%;

• Hydrogen bonds within the 2.75–3.5 Å range are substantially overpredicted;

• Key interactions such as halogen bonds are systematically underestimated.

图片2.png

Figure 1.Comparison between CrystalDataset and Crossdocked datasets in terms of molecular duplication, physicochemical properties, pocket matching, and interaction distribution. Figure sourced from the original paper.

When models repeatedly encounter simple, easily synthesizable scaffolds during training, they may boost metrics such as QED and SA by generating these structures repetitively, without genuinely learning pocket-specific 3D binding constraints.

Examples presented in the paper show that benzene rings appear 226 times in molecules generated by PocketFlow. Such “metric inflation” explains why some models achieve impressive computational scores yet fail to translate into tangible biological activity.


02丨EIP-Diff:Turning Interactions from “Latent Variables” into “Explicit Instructions”

The core idea of EIP-Diff is to move beyond full reliance on implicit induction. Instead, residue-level interactions valued by medicinal chemists are directly encoded as prompts and injected into the diffusion denoising process.

The model consists of two information streams: global self-attention and prompt-guided cross-attention, and maintains SE(3) equivariance throughout the entire coordinate update process.

图片3.png

Figure 2. Overview of the EIP-Diff framework, including global self-attention, interaction prompt embedding, prompt-guided cross-attention, and the diffusion generation process. Figure sourced from the original paper.

1.Global Geometric Modeling

EIP-Diff first performs global message passing on a heterogeneous graph containing protein–protein, protein–ligand, ligand–protein and ligand–ligand connections.

SE(3)-equivariant graph attention ensures that model predictions remain unaffected by the choice of coordinate system when the molecule undergoes overall rotation or translation. This serves as an essential foundation for high-fidelity 3D conformation generation.

2. Explicit Interaction Prompts

Interactions including hydrogen bonds, halogen bonds, cation–π and π–π stacking are encoded as learnable prompt vectors and fused with protein atomic features.

Cross-attention updates ligands only along the protein-to-ligand direction, enabling the generation process to explicitly respond to designated residues and interaction types.

• During training, interaction prompts act as explicit biological priors;

• During inference, interaction prompts serve as controllable instructions for molecular design.

For instance, researchers can instruct the model to generate molecules that form hydrogen bonds with a key residue, or introduce aromatic moieties at specific positions to achieve π–π stacking.

3. Two-Stage Learning

The team first pre-trains the model on the 3D ZINC dataset to equip it with fundamental chemical validity. The model is then fine-tuned on target datasets to learn authentic protein pocket conformations and interaction patterns.

This strategy appropriately separates “learning chemical grammar” from “learning complex binding geometries”, which helps improve training stability and model generalization.

03丨45,318 Crystal Complexes: Unleashing the Model’s Geometric Potential

To provide reliable supervision for the explicit architecture, the research team constructed CrystalDataset based on BioLiP2.

This dataset contains 45,318 experimentally resolved protein–ligand complexes and preserves:

• Native protein–ligand pairings;

• Experimentally determined native 3D conformations;

• Residue-level interaction annotations.

The merit of CrystalDataset lies not only in cleaner data, but more importantly, its provision of authentic geometric and interaction labels required by EIP-Diff.

The evaluation was conducted on 99 randomly selected crystal pockets, with 100 molecules generated per pocket, yielding a total of 9,900 samples.

Even when trained on the biased Crossdocked dataset, EIP-Diff exhibits strong resistance to mode collapse:

• Ratio of Repeated Molecules (RMR₁): 0.5%;

• Molecule Uniqueness: 99.7%;

• Struct-DCS: 0.951;

• Pharma-DCS: 0.920;

• Chemical Space Reconstruction rate (CSR): 99.8%;

• Chemical Space Exploration rate (CER): 38.7%.

图片4.png

Figure 3. Pharmacological property distribution, atomic composition distribution of different models, and generation success rates under varying numbers of interaction prompts. Figure sourced from the original paper.

When the training dataset is switched to CrystalDataset, the model’s 3D geometric capability is significantly unlocked:

• 3D ShapeSim Top-1 Dominance rises from 8.1% to 40.4%;

• Struct-DCS increases from 0.951 to 0.961;

• Pharma-DCS improves from 0.920 to 0.943.

This indicates that the low geometric scores obtained under Crossdocked training do not stem from insufficient structural capacity of the model. Instead, the model’s sensitivity to geometric information is constrained by erroneous supervision signals.

After incorporating authentic interaction prompts on the basis of crystal data:

• 3D ShapeSim Top-1 Dominance further reaches 47.5%;

• Struct-DCS reaches 0.968;

• Molecule Uniqueness reaches 99.9%;

• Real chemical space coverage reaches 100%;

• Chemical Space Exploration rate hits 41.9%.

For complexes with 11 known interactions, prompt-conditioned generation achieves up to a 64% higher success rate compared with prompt-free generation.

04丨From Recapturing Binding Modes to Optimizing Novel Interactions

Aggregate metrics cannot fully determine whether a model truly understands a specific protein pocket. Accordingly, the team further selected KAT6A and YTHDC1 for structural validation and molecular optimization.

KAT6A:Recovering Native Binding Modes

In the KAT6A case (PDB: 6OIO), EIP-Diff achieves a 3D structural similarity score of 0.467, ranking first among six benchmarked models.

The generated molecules retain the acyl-sulfonylhydrazide core of the original ligand and form stable hydrogen bonds with key residues including Ser690, Arg655, Gly659 and Arg660. Hydrophobic substituents can effectively fill the pocket.

YTHDC1:Optimization While Preserving Critical Interactions

In the YTHDC1 case (PDB: 8K2E), the team adopted the key interactions of the reference ligand as prompts to generate and screen approximately 30,000 candidate molecules.

The top design maintains the critical orientation of the original carboxyl group and exhibits superior MM/PBSA binding free energy in molecular dynamics simulations:

Designed molecule: −36.43 ± 1.41 kcal/mol

Reference molecule: −26.02 ± 1.08 kcal/mol

The newly designed molecule additionally forms stable hydrogen bonds with SER362 and SER378.

图片5.jpg

Figure 4. Recapitulation of KAT6A binding modes and molecular optimization for YTHDC1, including structural similarity, binding conformations, binding free energies and interaction analysis of key residues. Figure sourced from the original paper.

05丨Toward Experimental Validation: IDO1 Candidate Molecule Achieves 0.31 nM

IDO1 is a vital target for tumor immunotherapy. Its inhibitor BMS-986205 undergoes a dynamic binding process from the surface-extended state via a bent transition state to a deep stable state, imposing stringent requirements on the 3D conformational capability of generative models.

For this complex pocket, the team anchored the chlorobenzene moiety at the pocket entrance and performed generative scaffold hopping on the key tail region, yielding 30,000 unique molecules.

Following multi-parameter screening including molecular docking, synthetic accessibility and QED, researchers finally selected two candidate compounds for experimental validation.

In the IFN-γ-stimulated HeLa cell assay, the IDO1 inhibitory activities of the two candidates are as follows:

Candidate Compound 1: IC₅₀ = 0.75 nM;

Candidate Compound 2: IC₅₀ = 0.31 nM

Positive control: IC₅₀ = 0.29 nM

图片6.png

Figure 5. Scaffold hopping design for IDO1, candidate molecular structures, binding conformations, and cellular dose–response results. Figure sourced from the original paper.

This finding advances the research from investigating “whether generated molecules resemble real ones” to “whether generated molecules function in experiments”.

It demonstrates that explicit interaction prompts can not only improve computational metrics, but also empower models to tackle drug discovery-relevant challenges including dynamic binding pockets, scaffold hopping and constraints imposed by key residues.

06丨Where Do the Advantages of EIP-Diff Lie?

From Statistical Correlation to Biological Constraints

Instead of forcing the network to infer all binding rules independently, EIP-Diff introduces residue-level interactions as explicit conditions, aligning the generation logic of the model with real medicinal chemistry design principles.

Unlocking the Value of High-Quality Data

CrystalDataset provides not only an expanded set of samples, but also native pairings, experimentally determined conformations and authentic interaction annotations. Results reveal that high-fidelity data can only be translated into improved 3D geometric performance when supported by a suitable model architecture capable of utilizing such information.

Balancing Fidelity, Explorability and Controllability

A desirable model should faithfully reflect real pharmacological and structural distributions without repeatedly generating a limited set of simple scaffolds. It needs to cover known chemical space while exploring novel regions, and respond to specified residues and interaction requirements.

EIP-Diff achieves well-balanced performance across all these dimensions.

Enabling Closed-Loop Experimental Validation

The three validation cases of KAT6A, YTHDC1 and IDO1 correspond to binding mode recapitulation, structural optimization and prospective activity verification respectively, forming a complete evidence chain spanning algorithm evaluation to practical drug design tasks.

Models cannot replace compound synthesis and experimental testing, yet they can more efficiently translate design hypotheses into prioritized candidate molecules worthy of experimental validation.

Conclusion

This work highlights not merely a 3D molecular generative model with enhanced benchmark performance, but a paradigm of co-design for architecture and datasets:

Leveraging experimental crystallographic data for reliable supervision, adopting SE(3)-equivariant networks to preserve 3D geometric consistency, and employing explicit interaction prompts to translate medicinal chemistry knowledge into actionable generative conditions.

Shifting from “unconstrained molecular generation” to “interaction-driven generation centered on key binding interactions”, EIP-Diff delivers a more controllable technical route for target-specific molecular design, fragment optimization and scaffold hopping. It also offers new insights for generative AI to bridge the gap between computational metrics and tangible biological activity.


Paper and Code Information

Title: An Explicit Interaction-Prompted Diffusion Framework for High-Fidelity 3D Molecular Generation

Code repository: https://github.com/zephyrdhb/EIP-Diff


分享: