ArXiv · 2025
The increasing variety of molecular data creates a need for generative models that can flexibly incorporate heterogeneous constraints across modalities. However, existing SMILES-based diffusion models are typically designed for a fixed conditioning modality, and introducing new constraints often requires retraining the model. To address this limitation, we propose Cross-Modality Controlled Molecule Generation with Diffusion Language Model (CMCM-DLM), a modular framework that extends a pre-trained diffusion model to support heterogeneous molecular constraints without retraining the backbone. We demonstrate CMCM-DLM using two complementary modalities: molecular structure and chemical properties. Specifically, a Structure Control Module (SCM) guides early diffusion steps to establish the molecular scaffold, while a Property Control Module (PCM) subsequently steers generation toward target chemical properties. This staged design enables flexible integration of different molecular constraints within a unified generative framework. Experiments on multiple datasets demonstrate effective cross-modal controllability and strong adaptability, highlighting the potential of CMCM-DLM for heterogeneous molecular data modeling and data-driven drug discovery.
Try inveni