ArXiv · 2026
Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential for enabling large language models (LLMs) to both understand structure for chemical reasoning and generate molecules from natural-language design intent. However, the substantial cost of human annotation makes it infeasible to construct large-scale, high-quality datasets of structure-grounded descriptions. This work proposes a fully automated annotation framework for generating precise molecular descriptions at scale, such that the original molecule can be unambiguously reconstructed from the description alone. Our approach extends a rule-based chemical nomenclature parser to interpret IUPAC names and construct enriched, XML metadata that explicitly encodes molecular structure. This is then used to guide LLMs in producing accurate natural-language descriptions. Using this framework, we curate MolLangData, a dataset of approximately 163k molecule–description pairs. A rigorous validation protocol combining expert human and LLM-based reconstruction on a subset of 2,000 molecules demonstrates 98.6% description precision. Using the curated dataset, we train a 4B-parameter LLM via large-scale reinforcement learning for language-conditional molecule generation. The proposed framework and dataset provide a reliable foundation for molecule–language alignment, readily beneficial to broader chemical tasks.
Try inveni