ArXiv · 2025
Molecule generation from extremely limited training data is a key challenge in drug discovery. Existing fragment-based methods are more suitable than atom-based approaches in this regime, but typically optimize fragment selection separately from downstream generation. Expert feedback is also especially valuable with limited data, yet translating such feedback into model objectives usually requires AI engineering expertise. We introduce FRAGMENTA, an end-to-end framework for small-data drug lead optimization with two components: (1) LVSEF, a fragment-based generator that jointly optimizes fragmentation and generation through a tabular reward-update mechanism, and (2) an agentic system that converts conversational expert feedback into updated generative objectives. Across three small-data datasets (11–104 molecules), LVSEF outperforms state-of-the-art methods in the smallest-data settings, matches them at larger scales, and trains ∼16× faster. On three public protein targets, iterative closed-loop optimization improves final-round discovery yield by up to ∼16% over one-shot LVSEF-only on kinase, with gains depending on how well feedback matches target chemistry. In a real-world cancer drug-discovery deployment, Human-Agent FRAGMENTA identified nearly twice as many molecules with favorable docking scores (< -6) as baseline methods.
Try inveni