MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

2026-07-01Machine Learning

Machine LearningComputation and Language
AI summary

The authors point out that many AI models that create new molecules often ignore safety and might produce harmful or toxic chemicals. To fix this, they made MolSafeEval, a tool that checks how safe these AI-created molecules are by using lots of safety information organized in a knowledge graph. This tool helps explain why a molecule might be unsafe and tests different types of molecule generation tasks. Their work helps show where current AI methods can be risky and offers ways to design safer molecules in the future.

molecular generationtoxicity predictionsafety evaluationknowledge graphlarge language modelsproperty optimizationprotein-based designAI molecule designhazard rules
Authors
Tong Xu, Xinzhe Cao, Zhihui Zhu, Keyan Ding, Huajun Chen
Abstract
Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of AI-generated molecules. In practice, many generative models may produce molecules with toxic, reactive, or otherwise hazardous characteristics - posing hidden dangers that remain insufficiently addressed. To address this gap, we introduce MolSafeEval, a benchmark dedicated to evaluating and analyzing the safety risks of molecular generation. Unlike prior approaches that rely on narrow toxicity predictors, MolSafeEval integrates heterogeneous safety knowledge - ranging from toxicological databases to hazard rules - into a structured molecular safety knowledge graph. This graph serves as a foundation for large language model-based reasoning, enabling systematic detection and explanation of unsafe features in generated compounds. We further categorize molecular generative models into four representative task types - unconditional generation, property optimization, target protein-based design, and text-based generation - and provide standardized datasets and safety evaluation protocols for each. By systematically revealing the safety vulnerabilities of current generative approaches, MolSafeEval offers a new lens for benchmarking molecular models and provides essential guidance toward safer, more trustworthy molecular design.