DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery
2026-08-05 • Machine Learning
Machine LearningComputational Engineering, Finance, and ScienceSymbolic Computation
AI summaryⓘ
The authors developed a new method that uses a language model (LLM) to help find chemical reaction rate equations more quickly and accurately. Their method combines traditional symbolic regression with the LLM, which checks and suggests better models based on chemical knowledge. They tested this on simulated chemical and biological systems, showing it needed fewer tries to find the correct equations without losing accuracy. This approach could save time and effort in real experiments by reducing the number of experiments required. Overall, the authors show that language models can guide scientific discovery in chemistry by adding expert knowledge automatically.
kinetic modelsymbolic regressionrate expressionlarge language modelheterogeneous catalysisbioprocessmodel discoveryphysicochemical critiqueautomated modelingexperimental effort
Authors
Roberto Aliaga Medina, Paulina Quintanilla, Antonio del Rio Chanona
Abstract
Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models. Large language models (LLMs) offer a promising avenue for injecting domain expertise into this search. Here, we introduce an LLM-guided SR framework, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery. The LLM performs two roles at each iteration: (1) a qualitative physicochemical critique of the best SR candidates, and (2) the proposal of new candidate rate expressions guided by the SR-generated models and embedded chemical knowledge. Our framework is evaluated on four in silico case studies of increasing complexity, spanning heterogeneous catalysis and bioprocess systems. Results show the LLM-guided framework reduces iterations to identify the ground-truth model by $41.7-79.3\%$ versus a state-of-the-art SR framework, with the LLM directly proposing the correct model structure in over half of the guided runs. In practical settings, where each iteration typically requires a new wet-lab experiment, this translates into a substantial reduction in experimental effort. Predictive performance on an independent validation set is equivalent between both approaches, with $R^2>0.98$ in all case studies. Ablation studies indicate that both the SR component and the LLM scale contribute to this performance, with a reduced-size LLM largely retaining discovery efficiency. These findings demonstrate that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.