ArXiv

DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery

Authors
Roberto Aliaga Medina, Paulina Quintanilla, Antonio del Rio Chanona
Categories
cs.LG, cs.CE, cs.SC
arXiv
https://arxiv.org/abs/2608.05120v1
PDF
https://arxiv.org/pdf/2608.05120v1

Brief

DASyR-LLM introduces an LLM-integrated symbolic regression pipeline for automatic kinetic rate discovery: the LLM provides physicochemical critiques and proposes new candidate expressions each iteration. Evaluated on four in silico problems (heterogeneous catalysis to bioprocesses), it cut iterations by 41.7–79.3% versus state-of-the-art SR, often proposed correct models (>50% runs), and achieved R^2>0.98 on validation. Full text was not available; summary based on the abstract.

Why it matters

DASyR-LLM (Aliaga Medina, Quintanilla, del Rio Chanona; arXiv 2608.05120v1, published 2026-08-05) embeds an LLM into an iterative symbolic regression loop for kinetic model discovery and was evaluated on four in silico case studies spanning heterogeneous catalysis and bioprocess systems.

Key details

  • The LLM-guided framework reduced iterations to recover ground-truth models by 41.7–79.3% versus a state-of-the-art SR baseline, with the LLM directly proposing the correct model structure in over 50% of guided runs.
  • Predictive performance matched the baseline (R^2 > 0.98 on independent validation in all case studies); ablation shows both SR and LLM scale matter, and a reduced-size LLM largely retained discovery efficiency.
Source evidence

Abstract

Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models. Large language models (LLMs) offer a promising avenue for injecting domain expertise into this search. Here, we introduce an LLM-guided SR framework, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery. The LLM performs two roles at each iteration: (1) a qualitative physicochemical critique of the best SR candidates, and (2) the proposal of new candidate rate expressions guided by the SR-generated models and embedded chemical knowledge. Our framework is evaluated on four in silico case studies of increasing complexity, spanning heterogeneous catalysis and bioprocess systems. Results show the LLM-guided framework reduces iterations to identify the ground-truth model by $41.7-79.3\%$ versus a state-of-the-art SR framework, with the LLM directly proposing the correct model structure in over half of the guided runs. In practical settings, where each iteration typically requires a new wet-lab experiment, this translates into a substantial reduction in experimental effort. Predictive performance on an independent validation set is equivalent between both approaches, with $R^2>0.98$ in all case studies. Ablation studies indicate that both the SR component and the LLM scale contribute to this performance, with a reduced-size LLM largely retaining discovery efficiency. These findings demonstrate that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.