ArXiv

CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications

Authors
Brendan Smith, Susana Lopez-Moreno, Eric Dolores-Cuenca...
Categories
cs.LG, cond-mat.mtrl-sci, cond-mat.other, cs.AI, physics.chem-ph
arXiv
https://arxiv.org/abs/2608.04942v1
PDF
https://arxiv.org/pdf/2608.04942v1

Brief

CheMLFlow is an open-source platform designed to streamline scientific ML by providing modular, configuration-driven pipelines that cover data acquisition, curation, representation, model training, validation, screening, interpretation, and reporting. Based on the abstract, the system emphasizes reproducibility (deterministic splits, explicit artifacts), automation and agent integration, and reports benchmark results matching literature performance on quantum, physicochemical, and bioactivity prediction plus time-series use cases, showing broader applicability.

Why it matters

CheMLFlow is an open-source platform that assembles end-to-end, high-throughput and agentic workflows with modular components, ready-to-run reference pipelines, pluggable representations/models, deterministic data splits, explicit run artifacts, batch execution, and automated report generation to reduce orchestration overhead.

Key details

  • The authors report benchmarks that reach literature performance for quantum-mechanical, physicochemical, and bioactivity property prediction and demonstrate use cases on time-series datasets, indicating applicability beyond molecular chemistry datasets.
  • Paper metadata: Brendan Smith, Susana Lopez-Moreno, Eric Dolores-Cuenca et al.; published to arXiv 2026-08-05 (arXiv:2608.04942v1); abstract and PDF available at https://arxiv.org/abs/2608.04942v1.
Source evidence

Abstract

CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning development, where researchers often need to assemble data acquisition, curation, representation, model training, validation, screening, interpretation, and reporting into a reproducible pipeline, even when their primary research contribution concerns only one stage. CheMLFlow provides modular workflow components, ready-to-run reference pipelines, standardized artifacts, and evaluation outputs that reduce orchestration overhead and support benchmarking across methods and datasets. The platform is designed to be extensible, reproducible, and automation friendly, with pluggable representations and models, deterministic splits, explicit run artifacts, batch execution, and report generation. As scientific software increasingly moves toward agent assisted experimentation, CheMLFlow's configuration driven workflows and structured outputs also provide a practical interface for coding agents to help users construct experiments, inspect results, and summarize findings under human supervision. This article describes the system architecture, core workflows, and benchmarks that reach literature performance for quantum mechanical, physicochemical and bioactivity property prediction, and use cases involving time series datasets demonstrating applications beyond molecular chemistry datasets.