ArXiv

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

Authors
Yiting Zheng, Cheng Fang, Anthony Donofrio...
Categories
cs.LG
arXiv
https://arxiv.org/abs/2608.06259v1
PDF
https://arxiv.org/pdf/2608.06259v1

Brief

RxnCLF is a contrastive, transformation-aware reaction foundation model that encodes unified reactant–product information via condensed reaction graphs (CRGs). Pretrained on 1.7 million Pistachio reactions, the model learns a compact, chemically interpretable latent space and—when fine-tuned on several yield benchmarks (Buchwald–Hartwig, Pd-catalyzed BH, proprietary HTE C–N and amide datasets)—achieves improved R2 and best overall performance versus graph and sequence baselines. Full text was not available; summary is based on the abstract.

Why it matters

RxnCLF is a contrastive, self-supervised reaction foundation model pretrained on 1.7 million Pistachio reactions (authors: Yiting Zheng et al., published 2026-08-06) that uses condensed reaction graphs (CRGs) to encode unified reactant–product transformations.

Key details

  • After fine-tuning on yield benchmarks including Buchwald–Hartwig, Pd-catalyzed BH coupling, and proprietary HTE C–N coupling and amide-formation datasets, RxnCLF consistently outperforms graph- and sequence-based baselines, reporting improved R2 and the best overall yield-prediction performance.
  • The CRG-based representation produces a compact, transformation-aware and chemically interpretable latent space capturing reaction-center features plus side-chain context; authors propose extensions to regioselectivity, enantioselectivity, and reaction-condition optimization tasks.
Source evidence

Abstract

Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations. String-, fingerprint-, and graph-based reaction encodings only partially capture chemical transformations, making accurate prediction difficult for reactions with complex substrates. We propose reaction contrastive learning foundation (RxnCLF), a self-supervised contrastive framework for reaction representation learning. RxnCLF is built on a condensed reaction graph (CRG) that unifies reactant and product information into a single graph, enabling the model to learn explicit and enriched transformation structure rather than disconnected graphs. Pretrained on 1.7 million Pistachio reactions, RxnCLF learns a compact and continuous latent space that captures both reaction-center features and broader side chain contexts, making it transformation-aware and chemically interpretable. Fine-tuned on multiple yield prediction benchmarks, including Buchwald-Hartwig, Pd-catalyzed BH coupling, and proprietary HTE C-N coupling and amide formation datasets, RxnCLF consistently outperforms graph and sequence-based baselines, improving R2 and achieving the best performance overall. Our results highlight the promise of CRG-based RxnCLF as a scalable reaction foundation model, with the potential to generalize across broader reaction spaces and support diverse downstream reaction informatics tasks, including regioselectivity prediction, enantioselectivity prediction, and reaction condition optimization.

Comment: 8 pages, 6 figures