ArXiv

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

Authors
Brandon Wang, Andrei S. Tyrin, Daniil A. Boiko
Categories
cs.LG
arXiv
https://arxiv.org/abs/2608.02595v1
PDF
https://arxiv.org/pdf/2608.02595v1

Brief

onepot-Bench 0, introduced by Wang et al. (arXiv 2026-08-03), is a proprietary benchmark suite to evaluate language models on synthetic chemistry tasks relevant to wet‑lab execution. It comprises ChemAbacus (tool‑free cheminformatics and numerical reasoning), SynthRefusal (safety/refusal across benign, controlled, and designer‑drug targets), and SynthBench (reaction‑outcome and catalyst prediction using private lab experiments). Full text not available; summary based on the abstract.

Why it matters

onepot-Bench 0 (Brandon Wang, Andrei S. Tyrin, Daniil A. Boiko; arXiv 2026-08-03) is a proprietary benchmark suite for evaluating language models on wet‑lab synthetic chemistry tasks and comprises three evaluations: ChemAbacus (tool‑free cheminformatics literacy and numerical reasoning), SynthRefusal (safety/refusal across benign, controlled, and designer‑drug targets), and SynthBench (reaction‑outcome prediction and catalyst selection).

Key details

  • SynthBench uses private experimental data generated in the authors' laboratory to evaluate reaction outcome and catalyst choice, explicitly aiming to avoid public‑data leakage present in prior benchmarks and to probe basic competency, reliability, and deeper domain knowledge needed for reliable performance in physical labs.
Source evidence

Abstract

Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific intuition. Existing evaluations rarely measure the capabilities required to make reliable decisions in a physical laboratory and often rely on public data that may have appeared in model training corpora. We introduce onepot-Bench 0, a proprietary benchmark suite for evaluating language models on synthetic chemistry capabilities relevant to wet-lab execution. onepot-Bench 0 comprises three complementary evaluations: ChemAbacus measures tool-free cheminformatics literacy and numerical reasoning; SynthRefusal characterizes safety and refusal behavior across a variety of benign, controlled, and designer-drug targets; and SynthBench evaluates reaction-outcome prediction and catalyst selection using private experimental data generated in our laboratory. Together, these evaluations probe basic competency, reliability, and deeper knowledge, all skills which are required for reliable performance in the lab.