ArXiv

VEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester Design

Authors
Depeng Su, Yuyu Luo, Guobiao Hu
Categories
cs.CL, cs.SE
arXiv
https://arxiv.org/abs/2607.18181v1
PDF
https://arxiv.org/pdf/2607.18181v1

Brief

VEHBench is a diagnostic benchmark targeting LLM-assisted design of vibration energy harvesters: it defines 763 literature-grounded tasks and an analytical physical oracle to score designs across four stage-local roles (specification triage, verifier-guided search, corrupted-state recovery, policy-conditioned selection). Experiments (authors Su, Luo, Hu; arXiv 2026-07-20) find pronounced stage-dependent capability—no single model wins every role—supporting role-aware model selection and verifier-grounded workflows.

Why it matters

VEHBench (Depeng Su, Yuyu Luo, Guobiao Hu; arXiv 2026-07-20) is a stage-local diagnostic benchmark for LLM-assisted vibration energy harvester (VEH) design containing 763 literature-grounded tasks scored by an analytical physical oracle and released at https://huggingface.co/datasets/AnonymousVehbench/vehbench.

Key details

  • The benchmark explicitly evaluates four design roles—specification triage, verifier-guided search, corrupted-state recovery, and policy-conditioned selection—providing role-level diagnostics rather than only final-artifact validity.
  • Experimental results show LLM performance is strongly stage-dependent: no single model consistently dominates all four workflow roles, and models exhibit distinct response-control profiles that can guide role-specific selection and routing.
Source evidence

Abstract

Battery-free Internet of Things (IoT) requires iterative design of vibration energy harvesters (VEHs) under coupled physical constraints, while LLMs are emerging as interface layers for engineering workflows. However, existing engineering benchmarks primarily assess final artifact validity, offering limited insights into how LLMs behave across different stages of coupled physical design. We introduce VEHBench, an engineering-native diagnostic benchmark for LLM-assisted VEH design, featuring 763 literature-grounded tasks scored by an analytical physical oracle. VEHBench evaluates four design roles: specification triage, verifier-guided search, corrupted-state recovery, and policy-conditioned selection. Experimental results reveal that LLM capability is strongly stage-dependent: no single model consistently dominates the entire workflow, and response-control profiles expose distinct behavioral patterns across design roles. VEHBench thus provides a stage-aware foundation for evaluating, selecting, routing, and improving verifier-grounded engineering LLMs. The benchmark artifact is available at https://huggingface.co/datasets/AnonymousVehbench/vehbench