ArXiv

The authors built a benchmark with 1,600 human-written hallucination samples…

Authors
Timothee Mickus, Claudio Savelli, Eduardo Calò...
Categories
cs.CV, cs.AI, cs.CL
arXiv
https://arxiv.org/abs/2608.01021v1
PDF
https://arxiv.org/pdf/2608.01021v1

Brief

Human-written hallucination samples are proposed as a perennial alternative to model-generated benchmarks. The authors collect 1,600 human-written, span-level annotated examples in Chinese, English, French, and Italian and compare them to 18,400 outputs from five vision-and-language models. They find higher annotator agreement, finer dataset control, and distributional similarity, suggesting human data can validly substitute model-derived hallucination benchmarks. (Summary based on abstract; full text not checked.)

Why it matters

The authors built a benchmark with 1,600 human-written hallucination samples across four languages (Chinese, English, French, Italian) and 18,400 samples from five vision-and-language models, all annotated with a fine-grained span-level labeling scheme.

Key details

  • Human-written samples yielded higher annotator agreement and allowed greater control over dataset contents compared to model-generated hallucinations.
  • Human data was distributionally similar to model-derived samples and gave a reasonable portrayal of detection capabilities, supporting human-written benchmarks as a viable substitute for model-generated ones (paper published 2026-08-02).
Source evidence

Abstract

In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600 human-written samples, spanning four languages (Chinese, English, French, Italian), and 18,400 samples from five vision-and-language models, all annotated for hallucinations using a fine-grained span-level labeling scheme. We find that human-written samples result in higher agreement and allow greater control of dataset contents, while remaining distributionally similar to samples derived from vision-and-language samples and providing a reasonable portrayal of detection capabilities - suggesting that human data is a viable substitute for model-based hallucination benchmarks.