ArXiv

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

Authors
Yifan Zhou, Qihao Yang, Yan Li...
Categories
cs.AI
arXiv
https://arxiv.org/abs/2607.08758v1
PDF
https://arxiv.org/pdf/2607.08758v1

Brief

IG-Bench (IdeaGene-Bench) formalizes scientific ideas as minimal, typed, evidence-grounded Idea Genome objects with GenomeDiffs capturing inheritance, mutation, loss, import, and novel insertion across six evolutionary dynamics. The dataset (1,961 lineage traces, 1,085 genomes, 920 diffs, 10 domains) supports IG-Exam (42 task types, 1,029 instances) for reasoning and IG-Arena with a Population-Evolution Score for generation. Experiments on 14 LLMs expose a compositional bottleneck (best exact reasoning accuracy 27.3%), highlighting gaps in current benchmarks and the challenge of producing lineage-coherent, selectable research proposals. (Summary based on the provided abstract.)

Why it matters

IdeaGene-Bench (IG-Bench) introduces a formal IdeaGenome/GenomeDiff representation and contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records spanning 10 scientific domains.

Key details

  • IG-Bench supports two evaluations: IG-Exam (42 task types, 1,029 instances) for closed-form lineage reasoning, and IG-Arena which uses a lineage-conditioned Population-Evolution Score (PES) to evaluate lineage-grounded idea generation.
  • Benchmarks on 14 LLM-based 'scientists' reveal a compositional bottleneck: the best system attains only 27.3% exact accuracy on lineage reasoning, and providing structured lineage context reshuffles model rankings rather than uniformly helping models.
Source evidence

Abstract

Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.