ArXiv

Task Decomposition for Efficient Annotation

Authors
Nupoor Gandhi, Emma Strubell
Categories
cs.CL, cs.AI, cs.HC
arXiv
https://arxiv.org/abs/2606.24734v1
PDF
https://arxiv.org/pdf/2606.24734v1

Brief

Task Decomposition for Efficient Annotation (Gandhi & Strubell, arXiv:2606.24734v1, 2026-06-23) formalizes decomposing structured annotation into sub‑tasks to lower aggregate inferential load. Drawing on centering theory, it defines inferential load via degrees of freedom and shows that isolating 'centers'—salient anchor entities—reduces output-space complexity. The paper gives decomposition guidelines and an allocation procedure for mixing human and model annotators to maximize quality under a fixed budget; this summary is based on the abstract.

Why it matters

Gandhi and Strubell (arXiv:2606.24734v1, published 2026-06-23) propose decomposing structured annotation into sub-tasks to reduce aggregate inferential load and improve cost-efficiency compared to traditional end-to-end single-annotator workflows.

Key details

  • They introduce a formal model of inferential load based on degrees of freedom in the space of valid annotations and adapt centering theory to identify 'centers' (salient anchor entities) whose identification constrains output-space complexity.
  • The paper provides practical decomposition guidelines and a procedure to allocate sub-tasks across heterogeneous annotators (human and model) to maximize annotation quality under a fixed budget; examples cited show improved cost-efficiency (details in full text).
Source evidence

Abstract

High-quality annotations of structured representations are expensive to collect over large corpora. Manual annotation of structure is laborious, and model-based annotation, although cheaper to generate, requires expensive validation and potentially significant supervision to ensure that the annotation quality is strong enough to be useful downstream. In traditional annotation workflows, annotation of each complete example is performed end-to-end by a single annotator. However, structured annotation is complex, and each aspect of the task represents a unique challenge with an associated inferential load for a given annotator. Modern annotation projects can incorporate heterogeneous groups of annotators, including both models and human annotators with varying domain and linguistic expertise. It remains unclear, however, how to redesign annotation tasks in this setting, where efforts are discriminately allocated across heterogeneous annotators with respect to distinct annotation challenges. We propose to decompose annotation tasks into sub-tasks in order to reduce the aggregate inferential load of annotation projects. Inspired by the notion of centers from centering theory, we introduce a formal model of inferential load based on the degrees of freedom in the space of valid annotations. Using this model, we show that identifying these centers (i.e. salient anchor entities realized by annotation sub-tasks) constrains the output space complexity, and decompositions which isolate and advance center identification reduce the aggregate inferential load. We provide guidelines for decomposing complex structured annotation tasks, supported by examples demonstrating improved cost-efficiency from our prior work. Finally, we present a procedure for allocating sub-tasks across annotators to maximize quality under a fixed budget.