ArXiv

AI Exposure Scores: what they measure, what they miss, and what comes next

Authors
Campbell Lund, Thomas Euyang, Zanele Munyikwa...
Categories
cs.AI, econ.GN
arXiv
https://arxiv.org/abs/2606.23633v1
PDF
https://arxiv.org/pdf/2606.23633v1

Brief

AI exposure scores are evaluated by Lund, Euyang, Munyikwa, and Fadaee (2026), who show that static 2023 exposure metrics (Eloundou et al.) measure the fraction of occupational tasks LLMs can assist with but fail across time, place, and task ontologies. They catalogue five methodological improvements and argue the larger problem is weak researcher–policymaker coordination; better measurement helps but cannot substitute for participatory, policy-focused collaboration.

Why it matters

Lund et al. (2026) critique the widely used 2023 exposure scores from Eloundou et al. (the 'GPTs are GPTs' scores), noting these scores define exposure as the share of occupational tasks an LLM can assist with but suffer from temporal, geographic, and ontological limitations that distort policy analyses.

Key details

  • The paper surveys five research responses—dynamic/benchmark-based measures, ensemble methods, task-framework extensions, worker-centered metrics, and adoption/usage data—and highlights a coordination gap: policymakers must widen evidence, engage workers as epistemic partners, and shift from prediction to preparedness, while researchers should build data infrastructure, use participatory methods, and write with policymakers in mind.
Source evidence

Abstract

A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produced by Eloundou et al. (2023) and referred to here as the GPTs are GPTs scores, they define exposure as the share of occupational tasks a large language model can assist with. This work is a genuine methodological contribution, but as the scores travel from the time and place they were produced, the limitations the authors named do not always travel with them. Two gaps have widened as a result. The first is structural, between what static exposure scores measure and what policy questions actually require. Taking the diffusion of these scores as a case study, we show how their temporal, geographic, and ontological limitations compound in policy-facing analyses, and we survey five families of research responding to these limits: dynamic and benchmark-based measures, ensemble methods, task-framework extensions, worker-centered metrics, and adoption and usage data. The second gap is the one we argue needs more attention: the coordination between researchers and policymakers. The policy-relevant work which ask who is harmed, who benefits, how, and when, continues to reference the static GPTs are GPTs scores without engagement with the methodological updates that would let these questions be answered more reliably. We then ask what additional steps towards navigating uncertainty remain: ex-post frameworks and the deliberate, political work of reimagining what futures are worthy of building towards are. Closing the research-policy gap is a shared task: policymakers must widen their evidence base, engage workers as epistemic partners, and shift from prediction to preparedness; researchers must build data infrastructure, adopt participatory methods, and write with policymakers in mind. Better measurement matters, but it will not close the second gap alone.

Comment: 19 pages, 4 figures