Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Datasets are the foundation of modern artificial intelligence, yet the cost of building them, and whether that cost buys scientific value, is barely measured at scale. We develop a language-model-assisted pipeline that estimates machine and human-labor inputs separately, in dollars and in hours, and apply it to 173{,}369 datasets linked to arXiv papers from 2010 to 2025, yielding the first per-dataset cost record at population scale. The typical per-dataset cost follows a non-monotonic trajectory: broadly stable, declining through 2024, then rebounding in 2025. Yet a dataset's scholarly impact is largely uncoupled from its cost: neither machine nor human-hour cost is a meaningful predictor, whereas author count and topical prevalence are far stronger correlates. These findings inform how datasets are developed, resourced, and credited across the AI data economy. We release the pipeline and cost-estimation model as open weights for transparent, reproducible policy on dataset funding and attribution.</p>

Show More

Keywords

cost datasets scale pipeline machine

Related Articles

PORE

About

Connect