Abstract
<p>Behavioral science increasingly relies on complex, computational pipelines to infer latent psychological variables from time-series data. Because the same data can often be processed using multiple plausible approaches, analytic choices may substantially affect results, yet systematic method comparisons remain uncommon. Here, we adress this problem for pupillometric measurement of fear conditioning. We introduce M-BIDS, a constrained, machine-interpretable extension of the Brain Imaging Data Structure, together with a distributed dynamic database (DDDB) architecture for harmonizing source data stored across multiple repositories. Using this framework, we assembled PupilFear, comprising 361 pupillometry datasets from fear-acquisition experiments. We benchmarked more than 700 candidate measurement methods spanning alternative preprocessing and response-quantification choices. Performance was evaluated using retrodictive validity, and decision tree regression was used to identify the analytic decisions most strongly associated with performance. The distinction between condition- and trial-wise modelign was the dominant determinant of retrodictive validity. Condition-wise model-based approaches showed relatively stable performance across parameter settings, whereas trial-wise approaches were substantially more sensitive to preprocessing choices. Nevertheless, the best condition-wise and trial-wise methods achieved comparable retrodictive validity (d = 0.65). Methods also differed in their sensitivity to reinforced trials (US bias). Peak-scoring approaches showed lower retrodictive validity overall (d = 0.45) but were less influenced by reinforcement-related responses. These finding provide practical guidance for selecting pupillometry pipelines according to the research question. The publicly available PupilFear corpus and benchmarking workflow also enable researchers to reproduce the analyses and evaluate new methods against a common reference.</p>