Abstract
<title>Abstract</title> <p>Public repositories such as ProteomeXchange enable large-scale re-analysis, harmonisation, and reuse of proteomics datasets. Here we perform such a re-analysis, focusing on human serum Data-Independent Acquisition (DIA) datasets, to determine the feasibility of combining datasets and enabling "AI-ready" analysis of large data collections. Our aim was to explore methods for removing batch effects and performing metaanalysis, to determine whether statistical power could be gained in discovering associations between proteins and metadata variables (such as age and sex), without incurring false positives. Using eleven public DIA datasets and relevant metadata, we compared three data manipulation methods: native iBAQ intensities with no manipulation; parts per billion (ppb); and ranked values, each subjected or not to batch correction using ComBat and limma. Methods were evaluated, comparing associated proteins before and after batch correction. Overall, we were unable to effectively remove strong study batch effects, with false correlations often observed. Similarly, feature selection methods resulted in identifying false positives. In these cases, study signal was too strong to allow generation of a combined dataset for data mining or AI-driven analyses. We highlight the importance of treating datasets individually when performing data reanalysis, and the difficulties associated with combining datasets for large scale meta-analysis studies.</p>