Abstract
<sec> <title>UNSTRUCTURED</title> <p>Artificial Intelligence (AI) in Biomedicine is increasingly moving toward multi-modal architectures that integrate clinical epidemiological variables with multi-omics data using multi-modal architectures. Existing reporting frameworks, e.g. TRIPOD-AI, CONSORT-AI, DECIDE-AI, improved the understanding of model development and evaluation, but still, they show a model-centric approach. In multi-modal AI systems, a model does not receive variables as input, but complex upstream processes of data acquisition, transformation, and integration. Although these processes impact the behaviour of the model, they are rarely described. We claim that data-layer transparency is a critical, and often overlooked, aspect of reproducibility of AI in biomedicine. Therefore we propose the MERGE framework, which shows a minimal conceptual scaffold, and a starting point, to indicate the key areas in need of reporting. This approach shifts the focus to the upstream parts of the process, and highlights that in order to be able to interpret a model in a reasonable way, and to use the AI model in clinical practice, the processes for data generation have to be made transparent.</p> </sec>