Abstract
<jats:p>Recent developments in deep learning have advanced digital pathology and biomedical image analysis by enabling unprecedented performance in a variety of image analysis tasks, yet implementation can remain challenging due to data scarcity. Sufficiently large and diverse human imaging datasets are a requisite in the presence of biological heterogeneity, but are often unavailable due to disease rarity, limited specimen access, and practical constraints on data collection. Pancreatic neuroendocrine tumors (PNETs) provide one such example – a heterogeneous disease where access to human specimens can be limited and class-balanced datasets are difficult to assemble. While standard data augmentation methods can mitigate model convergence and overfitting issues, none introduce true biological diversity into the dataset, and cannot compensate for fundamental limits imposed by biological heterogeneity. Here, we present a novel framework for “biological data augmentation” using an unpaired animal-to-human image translation framework to improve deep learning classification by injecting biological diversity in data scarce contexts, and we demonstrate this with a dataset composed of label-free multiphoton microscopy (MPM) images of PNETs. Our framework maps animal images into the human imaging domain without requiring paired acquisitions, allowing translated images to be incorporated as biologically informed augmentation samples during classifier training. Unlike conventional augmentation, this strategy introduces additional structural and textural variability derived from real preclinical tissue data. Using patient-level splits and multiple cross-validation folds, we compared classifiers trained on the original human dataset alone with classifiers trained on datasets supplemented by translated images. Incorporating translated images improved classification performance, increasing mean accuracy from 74% to 79% with no loss of prediction stability. These results suggest that unpaired cross-species translation can expand the effective diversity of limited human datasets and generate augmentation samples that retain task-relevant information for downstream modeling. This work uses MPM images of PNETs as a representative test case, but more broadly, we introduce a generalizable strategy to mitigate data scarcity in digital pathology and to support deep learning development across rare-disease and small-cohort biomedical imaging applications.</jats:p>