Abstract
<jats:title>Abstract</jats:title> <jats:p>White matter lesion (WML) identification, assessment, and characterization using magnetic resonance imaging (MRI) are fundamental for diagnosis and monitoring of multiple sclerosis (MS). Portable ultra-low field (pULF) MRI at 64 millitesla (mT) has been shown to visualize WML with at least one dimension greater than 4 mm. An automated WML segmentation tool catered to pULF-MRI can provide standardized and accurate quantitative measurements of WML volume. In this study, we sought to investigate and compare the accuracy of machine-learning (ML) and deep-learning (DL) pULF MRI segmentation tools. Same-day paired pULF (64mT) and high-field (HF, 3T) MRI scans from 84 adults with MS or suspected-MS (mean age±SD: 48±13, 62 females) included T2-FLAIR and T1w images. Reference WML segmentations were manually annotated on pULF T2-FLAIR for all scans, with WML confirmed with registered HF T2-FLAIR. HF reference WML segmentations were created. Four automated segmentation methods were applied to pULF scans: Method for Inter-Modal Segmentation Analysis (MIMoSA), an ML algorithm trained on HF WML masks; WMH-SynthSeg, a convolutional neural network model with flexible segmentation capabilities across field strengths and resolution; nnU-Net, a DL algorithm trained on pULF reference WML masks; and Pseudo-Label Assisted nnU-Net (PLAn), a DL algorithm pre-trained on HF reference WML masks and refined with 64mT reference WML masks. Two models were trained with nnU-Net, one using T2-FLAIR images only (nnU-Net-FL) and one using T1w and T2-FLAIR images (nnU-Net-FL/T1). The same was done with PLAn, creating PLAn-FL and PLAn-FL/T1. The six automated WML segmentation outputs were compared to the manual segmentations to determine Dice Similarity Coefficient (DSC) scores.</jats:p> <jats:p>Associations of WML volume estimates with clinical measures were investigated.</jats:p> <jats:p>DSC scores with pULF reference WML masks from PLAn-FL (DSC mean±SD: 0.50±0.24) outperformed MIMoSA (0.24±0.20, p<0.0001), WMH-SynthSeg (0.30±0.18, p<0.0001), nnU-Net-FL (0.41±0.24, p<0.0001), and nnU-Net-FL/T1 (0.41 ± 0.26, p = 0.0004). Worse Expanded Disability Status Scale (EDSS) and Scripps Neurologic Rating Scale (SNRS) scores were correlated with higher WML volumes in the pULF and HF reference masks. They were also correlated with WML volumes derived from WHM-SynthSeg, nnU-Net-FL, nnU-Net-FL/T1, PLAn-FL, and PLAn-FL/T1, but not MIMoSA. After adjusting for age, WHM-SynthSeg, nnU-Net FL, nnU-Net-FL/T1, PLAn-FL, and PLAn-FL/T1 had significant associations with EDSS and SNRS scores.</jats:p> <jats:p>nnU-Net and PLAn performed best in segmenting WML on pULF-MRI at 64 mT, providing accurate quantitative estimates of WML burden. Moreover, WML volumes estimated by these algorithms were associated with clinical measures of disability, underscoring their utility for reflecting clinical and radiological disease severity. Given pULF-MRI’s mobility and lower cost, these findings highlight its relevance in clinical trials, particularly in involving more participants who face logistical constraints and barriers.</jats:p> <jats:sec> <jats:title>Highlights:</jats:title> <jats:list list-type="bullet"> <jats:list-item> <jats:p>Deep learning algorithms were trained on 20 paired, same-day pULF and HF MRI</jats:p> </jats:list-item> <jats:list-item> <jats:p>Automated pULF segmentation tools were evaluated using 84 paired, same day MRI</jats:p> </jats:list-item> <jats:list-item> <jats:p>Automated pULF WML segmentations align with reference pULF and HF segmentations</jats:p> </jats:list-item> <jats:list-item> <jats:p>Trained deep learning algorithms demonstrate strong pULF MRI lesion segmentation</jats:p> </jats:list-item> <jats:list-item> <jats:p>Automated pULF WML estimates are associated with clinical disability scores</jats:p> </jats:list-item> </jats:list> </jats:sec>