Abstract
<title>Abstract</title> <p>Multimodal medical image fusion combines structural and functional imaging to support clinical interpretation. Saliency-guided weighting is widely adopted as a simple, training-free way to emphasize informative image regions during fusion. However, most studies validate a saliency-guided method on a single modality pair and report only the metrics on which it performs best, leaving open whether the reported benefit generalizes across modalities and across the full space of evaluation metrics. This study addresses that gap through a controlled diagnostic framework, SAW-Fuse, which compares three baseline fusion rules, four individual saliency cues, and their combined multi-cue rule under an identical weighted fusion procedure. This procedure was applied to 24 registered image pairs from each of three modality pairs drawn from the Harvard Whole Brain Atlas: MRI-CT, MRI-PET, and MRI-SPECT. Every fused image was scored on nine established metrics and compared using the Friedman test, the Wilcoxon signed-rank test, and the Cliff delta effect size. The multi-cue rule reduced mutual information relative to a simple average baseline on all three modality pairs, by 15.4, 33.4, and 25.4 percent respectively, all with large effect sizes. Its effect on edge preservation reversed in sign with modality type, improving over the baseline by 39.4 percent on the structural pair, MRI-CT, but degrading it by 17.8 and 28.8 percent on the functional pairs, MRI-PET and MRI-SPECT. Combining four saliency cues did not exceed the strongest individual cue on any modality pair, and a rank spread analysis showed that the maximum fusion baseline, despite leading on five of nine metrics, simultaneously ranked lowest on entropy on every modality pair, exposing an internal inconsistency in the standard metric battery rather than in any single method. These findings indicate that the benefit of saliency-guided fusion is modality-conditioned rather than general, and that single-metric or narrow-metric reporting can mask this dependence. A small unblinded visual pilot gave an early, consistent signal in the same direction. The results discuss modality-specific validation and wide metric batteries as a minimum standard for evaluating fusion strategies intended for clinical use.</p>