Abstract
<title>Abstract</title> <p>Medical image anomaly detection and industrial visual inspection face sharedchallenges, including scarce abnormal samples, heterogeneous defect morphology,cross-domain shifts, and expensive pixel-level annotation. Existing distillationand reconstruction methods often rely on a single teacher or weak domain selection, limiting their ability to capture local texture deviations, blurred lesionboundaries, and global structural inconsistencies. This paper proposes SGMSDTDNet, a semantic-guided multi-scale dual-teacher distillation network forunified medical and industrial anomaly detection. A target-view teacher and asource-view teacher provide domain-specific semantics and transferable structural priors. The Dynamic Adaptive Multi-scale Enhancement module (DAME)strengthens fine-grained boundaries and contextual cues, the Diffusion Reconstruction Branch (DRB) models recoverable normal feature manifolds, and theSemantic-guided Domain-relevant Feature Selection module (S-DFS) suppressesdomain-irrelevant activations through adaptive semantic gating. Region connectivity analysis further converts noisy pixel-level responses into coherent anomalymaps and image-level scores. To evaluate methodological trends under controlled conditions, we conducted fixed-seed Monte Carlo experiments on fiveprotocol-defined medical and industrial domains. After stricter difficulty calibration, SGMS-DTDNetachieves average image-level AUROC, pixel-level AUROC,pixel-level AP, and image-level F1 scores of 92.41%, 93.06%, 57.38%, and 85.91%,respectively. The ablation experiments show that the proposed modules progressively improve localization robustness without producing near-saturatedscores.</p>