Abstract
<title>Abstract</title> <p>Comparisons of pretraining strategies for low-label cell instance segmentation can be confounded by transferred detector components, backbone adaptation strength, training budget, and an evaluation cap that is too small for dense microscopy images. We conduct a module-controlled study on the official 5\% LIVECell split using a fixed Mask R-CNN with a ResNet-50 feature pyramid. Four development-frozen transfer protocols are compared: COCO-full with a $1\times$ backbone learning-rate multiplier and three hybrid protocols with a $5\times$ multiplier. On 1,564 official test images and three training seeds, 20-epoch mask AP at 3,000 detections per image is $0.2413\pm0.0057$, $0.2222\pm0.0054$, $0.2082\pm0.0087$, and $0.2063\pm0.0057$ for COCO-full, ImageNet+COCO, SimCLR+COCO, and SupCon+COCO, respectively. With matched random detector components, ImageNet and SimCLR obtain close observed means, while their seed-wise difference changes direction. In a separate 100-epoch $2\times2$ control, COCO exceeds SimCLR at $1\times$, whereas SimCLR exceeds COCO at $5\times$, in all three observed seeds; the difference-in-differences ranges from $+0.0229$ to $+0.0433$. Raising the evaluation cap from 100 to 3,000 adds approximately 0.10 AP, although images containing more than 1,000 cells remain near 0.12--0.13 AP. The observed pretraining rankings depend on the explicit transfer-and-adaptation protocol and cannot be attributed to backbone weights alone; dense-cell evaluation likewise requires limits aligned with image density.</p>