Abstract
<p>How does the visual system construct object representations that are robust to the removal of color? I propose a two-stage scaffolding model. In Stage 1 (prenatal), retinal waves establish spectrally non-selective spatiotemporal contrast sensitivity, the necessary structural foundation of the visual system, before any visual experience occurs. Because the wave-generating circuitry encodes only contrasts in firing rate with no access to spectral information, this foundation is necessarily achromatic: whatever circuits retinal waves entrain, they are trained on a signal that carries no chromatic content. In Stage 2 (postnatal), the globally degraded neonatal visual signal, low spatial resolution, low contrast sensitivity, and immature chromatic processing, all coupled through shared photoreceptor immaturity, forces the first object representations to be built on luminance structure, the only dimension with sufficient fidelity under maximal front-end degradation. Together, these two stages explain why normal observers recognize objects in grayscale as readily as in color: luminance structure was the foundation of object representations from the beginning. The model also explains why late-sighted individuals who gain sight after dense congenital cataracts show disproportionate reliance on color: they retain the prenatal structural scaffold (intact retinas generating normal retinal waves) but lack the postnatal experiential scaffold (years of globally impoverished visual experience that normally grounds object representations in luminance). The model generates testable predictions that distinguish it from color-specific developmental accounts, and several can be evaluated on existing computational architectures without new patient data.</p>