Abstract
<p>Validating a scale requires spending a sample before knowing whether its structure will withstand response biases. This work presents a pre-empirical stress test that turns item wording into the parameters of a response-process simulator: a language model rates the desirability of each statement, embeddings detect near-paraphrased pairs, and the intended structure is subjected to increasing doses of six classical biases up to the breaking point, along with the items that fall first. On synthetic scales no alarms were fabricated (median clean structure = 1.00), and the seeded desirability halo was detected with a sensitivity of 0.97 and a specificity of 1.00, whereas two paraphrased pairs sank the probability from 0.90 to 0.01. The estimated desirability was stable across language models and correlated between 0.42 and 0.86 with the observed endorsement. Applied to nine public scales, the forecast separated those with real structural problems from the rest, and desirability was the only bias whose breaking point discriminated among scales. An additional calibration study with four Spanish couple scales showed that the real local dependences escaped both the lexical trigger and a paraphrase judge, and that half of them responded to positional adjacency, which grounds the scenario-based design of the engine. An optimistic forecast indicates the absence of wording-induced risk, not a warrant of good fit or of theoretical validity. Its use as an audit prior to fieldwork, together with its limits, is discussed.</p>