Abstract
<jats:p>We introduce \textbf{Intermediate-Evidence Substitution (IES)}, a benchmark design methodology that changes evidence granularity while preserving task semantics: decisive aggregate statistics in benchmark evidence are replaced with the raw intermediate values from which they are derived, while task labels, decision protocols, thresholds, and evaluation criteria remain unchanged. IES functions as a diagnostic probe\,---\,any change in model accuracy between the original and transformed evidence reveals how model performance depends on information representation, whether that change is a decrease, an increase, or neither. We instantiate IES on \textbf{ClimateTwinBench}, a 240-question deterministic hypothesis-verification benchmark built over India Meteorological Department (IMD) gridded climate data, spanning six verification categories. Comparing the original evidence design (V1) to its IES-transformed counterpart (V2), we observe statistically significant accuracy changes for Claude Sonnet~4.6 (100.00\%~$\rightarrow$~96.25\%, McNemar $p=0.008$, $\chi^2=7.111$) and Gemini~3.5 (32.50\%~$\rightarrow$~45.42\%, $p=0.008$, $\chi^2=7.087$), while the change for DeepSeek Chat is not statistically significant ($p=0.248$). ChatGPT improves markedly (+25.83~pp, $p\approx4.8\times10^{-11}$); all such outcomes\,---\,in either direction\,---\,are informative empirical results of the diagnostic. To separate numerical derivation from protocol execution, we conduct a structured derivation evaluation in which five models explicitly output predicted intermediate metrics, followed by oracle protocol recovery and cascade analysis distinguishing absorbed from propagated derivation errors. Claude Sonnet~4.6, Claude Opus~4.6, and Gemini~3.5 achieve 100.00\% oracle-recovered protocol accuracy; DeepSeek Chat achieves 99.58\%, with its single propagated failure attributable to connected-component counting in the Spatial Coherence category. GPT-OSS-120B exhibits failures in both metric derivation (4.17\%) and oracle-recovered protocol accuracy (42.92\%), preventing the clean separation between error sources that characterizes the frontier models. For the evaluated frontier models on ClimateTwinBench, numerical derivation\,---\,not protocol execution\,---\,is the dominant source of residual error, a finding that has direct implications for where benchmark evaluation effort should be directed.</jats:p>