Abstract
<title>Abstract</title> <p>Clinical question-answering first-token benchmarks can reward routes that appear fast by omitting provenance, using out-of-window evidence, hiding conflicts, or reusing stale dependencies. We introduce MedRouteGuard, which treats the complete answer-and-evidence route as the benchmark unit, admits evidence-capable candidates before execution, validates returned packets afterward, and retains one clock across failed attempts. In a retrospective evaluation under an every-successful-replay rule covering 2,541 routes for 847 clinician-authored questions, admission changed 61/847 (7.2%; 95% CI, 5.7–9.0%) request-to-first-generated-token winners and shifted retrospective-oracle P95 from 2.89 to 3.24 s. A public medication-safety workload reproduced 21/326 (6.4%) winner changes. Across 180 reviewed questions, admission-selected packets had 6.1 percentage points fewer material evidence-validity errors (95% CI, 1.9–10.4); 14 output-discordant pairs localized the association (2/14 versus 13/14 errors). Specialist validation reached 92.3% sensitivity and 93.3% specificity. This conservative estimand exposes a measurable first-token benchmark denominator error and changes the clinical evidence selected for review.</p>