Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> Graph neural network (GNN) policies trained with deep reinforcement learning (DRL) have been shown to generalize to job-shop instances larger than those seen during training and, separately, to be trainable for robustness against a single type of stochastic disruption. Whether these two properties — size generalization and multi-disruption robustness — co-emerge in the same policy has not previously been tested jointly under a statistically rigorous, pre-registered protocol. This study trains a relational graph-isomorphism-network (GIN) encoder with a proximal policy optimization (PPO) policy for the dynamic, stochastic flexible job-shop problem (FJSP) and pre-registers five hypotheses (H1–H5), with decision rules frozen before any experiment was run, evaluated against nine priority dispatching rules (PDR), a genetic algorithm (GA), and an exact constraint-programming solver (CP-SAT) across the Fisher–Thompson, Lawrence, and Brandimarte benchmark families. Four of five hypotheses are rejected. The policy does not match the best dispatching rule on static instances (H1: mean relative percentage deviation [RPD] 24.82% versus 18.96%, <italic>N</italic>  = 43, <italic>p</italic>  = 4.03 × 10⁻⁵); it is not robust under nine disruption regimes spanning machine breakdowns, stochastic processing times, and dynamic arrivals (H3: <italic>p</italic>  &lt; 1.7 × 10⁻¹¹ in all nine regimes); and augmenting the reward with a disruption-aware instability penalty fails to improve robustness without collapsing nominal performance (H4: nominal RPD increases by 46.6 to 106.0 percentage points). Decision latency is one to two orders of magnitude below the solver's on large, hard instances, but does not clear the strict pre-registered threshold on every instance (H5). Size generalization, by contrast, is accepted (H2): across four size tiers (1.5×–3.0× the training size), the policy's rank among ten methods is never significantly worse than the best dispatching rule's, even though its absolute optimality gap remains uncompetitive. We argue that this pattern — architectural transfer of rank without transfer of competitiveness, and no automatic transfer of robustness — is itself the study's central, defensible contribution: size generalization and multi-disruption robustness are separable properties of GNN–PPO scheduling policies, not a package deal, diagnosed here for the first time under a pre-registered, statistically powered, jointly tested protocol, with a concrete, actionable mechanism identified for every rejection. </p>

Show More

Keywords

robustness size policy instances stochastic

Related Articles

PORE

About

Connect