Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Background Chemotherapy dose adaptation requires sequential trade-offs between tumor control and physiological tolerance. Offline reinforcement learning can support preclinical evaluation of candidate dosing policies. For oncology applications, however, such evaluation requires explicit safety constraints, transparent simulation assumptions, and testing under clinically heterogeneous conditions before any patient-facing use. Methods A TCGA-informed digital twin framework was developed in which TCGA-STAD clinical variables parameterized mechanistic simulations rather than direct clinical predictions. Age, AJCC pathologic stage, and tumor grade informed initial conditions, safety thresholds, and selected kinetic parameters in a four-compartment chemotherapy model. Chemotherapy dosing was formulated as a constrained Markov decision process with four dose levels, a treatment reward, and a separate binary safety cost. SafeCQL was trained on offline simulated trajectories generated from training twins and evaluated on 93 held-out TCGA-informed digital twins. A 21-seed safety-budget analysis was then used to characterize training variability. Results In the prespecified representative checkpoint (epsilon = 0.1, seed = 15), SafeCQL achieved a return of -8.70 +/- 2.94. It selected an average dose of 1.01 and produced an empirical violation rate of 9.23% across 93 held-out twins. Compared with behavioral cloning and unconstrained CQL (alpha = 5), this checkpoint selected a lower average dose and reduced empirical safety-violation rates. Simulated return was lower than the unconstrained learned baselines and the high-dose fixed comparator. In the 21-seed safety-budget sweep, mean return was highest at epsilon = 0.5. The epsilon = 0.1 setting was retained as a prespecified safety-focused operating point because it provided conservative dose exposure and lower violations than BC/CQL in the representative comparison. Conclusions TCGA-informed digital twins provide a transparent preclinical testbed for evaluating safety-constrained offline reinforcement learning in chemotherapy policy triage. The framework is not a bedside dosing rule. Its clinical decision-support value lies in making efficacy-safety trade-offs explicit, stress-testing candidate policies in matched virtual patients, and prioritizing hypotheses for retrospective validation before prospective patient-facing evaluation.</p>

Show More

Keywords

dose chemotherapy twins offline evaluation

Related Articles

PORE

About

Connect