arXiv ScienceSearch

arXiv · 2506.10073

Patient-Specific Deep Reinforcement Learning for Automatic Replanning in Head-and-Neck Cancer Proton Therapy

Abstract

Anatomical changes during intensity-modulated proton therapy (IMPT) for head-and-neck cancer (HNC) can shift Bragg peaks, risking tumor underdosing and organ-at-risk overdosing. Treatment replanning is often required to maintain clinically acceptable treatment quality. However, current manual replanning processes are resource-intensive and time-consuming. We propose a patient-specific deep reinforcement learning (DRL) framework for automated IMPT replanning, with a reward-shaping mechanism based on a $150$-point plan quality score addressing competing clinical objectives. We formulate the planning process as a reinforcement learning problem where agents learn control policies to adjust optimization priorities, maximizing plan quality. Unlike population-based approaches, our framework trains agents for each patient using their planning Computed Tomography (CT) and augmented anatomies simulating anatomical changes (tumor progression and regression). This patient-specific approach leverages anatomical similarities along the treatment course, enabling effective plan adaptation. We implemented two DRL algorithms, Deep Q-Network and Proximal Policy Optimization, using dose-volume histograms (DVHs) as state representations and a $22$-dimensional action space of priority adjustments. Evaluation on eight HNC patients using actual replanning CT data showed that both agents improved initial plan scores from $120.78 \pm 17.18$ to $139.59 \pm 5.50$ (DQN) and $141.50 \pm 4.69$ (PPO), surpassing the replans manually generated by a human planner ($136.32 \pm 4.79$). Clinical validation confirms that improvements translate to better tumor coverage and OAR sparing across diverse anatomical changes. This work highlights DRL's potential in addressing geometric and dosimetric complexities of adaptive proton therapy, offering efficient offline adaptation solutions and advancing online adaptive proton therapy.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Malvern Madondo, Yuan Shao, Yingzi Liu, Jun Zhou, Xiaofeng Yang, Zhen Tian. 2025-08-11. Patient-Specific Deep Reinforcement Learning for Automatic Replanning in Head-and-Neck Cancer Proton Therapy. https://arxiv.org/abs/2506.10073

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Fail-closed conformal multiresolution error control for conservative three-dimensional dose remapping in synthetic phantoms

Conservative dose remapping can be formulated by transporting mass and deposited energy across nonmatching grids, but fixed high-sampling calculations use the same work for easy and difficult cases. We developed conformal multiresolution error control (CoMERC), a fail-closed multilevel quasi-Monte Carlo (QMC) controller for reference-relative target-cell allocation error in a fixed piecewise-constant source model with known deformation. Two randomized nested replicates were evaluated at a probe level of 4 and candidate levels of 8 and 16. Four endpoints were controlled jointly: reference-mass-weighted global, high-gradient, and density-gradient-threshold root-mean-square error, plus eligible-cell maximum absolute error. A split-conformal multiplier was calibrated from 240 synthetic cases, frozen, and evaluated on 400 independent cases from the same generator and 120 cases from a prespecified shifted generator. In the primary set, joint coverage was 394/400 (98.50%; one-sided 95% lower limit, 97.06%), 388/400 cases (97.00%; lower limit, 95.18%) received an output, and 0/388 released outputs exceeded any tolerance (one-sided 95% upper limit, 0.769%). Final actions were level 8 for 190 cases, level 16 for 198, and ABSTAIN for 12. The frozen policy implied mean relative production-sample work of 0.5844 compared with always running both randomized replicates through level 16; the prespecified 95th-percentile bootstrap upper limit was 0.6194. All seven primary gates passed. In the shifted-generator set, 113/120 cases were released and one released output exceeded a tolerance. CoMERC provided marginal finite-sample error control and lower modeled production-sample work in the locked synthetic population, but not patient-level, registration-level, conditional-on-release, or clinical safety validation.

physics.med-ph

Dosimetric equivalence of deep learning prostate contours after LDR brachytherapy: pre-declared margins, patient-level acceptance thresholds and the incremental predictive value of DVH indices

Background and purpose: Dose Volume Histogram (DVH) indices remain the main dose-effect metrics for toxicity prediction, but they depend on how contours were made. Inter-observer variability (IOV) is unavoidable and clinically accepted, so the question for automatic segmentation addresses equivalence: do automatic contours produce DVH errors comparable to human IOV, and can these indices predict patient-reported toxicity? Materials and methods: In 429 patients treated with iodine-125 LDR-BT monotherapy, indices from expert manual delineation were compared with a deterministic and a Bayesian nnUNet on a fixed dose distribution, by two-one-sided tests against margins set from CT contouring IOV. Logistic regression gave patient-level thresholds at 90\,\% probability of equivalence. In 380 patients, eleven DHV indices were added to a clinical baseline predicting change in International Prostate Symptom Score (IPSS) at six horizons over 5 years, with nested cross-validation, bootstrap intervals and corrected $t$-tests. Results: All cohort-level comparisons of DVH indices were declared equivalent, with no interval consuming more than 44\,\% of its equivalence margin. Individual agreement was weaker with equivalence rates from 55.9\,\% to 89.7\,\%, depending on the DVH index and automatic segmentation model. Thresholds ranged from 0.864 to 0.958 Dice. No DVH block improved IPSS prediction at any horizon. The largest improvement declared by bootstrap intervals was 0.12 IPSS points, well below the minimal clinically important difference, across all learners. Conclusion: Automatic contours matched expert dosimetry within human IOV on average, but individual equivalence requires a volume dependent quality metric threshold definition. Our DVH indices panel provides no significant predictive power for IPSS, irrespective of segmentation source and horizon.

physics.med-ph

The Electrodynamic Basis of Dichroism-Mediated Polarization Perception

Humans see the polarization of light through entoptic percepts arising from the macula's Henle fiber layer, where xanthophyll pigments absorb preferentially across the radiating fibers. From the layer's complex dielectric tensor alone, Maxwell's equations in Berreman $4\times4$ form yield its Mueller matrix, set by four scalars $\{A,B,C,D\}$. Its intensity channel depends only on the dichroic pair $A$ and $B$, which fix the percepts at a maximum contrast $|B|/A\approx0.05$. One relation generates the whole dichroism-mediated family: Haidinger's brushes under a uniform field, their dark arms perpendicular to the $\mathbf{E}$-vector; fractured brushes under spatially varying fields; and $N$-fold brushes under vector-vortex illumination.

physics.med-ph