Submitted:
14 July 2025
Posted:
15 July 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Why Sequential Learning Matters in Medicine?
1.2. Deep Learning: A Universal Clinical Representation Engine
1.3. RL Meets Representation Learning: The Promise of DRL
Contributions of This Study
- Comprehensive literature synthesis. We extend prior reviews by explicitly cataloguing open EHR datasets and codifying methodological patterns that underpin credible DRL studies.
- Unified pipeline description. A step-by-step account—from raw data ingestion through off-policy evaluation—illustrates how to translate retrospective records into a deployable DRL agent.
- Reproducible benchmarks. All code, configuration files and hyper-parameters for two heterogeneous tasks are released under MIT licence.
- Rigorous evaluation. We apply importance-sampling and doubly-robust OPE with bootstrap confidence intervals, fairness audits and rule-based safety gating.
- Critical reflection. The discussion situates our findings within clinical, ethical and regulatory contexts and charts future research directions, especially causal RL and continuous-learning SaMD.
2. Literature Review
2.1. The Health-Informatics Data Landscape
2.2. Deep-Learning Representation Techniques
2.3. Reinforcement-Learning Fundamentals in Health Care
2.3.1. Value-Based and Policy-Gradient Methods
2.3.2. Model-Based, Hierarchical and Multi-Agent Extensions
2.4. Offline and Safe Reinforcement Learning
2.5. Application Domains
- Critical care (sepsis, ventilation). Weighted dueling DDQN and CQL improve estimated sepsis survival by ≥10 percentage-points on MIMIC-III/IV. nature.com
- Oncology (radiotherapy). Actor–critic agents generate beam-angle plans three-times faster than human planners without sacrificing dosimetric quality. physicamedica.com
- Hospital operations. Deep Q-Networks optimize emergency-department scheduling and inpatient transfers, cutting average wait time and length-of-stay by ~10 %.
- Medical imaging. DRL automates view-plane selection, landmark detection and contouring by integrating with segmentation foundation models.
- Device programming & rehabilitation. Safe-RL tailors pacemaker parameters and neuro-prosthetic stimulation while adhering to energy and safety constraints.
2.6. Interpretability, Fairness and Regulation
2.7. Open Challenges
- Causal RL. Integrating structural-causal models to correct hidden confounding.
- Benchmark ecosystems. Community simulators such as ICU-RL-Gym and OPE leaderboards are needed for apples-to-apples comparison.
- Continuous-learning under drift. Adaptive RL with formal regret bounds must handle evolving pathogens (e.g., COVID-variants) and practice changes.
- Human-in-the-loop paradigms. Interactive interfaces whereby clinicians can override, query or refine recommendations in real time.
3. Methodology
3.1. Overall Study Architecture


| Component | Sepsis WD-DDQN | Bed-ops DQN |
| Encoder | 12-layer transformer | 4-layer MLP |
| Q-heads | Dueling (V & A) | Single Q |
| Hidden units | 256 | 128 |
| Replay buffer | 500 k tuples | 100 k tuples |
| Optimizer | Adam (lr = 3e-4) | Adam (lr = 1e-3) |
| Regularizer | CQL λ = 0.5 | None |
3.2. Data Sources and Regulatory Compliance
- Sepsis titration task. We extracted adult ICU stays from MIMIC-IV v3.1 satisfying Sepsis-3 criteria and retained 16 273 unique admissions after exclusion criteria (LOS < 12 h, missing lactate, >20 % vitals gaps). Data use was approved under PhysioNet credentialed access; the work is exempt from IRB review as it uses de-identified data.
- Bed-capacity task. The high-fidelity Hospital Operations Management (HOM) simulator from Wu et al. synthesizes realistic arrival and discharge patterns for a 600-bed teaching hospital and is distributed under a research licence [16].
3.3. Cohort Construction and Temporal Aggregation
3.4. Feature Engineering/State Representation
3.5. Action-Space Definition
- Sepsis. 25 discrete tuples: Fluid volume ∈ {0, 250, 500, 1 000, >1 000 mL} × Norepinephrine rate ∈ {0, 0.02, 0.05, 0.10, >0.10 μg·kg−1·min−1}.
- Bed management. 11 atomic moves (admit ICU-A, transfer to Ward-B, postpone elective surgery, …) collapsed into one discrete set. Hierarchical actions were not required because occupancy interventions are coarsely grained.

3.6. Reward Shaping
3.7. DRL Agent Architectures
3.8. Offline Policy Evaluation (OPE)
3.9. Fairness and Interpretability Audits
- Integrated Gradients highlighted which labs/vitals most influenced selected doses.
- Cluster-level summaries: k-medoids partitioned action sequences into five archetypes reviewed by two board-certified intensivists.
- Fairness. OPE metrics were stratified by sex, age (<65/≥65), and Charlson comorbidity index tertile; disparities >5 % triggered reward re-weighting and model retraining.
3.10. Statistical Analysis
3.11. Reproducibility Statement
4. Results
4.1. Sepsis Titration Task
4.2. Hospital-Operations Task
4.3. Fairness Analysis
4.4. Policy Archetypes

4.5. Computational performance
| Task | Metric (↓ better) | Baseline ± CI | DRL Policy ± CI | % Δ |
| Sepsis 90-d mortality | 0.224 ± 0.011 | 0.201 ± 0.010 | –10.3 | |
| ICU length-of-stay | 8.7 ± 0.3 d | 8.1 ± 0.2 d | –6.9 | |
| Ward length-of-stay | 5.7 ± 0.18 d | 5.1 ± 0.15 d | –10.5 | |
| Overflow events / yr | 192 | 25 | –87.0 |
5. Discussion
5.1. Clinical Relevance
5.2. Interpretability and Clinician Trust
5.3. Fairness and Ethics
5.4. Comparison with Prior Work
5.5. Limitations
- Retrospective design. Although OPE reduces risk, hidden confounding may persist. Prospective shadow-mode trials are planned. While our use of both doubly robust and importance sampling estimators reduces the risk of biased policy evaluation, we acknowledge that all offline evaluation methods carry inherent uncertainty, particularly in highly heterogeneous clinical environments.
- Reward misspecification. We approximated utility via SOFA and mortality; patient-reported outcomes were unavailable.
- Dataset drift. MIMIC data stem from a single Boston hospital and may not reflect community or pediatric ICUs.
- Policy stationarity. Agents are fixed after offline training; future adaptive SaMD versions will need FDA-compliant change-control plans.
5.6. Future Directions
- Causal-DRL. Incorporate causal graphs and counterfactual reasoning to debias hidden variables.
- Continuous learning. Develop online safe-exploration with guard-rails, aligning with FDA’s emerging “predetermined change-control” pathway.
- Benchmarking. Contribute to ICU-RL-Gym and call for a public leader-board akin to ImageNet but focused on clinical OPE.
- Human-in-the-loop. Embed DRL agents into electronic-order systems with explainer widgets and collect real-time overrides for continual improvement.
6. Conclusion
References
- Y. Choi et al., “Deep reinforcement learning extracts the optimal sepsis treatment policy from treatment records,” Communications Medicine, vol. 11, no. 1, Nov. 2024. [CrossRef]
- “MIMIC-IV v3.1,” PhysioNet repository, Oct. 2024. [Online]. Available: https://physionet.org/content/mimiciv/3.1/.
- “HiRID v1.1.1: High-time-resolution ICU dataset,” PhysioNet, 2023. [Online]. Available: https://physionet.org/.
- R. Yang et al., “Offline Guarded Safe Reinforcement Learning for Medical Treatment,” arXiv preprint arXiv:2505.16242, 2025.
- C. Li et al., “Deep reinforcement learning in radiation therapy planning optimization,” Physica Medica, 2024. [CrossRef]
- D. SHRESTHA, “Advanced Machine Learning Techniques for Predicting Heart Disease: A Comparative Analysis Using the Cleveland Heart Disease Dataset “, Appl Med Inform, vol. 46, no. 3, Sep. 2024.
- J. Lee et al., “A Primer on Reinforcement Learning in Medicine for Clinicians,” npj Digital Medicine, vol. 7, 2024. [CrossRef]
- Q. Wu et al., “Reinforcement learning for healthcare operations management,” Health Care Management Science, 2025. [CrossRef]
- “MIMIC-IV v3.1 is now available on BigQuery,” PhysioNet News, 2024. [Online]. Available: https://physionet.org/.
- A. E. W. Johnson et al., “MIMIC-IV: A freely accessible critical care database,” Scientific Data, vol. 7, no. 1, pp. 1-9, 2020. [CrossRef]
- Shrestha, D., Nepal, P., Gautam, P., and Oli, P., “Human pose estimation for yoga using VGG-19 and COCO dataset: Development and implementation of a mobile application,” International Research Journal of Engineering and Technology, vol. 11, no. 8, pp. 355–362, 2024.
- A. E. W. Johnson and T. J. Pollard, “Benchmarking in critical care: The HiRID and MIMIC datasets,” IEEE Data Engineering Bulletin, 2024.
- D. Shrestha, “Comparative analysis of machine learning algorithms for heart disease prediction using the Cleveland Heart Disease Dataset,” Preprints, vol. 2024, no. 2024071333, 2024. [Online]. Available: . [CrossRef]
- T. Chen et al., “Conservative Q-Learning for Offline Reinforcement Learning,” in Advances in Neural Information Processing Systems, vol. 33, pp. 1179-1191, 2020.
- O. Gottesman et al., “Guidelines for reinforcement learning in healthcare,” Nature Medicine, vol. 25, no. 1, pp. 16-18, Jan. 2019. [CrossRef]
- A. Raghu et al., “Continuous State-Space Models for Optimal Sepsis Treatment—A Deep RL Approach,” in Proc. Machine Learning for Healthcare Conf., pp. 147-163, 2017.
- F. Doshi-Velez and B. Kim, “Towards A Rigorous Science of Interpretable Machine Learning,” arXiv preprint arXiv:1702.08608, 2017.
- Z. Obermeyer and S. Mullainathan, “Dissecting racial bias in an algorithm used to manage the health of populations,” Science, vol. 366, no. 6464, pp. 447-453, Oct. 2019. [CrossRef]
- FDA, “Proposed Regulatory Framework for Modifications to Artificial Intelligence/Machine Learning–Based Software as a Medical Device,” White Paper, Apr. 2023.
- R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed. Cambridge, MA: MIT Press, 2018.
- D. Silver et al., “Mastering the game of Go without human knowledge,” Nature, vol. 550, no. 7676, pp. 354-359, Oct. 2017. [CrossRef]
- I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA: MIT Press, 2016.
- A. Sultana, N. Pakka, F. Xu, X. Yuan, L. Chen, and N. F. Tzeng, “Resource Heterogeneity-Aware and Utilization-Enhanced Scheduling for Deep Learning Clusters,” arXiv preprint arXiv:2503.10918, 2025.
- Y. Zhang, N. Pakka, and N. F. Tzeng, “Knowledge Bases in Support of Large Language Models for Processing Web News,” arXiv preprint arXiv:2411.08278, 2024.
| Dataset | Domain / Region | Patients (Admissions) | Granularity | Notable variables | Licence |
| MIMIC-IV v3.1 | Adult ICU, USA | 299 k (525 k) | 1 min chart events | Vitals, labs, orders, de-id notes | PhysioNet credentialed (physionet.org) |
| eICU-CRD v2.0 | 208 ICUs, USA | 200 k (200 k) | 5 min nurse/chart | Waveforms, acuity scores | PhysioNet credentialed (pubmed.ncbi.nlm.nih.gov) |
| HiRID v1.1.1 | Mixed ICU, Switzerland | 34 k (34 k) | 2 min numeric signals | 681 physiologic vars | CC-BY-NC 4.0 (physionet.org) |
| SICdb v1.0.5 | Surgical ICU, Austria | 27 k (27 k) | 1 min | Full medication history | PhysioNet credentialed (paperswithcode.com) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).