Forecast-Driven Doctor Scheduling in Primary Healthcare: Ensemble Learning, a Reinforcement Learning Agent, and an Exact Baseline

Authors

  • M. Agus Munandar Universitas Indo Global Mandiri
  • Terttiaavini Universitas Indo Global Mandiri, Indonesia

DOI:

https://doi.org/10.47709/cnahpc.v8i4.9575

Dimension:

Keywords:

Adaptive Schedulling, Deep Q-Network, Doctor Schedulling, Ensemble Learning, Healthcare Systems, Patient Visit Prediction

Abstract

Doctor scheduling in primary healthcare facilities is challenging because patient visit volume fluctuates due to seasonality, holidays, weather, and unplanned staff absences, while rule-based scheduling used in many healthcare systems remains static and unresponsive. This study builds and rigorously validates a forecast-driven scheduling pipeline for a primary healthcare center (Puskesmas Talang Ratu, Palembang, Indonesia), using five years (2019–2024) of visit, attendance, weather, and holiday records engineered into 46 temporal, operational, and contextual features. Six ensemble regressors were compared using a chronological split; Random Forest generalised best (MAE = 7.08, RMSE = 10.44, R² = 0.82, next-day visits). The forecast was embedded in a Deep Q-Network (DQN) scheduling agent and benchmarked, under a leak-free walk-forward forecasting procedure with the forecast-to-decision day alignment corrected, against three baselines: the facility's existing rule-based practice, a constraint-aware forecast-to-staffing rule, and a myopic cost-optimal integer program that solves the DQN's own objective exactly each day. Over a 301-day held-out period, the exact optimizer achieved the best capacity match (MAE 8.11) and the lowest objective cost, ahead of the constraint-aware rule (MAE 8.42) and the DQN policy (MAE 9.90–10.26 across reward-scaling and seed variants), all of which greatly outperformed the existing rule-based practice (MAE 20.61). The DQN did not surpass the exact optimizer even on its own training objective, a result that held across four random seeds, extended training, and reward rescaling. We trace this to a structural property of the environment: because staffing decisions do not influence future states or rewards, the task reduces to a per-day (contextual-bandit-like) optimization rather than a genuinely sequential one, removing the advantage reinforcement learning derives from learning delayed returns. The contribution is therefore a leak-free, day-aligned forecasting-to-scheduling pipeline and an empirically grounded account of when reinforcement learning is, and is not, warranted for demand-responsive capacity allocation in primary healthcare.

Downloads

Download data is not yet available.

References

Fatikasari, D., Dwi Pratama, Y. S., & Hozairi, H. (2024). Optimasi Penjadwalan Tenaga Kesehatan di Puskesmas Teja Kabupaten Pamekasan Menggunakan Solver Excel. TeknoIS?: Jurnal Ilmiah Teknologi Informasi Dan Sains, 14(2), 214–224. https://doi.org/10.36350/jbs.v14i2.257

Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5). https://doi.org/10.1214/aos/1013203451

Fu, Y., Wang, Y., Gao, K., & Huang, M. (2024). Review on ensemble meta-heuristics and reinforcement learning for manufacturing scheduling problems. Computers and Electrical Engineering, 120, 109780. https://doi.org/10.1016/j.compeleceng.2024.109780

Hendiawati, Y., Abdussalaam, F., & Gunawan, E. (2023). Tata Kelola Sistem Informasi Jadwal Praktik Dokter Berbasis Web dengan Menggunakan Framework Laravel. Jurnal Teknologi Sistem Informasi Dan Aplikasi, 6(4), 590–600. https://doi.org/10.32493/jtsi.v6i4.33690

Karami, F. H. A. (2018). Optimasi Penjadwalan Staf dengan Menggunakan Algortima Reinforcement Learning Hyper-Heuristics Studi Kasus Rumah Sakit Ibu dan Anak Kendangsari. Institut Teknologi Sepuluh Nopember.

Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T. Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 2017-Decem(Nips), 3147–3155.

Lee, S., & Lee, Y. H. (2020). Improving Emergency Department Efficiency by Patient Scheduling Using Deep Reinforcement Learning. Healthcare, 8(2), 77. https://doi.org/10.3390/healthcare8020077

Li, Y., Wang, H., Wang, N., & Zhang, T. (2022). Optimal scheduling in cloud healthcare system using Q-learning algorithm. Complex and Intelligent Systems, 8(6), 4603–4618. https://doi.org/10.1007/s40747-022-00776-9

Liu, K., Li, X., Zou, C. C., Huang, H., & Fu, Y. (2020). Ambulance Dispatch via Deep Reinforcement Learning. Proceedings of the 28th International Conference on Advances in Geographic Information Systems, 123–126. https://doi.org/10.1145/3397536.3422204

Liu, R., Piplani, R., & Toro, C. (2022). Deep reinforcement learning for dynamic scheduling of a flexible job shop. International Journal of Production Research, 60(13), 4049–4069. https://doi.org/10.1080/00207543.2022.2058432

Liu, X., Zheng, C., Chen, Z., Liao, Y., Chen, R., & Yang, S. (2025). Reinforcement Learning for Patient Scheduling with Combinatorial Optimisation (pp. 238–243). https://doi.org/10.1007/978-3-031-77918-3_18

Mahajan, P., Uddin, S., Hajati, F., & Moni, M. A. (2023). Ensemble Learning for Disease Prediction: A Review. Healthcare, 11(12), 1808. https://doi.org/10.3390/healthcare11121808

Mahariani, Y. R. (2023). Penjadwalan Ruang Operasi Rumah Sakit dengan Metode Non-Dominated Sorting Genetic Algorithm II (NSGA-II). JIPI (Jurnal Ilmiah Penelitian Dan Pembelajaran Informatika), 8(1), 338–344. https://doi.org/10.29100/jipi.v8i1.3989

Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533. https://doi.org/10.1038/nature14236

Pratiwi, E. W., & Siambaton, M. Z. (2022). Aplikasi Penjadwalan Dokter Pada Rumah Sakit Umum Kota Pinang dengan Menggunakan Algoritma Greedy. Hello World Jurnal Ilmu Komputer, 1(1), 1–9. https://doi.org/10.56211/helloworld.v1i1.4

Priatna, W., Warta, J., & Sulistiyo, D. (2023). Implementasi Algoritma Genetika untuk Aplikasi Penjadwalan Sistem Kerja Shift. Techno.Com, 22(1), 235–246. https://doi.org/10.33633/tc.v22i1.7049

Stevens, C. A., Lyons, A. R., Dharmayat, K. I., Mahani, A., Ray, K. K., Vallejo-Vaz, A. J., & Sharabiani, M. T. (2023). Ensemble machine learning methods in screening electronic health records: A scoping review. DIGITAL HEALTH, 9. https://doi.org/10.1177/20552076231173225

Wang, L., & Demeulemeester, E. (2023). Simulation optimization in healthcare resource planning: A literature review. IISE Transactions, 55(10), 985–1007. https://doi.org/10.1080/24725854.2022.2147606

Yu, C., Liu, J., Nemati, S., & Yin, G. (2023). Reinforcement Learning in Healthcare: A Survey. ACM Computing Surveys, 55(1), 1–36. https://doi.org/10.1145/3477600

Zuo, J., Jin, Y., & Liu, W. (2024). Outpatient scheduling problem in smart hospital with two-agent deep reinforcement learning algorithm. Discover Computing, 27(1), 41. https://doi.org/10.1007/s10791-024-09474-1

Downloads

Published

2026-10-05

How to Cite

Munandar, M. A., & Terttiaavini, T. (2026). Forecast-Driven Doctor Scheduling in Primary Healthcare: Ensemble Learning, a Reinforcement Learning Agent, and an Exact Baseline. Journal of Computer Networks, Architecture and High Performance Computing, 8(4), 527–539. https://doi.org/10.47709/cnahpc.v8i4.9575

Citation Tracker

Citation data is temporarily unavailable.

Source: OpenAlex