Implementation of K-Means Clustering for Grouping Post-Flood Disease Patterns in Affected Residential Settlements

Authors

  • Hotler Manurung Sekolah Tinggi Manajemen Informatika dan Komputer (STMIK) Kaputama, Binjai, Indonesia
  • Marto Sihombing Sekolah Tinggi Manajemen Informatika dan Komputer (STMIK) Kaputama, Binjai, Indonesia
  • Ratih Puspadini Sekolah Tinggi Manajemen Informatika dan Komputer (STMIK) Kaputama, Binjai, Indonesia

DOI:

https://doi.org/10.47709/cnahpc.v8i3.9210

Keywords:

Binjai City, Data Mining, K-Means Clustering, Post-Flood Diseases, Slum Settlements

Abstract

Flooding in Binjai City increases post-flood disease incidence in slum settlements due to inadequate sanitation and high environmental vulnerability, while disease records are often scattered, poorly organized, and late, slowing health interventions. This study groups post-flood disease data in affected residential areas using K-Means clustering based on disease type, village, and slum level. A total of 1,100 records were collected from five districts in Binjai City; after excluding incomplete records, 850 valid records were used in the 2-cluster scenario and 1,086 in the 3-cluster scenario, implemented in a Matlab-based application. In the 2-cluster scenario, Cluster 1 contained 536 records with centroid (3.64, 12.48, 2.41) and Cluster 2 contained 314 records with centroid (4.18, 16.25, 2.09). In the 3-cluster scenario, Cluster 1 contained 498 records with centroid (3.45, 11.62, 2.31), Cluster 2 contained 312 records with centroid (4.02, 16.47, 2.05), and Cluster 3 contained 276 records with centroid (2.91, 8.35, 2.76), all dominated by diarrhea in medium-to-high slum-level areas. Internal validity indices (Silhouette, Davies-Bouldin, Calinski-Harabasz), computed on a reference sample, support retaining the 3-cluster scheme for its finer, more actionable risk stratification. The results show that K-Means clustering groups affected areas by disease and slum characteristics, and a centroid-derived priority ranking of the clusters is proposed to support health priority setting, medical resource distribution, and data-driven post-flood disease mitigation.

Downloads

Download data is not yet available.

References

Adiputra, I. N. M. (2021). Clustering penyakit DBD pada Rumah Sakit Dharma Kerti menggunakan algoritma K-Means. INSERT: Information System and Emerging Technology Journal, 2(2), 99-105. https://ejournal.undiksha.ac.id/index.php/insert/article/view/41673

Agneresa, Hananto, A. L., Hilabi, S. S., Hananto, A., & Tukino. (2022). Strategi promosi penerapan data mining mahasiswa baru dengan metode K-Means clustering. Dirgamaya: Jurnal Manajemen dan Sistem Informasi, 2(2), 25-34.

Angin, S. P., & Sihombing, M. (2024). Clustering data on underage marriage. 3(November), 195-200.

Fahry, F., Miswaty, T. C., & Harun, H. (2025). Analyzing marketplace reviews using Word2Vec, CNN, and deep K-Means with sociolinguistic approaches. Jurnal Teknik Informatika, 6(6), 5489-5502. https://jutif.if.unsoed.ac.id/index.php/jurnal/article/view/5340

Joo, D., Na, R., Kim, H., Yoo, S., & Lee, S. (2025). Analysis of application of design standards for future climate change adaptive agricultural reservoirs using cluster analysis. 1-19.

Kiparisov, P., Lagutov, V., & Pflug, G. (2023). Quantification of loss of access to critical services during floods in Greater Jakarta: Integrating social, geospatial, and network perspectives. Remote Sensing, 15(21), 1-23.

Li, H., Huang, J., Zhang, X., Meng, Z., Fan, Y., & Wu, X. (2025). Flash flood risk classification using GIS-based fractional order k-means clustering method. Fractal and Fractional, 9(9), 1-18.

Li, J., Meng, Z., Zhang, J., Chen, Y., Yao, J., & Li, X. (2025). Prediction of seawater intrusion run-up distance based on K-Means clustering and ANN model. Journal of Marine Science and Engineering, 13(2), 1-18.

Liu, C., Feng, Q., Zhou, W., Zhang, C., & Zhang, X. (2025). Flow field evaluation method of high water-cut reservoirs based on K-Means clustering algorithm. Symmetry, 17(6), 1-18.

Manurung, H. (2024). Pengelompokan UMKM Kota Binjai menggunakan metode clustering K-Means untuk mengidentifikasi pola. 8(2), 93-99.

Mpakosi, A., Cholevas, V., Tzouvelekis, I., Passos, I., Kaliouli-Antonopoulou, C., & Mironidou-Tzouveleki, M. (2024). Autoimmune diseases following environmental disasters: A narrative review of the literature. Healthcare, 12(17), 1767. https://www.mdpi.com/2227-9032/12/17/1767/htm

Ordila, R., Wahyuni, R., Irawan, Y., & Sari, M. Y. (2020). Penerapan data mining untuk pengelompokan data rekam medis pasien berdasarkan jenis penyakit dengan algoritma clustering. Jurnal Ilmu Komputer, 9(2), 148-153. https://jik.htp.ac.id/index.php/jik/article/view/181

Prabowo, & Gibran. (2024). Asta Cita: Visi misi Prabowo-Gibran.

Pratama, R. Z., Sihombing, M., & Ambarita, I. (2024). Application of data mining to measure the level of satisfaction with public facilities and services at STMIK Kaputama Binjai using linear regression method. Journal of Artificial Intelligence and Engineering Applications, 4(1), 411-418.

Sembiring, C. S. D. B., Hanum, L., & Tamba, S. P. (2022). Penerapan data mining menggunakan algoritma K-Means untuk menentukan judul skripsi dan jurnal penelitian (studi kasus FTIK UNPRI). Jurnal Sistem Informasi dan Ilmu Komputer Prima (JUSIKOM PRIMA), 5(2), 80-85.

Song, H., Li, S., Yu, B., Sun, Y., Tao, Q., & Peng, L. (2025). Automatic text summary method based on optimized K-Means clustering algorithm with symmetry and maximal-marginal-relevance algorithm. 1-24.

Van Doorsselaere, T., Shariati, H., & Debosscher, J. (2019). Time series optimization on data mining. Journal of Physics: Conference Series, 1235(1), 012014. https://iopscience.iop.org/article/10.1088/1742-6596/1235/1/012014

Wala, J., Herman, H., & Umar, R. (2024). Implementasi K-Means clustering pada pengelompokan pasien penyakit jantung. JISKA (Jurnal Informatika Sunan Kalijaga), 9(3), 205-216. https://ejournal.uin-suka.ac.id/saintek/JISKA/article/view/4458

Zema, Maulita, Y., & Arliana, L. (2022). Penerapan data mining pengelompokan peserta BPJS Ketenagakerjaan berdasarkan program yang diambil menggunakan metode clustering. Jurnal Sistem Informasi Kaputama, 6(2), 152-164.

Downloads

Published

2026-07-22

How to Cite

Manurung, H., Sihombing, M., & Puspadini, R. (2026). Implementation of K-Means Clustering for Grouping Post-Flood Disease Patterns in Affected Residential Settlements. Journal of Computer Networks, Architecture and High Performance Computing, 8(3), 430–440. https://doi.org/10.47709/cnahpc.v8i3.9210