Implementation of K-Means Clustering for Grouping Post-Flood Disease Patterns in Affected Residential Settlements
DOI:
https://doi.org/10.47709/cnahpc.v8i3.9210Keywords:
Binjai City, Data Mining, K-Means Clustering, Post-Flood Diseases, Slum SettlementsAbstract
Flooding in Binjai City increases post-flood disease incidence in slum settlements due to inadequate sanitation and high environmental vulnerability, while disease records are often scattered, poorly organized, and late, slowing health interventions. This study groups post-flood disease data in affected residential areas using K-Means clustering based on disease type, village, and slum level. A total of 1,100 records were collected from five districts in Binjai City; after excluding incomplete records, 850 valid records were used in the 2-cluster scenario and 1,086 in the 3-cluster scenario, implemented in a Matlab-based application. In the 2-cluster scenario, Cluster 1 contained 536 records with centroid (3.64, 12.48, 2.41) and Cluster 2 contained 314 records with centroid (4.18, 16.25, 2.09). In the 3-cluster scenario, Cluster 1 contained 498 records with centroid (3.45, 11.62, 2.31), Cluster 2 contained 312 records with centroid (4.02, 16.47, 2.05), and Cluster 3 contained 276 records with centroid (2.91, 8.35, 2.76), all dominated by diarrhea in medium-to-high slum-level areas. Internal validity indices (Silhouette, Davies-Bouldin, Calinski-Harabasz), computed on a reference sample, support retaining the 3-cluster scheme for its finer, more actionable risk stratification. The results show that K-Means clustering groups affected areas by disease and slum characteristics, and a centroid-derived priority ranking of the clusters is proposed to support health priority setting, medical resource distribution, and data-driven post-flood disease mitigation.
Downloads
References
Adiputra, I. N. M. (2021). Clustering penyakit DBD pada Rumah Sakit Dharma Kerti menggunakan algoritma K-Means. INSERT: Information System and Emerging Technology Journal, 2(2), 99-105. https://ejournal.undiksha.ac.id/index.php/insert/article/view/41673
Agneresa, Hananto, A. L., Hilabi, S. S., Hananto, A., & Tukino. (2022). Strategi promosi penerapan data mining mahasiswa baru dengan metode K-Means clustering. Dirgamaya: Jurnal Manajemen dan Sistem Informasi, 2(2), 25-34.
Angin, S. P., & Sihombing, M. (2024). Clustering data on underage marriage. 3(November), 195-200.
Fahry, F., Miswaty, T. C., & Harun, H. (2025). Analyzing marketplace reviews using Word2Vec, CNN, and deep K-Means with sociolinguistic approaches. Jurnal Teknik Informatika, 6(6), 5489-5502. https://jutif.if.unsoed.ac.id/index.php/jurnal/article/view/5340
Joo, D., Na, R., Kim, H., Yoo, S., & Lee, S. (2025). Analysis of application of design standards for future climate change adaptive agricultural reservoirs using cluster analysis. 1-19.
Kiparisov, P., Lagutov, V., & Pflug, G. (2023). Quantification of loss of access to critical services during floods in Greater Jakarta: Integrating social, geospatial, and network perspectives. Remote Sensing, 15(21), 1-23.
Li, H., Huang, J., Zhang, X., Meng, Z., Fan, Y., & Wu, X. (2025). Flash flood risk classification using GIS-based fractional order k-means clustering method. Fractal and Fractional, 9(9), 1-18.
Li, J., Meng, Z., Zhang, J., Chen, Y., Yao, J., & Li, X. (2025). Prediction of seawater intrusion run-up distance based on K-Means clustering and ANN model. Journal of Marine Science and Engineering, 13(2), 1-18.
Liu, C., Feng, Q., Zhou, W., Zhang, C., & Zhang, X. (2025). Flow field evaluation method of high water-cut reservoirs based on K-Means clustering algorithm. Symmetry, 17(6), 1-18.
Manurung, H. (2024). Pengelompokan UMKM Kota Binjai menggunakan metode clustering K-Means untuk mengidentifikasi pola. 8(2), 93-99.
Mpakosi, A., Cholevas, V., Tzouvelekis, I., Passos, I., Kaliouli-Antonopoulou, C., & Mironidou-Tzouveleki, M. (2024). Autoimmune diseases following environmental disasters: A narrative review of the literature. Healthcare, 12(17), 1767. https://www.mdpi.com/2227-9032/12/17/1767/htm
Ordila, R., Wahyuni, R., Irawan, Y., & Sari, M. Y. (2020). Penerapan data mining untuk pengelompokan data rekam medis pasien berdasarkan jenis penyakit dengan algoritma clustering. Jurnal Ilmu Komputer, 9(2), 148-153. https://jik.htp.ac.id/index.php/jik/article/view/181
Prabowo, & Gibran. (2024). Asta Cita: Visi misi Prabowo-Gibran.
Pratama, R. Z., Sihombing, M., & Ambarita, I. (2024). Application of data mining to measure the level of satisfaction with public facilities and services at STMIK Kaputama Binjai using linear regression method. Journal of Artificial Intelligence and Engineering Applications, 4(1), 411-418.
Sembiring, C. S. D. B., Hanum, L., & Tamba, S. P. (2022). Penerapan data mining menggunakan algoritma K-Means untuk menentukan judul skripsi dan jurnal penelitian (studi kasus FTIK UNPRI). Jurnal Sistem Informasi dan Ilmu Komputer Prima (JUSIKOM PRIMA), 5(2), 80-85.
Song, H., Li, S., Yu, B., Sun, Y., Tao, Q., & Peng, L. (2025). Automatic text summary method based on optimized K-Means clustering algorithm with symmetry and maximal-marginal-relevance algorithm. 1-24.
Van Doorsselaere, T., Shariati, H., & Debosscher, J. (2019). Time series optimization on data mining. Journal of Physics: Conference Series, 1235(1), 012014. https://iopscience.iop.org/article/10.1088/1742-6596/1235/1/012014
Wala, J., Herman, H., & Umar, R. (2024). Implementasi K-Means clustering pada pengelompokan pasien penyakit jantung. JISKA (Jurnal Informatika Sunan Kalijaga), 9(3), 205-216. https://ejournal.uin-suka.ac.id/saintek/JISKA/article/view/4458
Zema, Maulita, Y., & Arliana, L. (2022). Penerapan data mining pengelompokan peserta BPJS Ketenagakerjaan berdasarkan program yang diambil menggunakan metode clustering. Jurnal Sistem Informasi Kaputama, 6(2), 152-164.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Hotler Manurung, Marto Sihombing, Ratih Puspadini

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.











