Comparison of K-Means and Hierarchical Clustering Using Silhouette Score on BPBD Data, Papua Barat

Authors

  • Najlah Putri Qudsiya Universitas Papua
  • Christian Dwi Suhendra Universitas Papua
  • Josua Josen A. Limbong Universitas Papua

DOI:

https://doi.org/10.32664/j-intech.v14i02.2315

Keywords:

Data Mining, Disaster Clustering, Hierarchical Clustering, K-Means Clustering, Silhouette Score

Abstract

This study compares K-Means Clustering and Hierarchical Clustering to group disaster-prone areas in Papua Barat Province using official data from UPTD Pusdalops BPBD for the period 2020 to November 11, 2025. The dataset covers seven regencies with three aggregated variables: total floods and tidal floods, other disasters, and earthquakes. The research procedure includes data preprocessing, z-score normalization, determining the optimal number of clusters using the Elbow Method and Silhouette Score, and visualizing results with scatter plots and dendrograms. Both algorithms produced two clusters, with Hierarchical Clustering achieving a higher Silhouette Score of 0.4274 compared to 0.339 for K-Means. The scores indicate moderate clustering quality, reflecting the exploratory nature of the study due to the small dataset. Manokwari Regency formed a distinct cluster due to differing disaster characteristics. The novelty lies in employing official BPBD data and Silhouette Score evaluation to compare the two clustering algorithms, supporting data-driven prioritization for disaster mitigation.

References

[1] S. Z. S. Abdalla, “Data-driven innovations in disaster risk management: Advancing resilience and sustainability through big data analytics,” Progress in Disaster Science, vol. 27, p. 100451, Oct. 2025, doi: 10.1016/J.PDISAS.2025.100451.

[2] A. Kondraganti, G. Narayanamurthy, and H. Sharifi, “A systematic literature review on the use of big data analytics in humanitarian and disaster operations,” Annals of Operations Research 2023 335:3, vol. 335, no. 3, pp. 1015–1052, Nov. 2022, doi: 10.1007/S10479-022-04904-Z.

[3] BPK Perwakilan Provinsi Papua Barat, “Provinsi Papua Barat.” Accessed: May 15, 2026. [Online]. Available: https://papuabarat.bpk.go.id/wilayah-pemeriksaan/provinsi-papua-barat/

[4] A. Jaeger and D. Banks, “Cluster analysis: A modern statistical review,” Wiley Interdisciplinary Reviews: Computational Statistics, vol. 15, no. 3, p. e1597, May 2023, doi: 10.1002/WICS.1597.

[5] T. T. Tin et al., “Natural Disaster Clustering Using K-Means, DBSCAN, SOM, GMM, and Mean Shift: An Analysis of Fema Disaster Statistics,” International Journal of Advanced Computer Science and Applications, vol. 15, no. 9, pp. 667–681, Sep. 2024, doi: 10.14569/IJACSA.2024.0150968.

[6] A. M. Ikotun, A. E. Ezugwu, L. Abualigah, B. Abuhaija, and J. Heming, “K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data,” Information Sciences, vol. 622, pp. 178–210, Apr. 2023, doi: 10.1016/J.INS.2022.11.139.

[7] T. Li, A. Rezaeipanah, and E. S. M. Tag El Din, “An ensemble agglomerative hierarchical clustering algorithm based on clusters clustering technique and the novel similarity measurement,” Journal of King Saud University - Computer and Information Sciences, vol. 34, no. 6, pp. 3828–3842, Jun. 2022, doi: 10.1016/J.JKSUCI.2022.04.010.

[8] S. Brina and W. Rizky, “PERBANDINGAN ALGORITMA K-MEANS DAN HIERARCHICAL CLUSTERING DALAM SEGMENTASI KABUPATEN/KOTA DI JAWA TIMUR BERDASARKAN DATA PERJALANAN DAN PERGERAKAN WISATAWAN,” JATI (Jurnal Mahasiswa Teknik Informatika), vol. 9, no. 5, pp. 8499–8506, Jul. 2025, doi: 10.36040/JATI.V9I5.15091.

[9] E. Choi and J. Song, “Clustering-based disaster resilience assessment of South Korea communities building portfolios using open GIS and census data,” International Journal of Disaster Risk Reduction, vol. 71, p. 102817, Mar. 2022, doi: 10.1016/J.IJDRR.2022.102817.

[10] G. Setapati, D. Hafidh Zulfikar, R. Fatah Palembang, P. Sistem Informasi, F. Sains dan Teknologi, and U. Raden Intan Lampung, “Algoritma K-Means untuk Mengelompokkan Daerah Rawan Bencana di Sumatera Selatan,” Journal of software enginering and computational intelligence, vol. 2, no. 2, p. 11, 2024.

[11] C. Wongoutong, “The impact of neglecting feature scaling in k-means clustering,” PLOS ONE, vol. 19, no. 12, p. e0310839, Dec. 2024, doi: 10.1371/JOURNAL.PONE.0310839.

[12] J. Ge and R. Tibshirani, “Optimal pruning of hierarchical clustering dendrograms,” Communications in Statistics - Theory and Methods, 2025, doi: 10.1080/03610926.2025.2543191;CSUBTYPE:STRING:AHEAD.

[13] M. Shutaywi, “Silhouette Analysis for Performance Evaluation in Machine Learning with Applications to Clustering,” pp. 1–17, 2021, doi: 10.3390/e23060759.

[14] J. Demšar and B. Zupan, “Hands-on training about data clustering with orange data mining toolbox,” PLOS Computational Biology, vol. 20, no. 12, p. e1012574, Dec. 2024, doi: 10.1371/JOURNAL.PCBI.1012574.

[15] Z. Dobesova, “Evaluation of Orange data mining software and examples for lecturing machine learning tasks in geoinformatics,” Computer Applications in Engineering Education, vol. 32, no. 4, p. e22735, Jul. 2024, doi: 10.1002/CAE.22735;ISSUE:ISSUE:DOI.

[16] J. C. N. Bittencourt, D. G. Costa, P. Portugal, and F. Vasques, “A data-driven clustering approach for assessing spatiotemporal vulnerability to urban emergencies,” Sustainable Cities and Society, vol. 108, p. 105477, Aug. 2024, doi: 10.1016/J.SCS.2024.105477.

[17] D. P. Dewi, I. H. Santi, and W. D. Puspitasari, “Perhitungan Penilaian Tingkat Kepuasan Pelanggan Dengan Menerapkan Algoritma K-Means,” J-INTECH, vol. 11, no. 2, pp. 257–265, Dec. 2023, doi: 10.32664/J-INTECH.V11I2.981.

[18] V. P. Ramadhan and G. M. Namung, “Klasterisasi Komentar Cyberbullying Masyarakat di Instagram berdasarkan K-Means Clustering,” J-INTECH, vol. 11, no. 1, pp. 32–39, Jul. 2023, doi: 10.32664/J-INTECH.V11I1.846.

[19] P. Rinekso Andriyanto, J. Santoso, and E. Setyati, “Identifikasi Viseme Untuk Fonem Bahasa Madura Berbasis Clustering Berdasarkan Facial Landmark Point,” J-INTECH, vol. 11, no. 1, pp. 73–82, Jul. 2023, doi: 10.32664/J-INTECH.V11I1.835.

[20] E. Yuniar, S. Ramadhania, P. S. Universitasari, M. Hermansyah, and A. B. Setiawan, “Segmentation and Prediction of Store Performance on the Shopee Marketplace Using a Hybrid Clustering Approach, Spatial Analysis, and Feature Importance,” J-INTECH, vol. 14, no. 01, pp. 121–129, Mar. 2026, doi: 10.32664/J-INTECH.V14I01.2256.

[21] I. T. Utami, F. Suryaningrum, D. Ispriyanti, I. T. Utami, F. Suryaningrum, and D. Ispriyanti, “K-MEANS CLUSTER COUNT OPTIMIZATION WITH SILHOUETTE INDEX VALIDATION AND DAVIES BOULDIN INDEX (CASE STUDY: COVERAGE OF PREGNANT WOMEN, CHILDBIRTH, AND POSTPARTUM HEALTH SERVICES IN INDONESIA IN 2020),” BAREKENG: Jurnal Ilmu Matematika dan Terapan, vol. 17, no. 2, pp. 0707–0716, Jun. 2023, doi: 10.30598/BAREKENGVOL17ISS2PP0707-0716.

[22] H. Herayati, A. Siregar, and H. Hariyanto, “K-Means Clustering Algorithm Measuring the Satisfaction Level of MNC TV Muslim I’murojaah Program Viewers,” J-INTECH, vol. 13, no. 01, pp. 179–191, Jul. 2025, doi: 10.32664/J-INTECH.V13I01.1926.

[23] D. Wintana, H. Sulaiman, R. S. Rohman, G. Gunawan, and M. A. Ghani, “Analyzing Students’ Interest in Mathematics Through the Implementation of the K-Means Clustering Algorithm,” J-INTECH, vol. 13, no. 01, pp. 71–77, Jun. 2025, doi: 10.32664/J-INTECH.V13I01.1861.

[24] J. J. A. Limbong and F. M. Likumahwa, “Implementation of the K-Means Algorithm for Clustering Students ’ Web Programming Course Grades Using Silhouette Score,” vol. 6, no. 3, pp. 271–280, 2025, doi: https://doi.org/10.61628/jsce.v6i3.2100.

[25] S. Wijitkosum, “Integrated spatial analysis of drought risk factors using agglomerative hierarchical clustering and correlation,” Environmental Advances, vol. 21, p. 100646, Oct. 2025, doi: 10.1016/J.ENVADV.2025.100646.

[26] D. Chicco, A. Campagner, A. Spagnolo, D. Ciucci, and G. Jurman, “The Silhouette coefficient and the Davies-Bouldin index are more informative than Dunn index, Calinski-Harabasz index, Shannon entropy, and Gap statistic for unsupervised clustering internal evaluation of two convex clusters,” PeerJ Computer Science, vol. 11, p. e3309, Nov. 2025, doi: 10.7717/PEERJ-CS.3309/TABLE-18.

[27] M. Gagolewski, M. Bartoszuk, and A. Cena, “Are cluster validity measures (in) valid?,” Information Sciences, vol. 581, pp. 620–636, Dec. 2021, doi: 10.1016/J.INS.2021.10.004.

[28] A. M. Ikotun, F. Habyarimana, and A. E. Ezugwu, “Cluster validity indices for automatic clustering: A comprehensive review,” Heliyon, vol. 11, no. 2, p. e41953, Jan. 2025, doi: 10.1016/J.HELIYON.2025.E41953.

[29] Q. Xin, “Self-Supervised Customer Representation Learning for Segmentation and Next-Purchase Prediction on UCI Online Retail,” J-INTECH, vol. 14, no. 01, pp. 20–37, Apr. 2026, doi: 10.32664/J-INTECH.V14I01.2229.

Downloads

Published

2026-06-26