ImpKmeans: An Improved Version of the K- Means Algorithm, by Determining Optimum Initial Centroids, based on Multivariate Kernel Density Estimation and Kd-Tree
Şenol, Ali
2025-09-16T10:27:20Z
2025-09-16T10:27:20Z
2024
1785-8860
hu_HU
http://hdl.handle.net/20.500.14044/33720
K-means is the best known clustering algorithm, because of its usage simplicity,
fast speed and efficiency. However, resultant clusters are influenced by the randomly selected
initial centroids. Therefore, many techniques have been implemented to solve the mentioned
issue. In this paper, a new version of the k-means clustering algorithm named as ImpKmeans
shortly (An Improved Version of K-Means Algorithm by Determining Optimum Initial
Centroids Based on Multivariate Kernel Density Estimation and Kd-tree) that uses kernel
density estimation, to find the optimum initial centroids, is proposed. Kernel density
estimation is used, because it is a nonparametric distribution estimation method, that can
identify density regions. To understand the efficiency of the ImpKmeans, we compared it with
some state-of-the-art algorithms. According to the experimental studies, the proposed
algorithm was better than the compared versions of k-means. While ImpKmeans was the most
successful algorithm in 46 tests of 60, the second-best algorithm, was the best on 34 tests.
Moreover, experimental results indicated that the ImpKmeans is fast, compared to the
selected k-means versions
hu_HU
dc.format
PDF
hu_HU
en
hu_HU
ImpKmeans: An Improved Version of the K- Means Algorithm, by Determining Optimum Initial Centroids, based on Multivariate Kernel Density Estimation and Kd-Tree
hu_HU
Open access
hu_HU
Óbudai Egyetem
hu_HU
Budapest
hu_HU
Óbudai Egyetem
hu_HU
Műszaki tudományok - multidiszciplináris műszaki tudományok