On the Consistency of k-means++ algorithm.

Saved in:
Bibliographic Details
Title: On the Consistency of k-means++ algorithm.
Authors: Kłopotek, Mieczysław A.1 (AUTHOR) klopotek@ipipan.waw.pl
Source: Fundamenta Informaticae. 2020, Vol. 172 Issue 4, p361-377. 17p.
Subjects: Big data, Expected returns, Databases, Algorithms
Abstract: We prove in this paper that the expected value of the objective function of the k-means++ algorithm for samples converges to population expected value. As k-means++, for samples, provides with constant factor approximation for k-means objectives, such an approximation can be achieved for the population with increase of the sample size. This result is of potential practical relevance when one is considering using subsampling when clustering large data sets (large data bases). [ABSTRACT FROM AUTHOR]
Copyright of Fundamenta Informaticae is the property of Polskie Towarzystwo Matematyczne and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:We prove in this paper that the expected value of the objective function of the k-means++ algorithm for samples converges to population expected value. As k-means++, for samples, provides with constant factor approximation for k-means objectives, such an approximation can be achieved for the population with increase of the sample size. This result is of potential practical relevance when one is considering using subsampling when clustering large data sets (large data bases). [ABSTRACT FROM AUTHOR]
ISSN:01692968
DOI:10.3233/FI-2020-1909