Fast and Robust General Purpose Clustering Algorithms
File version
Author(s)
Yang, J
Griffith University Author(s)
Primary Supervisor
Other Supervisors
Editor(s)
Heikki Mannila
Date
Size
905036 bytes
61987 bytes
File type(s)
application/pdf
text/plain
Location
License
Abstract
General purpose and highly applicable clustering methods are usually required during the early stages of knowledge discovery exercises. k-MEANS has been adopted as the prototype of iterative model-based clustering because of its speed, simplicity and capability to work within the format of very large databases. However, k-MEANS has several disadvantages derived from its statistical simplicity. We propose an algorithm that remains very efficient, generally applicable, multidimensional but is more robust to noise and outliers. We achieve this by using medians rather than means as estimators for the centers of clusters. Comparison with k-MEANS, EXPECTATION and MAXIMIZATION sampling demonstrates the advantages of our algorithm.
Journal Title
Data Mining and Knowledge Discovery
Conference Title
Book Title
Edition
Volume
8
Issue
2
Thesis Type
Degree Program
School
Publisher link
Patent number
Funder(s)
Grant identifier(s)
Rights Statement
Rights Statement
© Springer 2004. This is the author-manuscript version of this paper. The original publication is available at www.springerlink.com
Item Access Status
Note
Access the data
Related item(s)
Subject
Data management and data science
Information systems