Sparse computation for large-scale binary classification
Abstract
Well-known data mining algorithms rely on inputs in the form of pairwise similarities between objects. For large datasets it is computationally impossible to perform all pairwise comparisons. We therefore propose a novel approach that uses approximate Principal Component Analysis to efficiently identify groups of similar objects. The effectiveness of the approach is demonstrated in the context of binary classification using the supervised normalized cut as a classifier. For large datasets from the UCI repository, the approach significantly improves run times with minimal loss in accuracy.
Date Issued
2014
Publication Type
Conference Item
Subject(s)
Language(s)
en
Additional Credits
Access(Rights)
metadata.only