Online Feature Selection and Its Applications

Feature selection is an important technique for data mining before a machine learning algorithm is applied. Despite its importance, most studies of feature selection are restricted to batch learning. Unlike traditional batch learning methods, online learning represents a promising family of efficien...

Full description

Saved in:
Bibliographic Details
Main Authors: HOI, Steven, WANG, Jialei, ZHAO, Peilin, JIN, Rong
Format: text
Published: Institutional Knowledge at Singapore Management University 2012
Subjects:
Online Access:https://ink.library.smu.edu.sg/researchdata/13
http://ofs.stevenhoi.org/
Tags: Add Tag
No Tags, Be the first to tag this record!
Institution: Singapore Management University
id sg-smu-ink.researchdata-1012
record_format dspace
spelling sg-smu-ink.researchdata-10122015-11-26T09:58:39Z Online Feature Selection and Its Applications HOI, Steven WANG, Jialei ZHAO, Peilin JIN, Rong Feature selection is an important technique for data mining before a machine learning algorithm is applied. Despite its importance, most studies of feature selection are restricted to batch learning. Unlike traditional batch learning methods, online learning represents a promising family of efficient and scalable machine learning algorithms for large-scale applications. Most existing studies of online learning require accessing all the attributes/features of training instances. Such a classical setting is not always appropriate for real-world applications when data instances are of high dimensionality or it is expensive to acquire the full set of attributes/features. To address this limitation, we investigate the problem of Online Feature Selection (OFS) in which an online learner is only allowed to maintain a classifier involved only a small and fixed number of features. The key challenge of Online Feature Selection is how to make accurate prediction using a small and fixed number of active features. This is in contrast to the classical setup of online learning where all the features can be used for prediction. We attempt to tackle this challenge by studying sparsity regularization and truncation techniques. Specifically, this article addresses two different tasks of online feature selection: (1) learning with full input where an learner is allowed to access all the features to decide the subset of active features, and (2) learning with partial input where only a limited number of features is allowed to be accessed for each instance by the learner. We present novel algorithms to solve each of the two problems and give their performance analysis. We evaluate the performance of the proposed algorithms for online feature selection on several public datasets, and demonstrate their applications to real-world problems including image classification in computer vision and microarray gene expression analysis in bioinformatics. The encouraging results of our experiments validate the efficacy and efficiency of the proposed techniques. 2012-01-01T08:00:00Z text https://ink.library.smu.edu.sg/researchdata/13 http://ofs.stevenhoi.org/ SMU Research Data Institutional Knowledge at Singapore Management University Feature selection online learning large-scale data mining classification big data analytics Computer Sciences Databases and Information Systems Numerical Analysis and Scientific Computing
institution Singapore Management University
building SMU Libraries
country Singapore
collection InK@SMU
topic Feature selection
online learning
large-scale data mining
classification
big data analytics
Computer Sciences
Databases and Information Systems
Numerical Analysis and Scientific Computing
spellingShingle Feature selection
online learning
large-scale data mining
classification
big data analytics
Computer Sciences
Databases and Information Systems
Numerical Analysis and Scientific Computing
HOI, Steven
WANG, Jialei
ZHAO, Peilin
JIN, Rong
Online Feature Selection and Its Applications
description Feature selection is an important technique for data mining before a machine learning algorithm is applied. Despite its importance, most studies of feature selection are restricted to batch learning. Unlike traditional batch learning methods, online learning represents a promising family of efficient and scalable machine learning algorithms for large-scale applications. Most existing studies of online learning require accessing all the attributes/features of training instances. Such a classical setting is not always appropriate for real-world applications when data instances are of high dimensionality or it is expensive to acquire the full set of attributes/features. To address this limitation, we investigate the problem of Online Feature Selection (OFS) in which an online learner is only allowed to maintain a classifier involved only a small and fixed number of features. The key challenge of Online Feature Selection is how to make accurate prediction using a small and fixed number of active features. This is in contrast to the classical setup of online learning where all the features can be used for prediction. We attempt to tackle this challenge by studying sparsity regularization and truncation techniques. Specifically, this article addresses two different tasks of online feature selection: (1) learning with full input where an learner is allowed to access all the features to decide the subset of active features, and (2) learning with partial input where only a limited number of features is allowed to be accessed for each instance by the learner. We present novel algorithms to solve each of the two problems and give their performance analysis. We evaluate the performance of the proposed algorithms for online feature selection on several public datasets, and demonstrate their applications to real-world problems including image classification in computer vision and microarray gene expression analysis in bioinformatics. The encouraging results of our experiments validate the efficacy and efficiency of the proposed techniques.
format text
author HOI, Steven
WANG, Jialei
ZHAO, Peilin
JIN, Rong
author_facet HOI, Steven
WANG, Jialei
ZHAO, Peilin
JIN, Rong
author_sort HOI, Steven
title Online Feature Selection and Its Applications
title_short Online Feature Selection and Its Applications
title_full Online Feature Selection and Its Applications
title_fullStr Online Feature Selection and Its Applications
title_full_unstemmed Online Feature Selection and Its Applications
title_sort online feature selection and its applications
publisher Institutional Knowledge at Singapore Management University
publishDate 2012
url https://ink.library.smu.edu.sg/researchdata/13
http://ofs.stevenhoi.org/
_version_ 1681132346675298304