Determining appropriate approaches for using data in feature selection

Aldehim, Ghadah and Wang, Wenjia (2017) Determining appropriate approaches for using data in feature selection. International Journal of Machine Learning and Cybernetics, 8 (3). 915–928. ISSN 1868-808X

[thumbnail of PublishedPrint]
PDF (PublishedPrint) - Published Version
Download (2MB) | Preview


Feature selection is increasingly important in data analysis and machine learning in big data era. However, how to use the data in feature selection, i.e. using either ALL or PART of a dataset, has become a serious and tricky issue. Whilst the conventional practice of using all the data in feature selection may lead to selection bias, using part of the data may, on the other hand, lead to underestimating the relevant features under some conditions. This paper investigates these two strategies systematically in terms of reliability and effectiveness, and then determines their suitability for datasets with different characteristics. The reliability is measured by the Average Tanimoto Index and the Inter-method Average Tanimoto Index, and the effectiveness is measured by the mean generalisation accuracy of classification. The computational experiments are carried out on ten real-world benchmark datasets and fourteen synthetic datasets. The synthetic datasets are generated with a pre-set number of relevant features and varied numbers of irrelevant features and instances, and added with different levels of noise. The results indicate that the PART approach is more effective in reducing the bias when the size of a dataset is small but starts to lose its advantage as the dataset size increases.

Item Type: Article
Uncontrolled Keywords: features selection , reliability ,big data,cross-validation,classification,similarity measure
Faculty \ School: Faculty of Science > School of Computing Sciences

UEA Research Groups: Faculty of Science > Research Groups > Data Science and Statistics
Related URLs:
Depositing User: Pure Connector
Date Deposited: 08 Jan 2016 10:00
Last Modified: 21 Oct 2022 00:34
DOI: 10.1007/s13042-015-0469-8

Actions (login required)

View Item View Item