Analyzing high dimensional correlated data using feature ranking and classifiers Academic Article uri icon

abstract

  • Abstract The Illumina Infinium HumanMethylation27 (Illumina 27K) BeadChip assay is a relatively recent high-throughput technology that allows over 27,000 CpGs to be assayed. The Illumina 27K methylation data is less commonly used in comparison to gene expression in bioinformatics. It provides a critical need to find the optimal feature ranking (FR) method for handling the high dimensional data. The optimal FR method on the classifier is not well known, and choosing the best performing FR method becomes more challenging in high dimensional data setting. Therefore, identifying the statistical methods which boost the inference is of crucial importance in this context. This paper describes the detailed performances of FR methods such as fisher score, information gain, chi-square, and minimum redundancy and maximum relevance on different classification methods such as Adaboost, Random Forest, Naive Bayes, and Support Vector Machines. Through simulation study and real data applications, we show that the fisher score as an FR method, when applied on all the classifiers, achieved best prediction accuracy with significantly small number of ranked features.

published proceedings

  • Computational and Mathematical Biophysics

author list (cited authors)

  • Patil, A. R., Chang, J., Leung, M., & Kim, S.

citation count

  • 3

complete list of authors

  • Patil, Abhijeet R||Chang, Jongwha||Leung, Ming-Ying||Kim, Sangjin

publication date

  • January 2019