A new feature selection method for text categorization
Xin-Shun Xu
Abstract
Xin-Shun Xu
Abstract
How to reduce feature dimension while maintaining categorization accuracy is a key issue of text categorization.A new method based on information theory was proposed to solve this problem.This approach aims to eliminate sparsely distributed features and find features useful for categorization.Working with these feature reduction methods,it could further reduce the feature dimension.The performance of this proposed method was tested on benchmark text classification problems.The results showed that it could not only reduce the feature dimension to hundreds but also improve the performance.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
How to reduce feature dimension while maintaining categorization accuracy is a key issue of text categorization.A new method based on information theory was proposed to solve this problem.This approach aims to eliminate sparsely distributed features and find features useful for categorization.Working with these feature reduction methods,it could further reduce the feature dimension.The performance of this proposed method was tested on benchmark text classification problems.The results showed that it could not only reduce the feature dimension to hundreds but also improve the performance.
Key concepts: Categorization, Feature selection, Text categorization, Feature (linguistics), Benchmark (surveying), Computer science, Dimension (graph theory), Dimensionality reduction