Development and Design of General Data Mining System
Baowen Chen
Abstract
Open-access reader
Baowen Chen
Abstract
Open-access reader
In this paper, we focus on top-down discretization methods and propose a new method for supervised discretization based on class-feature correlation by defining a class-feature contingency factor.The proposed method takes into consideration the distribution of all samples to generate an ideal discretization scheme.The method maintains a high interdependence between the target class and the discretized attribute, and avoids overfitting.Empirical evaluation of seven discretization algorithms on UCI real datasets show that the novel algorithm can yield a better discretization scheme that improves the accuracy of decision tree classification.As to the execution time of discretization and the number of generated rules, our approach also achieves promising results.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In this paper, we focus on top-down discretization methods and propose a new method for supervised discretization based on class-feature correlation by defining a class-feature contingency factor.The proposed method takes into consideration the distribution of all samples to generate an ideal discretization scheme.The method maintains a high interdependence between the target class and the discretized attribute, and avoids overfitting.Empirical evaluation of seven discretization algorithms on UCI real datasets show that the novel algorithm can yield a better discretization scheme that improves the accuracy of decision tree classification.As to the execution time of discretization and the number of generated rules, our approach also achieves promising results.
Key concepts: Computer science, Data science, Data mining