2021•Unpublished venueRequires access

Research on C4.5 Algorithm Optimization For User Churn

Chao Deng, Zhaohui Ma

Open publisher page 2 citations

Abstract

Decision tree is a kind of machine learning method which can decide a corresponding result according to the probability of different eigenvalues, The effective decision number constructed can provide help for our data analysis. The generation of decision tree is a recursive process, It mainly uses the optimal partition attribute as the corresponding tree node, and then uses various values of the attribute to construct branches. In this way, until the data reaches a certain purity, the leaf nodes are obtained, and a decision tree in accordance with the rules is constructed. Among the traditional decision tree algorithms, C4.5 algorithm has a gain rate because of its attribute division. This leads to another obvious disadvantage, that is, it has a preference for the attributes with a small number of values, so that the accuracy of the decision tree is often not particularly ideal. In view of this, this paper proposes an improved E-C4.5 algorithm, which combines information gain and information gain rate to generate a new attribute partition criterion. The attribute partition method greatly eliminates the shortcoming of C4.5 algorithm which has a preference for the attributes with a small number of values, and further improves the decision accuracy of decision tree generation. In this paper, the actual data sets are used to verify the accuracy of the decision tree generated by the improved algorithm compared with the traditional C4.5 algorithm.

About this research paper

What this paper is about

Decision tree is a kind of machine learning method which can decide a corresponding result according to the probability of different eigenvalues, The effective decision number constructed can provide help for our data analysis. The generation of decision tree is a recursive process, It mainly uses the optimal partition attribute as the corresponding tree node, and then uses various values of the attribute to construct branches. In this way, until the data reaches a certain purity, the leaf nodes are obtained, and a decision tree in accordance with the rules is constructed. Among the traditional decision tree algorithms, C4.5 algorithm has a gain rate because of its attribute division. This leads to another obvious disadvantage, that is, it has a preference for the attributes with a small number of values, so that the accuracy of the decision tree is often not particularly ideal. In view of this, this paper proposes an improved E-C4.5 algorithm, which combines information gain and information gain rate to generate a new attribute partition criterion. The attribute partition method greatly eliminates the shortcoming of C4.5 algorithm which has a preference for the attributes with a small number of values, and further improves the decision accuracy of decision tree generation. In this paper, the actual data sets are used to verify the accuracy of the decision tree generated by the improved algorithm compared with the traditional C4.5 algorithm.

Why it matters

OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Decision tree is a kind of machine learning method which can decide a corresponding result according to the probability of different eigenvalues, The effective decision number constructed can provide help for our data analysis. The generation of decision tree is a recursive process, It mainly uses the optimal partition attribute as the corresponding tree node, and then uses various values of the attribute to construct branches. In this way, until the data reaches a certain purity, the leaf nodes are obtained, and a decision tree in accordance with the rules is constructed. Among the traditional decision tree algorithms, C4.5 algorithm has a gain rate because of its attribute division. This leads to another obvious disadvantage, that is, it has a preference for the attributes with a small number of values, so that the accuracy of the decision tree is often not particularly ideal. In view of this, this paper proposes an improved E-C4.5 algorithm, which combines information gain and information gain rate to generate a new attribute partition criterion. The attribute partition method greatly eliminates the shortcoming of C4.5 algorithm which has a preference for the attributes with a small number of values, and further improves the decision accuracy of decision tree generation. In this paper, the actual data sets are used to verify the accuracy of the decision tree generated by the improved algorithm compared with the traditional C4.5 algorithm.

Key concepts: Decision tree, Incremental decision tree, Computer science, ID3 algorithm, Partition (number theory), Data mining, Algorithm, Information gain ratio

Related papers

Back to paper searchBrowse research topicsOriginal source
Research on C4.5 Algorithm Optimization For User Churn — Research Paper | ScholarLens