A Novel Method for Mining Sequential Patterns in Datasets
Xiaoyu Chang, Chunguang Zhou, Zhe Wang, Ping Hu
Abstract
Xiaoyu Chang, Chunguang Zhou, Zhe Wang, Ping Hu
Abstract
Sequential pattern mining is one of the most important fields in data mining. In this paper, we propose a novel algorithm FSPAN (Fast Sequential Pattern mining algorithm) to do the sequence mining. FSPAN can mine all the frequent sequential patterns in large datasets and it integrates a depth-first traversal approach with an effective pruning mechanism. This pruning mechanism solves the problem of searching frequent sequences in a sequence database by searching frequent items or frequent itemsets, which makes this method very efficient. Moreover, the databases scanned via FSPAN keep shrinking quickly, which makes the algorithm more efficient when the sequential patterns are longer. Experiments on standard test data show that FSPAN is very effective
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Sequential pattern mining is one of the most important fields in data mining. In this paper, we propose a novel algorithm FSPAN (Fast Sequential Pattern mining algorithm) to do the sequence mining. FSPAN can mine all the frequent sequential patterns in large datasets and it integrates a depth-first traversal approach with an effective pruning mechanism. This pruning mechanism solves the problem of searching frequent sequences in a sequence database by searching frequent items or frequent itemsets, which makes this method very efficient. Moreover, the databases scanned via FSPAN keep shrinking quickly, which makes the algorithm more efficient when the sequential patterns are longer. Experiments on standard test data show that FSPAN is very effective
Key concepts: Pruning, Tree traversal, Computer science, Sequential Pattern Mining, Data mining, Sequence (biology), GSP Algorithm, Sequence database