2013•Unpublished venueRequires access

Mining Sequential Pattern with Time-Constraint

Anita Zala, Mehul Barot

Open publisher page 1 citations

Abstract

Sequential pattern mining is an important data mining task, and different algorithms have been proposed to perform this task efficiently. The problem is to find all sequential patterns with higher or equal support to a predefined minimum support threshold in a data sequence database. Here we present a new methodology to mine a sequential pattern with time- constraint. Our study shows that constraints can be effectively and efficiently pushed deep into sequential pattern mining under this new framework. I. INTRODUCTION Data mining is to find valid, novel, potentially useful, and ultimately understandable patterns in data. Sequential pattern mining, which extracts frequent subsequences from a sequence database, has attracted a great deal of interest during the recent surge in data mining research because it is the basis of many applications, such as customer behavior analysis, stock trend prediction, and DNA sequence analysis. The sequential mining problem was first introduced in; two sequential patterns examples are: ―80% of the people who buy a television also buy a video camera within a day‖, and ―Every time Microsoft stock drops by 5%, then IBM stock will also drop by at least 4% within three days‖. The above patterns can be used to determine the efficient use of shelf space for customer convenience, or to properly plan the next step during an economic crisis. Sequential pattern mining is also very important for analyzing bio- logical data, in which a very small alphabet (i.e., 4 for DNA sequences and 20 for protein sequences) and long patterns with a typical length of few hundreds or even thousands frequently appear. Sequence discovery can be thought of as essentially an association discovery over a temporal database. While association rules discern only intra-event patterns (item sets), sequential pattern mining discerns inter-event patterns (sequences). There are many other important tasks related to association rule mining, such as correlations, causality, episodes, multi dimensional patterns, maximal patterns, partial periodicity, and emerging patterns. Elaborate exploration of sequential pattern mining issue will be beneficial to the other research problems shown above a lot. Thus, effective and efficient sequential pattern mining is an important and interesting research problem. Efficient sequential pattern mining methodologies have been studied extensively in many related problems, including the general sequential pattern mining, constraint-based sequential pattern mining, incremental sequential pat- tern mining, frequent episode mining, approximate sequential pattern mining, partial periodic pattern mining, temporal pattern mining in data stream, maximal and closed sequential pattern mining. Although there are so many problems related to sequential pattern mining explored, we realize that the general sequential pattern mining algorithm development is the most basic one because all the others can benefit from the strategies it employs, i.e., Apriori heuristic and projection-based pattern growth. Hence we aim to develop an efficient general sequential pattern mining algorithm. All of these works suffer from the problems of having a large search space and the ineffectiveness in handling dense data sets, i.e., biological data. In this work, we propose new strategies to reduce the space necessary to be searched. Instead of searching the entire projected database for each item, as Prefix Span (1) does, we only search a small portion of the database by recording the last position of each item in each sequence.

About this research paper

What this paper is about

Sequential pattern mining is an important data mining task, and different algorithms have been proposed to perform this task efficiently. The problem is to find all sequential patterns with higher or equal support to a predefined minimum support threshold in a data sequence database. Here we present a new methodology to mine a sequential pattern with time- constraint. Our study shows that constraints can be effectively and efficiently pushed deep into sequential pattern mining under this new framework. I. INTRODUCTION Data mining is to find valid, novel, potentially useful, and ultimately understandable patterns in data. Sequential pattern mining, which extracts frequent subsequences from a sequence database, has attracted a great deal of interest during the recent surge in data mining research because it is the basis of many applications, such as customer behavior analysis, stock trend prediction, and DNA sequence analysis. The sequential mining problem was first introduced in; two sequential patterns examples are: ―80% of the people who buy a television also buy a video camera within a day‖, and ―Every time Microsoft stock drops by 5%, then IBM stock will also drop by at least 4% within three days‖. The above patterns can be used to determine the efficient use of shelf space for customer convenience, or to properly plan the next step during an economic crisis. Sequential pattern mining is also very important for analyzing bio- logical data, in which a very small alphabet (i.e., 4 for DNA sequences and 20 for protein sequences) and long patterns with a typical length of few hundreds or even thousands frequently appear. Sequence discovery can be thought of as essentially an association discovery over a temporal database. While association rules discern only intra-event patterns (item sets), sequential pattern mining discerns inter-event patterns (sequences). There are many other important tasks related to association rule mining, such as correlations, causality, episodes, multi dimensional patterns, maximal patterns, partial periodicity, and emerging patterns. Elaborate exploration of sequential pattern mining issue will be beneficial to the other research problems shown above a lot. Thus, effective and efficient sequential pattern mining is an important and interesting research problem. Efficient sequential pattern mining methodologies have been studied extensively in many related problems, including the general sequential pattern mining, constraint-based sequential pattern mining, incremental sequential pat- tern mining, frequent episode mining, approximate sequential pattern mining, partial periodic pattern mining, temporal pattern mining in data stream, maximal and closed sequential pattern mining. Although there are so many problems related to sequential pattern mining explored, we realize that the general sequential pattern mining algorithm development is the most basic one because all the others can benefit from the strategies it employs, i.e., Apriori heuristic and projection-based pattern growth. Hence we aim to develop an efficient general sequential pattern mining algorithm. All of these works suffer from the problems of having a large search space and the ineffectiveness in handling dense data sets, i.e., biological data. In this work, we propose new strategies to reduce the space necessary to be searched. Instead of searching the entire projected database for each item, as Prefix Span (1) does, we only search a small portion of the database by recording the last position of each item in each sequence.

Why it matters

OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Sequential pattern mining is an important data mining task, and different algorithms have been proposed to perform this task efficiently. The problem is to find all sequential patterns with higher or equal support to a predefined minimum support threshold in a data sequence database. Here we present a new methodology to mine a sequential pattern with time- constraint. Our study shows that constraints can be effectively and efficiently pushed deep into sequential pattern mining under this new framework. I. INTRODUCTION Data mining is to find valid, novel, potentially useful, and ultimately understandable patterns in data. Sequential pattern mining, which extracts frequent subsequences from a sequence database, has attracted a great deal of interest during the recent surge in data mining research because it is the basis of many applications, such as customer behavior analysis, stock trend prediction, and DNA sequence analysis. The sequential mining problem was first introduced in; two sequential patterns examples are: ―80% of the people who buy a television also buy a video camera within a day‖, and ―Every time Microsoft stock drops by 5%, then IBM stock will also drop by at least 4% within three days‖. The above patterns can be used to determine the efficient use of shelf space for customer convenience, or to properly plan the next step during an economic crisis. Sequential pattern mining is also very important for analyzing bio- logical data, in which a very small alphabet (i.e., 4 for DNA sequences and 20 for protein sequences) and long patterns with a typical length of few hundreds or even thousands frequently appear. Sequence discovery can be thought of as essentially an association discovery over a temporal database. While association rules discern only intra-event patterns (item sets), sequential pattern mining discerns inter-event patterns (sequences). There are many other important tasks related to association rule mining, such as correlations, causality, episodes, multi dimensional patterns, maximal patterns, partial periodicity, and emerging patterns. Elaborate exploration of sequential pattern mining issue will be beneficial to the other research problems shown above a lot. Thus, effective and efficient sequential pattern mining is an important and interesting research problem. Efficient sequential pattern mining methodologies have been studied extensively in many related problems, including the general sequential pattern mining, constraint-based sequential pattern mining, incremental sequential pat- tern mining, frequent episode mining, approximate sequential pattern mining, partial periodic pattern mining, temporal pattern mining in data stream, maximal and closed sequential pattern mining. Although there are so many problems related to sequential pattern mining explored, we realize that the general sequential pattern mining algorithm development is the most basic one because all the others can benefit from the strategies it employs, i.e., Apriori heuristic and projection-based pattern growth. Hence we aim to develop an efficient general sequential pattern mining algorithm. All of these works suffer from the problems of having a large search space and the ineffectiveness in handling dense data sets, i.e., biological data. In this work, we propose new strategies to reduce the space necessary to be searched. Instead of searching the entire projected database for each item, as Prefix Span (1) does, we only search a small portion of the database by recording the last position of each item in each sequence.

Key concepts: Sequential Pattern Mining, Computer science, Data mining, Sequence database, Constraint (computer-aided design), Sequence (biology), Task (project management), Engineering

Related papers

Back to paper searchBrowse research topicsOriginal source
Mining Sequential Pattern with Time-Constraint — Research Paper | ScholarLens