Concept Drift Detection on Streaming Data under Limited Labeling
Young In Kim, Cheong Hee Park
Abstract
Young In Kim, Cheong Hee Park
Abstract
In data stream analyses, detecting the concept drift accurately is important to maintain the classification performance. Most drift detection methods assume that the class labels become available immediately after a data sample arrives. However, this assumption is overly optimistic, as labeling costs are high and much time is needed to obtain the label of data samples. Therefore, it is un-realistic to attempt to acquire all of the labels when processing the data streams. In this paper, we propose a concept drift detection method under the assumption that there is limited access to labels. The proposed method detects concept drift on unlabeled data streams based on the class label information which is predicted by the classifier trained with the limited number of labeled data samples. Experimental results on synthetic and real streaming data show that the proposed method is competent to detect the concept drift using only a small amount of labeled data.
OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In data stream analyses, detecting the concept drift accurately is important to maintain the classification performance. Most drift detection methods assume that the class labels become available immediately after a data sample arrives. However, this assumption is overly optimistic, as labeling costs are high and much time is needed to obtain the label of data samples. Therefore, it is un-realistic to attempt to acquire all of the labels when processing the data streams. In this paper, we propose a concept drift detection method under the assumption that there is limited access to labels. The proposed method detects concept drift on unlabeled data streams based on the class label information which is predicted by the classifier trained with the limited number of labeled data samples. Experimental results on synthetic and real streaming data show that the proposed method is competent to detect the concept drift using only a small amount of labeled data.
Key concepts: Concept drift, Streaming data, Computer science, Data stream, Data stream mining, Classifier (UML), Labeled data, Data mining