Concept Drift Visualization Using Feature Importance on the Streaming Data
Martin Sarnovský
Abstract
Martin Sarnovský
Abstract
Currently, online processing of the data streams is a very active research topic. Streaming data are usually dynamic, where the underlying data distributions evolve during the time. Predictive data analytical tasks, such as classification, must be able to reflect such dynamics. This phenomenon is called a concept drift, and multiple adaptive classification methods have been proposed to handle drifting streams. To understand how the adaptive models work, it is necessary to use the techniques able to visualize how the model performs, as well as to provide explanations of the drift occurrence. In this paper, we present the visualization technique based on feature importance. In this case, we want to provide information about the continuous importance of the input features and use it to explain the possible drifts in the data. We used the commonly used ADWIN adaptive streaming classifier and evaluated the technique on the two real-world data streams with concept drift.
OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Currently, online processing of the data streams is a very active research topic. Streaming data are usually dynamic, where the underlying data distributions evolve during the time. Predictive data analytical tasks, such as classification, must be able to reflect such dynamics. This phenomenon is called a concept drift, and multiple adaptive classification methods have been proposed to handle drifting streams. To understand how the adaptive models work, it is necessary to use the techniques able to visualize how the model performs, as well as to provide explanations of the drift occurrence. In this paper, we present the visualization technique based on feature importance. In this case, we want to provide information about the continuous importance of the input features and use it to explain the possible drifts in the data. We used the commonly used ADWIN adaptive streaming classifier and evaluated the technique on the two real-world data streams with concept drift.
Key concepts: Concept drift, Streaming data, Computer science, Data stream mining, Visualization, Data mining, Classifier (UML), Feature (linguistics)