Determination of loss of information during data anonymization procedure
Alexey Pervushin, Viktoria Ermachkova, Anton Spivak
Abstract
Alexey Pervushin, Viktoria Ermachkova, Anton Spivak
Abstract
One of the effective approaches to the protection of personal data is their depersonalization, as it reduces the requirements for the level of data protection. Therefore anonymization procedure is widely used in practice. As a result of the application of such methods discards some of the information, which entails a loss of information content of the personal data. The problem of determining the loss of information during data anonymization procedure. It is often easier to allocate a group of similar objects, to analyze the loss for each group, than analyze all the data at once. A large amount of data and the lack of pre-known cluster, prompted the decision of one of the tasks of Data Mining - Data Clustering. Realize the data by identifying the cluster structure allows us to understand how the data looked at the beginning, and how have they changed after anonymization procedure. Analyzing the structure of the cluster can be seen on the data identity.
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
One of the effective approaches to the protection of personal data is their depersonalization, as it reduces the requirements for the level of data protection. Therefore anonymization procedure is widely used in practice. As a result of the application of such methods discards some of the information, which entails a loss of information content of the personal data. The problem of determining the loss of information during data anonymization procedure. It is often easier to allocate a group of similar objects, to analyze the loss for each group, than analyze all the data at once. A large amount of data and the lack of pre-known cluster, prompted the decision of one of the tasks of Data Mining - Data Clustering. Realize the data by identifying the cluster structure allows us to understand how the data looked at the beginning, and how have they changed after anonymization procedure. Analyzing the structure of the cluster can be seen on the data identity.
Key concepts: Computer science, Data anonymization, Data mining, Cluster analysis, Information loss, Data loss, Cluster (spacecraft), Information privacy