Extracting Topological Features from Big Data Using Persistent Density Entropy
Jinzhong Xu, Xuzhi Li, Hongfei Wang
Abstract
Open-access reader
Jinzhong Xu, Xuzhi Li, Hongfei Wang
Abstract
Open-access reader
Topological data analysis is a method of extracting shape information of big data by means of algebraic topology in mathematics. Persistent homology is a very important method in topological data analysis. It constructs multi-scale simplicial complexes (also called filtration) to approximate the underlying space of the data set. By studying these simplicial complexes, the topological features of each dimension of big data are summarized. However, it does not give us the uncertainty of each simplicial complex to approximate the underlying space of the data set. This paper defines an entropy called persistent density entropy, which gives the uncertainty of each simplicial complex approximating the underlying space. The examples demonstrate that it is able to find the best simplicial complex that approximates the underlying space and can be used to detect outliers to a certain extent.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Topological data analysis is a method of extracting shape information of big data by means of algebraic topology in mathematics. Persistent homology is a very important method in topological data analysis. It constructs multi-scale simplicial complexes (also called filtration) to approximate the underlying space of the data set. By studying these simplicial complexes, the topological features of each dimension of big data are summarized. However, it does not give us the uncertainty of each simplicial complex to approximate the underlying space of the data set. This paper defines an entropy called persistent density entropy, which gives the uncertainty of each simplicial complex approximating the underlying space. The examples demonstrate that it is able to find the best simplicial complex that approximates the underlying space and can be used to detect outliers to a certain extent.
Key concepts: Persistent homology, Topological data analysis, Simplicial complex, Outlier, Abstract simplicial complex, Mathematics, Simplicial homology, Entropy (arrow of time)