XML Structure Extraction from Plain Texts with Hidden Markov Model
Piao Yong, Zou Sha-sha, Wang Xiu-Kun
Abstract
Piao Yong, Zou Sha-sha, Wang Xiu-Kun
Abstract
Information extraction is one of the ways to convert unstructured text into structured records. Most of the previous work in this field are devoted to add semantic tags to specific textual content, so their structures are often plain which cannot illustrate relationships among semantic features. A novel approach, Structure Information Extraction System based on Hidden Markov Model (SIEHMM), for the task of extracting structure from plain texts is proposed in these papers, which utilizes path information for HMM training and automatically generate XML. Experiments on a real life dataset show SIEHMM has a high precision and recall ratio and can not only help solve problems of structural storage and text information retrieval, but also take advantages of XML to meet the future trends.
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Information extraction is one of the ways to convert unstructured text into structured records. Most of the previous work in this field are devoted to add semantic tags to specific textual content, so their structures are often plain which cannot illustrate relationships among semantic features. A novel approach, Structure Information Extraction System based on Hidden Markov Model (SIEHMM), for the task of extracting structure from plain texts is proposed in these papers, which utilizes path information for HMM training and automatically generate XML. Experiments on a real life dataset show SIEHMM has a high precision and recall ratio and can not only help solve problems of structural storage and text information retrieval, but also take advantages of XML to meet the future trends.
Key concepts: Computer science, XML, Information extraction, Hidden Markov model, Information retrieval, Plain text, Precision and recall, Document Structure Description