2010Unpublished venueRequires access

XML Structure Extraction from Plain Texts with Hidden Markov Model

Piao Yong, Zou Sha-sha, Wang Xiu-Kun

Open publisher page 2 citations

Abstract

Information extraction is one of the ways to convert unstructured text into structured records. Most of the previous work in this field are devoted to add semantic tags to specific textual content, so their structures are often plain which cannot illustrate relationships among semantic features. A novel approach, Structure Information Extraction System based on Hidden Markov Model (SIEHMM), for the task of extracting structure from plain texts is proposed in these papers, which utilizes path information for HMM training and automatically generate XML. Experiments on a real life dataset show SIEHMM has a high precision and recall ratio and can not only help solve problems of structural storage and text information retrieval, but also take advantages of XML to meet the future trends.

About this research paper

What this paper is about

Information extraction is one of the ways to convert unstructured text into structured records. Most of the previous work in this field are devoted to add semantic tags to specific textual content, so their structures are often plain which cannot illustrate relationships among semantic features. A novel approach, Structure Information Extraction System based on Hidden Markov Model (SIEHMM), for the task of extracting structure from plain texts is proposed in these papers, which utilizes path information for HMM training and automatically generate XML. Experiments on a real life dataset show SIEHMM has a high precision and recall ratio and can not only help solve problems of structural storage and text information retrieval, but also take advantages of XML to meet the future trends.

Why it matters

OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Information extraction is one of the ways to convert unstructured text into structured records. Most of the previous work in this field are devoted to add semantic tags to specific textual content, so their structures are often plain which cannot illustrate relationships among semantic features. A novel approach, Structure Information Extraction System based on Hidden Markov Model (SIEHMM), for the task of extracting structure from plain texts is proposed in these papers, which utilizes path information for HMM training and automatically generate XML. Experiments on a real life dataset show SIEHMM has a high precision and recall ratio and can not only help solve problems of structural storage and text information retrieval, but also take advantages of XML to meet the future trends.

Key concepts: Computer science, XML, Information extraction, Hidden Markov model, Information retrieval, Plain text, Precision and recall, Document Structure Description

Related papers

Back to paper searchBrowse research topicsOriginal source
XML Structure Extraction from Plain Texts with Hidden Markov Model — Research Paper | ScholarLens