XML Schema Directory: a data structure for XML data processing
E. Kotsakis, Klemens Böhm
Abstract
E. Kotsakis, Klemens Böhm
Abstract
The problem addressed in this paper is the execution of XML queries over a large collection of XML documents. This paper concentrates on how to develop the necessary infrastructure to effectively manipulate XML data and it proposes a data structure, named the XML Schema Directory (XSD), as an access means to XML repositories. The aim of XSD is to accelerate query processing by quickly finding the relevant set of XML documents for a given query. This is obtained by considering only a small number of relative XML schemata and consequently a limiting number of XML documents, rather than the entire corpus of XML documents. XML schema similarity is introduced as a way to determine the relevance among XML documents which belong to the same knowledge category. The proposed algorithms for maintaining the XSD structure do not require reorganisation and they may be efficiently used in practice. An alternative advantage of the XSD structure is that it may also be used as a method for facilitating browsing.
OpenAlex reports 18 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The problem addressed in this paper is the execution of XML queries over a large collection of XML documents. This paper concentrates on how to develop the necessary infrastructure to effectively manipulate XML data and it proposes a data structure, named the XML Schema Directory (XSD), as an access means to XML repositories. The aim of XSD is to accelerate query processing by quickly finding the relevant set of XML documents for a given query. This is obtained by considering only a small number of relative XML schemata and consequently a limiting number of XML documents, rather than the entire corpus of XML documents. XML schema similarity is introduced as a way to determine the relevance among XML documents which belong to the same knowledge category. The proposed algorithms for maintaining the XSD structure do not require reorganisation and they may be efficiently used in practice. An alternative advantage of the XSD structure is that it may also be used as a method for facilitating browsing.
Key concepts: Computer science, XML validation, XML Schema Editor, Document Structure Description, Efficient XML Interchange, Streaming XML, XML Encryption, XML database