Synopses Generation for Specialized Document-Element Search Engines
Sumit Bhatia, Prasenjit Mitra
Abstract
Sumit Bhatia, Prasenjit Mitra
Abstract
Scientists often want to search for document-elements like tables and gures in digital documents. Using a documentelement search engine helps them to retrieve a set of documentelements using keyword queries. Consequently, they need to decide whether the returned document-element is useful and then determine what information is contained in it. The last step is typically done by downloading the paper and reading it. In this paper, we investigate how to extract information (synopsis) related to document-elements from documents automatically. The extracted information can be indexed and provided along with the search results, enabling the end-user to quickly nd the related information. Thus, this work has signicant potential to facilitate easeof-use for a document-element search engine, consequently increasing the productivity of the end-user. We propose a novel method to extract synopses, investigate the optimum synopsis-size and demonstrate the utility of our extracted synopsis in document-element understanding with a user study.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Scientists often want to search for document-elements like tables and gures in digital documents. Using a documentelement search engine helps them to retrieve a set of documentelements using keyword queries. Consequently, they need to decide whether the returned document-element is useful and then determine what information is contained in it. The last step is typically done by downloading the paper and reading it. In this paper, we investigate how to extract information (synopsis) related to document-elements from documents automatically. The extracted information can be indexed and provided along with the search results, enabling the end-user to quickly nd the related information. Thus, this work has signicant potential to facilitate easeof-use for a document-element search engine, consequently increasing the productivity of the end-user. We propose a novel method to extract synopses, investigate the optimum synopsis-size and demonstrate the utility of our extracted synopsis in document-element understanding with a user study.
Key concepts: Computer science, Information retrieval, Search engine, Upload, Element (criminal law), Set (abstract data type), Document retrieval, World Wide Web