Auto-Markup BenchMark: Towards an Industry-standard Benchmark for Evaluating Automatic Document Markup
Paul Prescod, B. Feuer, Andrii Hladkyi, Sean Paulk, A.Mallikarjuna Prasad
Abstract
Open-access reader
Paul Prescod, B. Feuer, Andrii Hladkyi, Sean Paulk, A.Mallikarjuna Prasad
Abstract
Open-access reader
Recent large language models (LLMs) such as GPT, Bard and Claude have achieved impressive progress in comprehending and generating human language. We explore the potential for LLMs to automate markup of semi-structured texts into HTML, DITA, and other related languages. Automated markup could make it possible to move documents into XML at much lower cost. Currently, there is no standard benchmark for evaluating automatic markup systems. Establishing a benchmark for the understanding and generation of structured documents and markup languages can drive innovation, standardize evaluations, identify algorithm strengths and weaknesses, clarify the State of the Art and foster interdisciplinary collaborations. This paper introduces an early version of a benchmark for this purpose.
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Recent large language models (LLMs) such as GPT, Bard and Claude have achieved impressive progress in comprehending and generating human language. We explore the potential for LLMs to automate markup of semi-structured texts into HTML, DITA, and other related languages. Automated markup could make it possible to move documents into XML at much lower cost. Currently, there is no standard benchmark for evaluating automatic markup systems. Establishing a benchmark for the understanding and generation of structured documents and markup languages can drive innovation, standardize evaluations, identify algorithm strengths and weaknesses, clarify the State of the Art and foster interdisciplinary collaborations. This paper introduces an early version of a benchmark for this purpose.
Key concepts: Markup language, Benchmark (surveying), XML, Computer science, RuleML, SGML, Strengths and weaknesses, XHTML