2023Balisage series on markup technologiesOpen access

Auto-Markup BenchMark: Towards an Industry-standard Benchmark for Evaluating Automatic Document Markup

Paul Prescod, B. Feuer, Andrii Hladkyi, Sean Paulk, A.Mallikarjuna Prasad

Open full text 2 citations

Abstract

Recent large language models (LLMs) such as GPT, Bard and Claude have achieved impressive progress in comprehending and generating human language. We explore the potential for LLMs to automate markup of semi-structured texts into HTML, DITA, and other related languages. Automated markup could make it possible to move documents into XML at much lower cost. Currently, there is no standard benchmark for evaluating automatic markup systems. Establishing a benchmark for the understanding and generation of structured documents and markup languages can drive innovation, standardize evaluations, identify algorithm strengths and weaknesses, clarify the State of the Art and foster interdisciplinary collaborations. This paper introduces an early version of a benchmark for this purpose.

Open-access reader

About this research paper

What this paper is about

Recent large language models (LLMs) such as GPT, Bard and Claude have achieved impressive progress in comprehending and generating human language. We explore the potential for LLMs to automate markup of semi-structured texts into HTML, DITA, and other related languages. Automated markup could make it possible to move documents into XML at much lower cost. Currently, there is no standard benchmark for evaluating automatic markup systems. Establishing a benchmark for the understanding and generation of structured documents and markup languages can drive innovation, standardize evaluations, identify algorithm strengths and weaknesses, clarify the State of the Art and foster interdisciplinary collaborations. This paper introduces an early version of a benchmark for this purpose.

Why it matters

OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Recent large language models (LLMs) such as GPT, Bard and Claude have achieved impressive progress in comprehending and generating human language. We explore the potential for LLMs to automate markup of semi-structured texts into HTML, DITA, and other related languages. Automated markup could make it possible to move documents into XML at much lower cost. Currently, there is no standard benchmark for evaluating automatic markup systems. Establishing a benchmark for the understanding and generation of structured documents and markup languages can drive innovation, standardize evaluations, identify algorithm strengths and weaknesses, clarify the State of the Art and foster interdisciplinary collaborations. This paper introduces an early version of a benchmark for this purpose.

Key concepts: Markup language, Benchmark (surveying), XML, Computer science, RuleML, SGML, Strengths and weaknesses, XHTML

Related papers

Back to paper searchBrowse research topicsOriginal source
Auto-Markup BenchMark: Towards an Industry-standard Benchmark for Evaluating Automatic Document Markup — Research Paper | ScholarLens