2018RePEc: Research Papers in EconomicsRequires access

NGRAM: Stata module to provide n-gram feature extractor

Matthias Schonlau

Open publisher page 0 citations

Abstract

ngram extracts n-gram variables containing counts of how often the n-grams occur in a given text. An n-gram is an n-long sequence of words. For example, "sheep" is a unigram (1-gram), "black sheep" is a bigram (2-gram), and "the black sheep is happy" is a 5-gram. This is useful for text mining applications.

About this research paper

What this paper is about

ngram extracts n-gram variables containing counts of how often the n-grams occur in a given text. An n-gram is an n-long sequence of words. For example, "sheep" is a unigram (1-gram), "black sheep" is a bigram (2-gram), and "the black sheep is happy" is a 5-gram. This is useful for text mining applications.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

ngram extracts n-gram variables containing counts of how often the n-grams occur in a given text. An n-gram is an n-long sequence of words. For example, "sheep" is a unigram (1-gram), "black sheep" is a bigram (2-gram), and "the black sheep is happy" is a 5-gram. This is useful for text mining applications.

Key concepts: Bigram, n-gram, Gram, Extractor, Feature (linguistics), Computer science, Sequence (biology), Statistics

Related papers

Back to paper searchBrowse research topicsOriginal source
NGRAM: Stata module to provide n-gram feature extractor — Research Paper | ScholarLens