An algorithm for finding tandem repeats of unspecified pattern size
Gary Benson
Abstract
Open-access reader
Gary Benson
Abstract
Open-access reader
A tandem repeat is two or more contiguous, approtimate copies of a pattern of nucleotides.Tandem repeats occur frequently in the human genome.They have been shown to cause human disease, may play a variety of regulatory and evolutionary roles, and are important laboratory tools.Extensive knowledge about pattern sizes, copy number, mutational history, etc. for t.andem repeats has been limited because of the difficulty of detecting them in genomic sequence data.In this paper, me present a new algorithm for finding tandem repeats in DNA sequences without the need to specify either the pattern or pattern size.The algorithm is based on the detection of k-tuple matches.It uses a probabiitic model of tandem repeats and a collection of statistical criteria based on that modeL We demonstrate the algorithm's speed and its abiity to detect tandem repeats that have undergone extensive mutational change by analyzing 4 sequences in the 2OOKb to 700Kb range.
OpenAlex reports 26 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
A tandem repeat is two or more contiguous, approtimate copies of a pattern of nucleotides.Tandem repeats occur frequently in the human genome.They have been shown to cause human disease, may play a variety of regulatory and evolutionary roles, and are important laboratory tools.Extensive knowledge about pattern sizes, copy number, mutational history, etc. for t.andem repeats has been limited because of the difficulty of detecting them in genomic sequence data.In this paper, me present a new algorithm for finding tandem repeats in DNA sequences without the need to specify either the pattern or pattern size.The algorithm is based on the detection of k-tuple matches.It uses a probabiitic model of tandem repeats and a collection of statistical criteria based on that modeL We demonstrate the algorithm's speed and its abiity to detect tandem repeats that have undergone extensive mutational change by analyzing 4 sequences in the 2OOKb to 700Kb range.
Key concepts: Computer science, Algorithm, Tandem repeat, Tandem, Genetics, Engineering, Biology, Genome