Phylogenetic trees based on k-nucleotide frequencies
Markus T. Friberg, Gastón H. Gonnet, Peter von Rohr, Kurt Tobler
Abstract
Open-access reader
Markus T. Friberg, Gastón H. Gonnet, Peter von Rohr, Kurt Tobler
Abstract
Open-access reader
We investigate an alternative approach to phylogenetic tree construction based on -nucleotide frequencies.Compared to traditional phylogenetic analysis, where trees are produced from multiple sequence alignments, the -nucleotide approach has the advantage that it uses all the available information.It is also almost insensitive to gene rearrangements, which are common in viruses and cause havoc to MSA methods.There are several ways of comparing the -nucleotide frequencies to compute the phylogenetic distance between two organisms.In previous work, different methods have been used producing different results.We study in detail the method of going from -nucleotide counts to phylogenetic trees.The different methods are evaluated using a tree quality index.We analyze an extensive number of new and already reported methods to see which produce the most accurate phylogenetic trees.According to our analysis, the method that works best in general is the Euclidean distance of odds ratios.This paper was motivated by the good results obtained for the classification of the SARS virus in the context of other coronaviruses and RNA viruses.As an additional example of how the method can be used, we study 11 vertebrate species for which the phylogenetic tree is known.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
We investigate an alternative approach to phylogenetic tree construction based on -nucleotide frequencies.Compared to traditional phylogenetic analysis, where trees are produced from multiple sequence alignments, the -nucleotide approach has the advantage that it uses all the available information.It is also almost insensitive to gene rearrangements, which are common in viruses and cause havoc to MSA methods.There are several ways of comparing the -nucleotide frequencies to compute the phylogenetic distance between two organisms.In previous work, different methods have been used producing different results.We study in detail the method of going from -nucleotide counts to phylogenetic trees.The different methods are evaluated using a tree quality index.We analyze an extensive number of new and already reported methods to see which produce the most accurate phylogenetic trees.According to our analysis, the method that works best in general is the Euclidean distance of odds ratios.This paper was motivated by the good results obtained for the classification of the SARS virus in the context of other coronaviruses and RNA viruses.As an additional example of how the method can be used, we study 11 vertebrate species for which the phylogenetic tree is known.
Key concepts: Phylogenetic tree, Phylogenetic network, Tree (set theory), Context (archaeology), Biology, Computational phylogenetics, Phylogenetics, Tree rearrangement