Implementation of some similarity coefficients in conjunction with multiple UPGMA and neighbor-joining algorithms for enhancing phylogenetic trees.
Tarik Rabie
Abstract
Tarik Rabie
Abstract
Random Amplified Polymorphic DNA (RAPD) markers was used to analyze the genetic structure of five Indigenous Egyptian's chicken populations including Fayoumi, Dokki-4, Golden Montazah, Silver Montazah, and El- Salam, based on the taxa generated by the analysis of ten RAPD markers. The population genetic distances were estimated by using two cluster algorithms (UPGMA & NJ neighbor-joining) accompanied with ten similarity coefficients comprising Jaccard, Sorensen-Dice, Russel& Rao, Rogers & Tanimoto, Simple Matching, Pearson Phi, Lance &Williams, Mountford, Michael, and Kulchenzky-1. The results demonstrated that for almost all methodologies, the Jaccard and Sorensen-Dice followed by Simple Matching coefficients revealed extremely close results, because both of them exclude negative co-occurrences. Due to the fact that there is no guarantee that the DNA regions with negative co-occurrences between two strains are indeed identical, the use of coefficients such as Jaccard and Sorensen-Dice that do not include negative co- occurrences was imperative for closely related organisms along with the NJ neighbor-joining cluster algorithm.
OpenAlex reports 7 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Random Amplified Polymorphic DNA (RAPD) markers was used to analyze the genetic structure of five Indigenous Egyptian's chicken populations including Fayoumi, Dokki-4, Golden Montazah, Silver Montazah, and El- Salam, based on the taxa generated by the analysis of ten RAPD markers. The population genetic distances were estimated by using two cluster algorithms (UPGMA & NJ neighbor-joining) accompanied with ten similarity coefficients comprising Jaccard, Sorensen-Dice, Russel& Rao, Rogers & Tanimoto, Simple Matching, Pearson Phi, Lance &Williams, Mountford, Michael, and Kulchenzky-1. The results demonstrated that for almost all methodologies, the Jaccard and Sorensen-Dice followed by Simple Matching coefficients revealed extremely close results, because both of them exclude negative co-occurrences. Due to the fact that there is no guarantee that the DNA regions with negative co-occurrences between two strains are indeed identical, the use of coefficients such as Jaccard and Sorensen-Dice that do not include negative co- occurrences was imperative for closely related organisms along with the NJ neighbor-joining cluster algorithm.
Key concepts: Jaccard index, UPGMA, RAPD, Dice, Similarity (geometry), Phylogenetic tree, Mathematics, Matching (statistics)