A MIXED METHOD FOR PROVIDING EXACT SNP IDs FROM SEQUENCES
Li‐Yeh Chuang, Cheng‐Hong Yang, Yu‐Huei Cheng, Hsueh‐Wei Chang
Abstract
Li‐Yeh Chuang, Cheng‐Hong Yang, Yu‐Huei Cheng, Hsueh‐Wei Chang
Abstract
Single nucleotide polymorphisms (SNPs) are the most frequently occurring genetic variations. Biologists use identified SNPs to investigate genetic diseases and heredity markers. They are also used to prevent side effects of medication. Thus, SNPs play an important role in personalized medicine. However, many association studies provide only the relationship among SNPs, diseases and cancers, without giving an SNP ID. In order to identify SNPs in a sequence, this research built dbSNP, SNP fasta and SNP flanking marker databases for the rat, mouse and human genome from the NCBI database. The proposed method utilizes SNP flanking markers that are extracted from a SNP fasta sequence and combines a Boyer–Moore algorithm with a dynamic programming method. The Boyer–Moore algorithm helps to select possible SNPs from the SNP fasta database using unknown sequences, and the dynamic programming method will then validate these SNPs. This method is very reliable retrieving SNP IDs from an unknown sequence. The experimental results show that this method is indeed able to determine exact SNP IDs from a sequence. It constitutes a novel application for the identification of SNP IDs from the literature and can be used in systematic association studies.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Single nucleotide polymorphisms (SNPs) are the most frequently occurring genetic variations. Biologists use identified SNPs to investigate genetic diseases and heredity markers. They are also used to prevent side effects of medication. Thus, SNPs play an important role in personalized medicine. However, many association studies provide only the relationship among SNPs, diseases and cancers, without giving an SNP ID. In order to identify SNPs in a sequence, this research built dbSNP, SNP fasta and SNP flanking marker databases for the rat, mouse and human genome from the NCBI database. The proposed method utilizes SNP flanking markers that are extracted from a SNP fasta sequence and combines a Boyer–Moore algorithm with a dynamic programming method. The Boyer–Moore algorithm helps to select possible SNPs from the SNP fasta database using unknown sequences, and the dynamic programming method will then validate these SNPs. This method is very reliable retrieving SNP IDs from an unknown sequence. The experimental results show that this method is indeed able to determine exact SNP IDs from a sequence. It constitutes a novel application for the identification of SNP IDs from the literature and can be used in systematic association studies.
Key concepts: dbSNP, SNP, Single-nucleotide polymorphism, Tag SNP, Molecular Inversion Probe, SNP genotyping, SNP array, Computational biology