A Novel Method Providing Exact SNP IDs from Sequences
Yu‐Huei Cheng, Cheng‐San Yang, Hsueh‐Wei Chang, Li‐Yeh Chuang, Cheng‐Hong Yang
Abstract
Yu‐Huei Cheng, Cheng‐San Yang, Hsueh‐Wei Chang, Li‐Yeh Chuang, Cheng‐Hong Yang
Abstract
Single-nucleotide polymorphisms (SNPs) are the most common type of DNA sequence variation. An SNP is the substitution of a single base in the sequence for one that is different from that present in the majority of the population. SNPs were very important for personalized medicine, especially for association studies. Each SNP has an ID number (rs#) in dbSNP of NCBI, providing the information for SNP genotype and frequency of many populations. However, many previous association studies provide only the SNP nucleotide position or primer sequences, without giving an SNP ID of NCBI. In this study, we built the dbSNP, SNP fasta and SNP flanking marker databases for the rat, mouse and human organisms from the NCBI databases. Boyer-Moore algorithm, dynamic programming method and database technologies were applied and integrated to identify the SNP IDs within input sequences. Therefore, we proposed a novel method to provide efficient, exact and stable output for SNP IDs discovery from a sequence. It also constitutes a novel application to identify SNP IDs from the literatures for systematic association studies.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Single-nucleotide polymorphisms (SNPs) are the most common type of DNA sequence variation. An SNP is the substitution of a single base in the sequence for one that is different from that present in the majority of the population. SNPs were very important for personalized medicine, especially for association studies. Each SNP has an ID number (rs#) in dbSNP of NCBI, providing the information for SNP genotype and frequency of many populations. However, many previous association studies provide only the SNP nucleotide position or primer sequences, without giving an SNP ID of NCBI. In this study, we built the dbSNP, SNP fasta and SNP flanking marker databases for the rat, mouse and human organisms from the NCBI databases. Boyer-Moore algorithm, dynamic programming method and database technologies were applied and integrated to identify the SNP IDs within input sequences. Therefore, we proposed a novel method to provide efficient, exact and stable output for SNP IDs discovery from a sequence. It also constitutes a novel application to identify SNP IDs from the literatures for systematic association studies.
Key concepts: dbSNP, SNP, Single-nucleotide polymorphism, Tag SNP, Molecular Inversion Probe, SNP genotyping, Genetics, Computational biology