Contributions to Methods Useful for Optimising Animal Breeding Plans
Vinzent Börner
Abstract
Vinzent Börner
Abstract
The goal of breeding activities in commercial livestock populations is the increase of the mean of the genetically based performance capacity concerning one or numerous traits being summarised the in aggregate genotype via weighting factors if a certain breeding scheme is applied. Given the breeding scheme, the extent of this increase depends on the accuracy of breeding value estimation which is a function of a) the amount of information available about the selection candidates and its correlation structure to the aggregate genotype, and b) the usefulness of the statistical model in order to regress the genotype of the selection candidate on these available information. Recent molecular-genetical findings concern both, the amount of available information as well as the statistical model. The latter is affected by a ``genomic imprinting'' called mechanism, leading to an alteration of a genes effect on the phenotype of offspring due to a sex specific DNA methylation during gametogenesis in parents. Genomic imprinting can be regarded in breeding value estimation due to the calculation of two breeding values for each individual, one if it acts as sire and the other if it acts as a dam. Weighting factors to summarise this breeding values can be derived by an extension of the gene flow method. This extension is developed in the first part of this thesis and allows for tracing the flow of genes of a certain founder or group of founders within a population across tiers (e.g. nucleus, multiplier, production) and generations with special regard to the sex of the direct parent of an individual carrying these genes. Thus, it allows to assess the probability that a gene of a founder is inherited to its descendants via their direct sires or dams. The discounted and summarised trait realisations out of the genes inherited by the sire and the dam can be used as weighting coefficients for summarising the breeding values of an individual as a sire and a dam. The extended gene flow method is applied to a hypothetical pig breeding program showing that the weights for the breeding values as a dam and as a sire can differ according to the chosen breeding scheme and the planning horizon. Furthermore, it is shown that depending the breeding scheme the breeding value of a dam when acting as a sire might be weighted higher than when acting as a dam. Additionally, a possibility to predict the increase in inbreeding due to one round of selection inherent in the method is presented. The above mentioned amount of available information about a selection candidate is affected by the discovery of hundreds of thousands DNA markers in form of single nucleotide polymorphisms (SNP marker) being in strong linkage disequilibrium with neighbouring trait affecting genes or quantitative trait loci (QTL). The sum over all estimated marker effects on the phenotype, the genomically estimated breeding value (GEBV), allow for the explanation of a certain proportion of the additive genetic variance dependent on the trait, and can be used as an additional information about the selection candidate for estimating breeding values. The application of genomic selection (GS) as the selection on the basis of GEBVs may lead to multistage selection schemes especially in dairy cattle, using GS as a preselection stage in order to reduce the number of test bulls in breeding schemes using progeny testing, or to replace this information source. A major problem of multistage selection is to choose the combination of stages and selection intensities maximising the genetic gain. Approaches of optimisation research may be applied, but since the selection indices of successive stages are correlated, multidimensional integration for deriving the selection intensities at selection stages is necessary, which might be unstable and time consuming according to the correlation structure and number of stages. The second part of this thesis compares the optimisation results of multistage breeding schemes regarding genomic selection, where two different approaches for deriving the selection intensity and the genetic gain are used and the accuracy and cost of GEBVs are varied. The first approach derives the stage dependent breeding values such that the correlation between stages is zero allowing for the calculation of the stage selection intensity via one dimensional integration and, therefore, a fast optimisation of breeding schemes containing even an unlimited number of selection stages. A disadvantage of this approach is a loss in variance of stage breeding values and the genetic gain. The second approach uses new developments for the integration of multivariate normal distributions and calculates an exact solution for the selection intensity and the genetic gain after a certain number of selection stages. The results clearly show that the integration algorithm is fast and stable enough to compare even a large number of possible breeding schemes. Furthermore, the loss in breeding value variance is unpredictable when using the decorrelated selection indices, and a proper consideration of the interaction between selection paths due to cost limitation and paths specific selection strategies will lead to illogical suggestions concerning the breeding scheme structure. As the accuracies and costs of GEBVs were varied in a certain range, the results also show that GS is competitive to conventional progeny testing in dairy cattle breeding even if the accuracy of GEBVs is decreased to 0.45. GS will increase the breeding costs linear due to the number of genotyped individuals. Thus, genotyping large proportions of a population might lead to uneconomical breeding schemes. This is especially the case for bull dam selection in dairy cattle breeding because the cow population size is equal to the number of potential selection candidates. Additionally, the number of selected bull dams is dictated by the demand for potential sires. Therefore, decreasing the number of genotyped selection candidates in order to fulfil economical limitations might lead to a very small selection intensity making the financial efforts difficult to justify concerning the genetic gain. A possible way out is the usage of inexpensive SNP chips containing only a minor number of SNPs for genotyping huge proportions of the selection candidates population and estimate less accurate GEBVs on this basis by using imputation algorithms. The third part of this thesis investigates multistage dairy cattle breeding schemes regrading the possibility of using a low density and high density SNP chip in each selection path. The costs of each chip and the accuracy of the subsequently estimated GEBVs were varied within a certain parameter space, where it was assured that the costs of the low density SNP chip and the subsequent accuracy of the GEBVs were always lower than those for the high density SNP chip. The results underline the potential of low density SNP chips for selecting bull dams from large cow populations, but also draw the attention to the non-linearity of the genetic gain as a function of the selection intensity. Thus, there exist combinations of cost and accuracies were it was found to be economical to limit the number of low density genotyped bull dams and include a further selection stage using high density SNP chips in that path. Furthermore, the results also show that the genetic gain is much more influenced by the cost and accuracy of the GEBV out of a high density chip, but the breeding scheme structure reacts more sensible to a change of this parameter concerning the low density chip.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The goal of breeding activities in commercial livestock populations is the increase of the mean of the genetically based performance capacity concerning one or numerous traits being summarised the in aggregate genotype via weighting factors if a certain breeding scheme is applied. Given the breeding scheme, the extent of this increase depends on the accuracy of breeding value estimation which is a function of a) the amount of information available about the selection candidates and its correlation structure to the aggregate genotype, and b) the usefulness of the statistical model in order to regress the genotype of the selection candidate on these available information. Recent molecular-genetical findings concern both, the amount of available information as well as the statistical model. The latter is affected by a ``genomic imprinting'' called mechanism, leading to an alteration of a genes effect on the phenotype of offspring due to a sex specific DNA methylation during gametogenesis in parents. Genomic imprinting can be regarded in breeding value estimation due to the calculation of two breeding values for each individual, one if it acts as sire and the other if it acts as a dam. Weighting factors to summarise this breeding values can be derived by an extension of the gene flow method. This extension is developed in the first part of this thesis and allows for tracing the flow of genes of a certain founder or group of founders within a population across tiers (e.g. nucleus, multiplier, production) and generations with special regard to the sex of the direct parent of an individual carrying these genes. Thus, it allows to assess the probability that a gene of a founder is inherited to its descendants via their direct sires or dams. The discounted and summarised trait realisations out of the genes inherited by the sire and the dam can be used as weighting coefficients for summarising the breeding values of an individual as a sire and a dam. The extended gene flow method is applied to a hypothetical pig breeding program showing that the weights for the breeding values as a dam and as a sire can differ according to the chosen breeding scheme and the planning horizon. Furthermore, it is shown that depending the breeding scheme the breeding value of a dam when acting as a sire might be weighted higher than when acting as a dam. Additionally, a possibility to predict the increase in inbreeding due to one round of selection inherent in the method is presented. The above mentioned amount of available information about a selection candidate is affected by the discovery of hundreds of thousands DNA markers in form of single nucleotide polymorphisms (SNP marker) being in strong linkage disequilibrium with neighbouring trait affecting genes or quantitative trait loci (QTL). The sum over all estimated marker effects on the phenotype, the genomically estimated breeding value (GEBV), allow for the explanation of a certain proportion of the additive genetic variance dependent on the trait, and can be used as an additional information about the selection candidate for estimating breeding values. The application of genomic selection (GS) as the selection on the basis of GEBVs may lead to multistage selection schemes especially in dairy cattle, using GS as a preselection stage in order to reduce the number of test bulls in breeding schemes using progeny testing, or to replace this information source. A major problem of multistage selection is to choose the combination of stages and selection intensities maximising the genetic gain. Approaches of optimisation research may be applied, but since the selection indices of successive stages are correlated, multidimensional integration for deriving the selection intensities at selection stages is necessary, which might be unstable and time consuming according to the correlation structure and number of stages. The second part of this thesis compares the optimisation results of multistage breeding schemes regarding genomic selection, where two different approaches for deriving the selection intensity and the genetic gain are used and the accuracy and cost of GEBVs are varied. The first approach derives the stage dependent breeding values such that the correlation between stages is zero allowing for the calculation of the stage selection intensity via one dimensional integration and, therefore, a fast optimisation of breeding schemes containing even an unlimited number of selection stages. A disadvantage of this approach is a loss in variance of stage breeding values and the genetic gain. The second approach uses new developments for the integration of multivariate normal distributions and calculates an exact solution for the selection intensity and the genetic gain after a certain number of selection stages. The results clearly show that the integration algorithm is fast and stable enough to compare even a large number of possible breeding schemes. Furthermore, the loss in breeding value variance is unpredictable when using the decorrelated selection indices, and a proper consideration of the interaction between selection paths due to cost limitation and paths specific selection strategies will lead to illogical suggestions concerning the breeding scheme structure. As the accuracies and costs of GEBVs were varied in a certain range, the results also show that GS is competitive to conventional progeny testing in dairy cattle breeding even if the accuracy of GEBVs is decreased to 0.45. GS will increase the breeding costs linear due to the number of genotyped individuals. Thus, genotyping large proportions of a population might lead to uneconomical breeding schemes. This is especially the case for bull dam selection in dairy cattle breeding because the cow population size is equal to the number of potential selection candidates. Additionally, the number of selected bull dams is dictated by the demand for potential sires. Therefore, decreasing the number of genotyped selection candidates in order to fulfil economical limitations might lead to a very small selection intensity making the financial efforts difficult to justify concerning the genetic gain. A possible way out is the usage of inexpensive SNP chips containing only a minor number of SNPs for genotyping huge proportions of the selection candidates population and estimate less accurate GEBVs on this basis by using imputation algorithms. The third part of this thesis investigates multistage dairy cattle breeding schemes regrading the possibility of using a low density and high density SNP chip in each selection path. The costs of each chip and the accuracy of the subsequently estimated GEBVs were varied within a certain parameter space, where it was assured that the costs of the low density SNP chip and the subsequent accuracy of the GEBVs were always lower than those for the high density SNP chip. The results underline the potential of low density SNP chips for selecting bull dams from large cow populations, but also draw the attention to the non-linearity of the genetic gain as a function of the selection intensity. Thus, there exist combinations of cost and accuracies were it was found to be economical to limit the number of low density genotyped bull dams and include a further selection stage using high density SNP chips in that path. Furthermore, the results also show that the genetic gain is much more influenced by the cost and accuracy of the GEBV out of a high density chip, but the breeding scheme structure reacts more sensible to a change of this parameter concerning the low density chip.
Key concepts: Engineering, Computer science