Better Data for Better Health

The Role of Whole Genome and Exome Sequencing in Uncovering Variants in Rare Diseases 

Written by By Dr. Ali M Tabish, PhD.

Rare diseases impact the lives of around 300 million people in the world with an estimated affected population frequency of 10%. With current standard clinical diagnostic practices, it takes a long time to diagnose rare diseases, and in some cases, up to 30 years. Approximately 80% of rare diseases are believed to have a genetic cause. 

The recent advancements in genomic and bioinformatic technologies have enabled the discovery of genetic etiologies in 20-30% of rare disease. Next-generation sequencing techniques (NGS) i.e., Exome Sequencing (ES) and Whole Genome Sequencing (WGS), are increasingly used in the study of rare genetic diseases. Studies have shown in the clinical perspective of rare diseases that with NGS techniques this is a ~40% diagnostic rate in rare genetic diseases compared to previously employed traditional genetic methodologies which yielded a 10% diagnostic rate. 

Navigating the Complexities of Genetic Data in Rare Disease Diagnosis

A single human genome contains 3 billion bases. Thus, studies of rare diseases by Exome Sequencing and Whole Genome Sequencing generate enormous amounts of genetic data from a single individual, much of which, in the context of a disease study, is not relevant. The complexity of genetic diagnosis is multi-layered. There is uncertainty on the number of coding genes in the human genome and their functional effect at a cellular level, which is further dwarfed by variants in noncoding DNA, the understanding of which is in its infancy. Genes exert their functions depending on the developmental stage, tissue specificity, and external and internal stimulus. Translating these complexities to disease level, a single genetic variant can have a myriad of effects on human health. 

 

Compared to single gene analysis, the interpretation of the large number of variants from exome sequencing and whole genome sequencing is obviously quite different. This requires not only in‐depth knowledge of the technique to assess the quality of data and identified variants but also new approaches for variant interpretation. Exome sequencing or whole genome sequencing from a patient undergoing genetic assessment will yield tens of thousands of variants. And with extensive data analysis to sift through these variants to find the causal variant for the disease, via filtering out common variants in the population using resources such as the Genome Aggregation Database, geneticists are still left with 100s of variants to review.  

 

Further ultra-fine analysis of remaining variants relies heavily on the phenotype of the referred patient and uses Human Phenotype Ontology (HPO) terms that match phenotype to gene to help with the analysis. However, even with these resources, a geneticist may still have way too many variants to manually look at one by one to find the causal variant for the disease. Although there are consistent improvements in technology platforms, reference data sets, software, and analysis pipelines, however, standards and guidelines for prioritizing genetic variant candidates likely to be relevant in disease are still in their infancy.

  

Overcoming Bottlenecks in Genetic Analysis for Rare Disease Diagnosis

The American College of Medical Genetics and Genomics (ACMG) has published standard guidelines to implement and standardize the identification and reporting of genetic variants in the clinical context. Since the individual’s genetic data do not fluctuate over the period of a lifetime, it thus provides the possibility to re-evaluate the genetic information of the patient in the context of recent knowledge, where a previous analysis did not yield genetic etiology for the patient disease.

The bottleneck here is that the primary analysis or re-evolution of the genetic data is time-consuming and requires dedicated personnel to accomplish the task, which even with all the advancement takes much time and effort. Not just that, there is a great risk for human errors even from the field experts owing to the complexity and enormity of the genetic data. Several large-scale pan-continental efforts have been initiated in order to unravel the previously undiagnosed rare genetic disorder e.g., solve-RD is a Horizon 2020-supported EU study bringing together more than 300 clinicians, scientists, and patient representatives from 51 sites in over 15 European countries. Solve-RD consortium is expected to analyze 19,000 exomes and genomes of genetically unsolved cases in several batches over 5 years. Such analysis or re-analysis of genetic data is error-prone.

In fact, a study has shown that nearly half of previously reported deleterious de-novo variants were in fact nearly neutral mutations, based on data from functional genomics studies. Another recent study collected variants of uncertain significance (VUS) from over 1.5 million sequencing test results from 19 clinical laboratories in North America from 2020 to 2021. This study concluded that high rates of VUS observed in diagnostic testing warrant examining current variant reporting practices. VUS poses major challenges in the interpretation of genetic results and current guidelines require an update in order to increase the diagnostic yield of the genetic testing. 

 

Current NGS bioinformatic pipelines and methods have been instrumental in detecting single nucleotide variants (SNV) within the coding region of genes and splice sites. However, there are several other types of variants such structural variants, intronic variants, uniparental disomies, mitochondrial variants and repeat expansion, where existing pipelines have severe technological limitations in order to reliably detect them, which otherwise could substantially increase the diagnostic yield of genetic testing. In many genetic cases where a recessive inheritance is expected, there could initially be only a single variant identified in a recessive gene. In such cases, the second variant may be a different type of mutation, may not meet the quality standards, or may seem less likely to be pathogenic; all these factors potentially lead to the exclusion of the second variant even in the eyes of an expert geneticist. This necessitates the development and implementation of new analytical algorithms in order not to miss such variants. Existing algorithms are not trained enough to analyze the chromosomal changes from the WES data, large copy number variants (CNVs) and mosaic nature of certain SNVs and CNVs which are filtered out in the bioinformatic pipeline owing to the low coverage or deemed as potential artifacts.  

AI Brings Promise to Incomplete Phenotypic Data That Hinders Genetic Analysis

Another common step in the bioinformatic pipeline is that the filtering eliminates all data with an allele frequency >1% or based on the frequency and inheritance patterns of the disease. A frequent pathogenic variant may be due to homopolymeric stretches or a founder variant in a population; however, current algorithms are not perfect in identifying such high-frequency pathogenic variants. Another practical problem faced by molecular geneticists is the incomplete phenotypic information of the patient, which significantly hinders the proper genetic work-up in the genetic laboratories. It has been reported that the genetic yield is proportional to clinical work and provides phenotype of the patients. However, patient phenotype might also evolve as the patient ages and the full phenotypic spectrum might not be available at the time of testing. This further necessitates building and training the analytical algorithms taking an unbiased genome-wide approach in prioritizing the genetic variants for potential clinical correlation.

There are several other practical issues e.g., pseudogene, gene transcripts, digenic and complex inheritance where existing bioinformatic pipelines do not sufficiently overcome the diagnostic hurdles, and in such scenarios, geneticists are prone to misreporting the genetic findings. With the recent advancement in the field of artificial intelligence (AI), it has become possible now to implement and train AI algorithms to overcome the above-discussed caveats in the interpretation and reporting of genetic results. 

Read an interesting artcile:

Epigenetics: Impact, Resources, and Technology in DNA Methylation Analysis by Eli Sward PhD » Geneyx

 

+

Selected Videos

Schedule Demo

Contact us to set a live demo


Contact Us

Whether you have general questions about our solutions or would like to schedule a demo or to suggest collaboration – our team is on hand for you.