Geneyx Analysis is a powerful diagnostic analysis software specifically
designed for in vitro laboratory developed tests (LDTs) within health
institutions. It serves as an automated solution for conducting
secondary and tertiary analysis of variants identified in gene panels,
exomes, and whole genomes, primarily generated through next-generation
sequencing (NGS) assays.
The primary objective of Geneyx Analysis is to facilitate the
identification and classification of causal variants that have clinical
significance. By leveraging advanced algorithms and annotation tools,
the software assists healthcare professionals and researchers in making
informed decisions related to targeted diagnoses and targeted therapy
applications in both hereditary disorders and cancer.
The software offers comprehensive features for variant analysis,
including filtering, annotation, variant classification, and report
generation. It streamlines the analysis process, providing a
standardized and efficient workflow for variant interpretation.
Geneyx Analysis is specifically developed for use by healthcare
professionals and researchers who possess the necessary expertise and
domain knowledge in genetics, genomics, and molecular diagnostics. It is
intended to be utilized within healthcare institutions for diagnostic
purposes, aiding in the accurate and timely identification of clinically
relevant variants.
It is worth noting that Geneyx Analysis is a specialized software
solution and should be used as part of a broader clinical workflow and
decision-making process. It is crucial to adhere to established
guidelines and best practices in genetic analysis, interpretation, and
reporting when utilizing the software.
For specific details on the features, functionalities, and regulatory
aspects of Geneyx Analysis, it is recommended to refer to the official
documentation provided by the software vendor or consult with their
representatives for a comprehensive understanding of its capabilities
and implementation requirements.
1.2 Indications for Use
Geneyx Analysis is indeed tailored for research and laboratory settings,
specifically for the analysis of human samples in the context of
diagnosing rare Mendelian diseases and providing prognostic, diagnostic,
and therapeutic guidance. The software is intended to be utilized by
trained medical professionals who have expertise in genetics and
genomics.
The primary purpose of Geneyx Analysis is to analyze and interpret NGS
(Next-Generation Sequencing) aligned reads and variant calls obtained
from sequencing data. It is not involved in the sample collection,
preparation for sequencing, or primary sequence analysis, such as
base-calling from raw instrument data. The software assumes that the
input data, including aligned reads and variant calls, have already been
generated through established laboratory protocols.
The analysis process within Geneyx Analysis involves applying
sophisticated algorithms, variant filtering, annotation, and
classification methods to identify potentially causative variants in the
context of rare Mendelian diseases. The generated interpretations and
reports are designed to be interpreted by healthcare professionals, such
as physicians, genetic counselors, or other trained medical experts. The
insights and recommendations provided by the software are intended to
guide clinical decision-making and facilitate patient management.
Geneyx Analysis can be deployed either on cloud infrastructures or
within the infrastructure of clinical laboratories, offering flexibility
in terms of implementation options. This allows the software to adapt to
different laboratory environments and data management strategies
It is important to note that Geneyx Analysis is a specialized tool that
should be used as part of a comprehensive clinical workflow and in
conjunction with other relevant clinical and laboratory information. The
software is not intended for direct use by patients, as its outputs
require interpretation and expertise from healthcare professionals.
For specific details regarding the technical requirements, deployment
options, and regulatory considerations of Geneyx Analysis, it is
recommended to consult the official documentation provided by the
software vendor or engage with their representatives to ensure proper
utilization and compliance within the intended laboratory and clinical
contexts.
1.3 Contraindications
There are no specific contraindications associated with Geneyx Analysis.
When utilizing Geneyx Analysis or any similar diagnostic analysis
software, it is crucial to adhere to standard protocols and best
practices in the field of genetics and genomics. This includes
appropriate validation and quality control processes, compliance with
data protection and privacy regulations, and the involvement of trained
medical professionals who are knowledgeable in genetic analysis and
interpretation.
1.4 Warnings
Geneyx Analysis is specifically designed for use in research and
clinical laboratory settings, and it is intended for the analysis of
genetic variants in the context of diagnosing rare Mendelian diseases
and providing prognostic, diagnostic, and therapeutic guidance. It is
primarily used by trained medical professionals, such as physicians,
geneticists, and genetic counselors, who have the necessary expertise to
interpret and analyze genetic data.
The software’s purpose is not for personal genomics or for diagnosing or
interpreting variants in healthy individuals. Its use is focused on
identifying and classifying causal variants in patients with suspected
genetic disorders or cancer, with the aim of supporting informed
decision-making in personalized medicine.
It is important to use Geneyx Analysis or any similar genetic analysis
software within its intended scope and in accordance with applicable
regulations, guidelines, and ethical considerations. This ensures that
the results and interpretations are accurately understood and
appropriately utilized by healthcare professionals for patient care.
1.5 Precautions
Important precautions to follow when implementing Geneyx Analysis or any
similar laboratory testing software. Customization, calibration, and
validation are crucial steps in adapting the software to the specific
test environment and ensuring accurate and reliable results. Here are
some further explanations and considerations for the precautions you
mentioned:
Customize and calibrate workflows: It is essential to customize the provided workflows (project protocols) to match the specific requirements of the genetic test being performed. This includes adjusting QC thresholds, analysis parameters, and incorporating test-specific annotations and quality filters. The customization should be based on comprehensive validation procedures that encompass the entire sample preparation, sequencing, secondary analysis, and tertiary analysis pipeline. By validating the complete workflow, you can ensure the software’s performance is optimized for accurate variant detection and interpretation.
Implement cybersecurity measures: To protect the integrity and confidentiality of the data processed by Geneyx Analysis, it is important to implement robust cybersecurity measures. This includes anti-malware software, system firewalls, and other security best practices to prevent cybersecurity breaches. Following cybersecurity guidance and staying up to date with industry standards can help safeguard against unauthorized access, data breaches, and potential risks associated with the software.
Secure system access: To maintain the security and privacy of the software and the data it processes, it is recommended to avoid running the software or server components on directly accessible internet-facing hosts. Instead, deploy them within a secure network environment with restricted access. Physical and system-level access should be limited to authorized individuals only, reducing the risk of unauthorized access through external networks or physical intrusions.
By adhering to these precautions, laboratories can mitigate potential
risks associated with implementing Geneyx Analysis or similar software
as part of their testing process. Following proper validation,
cybersecurity measures, and access control protocols helps ensure the
accuracy, security, and confidentiality of the genetic data and the
results generated for clinical and research purposes.
1.6 Description
Geneyx Analysis is an advanced decision support software and in vitro
medical device specifically designed to assist in the diagnosis of rare
diseases and provide valuable insights into the prognostic and
therapeutic implications of genomic biomarkers. It achieves this by
analyzing genomic mutations (variants) detected through Next Generation
Sequencing (NGS) instruments.
Laboratories can select from a range of NGS-based assay products to
sequence either the entire human genome or a targeted set of genes (gene
panels), exomes, or genomes. This sequencing process generates raw
variant calls for each patient. Geneyx Analysis seamlessly handles these
raw variant calls and offers the following key features:
Reading and processing input in standard VCF files, which contain both small variants (detectable using short-read NGS data) and CNVs (detected through coverage profile analysis of NGS data).
Annotating imported data with comprehensive information from numerous public and licensed databases, enhancing the understanding of variant characteristics.
Evaluating the quality of samples and variants through a robust set of metrics, derived statistics, and intuitive visualizations.
Utilizing advanced algorithms to conduct in silico analyses, providing valuable insights into the impact of variants on the patient’s genes and their respective functions.
Adhering to industry-standard guidelines for scoring and classifying germline variants and CNVs, employing guided workflows and auto-scoring algorithms.
Reviewing relevant clinical evidence and literature to assess whether a variant meets the criteria for reporting and further analysis.
Creating customizable clinical reports that include all selected variants, along with patient and sample-level data, facilitating effective communication of findings.
Storing variants from previously sequenced samples in a secure warehouse, enabling the computation of cohort-level statistics and per-variant frequencies for future analyses and research.
Although Geneyx Analysis comprises several modules catering to diverse
market needs, it is developed, deployed, and distributed as a unified
software solution, ensuring a seamless and comprehensive user
experience.
1.7 System Requirements
Geneyx Analysis offers multiple Cloud installations on the Azure Cloud
platform, providing users with a scalable and reliable infrastructure.
When deployed on Azure Cloud, the following components are utilized:
Web Apps: Geneyx Analysis leverages Azure Web Apps to provide a robust and accessible user interface. This enables users to access the software securely through web browsers without the need for complex local installations.
Storage Account: Azure Storage Account is utilized to store and manage data associated with Geneyx Analysis. This includes input files, variant data, annotations, and other relevant information. The storage account ensures efficient data management and retrieval.
Batch Service: Azure Batch Service is employed to handle computationally intensive tasks within Geneyx Analysis. This service enables efficient parallel processing and scaling of computational workloads, optimizing the performance of data analysis and variant processing.
Azure SQL Hyperscale DB: Geneyx Analysis utilizes Azure SQL Hyperscale as the database solution to store and manage structured data. This highly scalable and fully managed database service ensures reliable data storage, retrieval, and query processing for efficient analysis.
VMs: Virtual Machines (VMs) on Azure Cloud are utilized to provide the necessary computing resources for running Geneyx Analysis. These VMs host the software and handle the computational tasks required for variant analysis, annotation, and interpretation.
By leveraging the capabilities of Azure Cloud, Geneyx Analysis benefits
from a robust and scalable infrastructure that ensures efficient data
management, computational power, and storage. The use of Azure Cloud
allows for flexibility, scalability, and automated deployment of
updates, providing users with a reliable and up-to-date environment for
their genetic analysis needs.
1.8 Cybersecurity Guidance
Establishing a secure network and system environment is crucial to
safeguard the sample, genomic data, and patient information processed by
Geneyx Analysis. To minimize cybersecurity risks and ensure compliance
with privacy regulations like GDPR, the following best practices are
recommended:
Install Institutional Firewalls: Implement institutional firewalls to restrict unauthorized incoming connections to the hosts running Geneyx Analysis. This helps protect the system from external threats and unauthorized access.
Use Anti-Malware and Endpoint Security Software: Install and regularly update anti-malware and endpoint security software to prevent system-level exploits that could grant access to the locally installed software and data. This helps protect against malware and other security threats.
Regularly Backup Data: Back up all input, project, and user data to non-network attached mediums. This ensures that data can be recovered in the event of ransomware attacks or data loss incidents. Regular backups help maintain data integrity and availability.
Implement User Authentication and Access Controls: Provide system-level user authentication with a strong password policy and consider implementing two-factor authentication for an added layer of security. Restrict access to the installed software and data only to authorized personnel.
Ensure Physical Security: Implement physical security measures to prevent unauthorized access to workstations and servers running Geneyx Analysis. This includes securing the physical premises, implementing access control systems, and enforcing policies for workstation auto-locking after periods of inactivity.
GDPR Compliance: Adhere to GDPR privacy rights when using Geneyx Analysis. Create separate analyses for each sample to facilitate precise compliance with requests to delete personal data. Update or modify personal data within the analysis itself and generate new reports when necessary. It’s important to understand and implement the full range of responsibilities, policies, and procedures required by GDPR to ensure compliance. Consult with network administrators, IT specialists, or legal experts to ensure adherence to institutional policies and GDPR requirements.
By following these best practices, laboratories can establish a secure
and compliant environment for using Geneyx Analysis, safeguarding
sensitive data and protecting the privacy rights of individuals.
1.9 Maintenance and Upgrade Guidance
Geneyx Analysis is designed to undergo validation in coordination with
the cycle of re-validation of laboratory processes, ensuring its
reliability and accuracy. When using the software, it is important to
consider the following points:
Stay Updated with Release Notes: Regularly monitor notifications of release notes and known issues to stay informed about software updates. This helps ensure that any new features or bug fixes do not impact production workflows. Keeping track of updates allows for timely implementation of improvements and bug fixes.
Update Annotations: Annotations used during analysis can be updated by creating a copy of the analysis. This enables users to stay up to date with monthly updates of annotation sources like ClinVar and OMIM. However, it is essential to evaluate the impact of these changes on the specific test being performed. Consider how updated annotations may affect the interpretation and reporting of variants.
Re-validation of Workflows: When choosing to update the software, it is important to re-validate all production workflows. Use the same benchmark samples that were initially used for calibration and validation. By comparing the results with expected changes, such as updated annotations, algorithm improvements, or changes in the laboratory-specific knowledge base, you can ensure the reliability and accuracy of the analysis.
System Monitoring and Maintenance: Implement standard system monitoring and maintenance practices for the hosts running Geneyx Analysis. Regularly monitor volumes containing user data, shared application data, and performance metrics. This helps detect any issues or degraded performance in a timely manner. Additionally, maintain and update security measures and system software patches to minimize the risk of exploitation.
By following these guidelines, laboratories can ensure that Geneyx
Analysis remains updated, validated, and secure, enabling reliable and
accurate analysis of genomic data.
1.9.1 Clinical Performance & Risk/Benefit Information
Geneyx Analysis has undergone clinical performance evaluation studies
conducted with users in their own laboratory settings. These studies
have shown that Geneyx Analysis offers several benefits over manual
variant interpretation and reporting, including improved efficiency and
decision support. Some key points to consider are:
Improved Efficiency: The use of Geneyx Analysis automates the analysis process, resulting in improved efficiency and reduced time to completion compared to manual variant interpretation. This allows laboratory professionals to analyze and interpret variants more quickly, facilitating faster turnaround times for test results.
Workflow and Decision Support: Geneyx Analysis provides predefined workflows and decision support tools, which guide users through the analysis process and assist in variant scoring. These workflows help streamline the interpretation process and ensure consistent and standardized analysis across different users and cases.
Maintaining Quality: The clinical performance evaluation studies have demonstrated that the use of Geneyx Analysis does not compromise the quality of variant interpretation. The studies have not identified any significant drop in quality or unexplained discordance in outcomes when compared to manual variant interpretation. This indicates that Geneyx Analysis can maintain high-quality results while improving efficiency.
Risk-Benefit Assessment: It is important to consider the overall risk-benefit ratio when adopting automated software like Geneyx Analysis. The demonstrated improvements in efficiency, coupled with the maintained quality of variant interpretation, outweigh the potential risks associated with using automated software.
By leveraging Geneyx Analysis, laboratories can benefit from improved
efficiency, standardized workflows, and reliable decision support,
ultimately enhancing their clinical variant interpretation processes.
1.9.2 Storage Specifications
FASTQ + BAM files:
Our FASTQ and BAM files are securely stored on AWS S3 in the Ireland datacenter. As part of AWS’s redundancy and disaster recovery policy, the data is persisted in three availability zones, ensuring data redundancy and high availability. It provides an extra layer of protection for your files against potential failures or outages. Additionally, please note that FASTQ files are archived three days after being uploaded, ensuring efficient data management practices.
VCF files:
For VCF files, we utilize Azure’s Netherlands datacenter. The files are stored in a read-access geo-redundant storage (RA-GRS) system across two availability zones. This setup ensures data replication and redundancy, maximizing data availability and durability. Moreover, the storage system maintains six days’ worth of restore points, enabling you to access previous versions of the VCF files if needed.
Database:
Our database is hosted in Azure’s Netherlands datacenter, providing a robust and reliable infrastructure for data storage. The database is equipped with built-in resilience measures, including seven days’ worth of restore points. This ensures that in case of any data-related issues or accidental modifications, we can restore the database to a previous state within the specified timeframe.
By leveraging these cloud providers and their respective datacenters, we
can ensure the security, redundancy, and disaster recovery capabilities
necessary to safeguard your data. If you have any further questions or
concerns about our data storage and backup practices, please feel free
to let us know.
1.10 Technical Support and Turnaround Times
1. Technical Support Availability
Our technical support team is available to assist you during our regular business hours, Monday to Friday, from 9:00 AM to 5:00 PM (local time).
Support requests can be submitted through our designated channels, including email, or our online support portal. The online support portal can be accessed by clicking on the Help icon in the upper right corner. This will connect you with a live representative or will be addressed via email.
Upon receiving a support request, our team will provide an initial response within 24 hours during business days.
For requests submitted outside of business hours or on weekends, the initial response will be provided on the next business day.
3. Issue Resolution Timeframes
We strive to resolve technical issues as quickly as possible, based on their severity and complexity.
The resolution timeframes are categorized into four priority levels:
i. Priority 1 (P1): Critical
Definition: Critical system failure or significant impact on business operations.
Response Time: Within 2 hours.
Resolution Time: Within 24 hours or as mutually agreed upon based on the nature of the issue.
ii. Priority 2 (P2): High
Definition: Major issue affecting system functionality or performance.
Response Time: Within 4 hours.
Resolution Time: Within 48 hours or as mutually agreed upon based on the complexity of the issue.
iii. Priority 3 (P3): Medium
Definition: Issue causing moderate impact on system functionality or performance.
Response Time: Within 8 hours.
Resolution Time: Within 5 business days or as mutually agreed upon based on the nature of the issue.
iv. Priority 4 (P4): Low
Definition: Minor issue or general inquiries.
Response Time: Within 24 hours.
Resolution Time: Within 10 business days or as mutually agreed upon based on the nature of the issue.
4. Escalation Process
If you are not satisfied with the progress or resolution of your support request, you have the option to escalate the issue.
Escalation contacts and procedures will be provided upon request.
Please note that the above turnaround times are general guidelines and
may vary depending on the specific circumstances and agreement. Our goal
is to provide timely and effective technical support to ensure the
smooth operation of your systems and minimize any disruptions to your
business.
2.0 Introduction to Geneyx Analysis
Geneyx Analysis is a powerful diagnostic analysis software designed for
in vitro laboratory developed testing. It serves as an essential tool
for automating the analysis of genetic variants detected in gene panels,
exomes, and whole genomes obtained through next-generation sequencing
(NGS) assays. The primary goal of this analysis is to accurately
identify and classify causal variants, enabling healthcare professionals
to make well-informed decisions, provide precise diagnoses, and
facilitate targeted therapy applications in hereditary disorders.
To cater to the diverse requirements of the clinical testing market, an
efficient workflow engine is essential. Geneyx Analysis excels in this
aspect by supporting multiple guidelines and accommodating various
levels of laboratory-specific customizations within a single solution.
This flexibility ensures that the software can adapt to the specific
needs and preferences of different laboratories, enabling them to
streamline their analysis processes effectively.
One of the key strengths of Geneyx Analysis lies in its automated
workflow, which incorporates the guidelines set forth by the American
College of Medical Genetics and Genomics (ACMG) and the Association for
Molecular Pathology (AMP) for interpreting NGS results. By automating
the interpretation process, the software significantly reduces the
complexities and challenges associated with manual interpretation. This
standardized and repeatable approach to variant interpretation in a
clinical context can be further customized to align with the practices
and protocols of individual laboratories.
In summary, Geneyx Analysis offers a comprehensive solution for genetic
variant analysis in a clinical setting. Its advanced workflow engine,
adherence to industry guidelines, and ability to accommodate
laboratory-specific customizations make it an invaluable tool for
clinicians and geneticists, providing them with accurate and efficient
variant interpretation to support patient care and decision-making.
2.1 QuickStart Guide
2.1.1 Upload Local Fastq Files
Geneyx Analysis simplifies the process of implementing a secondary
pipeline for converting fastq files to vcf (variant call format) through
its user-friendly interface, eliminating the need for command line
expertise. This feature is available as part of the licensed package,
and if you wish to add it to your license, you can send an email request
to support@geneyx.com.
In addition to its pipeline capabilities, Geneyx Analysis offers the
convenience of automatically retrieving fastq files from a cloud
infrastructure, known as the 3.10.6 Data Sources. This enables seamless
integration with cloud-based storage solutions for effortless batch
uploads. You can find detailed information on this feature in the
platform’s documentation.
For the purpose of explaining the process of importing a single sample
from a local directory, follow these steps:
Navigate to the Data Management view within Geneyx Analysis.
Hover over the “Seq. Samples” section, and a menu will appear.
Select “Upload Single Sample” from the menu.
By choosing this option, you can initiate the upload process for fastq
files associated with a single sample. Geneyx Analysis provides a
straightforward and streamlined experience for importing and managing
your sequencing data, empowering you to effortlessly analyze genetic
variants with ease and efficiency.
Upload Single Sample
After initiating the upload process for a single sample in Geneyx
Analysis, the next dialog will prompt you to enter a Subject ID. The
Subject ID serves as an identifier that is applied at the subject level,
allowing you to uniquely identify and track specific individuals or
samples within the analysis.
To proceed with the upload, follow these steps:
Enter the desired Subject ID in the provided field. Choose an identifier that is meaningful and helps you easily identify the sample associated with it. It could be a patient ID, sample code, or any other relevant identifier.
Once you have entered the Subject ID, click on the “Next” button to proceed to the next step in the upload process.
By providing a Subject ID, you ensure that the uploaded sample is
associated with a specific identifier within the Geneyx Analysis
platform, facilitating efficient tracking and management of the data
throughout the analysis workflow.
Entering Fastq Subject ID
In the subsequent dialog of the upload process, you will be prompted to
enter the Serial Number, Expected # Fastq Files, and the Sequencing
Target. These details help Geneyx Analysis accurately process and
analyze the uploaded fastq files. Here’s how you can provide the
required information:
Serial Number: The Serial Number refers to the prefix of the fastq file names for Illumina FASTQ. In this example, the fastq files are named MQ21111B_005.R1.fq.gz and MQ21111B_005.R1.fq.gz. In this case, the Serial Number would be “MQ21111B.”
For MGI FASTQ, the Serial Number refers to the sample ID: In this
example, the FASTQ files are named V450272805_L01_D12345_1.fq.gz. In
this case, the serial Number would be “D12345”.
Enter the appropriate Serial Number in the provided field.
Expected # Fastq Files: Specify the number of fastq files associated with the sample you are uploading. In your example, there are two fastq files. Enter “2” as the Expected # Fastq Files.
Sequencing Target: The Sequencing Target refers to the specific genomic regions or targets that were sequenced. If you are working with gene panel or exome data, the Enrichment Kit & SMART Filtering option allows you to define your target capture regions. However, if the sample represents whole genome data, you can leave the default setting as it is.
Enrichment Kit & SMART Filtering (for gene panel and exome data): If you are analyzing gene panel or exome data, you can specify the Enrichment Kit & SMART Filtering options to define your target capture regions. This helps in focusing the analysis on specific genomic regions of interest. If this is not applicable to your data, you can disregard this field.
Once you have entered the required information, click on the “Next”
button to proceed with the upload process. These details will assist
Geneyx Analysis in correctly processing and analyzing the fastq files
associated with the sample.
Uploading sample information
After entering the necessary sample information, click on the “Next”
button to proceed. In the following dialog, you will have the option to
select “Upload Local Files.” This selection enables you to browse and
select the samples from a local directory on your computer.
Here’s how you can upload the samples from a local directory:
Click on the “Upload Local Files” option.
A file browser window will appear, allowing you to navigate to the directory where your samples are located.
Browse through your local directories and select the relevant sample files that you want to upload to Geneyx Analysis. You can select multiple files at once by holding down the Ctrl key (or Command key on Mac) while clicking on the files.
Once you have selected the sample files, click on the “Open” or “Choose” button (depending on your operating system) in the file browser window. This will initiate the upload process.
The selected sample files will be uploaded to Geneyx Analysis.
After the upload is complete, click on the “Next” button to proceed.
By following these steps, you can upload your samples from a local
directory and move forward with the analysis
process.
Upload local fastq files
In the final step, you will need to select the Secondary Pipeline for
your analysis. Geneyx Analysis offers several options including:
“Dragen – hg19” and “Dragen – hg38.” As well as Sentieon hg19 and
Sentieon hg38 (NB. Sentieon is for SNV only)
The choice of Secondary Pipeline depends on the reference genome version
you want to use for your analysis.
Choose the appropriate Secondary Pipeline based on your specific
analysis requirements and reference genome preference. Once you have
made your selection, proceed to the next step. continue with the
analysis process.
Selecting the secondary pipeline
After clicking “Save,” the system will populate with details related to
the sequencing process. This information can be accessed through various
sections:
Processing Tasks: The Processing Tasks section provides insights into the progress of your sample within the processing pipeline. You can monitor the different stages of analysis, such as alignment, variant calling, and annotation. This allows you to track the status of your sample and observe any potential issues or delays.
QC Data: The QC Data section allows you to view sample statistics and coverage profiles. It provides valuable information about the quality of the sequencing data and the depth of coverage across the target regions. This data helps in assessing the reliability and accuracy of the analysis results.
Additional QC Metrics: By utilizing the Geneyx Analysis APIs (Application Programming Interfaces), you can extract additional QC metrics beyond what is displayed in the QC Data section. These APIs provide programmatic access to the data, allowing you to retrieve specific metrics or integrate them into other workflows or systems.
VCF Samples: Once the variant calling and alignment processes have been completed, the VCF Samples section will be populated with the relevant files. These files contain the variant information, such as the genomic variants identified in the sample and their associated annotations. You can access and explore these files to perform downstream analysis or interpretation of the genetic variants.
These sections provide valuable insights and data related to the
sequencing and analysis process, allowing you to monitor progress,
assess data quality, and access variant information for further analysis
and interpretation.
Summary table of the secondary pipeline.
Once the analysis is completed, the output will include several files
that are relevant for tertiary analysis:
SNV (Single Nucleotide Variant) File: This file contains information about single nucleotide variants, also known as point mutations, identified in the sample. It provides details such as the genomic position, reference and alternate alleles, quality scores, and annotations for each variant.
SV (Structural Variant) File: This file contains information about structural variants, which are larger genomic alterations such as insertions, deletions, duplications, inversions, or translocations. It provides details about the structural variants detected in the sample, including their genomic coordinates, size, and other relevant annotations.
BAM (Binary Alignment Map) File: This file represents the alignment of the sequencing reads to a reference genome. It contains the aligned reads, along with their qualities and mapping information, allowing for visualization and further analysis of the sequencing data.
BAI (BAM Index) File: This is the index file associated with the BAM file. It enables efficient random access to specific regions of the BAM file, facilitating quick retrieval and analysis of specific genomic regions.
To access these files, you can either click on the corresponding file
name under the VCF Samples section, which will open a preview or allow
you to download the file, or you can navigate to the VCF Samples section
in the navigation pane on the left-hand side of the interface. From
there, you can select the desired file and access or download it as
needed.
These files serve as the foundation for tertiary analysis, where further
interpretation, filtering, and annotation of the variants can be
performed using specialized software or
tools.
VCF Sample output from the secondary pipeline
2.2 Apply VCF from fastq to Protocol
By importing the output files into a protocol for tertiary analysis, you
can perform advanced analysis and interpretation of the variants,
including filtering, annotation, prioritization, and identification of
potential disease-causing variants. These steps allows for a deeper
exploration of the genomic data and aids in making informed decisions,
diagnoses, and targeted therapy applications in the context of
hereditary disorders.
To apply the output files from the previous analysis to a protocol for
tertiary analysis, follow these steps:
Navigate to the Dashboard view in the Geneyx Analysis interface.
Locate the VCF sample that was generated under the VCF Samples section. Next to the sample, you should see an option labeled “New Analysis.” Click on this option to initiate the import process for tertiary analysis.
A dialog or pop-up window will appear, guiding you through the import process. This dialog will allow you to configure the settings and parameters for the tertiary analysis.
Follow the prompts and provide the necessary information in each step of the import process. This will include selecting the appropriate protocol, and setting any additional analysis parameters or options.
Selecting New Analysis for the output VCF created by the secondary
After selecting “New Analysis,” you will be presented with a dialog
displaying the available protocols to choose from. A protocol serves as
a template for analysis, containing predefined filters and annotations.
Default protocols are provided in the categories of Germline, Somatic,
and Health Screening. Here are some of the default protocols available:
Germline:
Single Sample Analysis: This workflow is used to analyze affected germline samples that do not have associated maternal or paternal VCF files.
Trio Analysis: The trio workflow utilizes inheritance models to analyze rare Mendelian variants using mother, father, and proband samples.
Single Sample SNV/SV: This template is used to analyze affected germline samples for single nucleotide polymorphisms and structural variations.
CNV Only Analysis: The CNV workflow is designed to detect large structural variations using a specific file format that includes information such as chromosome, start, end, effect, score, and copy number.
Mitochondria Analysis: This protocol is specifically tailored for the analysis of mitochondrial variants.
Somatic:
Tumor Only Analysis: This template is used to analyze tumor-only samples without paired normal/germline samples.
Tumor Normal Analysis: This protocol is commonly used to analyze somatic samples along with their corresponding germline counterparts.
Cancer Screening: This template is used to detect the presence of somatic variants that are not isolated from tumor tissue.
Health Screening:
Carrier Screening: This protocol is used for the analysis of variants that could be inherited and passed on as genetic conditions to offspring
When selecting a protocol, consider the specific type of analysis you
wish to perform based on your sample and data. These predefined
protocols provide a starting point for analysis, but they can also be
customized to fit your specific requirements using the 3.10.9 Protocols
feature.
Choose the appropriate protocol that aligns with your analysis goals,
and then proceed to the next step of the analysis workflow.
Selecting protocol
Once you have selected a protocol, you can enter the subject information
in the Subject dialog. This information helps provide context to the
analysis and is available throughout the process for reference. It will
also be automatically included in the clinical report. When entering
details for a new subject, you can provide the following information:
SUBJECT ID (required): This is a unique identifier for the subject. It can be the patient identifier in the EMR or LIMS system, or any other identifier that you choose.
NAME (optional): The name of the subject.
DATE OF BIRTH (optional): The date of birth of the subject.
GENDER (optional): Specify the gender as Male or Female. Leave blank if the gender is unknown or unspecified.
CONSENT – PERSONAL DATA (optional): Specify whether the subject has given consent for the use of their personal data, with options such as Yes or No. Leave blank if consent status is unknown or unspecified.
CONSENT – CLINICAL DATA (optional): Specify whether the subject has given consent for the use of their clinical data, with options such as Yes or No. Leave blank if consent status is unknown or unspecified.
CONSANGUINITY (optional): Use this field to indicate parental consanguinity, particularly in cases of rare disease analysis.
ETHNICITY (optional): Enter the ethnicity of the subject.
PATERNAL ANCESTRY (optional): Specify the ancestry of the subject’s father.
MATERNAL ANCESTRY (optional): Specify the ancestry of the subject’s mother.
FAMILY HISTORY (optional): This is a text box where you can add notes regarding the family history of the subject.
By providing these subject details, you enhance the analysis with
important contextual information. Once entered, the information will be
associated with the subject throughout the analysis process and can be
accessed as needed.
Entering in Subject information
After clicking “Next” on the Subject dialog, you will be taken to the
Samples page. Since you started a new analysis from an existing sample,
the associated information is automatically populated. However, if you
need to update the VCF information, you can do so by clicking on the
edit icon next to the VCF sample. The available fields for editing
include:
Serial Number: Specify a unique identifier for this analysis. By default, a value is generated, but you can replace it with your own identifier, such as the ID of the case in your system.
RELATION: Specify the relation of the sample to the proband (e.g., mother, father, sibling, etc.).
Sequencing Machine: Specify the machine used for automating the DNA sequencing process.
Enrichment Kit & Filtering: Define the target capture regions using an enrichment kit. You can also implement your own BED files in the Settings dialog.
Genome Build: Select the reference genome assembly build used for alignment and mapping.
Taken Date: Enter the date when the sample was obtained.
Sequence Date: Enter the date when the sample was sequenced.
Received Date: Enter the date when the sequencing data was received.
Sequencing Target: Specify the NGS data type (e.g., gene panel, exome, whole genome).
Sample Source: Describe the method by which the DNA sample was collected.
Notes: Add any additional comments or notes related to the sample.
BAM file URL: Provide the URL of the BAM file associated with the sample.
files must be bgzipped with corresponding index file (tbi)
eg. xxx.bam.gz and xxx.bam.gz.tbi
Methylation file URL: Provide the URL of the methylation file, if applicable.
files must be bgzipped with corresponding index file (tbi)
xxx.bed.gz and xxx.bed.gz.tbi
Exclude sample from local allele frequency: Choose whether to exclude the sample from local allele frequency calculations. In the event where artifact or low quality samples are uploaded, this feature will exclude the sample from the local allele frequency calculation. This will prevent inflation of local variant allele frequencies.
By updating these fields, you can ensure that the analysis is performed
accurately and with the relevant sample information.
Sample being pulled automatically
After clicking “Next” on the Samples page, you will reach the Clinical
Information section. This section allows you to provide additional
context about the analysis and subject. Here, you can set the phenotype
terms that will be used to score and rank candidate variants during the
analysis stage. The following information can be entered:
Serial Number (required): Specify a unique identifier for this analysis. By default, a value is automatically generated, but you can replace it with your own identifier, such as the ID of the case in your system.
Name (optional): Enter the name of the protocol used for the analysis.
Owner (Doctor) (optional): Specify the name of the referring clinician or the owner of the analysis.
Description (optional): Use free text to describe the clinical aspects of the case, providing any relevant information.
Disease frequency (required): Enter the expected frequency of this disease occurring in the population. This information will be applied in the filtering process to reduce the number of observed variants above the defined threshold.
Department: Specify the name of the department associated with the analysis.
Phenotypes (optional): Enter any biological terms (phenotypes) that will be used to score and rank the candidate variants. You can enter multiple phenotypes using the Advanced option.
To ensure compound terms are recognized properly, use quotation marks around the term. For example, if the phenotype is “Developmental Delay,” enter it exactly as “Developmental Delay” to ensure the entire term is used.
After entering the phenotypes, click on Keywords. This will display which terms are highly specific, indicating how many genes are associated with the given condition. The HPO option allows you to enter a list of HPO terms and select all to import.
By providing this clinical information, you can enhance the analysis
process and improve the accuracy of candidate variant ranking based on
the specified phenotypes.
Entering in Clinical Information
After clicking “Save” on the Clinical Information page, the VCF sample
will be applied to the selected protocol, incorporating the information
entered during the import process. This action will lead you to the
Analysis Details page, where you can access and review the details of
the analysis. The Analysis Details page provides comprehensive
information about the analysis, including the applied protocol, subject
information, sample details, and any additional clinical information
provided. It serves as a hub for monitoring and managing the analysis
process, allowing you to track the progress and view the results of the
analysis. Further information about the (2.5. Analysis Details page can
be found in the section dedicated to it in the manual.
2.3 Upload VCF to Protocol
One the VCF file has been obtained the next step in using Geneyx
Analysis is creating an analysis. In Geneyx, an analysis is a mechanism
for analyzing a set of data from one or more samples using a template
for filtering and annotation to classification and reporting.
To initiate the process of creating a new analysis in Geneyx Analysis,
you can follow these steps:
Access the Geneyx Dashboard by navigating to the platform’s interface.
Look for the option labeled “Click here” or a similar indication that prompts you to analyze your own data. This option is typically located prominently on the Dashboard page. If samples are already loaded, you can select Start New Analysis.
Click on the provided link or button to proceed with creating a new analysis.
By following these steps, you will initiate the process of setting up an
analysis in Geneyx Analysis, allowing you to utilize the platform’s
filtering, annotation, classification, and reporting capabilities on
your data.
When you open the Analyses dialog in Geneyx Analysis, you will have the
option to choose a protocol for your analysis. A protocol serves as a
template that includes predefined filters and annotations to guide the
analysis process. Geneyx provides default protocols that are categorized
into Germline, Somatic, and Health Screening. Here are some examples of
the default protocols available:
Germline:
Single Sample Analysis: This protocol is used to analyze affected germline samples that do not have associated maternal or paternal VCF files.
Trio Analysis: The trio workflow is utilized for analyzing rare Mendelian variants by incorporating mother, father, and proband samples.
Single Sample SNV/SV: This template is specifically designed for analyzing affected germline samples to detect single nucleotide polymorphisms and structural variations.
CNV Only Analysis: The CNV workflow focuses on detecting large structural variations that follow a specific file format, including information such as chromosome (chr), start and end positions, effect, score, and copy number.
– Mitochondria Analysis: This protocol is tailored for the analysis of
mitochondrial variants.
Somatic:
Tumor Only Analysis: This template is used to analyze tumor-only samples that do not have a paired normal or germline sample for comparison.
Tumor Normal Analysis: The tumor normal workflow is commonly employed to analyze somatic samples along with their corresponding germline or normal samples.
Cancer Screening: This protocol is designed to detect the presence of somatic variants that are not isolated from tumor tissue.
Health Screening:
Carrier Screening: This template is used for the analysis of variants that have the potential to be passed on as inherited genetic conditions to offspring during carrier screening.
By selecting the appropriate protocol based on your specific analysis
requirements, you can streamline the analysis process and focus on the
relevant aspects of your
data.
When creating a new analysis in Geneyx Analysis, after selecting the
protocol, you will be prompted to enter the subject information in the
Subject dialog. This information is important for reference throughout
the analysis process and will also be included in the generated clinical
report. Here are the details that can be entered for a new subject:
SUBJECT ID (required): This field requires a unique identifier for the subject. Typically, it would be the patient identifier used in the EMR (Electronic Medical Record) or LIMS (Laboratory Information Management System) system. However, it can be any identifier that uniquely represents the subject.
NAME (optional): You can enter the name of the subject if available.
DATE OF BIRTH (optional): This field allows you to provide the date of birth of the subject if known.
GENDER (optional): You can specify the gender of the subject as Male or Female. If the gender is unknown or unspecified, you can leave this field blank.
CONSENT – PERSONAL DATA (optional): This field allows you to indicate whether the subject has given consent for the use of their personal data. You can select Yes or No, or leave it blank if the consent status is unknown or unspecified.
CONSENT – CLINICAL DATA (optional): Similarly, you can specify whether the subject has given consent for the use of their clinical data. You can select Yes or No, or leave it blank if the consent status is unknown or unspecified.
CONSANGUINITY (optional): This field is commonly used to specify parental consanguinity in cases of rare disease analysis. You can enter relevant information if applicable.
ETHNICITY (optional): Here, you can provide the ethnicity of the subject if known.
PATERNAL ANCESTRY (optional): You can enter the ancestry or ethnic background of the subject’s father.
MATERNAL ANCESTRY (optional): Similarly, you can enter the ancestry or ethnic background of the subject’s mother
FAMILY HISTORY (optional): This text box allows you to add any relevant notes or details regarding the family history of the subject.
By entering these details, you can provide important context for the
analysis and ensure that the generated clinical report includes accurate
and comprehensive information about the subject.
Entering Subject Information
After saving the subject information, the next step in Geneyx Analysis
is to upload the VCF files. In the SAMPLES dialog, you can click on the
“Select Sample” button, which will allow you to browse your local
directory to locate and upload the VCF file. Once the VCF file is
uploaded, you can define additional parameters specific to the sample.
Here are the parameters that can be configured:
Serial Number: This field allows you to specify a unique identifier for this analysis. By default, a value is automatically generated, but you can replace it with your own identifier, such as the ID of the case in your system.
RELATION: This field indicates the relation of the sample to the proband (the subject of the analysis).
Sequencing Machine: You can specify the machine used to automate the DNA sequencing process for this sample.
Enrichment Kit & Filtering: This field defines the target capture regions and allows you to specify the enrichment kit used. You can also implement your own BED files in the Settings dialog.
Genome Build: Specify the reference genome assembly build that was used for alignment and mapping of the sequencing data.
Taken Date: Enter the date on which the sample was obtained.
Sequence Date: Provide the date on which the sample was sequenced.
Received Date: Enter the date on which the sequencing data was received.
Sequencing Target: Specify the type of NGS (Next-Generation Sequencing) data for this sample.
Sample Source: Indicate the method by which the DNA was collected for this sample.
Notes: This field allows you to add any additional comments or notes relevant to this sample.
BAM file URL: If available, you can provide the URL of the BAM file associated with this sample.
Users can also provide http address for CRAM files and Geneyx will support visualization in IGV.
files must be bgzipped with corresponding index file (tbi)
eg. xxx.bam.gz and xxx.bam.gz.tbi
Methylation file URL: If applicable, you can enter the URL of the methylation file associated with this sample.
files must be bgzipped with corresponding index file (tbi)
xxx.bed.gz and xxx.bed.gz.tbi
Exclude sample from local allele frequency: This option allows you to exclude the sample from local allele frequency calculations if desired.
By providing these parameters, you can ensure that the analysis is
performed accurately and that the specific details of the sample are
taken into account during the analysis process.
Uploading VCF files
In the Clinical Information section of Geneyx Analysis, you can
incorporate additional context about the analysis and the subject being
analyzed. This section allows you to set phenotype terms that will be
used to score and rank candidate variants during the analysis stage.
Here are the components of the Clinical Information section:
Serial Number (required): Specify a unique identifier for this analysis. A default value is automatically generated, but you can replace it with your own identifier, such as the ID of the case in your system.
Name (optional): You can provide a name or title for the analysis, which helps in identifying and organizing the analysis.
Owner (Doctor) (optional): Enter the name of the referring clinician or the doctor responsible for the analysis.
Description (optional): This field allows you to provide free text to describe the clinical aspects of the case, providing additional context and information.
Disease frequency (required): Specify the expected frequency of the disease occurring in the population. This information will be applied as a filter to reduce the number of variants observed above the defined threshold.
Department: Enter the name of the department associated with the analysis.
Phenotypes (optional): Here, you can enter any biological term (phenotype) that will be used to score and rank the candidate variants during the analysis. You can enter multiple phenotypes, and for compound terms to be recognized properly, you should enclose them in quotation marks. For example, if you want to enter the term “Developmental Delay,” you should type it as “Developmental Delay” to ensure the entire term is used.
Keywords: After entering the phenotypes, you can select the “Keywords” option, which will display which terms are highly specific and reflect how many genes are associated with the given condition. This information can help in refining the analysis and focusing on specific aspects.
HPO (Human Phenotype Ontology) option: If you have a list of HPO terms, you can enter them here and select all to import them.
Entering Clinical Information
By providing clinical information and setting phenotype terms, you can
enhance the analysis process by incorporating relevant clinical context
and enabling the system to prioritize candidate variants based on their
association with the specified phenotypes. Once you click Save, the
VCF sample will be applied to the protocol with the information entered
during import.
2.4 Analysis Details
The Analysis Details page in Geneyx Analysis provides an audit trail and
displays important information about the completed analysis. Here are
the details that you can find on this page:
Creator and Creation Time: This shows the user who created the case and
the timestamp of when it was created.
Description: If a description was provided during the import process, it will be displayed here. This allows for additional information or context about the case to be included.
Protocol: The selected protocol for the case analysis will be shown. It represents the template that was used for filtering and annotation during the analysis.
Genetic Models: This section displays the genetic models that were applied to the sample. Genetic models are integrated at the protocol level and help in analyzing and interpreting rare Mendelian variants.
Disease Frequency (%): The disease frequency indicates the expected frequency of the disease occurring in the population. It is used as a threshold to capture rare variants present in the sample.
Phenotype: The phenotype terms that were entered during the import process will be shown here. These terms were used to score and rank candidate variants during the analysis.
Version Set: This reflects the version of Geneyx Analysis that was used for the analysis.
Local Frequency Date: The date of the in-house allele frequency database used for the analysis is displayed here. It indicates when the local frequency data was last updated.
#SNV Samples in Local Frequency: This shows the number of single nucleotide variant (SNV) samples present in the in-house allele frequency database at the time of analysis.
#SV Samples in Local Frequency: This indicates the number of structural variant (SV) samples present in the in-house allele frequency database at the time of analysis.
Case ACMG Date: The date of the ACMG (American College of Medical Genetics and Genomics) guidelines annotation that was used in the analysis is shown here.
Last Modified: This displays the user who made the last modification to the analysis and the timestamp of when the modification occurred. It helps in tracking any recent changes made to the analysis.
Associated samples: Associated samples defined in the protocol will be present here. There is also the option to include additional associated samples at this stage using the ‘+Add button’
Gene panels: Gene panels defined in the protocol will be present here. There is also the option to include gene panels at this stage using the ‘+Add button’
The Analysis Details page provides a comprehensive overview of the
metrics, settings, and information associated with the completed
analysis, allowing users to review and reference these details as
needed.
Analysis Details
Clicking on the Analyze icon in the upper right corner of the Analysis
Details or accessing it from the Dashboard view will take you to the
variant analysis window in Geneyx Analysis. In this window, you can
further explore and analyze the variants identified in the analysis. The
variant analysis window provides tools and features to investigate and
interpret the genetic variants detected in the sample. Here, you can
perform various tasks such as:
Filtering Variants: Apply filters based on variant properties such as allele frequency, variant type, impact, gene region, and more to narrow down the list of variants.
Sorting and Ranking Variants: Sort and rank the variants based on different criteria such as pathogenicity, predicted impact, allele frequency, and functional annotations to prioritize the most relevant variants.
Variant Visualization: View detailed information about each variant, including genomic position, variant type, allele frequencies in databases, predicted functional impact, and associated gene annotations.
Variant Annotation: Access comprehensive annotations for each variant, such as gene information, functional consequences, conservation scores, and predicted pathogenicity.
Gene Analysis: Explore the genes associated with the variants, including gene function, known disease associations, pathways, and relevant literature.
Variant Filtering and Comparison: Apply additional filters and compare variants across different samples or datasets to identify shared or unique variants.
Pathway and Functional Analysis: Investigate the functional impact of variants on biological pathways and perform enrichment analysis to identify affected biological processes.
Visualization Tools: Utilize interactive visualizations, such as variant allele frequency plots, gene diagrams, and protein domain annotations, to gain insights into the genetic data.
By navigating through the variant analysis window, users can delve into
the genetic variants, examine their potential significance, and gain a
deeper understanding of their role in the analyzed sample.
Analyze icon in the upper right corner.
2.5 Variant Analysis Workflow
In Geneyx Analysis, you can effectively manage and filter the large
number of variants in a sample using the Filter Chain function. This
allows you to apply specific criteria to refine the variants based on
selected annotations. Here’s how you can use the Filter Chain:
In the variant analysis window, locate the column header for the annotation you want to filter by. Click on the filter icon next to the column header.
A dropdown menu will appear, showing various filtering options based on the selected annotation. Select the desired filtering criteria from the available options. For example, you may choose to filter variants based on variant type, impact, allele frequency, or any other relevant annotation.
After selecting the filtering criteria, Geneyx Analysis will automatically apply the filter and display the filtered variants in the variant table.
You can add additional filters by repeating the above steps for other annotations or criteria. Each filter will be listed on the left side of the screen under Column Filters.
To remove or modify a filter, you can click on the “X” icon next to the filter in the Column Filters section. This will remove the filter and update the variant table accordingly.
Additionally, if you need to create customized filters with specific
combinations of criteria, you can refer to the (3.10.10 Filters section
in the Settings window of Geneyx Analysis. The Filter section allows you
to build complex filters using logical operators (AND, OR) and define
multiple conditions for variant filtering.
By using the Filter Chain function and custom filters, you can
effectively narrow down the variants of interest based on your specific
criteria and focus on the most relevant variants for further analysis
and interpretation.
2.5.1 Visualize in GenomeBrowse
To display the GenomeBrowse visualization tool in Geneyx Analysis, you
can follow these steps:
In the variant analysis window, locate the variant of interest in the variant table.
Click on the location of the variant, which is typically represented as a hyperlink or a clickable element in the variant’s genomic position.
Upon clicking the variant’s location, Geneyx Analysis will open the GenomeBrowse visualization tool. This tool provides a graphical representation of the genomic region surrounding the variant, allowing you to explore the nearby genes, transcripts, and genomic features.
In the GenomeBrowse tool, you can navigate through the genomic region using zoom and pan controls to adjust the level of detail and explore the surrounding genomic context.
The GenomeBrowse tool may display various genomic annotations, such as gene models, known variants, regulatory elements, and other relevant information depending on the available data and settings.
You can interact with the GenomeBrowse visualization to explore the genomic features and their relationships to the variant of interest. For example, you can hover over a gene to view its details, click on a known variant to see its information, or adjust the view to focus on specific regions of interest.
The GenomeBrowse tool provides a visual representation of the genomic
landscape surrounding a variant, aiding in the interpretation and
analysis of the variant within its genomic
context.
In Geneyx Analysis, when you open the GenomeBrowse visualization tool to
view a selected variant in the context of the genome, chromosome, gene,
and protein, you will have the option to customize the annotation view
using the settings icon. Here’s how you can access and customize the
annotation view:
After opening the GenomeBrowse tool by clicking on the location of a variant, you will see the schematic view of the variant and its genomic context.
On the right side of the GenomeBrowse dialog, you will find a settings icon (usually represented by a gear or cogwheel icon). Click on this settings icon to open the customization options for the annotation view.
The settings dialog allows you to customize various aspects of the annotation view. You may have options to enable or disable specific annotation sources, adjust the visualization settings, and modify the display of genomic features.
Depending on the available customization options in Geneyx Analysis, you can configure the annotation view to suit your preferences and analysis requirements. For example, you may choose to show or hide specific types of annotations, adjust the color scheme or track display, and modify the level of detail or zoom level.
Explore the settings dialog and experiment with different customization options to optimize the annotation view for your analysis needs. The specific options and features available may vary depending on the version of Geneyx Analysis and the data sources used for annotation.
By utilizing the settings icon in the GenomeBrowse view, you can tailor
the annotation view to focus on the specific information and features
that are most relevant to your analysis, enhancing your understanding of
the variant within its genomic context.
Geneyx Analysis now provides direct integration with IGV (Integrative
Genomics Viewer) Desktop, enabling users to visualize local BAM and
methylation BED files within their native IGV environment. This feature
addresses the common challenge faced by users with local data storage,
offering an efficient solution to the bottleneck associated with
visualizing BAM files when a cloud environment is not utilized.
Key Functionalities and Operational Details:
Seamless Navigation and Dynamic View Management: Upon selecting a variant location within Geneyx Analysis, the system automatically facilitates navigation to the corresponding genomic locus in IGV. Subsequent selections of new variant locations perform a reset of the previous IGV view, ensuring the relevant data for the newly chosen variant is loaded dynamically.
Optimized Local Data Access: This integration is particularly beneficial for users operating in local or restricted environments. It ensures fast, secure, and flexible access to raw read data without the necessity of remote upload.
Remote Control Configuration: The integration leverages IGV’s remote control capabilities via a WebSocket interface, typically on port 60151.
How to Set Up and Use Desktop IGV Integration:
To utilize the Desktop IGV Integration for local BAM visualization,
follow these steps:
Enable IGV Desktop in Geneyx Analysis:
Click on your user name in the upper right corner of the Geneyx Analysis interface.
Select “Preferences” from the dropdown menu.
Tick the checkbox labeled “Enable IGV Desktop”.
Enabling IGV Desktop
Install IGV Desktop: Ensure you have the Integrative Genomics Viewer (IGV) Desktop application and your BAM and/or methylation files are prepared, navigate to a variant location within the Geneyx Analysis platform. Upon clicking a variant, the system will give an option to automatically navigate IGV to the corresponding genomic locus.
Navigating to Desktop IGV
This setup enables seamless visualization of your local BAM and
methylation BED files directly in IGV, streamlining your analysis
workflow.
2.6 Interpret Results
In Geneyx Analysis, you can specify the variants you wish to report for
a sample using the Relevance column. The Relevance column allows you to
indicate the significance or relevance of each variant based on your
analysis criteria.
Here’s how you can utilize the Relevance column to select variants for
reporting:
In the variant analysis window, locate the Relevance column. This column is typically displayed alongside other annotation columns that provide information about the variants.
The Relevance column contains values or options that allow you to assign a relevance score or category to each variant. The specific options available may depend on the configuration of your analysis and the criteria you are using to assess variant significance.
Review the annotations, genomic context, and any other relevant information for each variant in the analysis. Based on your evaluation, assign a relevance value or category to indicate the significance of the variant. This can be done by selecting the appropriate option from a drop-down menu or using a predefined scoring system.
The relevance values or categories can be used to prioritize or filter the variants for reporting. You can choose to report only the variants that meet a certain relevance threshold or fall into specific categories, depending on your analysis goals and criteria.
By assigning relevance values to the variants, you can identify and focus on the most relevant or clinically significant variants in the analysis results. This helps streamline the reporting process and ensures that the reported variants align with your analysis objectives.
Note that the specific options and criteria for assigning relevance
values may vary depending on your analysis settings and the annotation
sources used in Geneyx Analysis.
Annotating a variant for reporting
When you click “Save” after applying the relevant information and
annotations to the variants in Geneyx Analysis, the system will store
this annotation for the analysis. The saved annotation will be
associated with the specific variants you have selected and will be
automatically included in the generated report.
By saving the annotations, you ensure that the selected variants and
their associated information, including the relevance values, are
recorded and incorporated into the final report. This helps maintain a
comprehensive and accurate documentation of the analysis process and the
variants of interest.
The report generated by Geneyx Analysis will include the saved
annotations, allowing you to present the findings and relevant details
to clinicians, researchers, or other stakeholders. The report may
include information such as variant classifications, functional impact,
gene annotations, genomic context, and any additional annotations or
customizations you have applied.
Remember to review the saved annotations and ensure their accuracy
before generating the final report. This step helps maintain the quality
and reliability of the reported variants and their associated
information.
It’s important to note that the specific report generation process and
formatting may vary based on the settings and configurations of your
Geneyx Analysis platform. It is recommended to consult the platform’s
documentation or reach out to their support team for further guidance on
generating and customizing reports.
2.7 Patient Report
Once you click the green Report Preview icon in the upper right corner
of the Geneyx Analysis interface, the Report Editor will be displayed.
The Report Editor allows you to make final modifications and
customizations to the generated report before finalizing and exporting
it.
In the Report Editor, you can perform various actions such as:
Adding and arranging sections: You can add or remove sections in the report and adjust their order according to your preferences. Sections can include general information, analysis details, variant tables, genomic context, clinical interpretations, and more
Customizing content: You can modify the content of each section, including adding or removing text, tables, images, and other elements. This allows you to tailor the report to your specific requirements and include relevant information.
Formatting and styling: The Report Editor provides options to format and style the report, such as adjusting font styles, sizes, colors, and alignments. You can also add headers, footers, page numbers, and other formatting elements to enhance the visual presentation of the report.
Adding comments and annotations: The Report Editor enables you to add comments, notes, and annotations to specific sections or variants in the report. This can be useful for providing additional context, explanations, or recommendations related to the findings.
Previewing and reviewing: You can preview the report in real-time as you make modifications to ensure that it reflects the desired format and content. This allows you to review and proofread the report before finalizing it.
Once you are satisfied with the modifications made in the Report Editor,
you can proceed to finalize and export the report in the desired format,
such as PDF, JSON, or HTML. The finalized report can then be shared with
relevant stakeholders, such as clinicians, researchers, or patients, to
communicate the analysis results and findings effectively.
It’s important to note that the specific functionalities and features of
the Report Editor may vary depending on the version and configuration of
Geneyx Analysis you are using. It is recommended to refer to the
platform’s documentation or consult their support team for detailed
instructions on using the Report Editor and generating customized
reports.
Generating a report using the Report Preview
icon¶
Report Editor
Once you have made the desired modifications to the report using the
Report Editor in Geneyx Analysis, you can select the “Save” option. This
action will generate a PDF document and TSV (Tab-Separated Values) files
containing the filtered variants for each genetic model.
The PDF document will contain the finalized report with all the
customized sections, content, and formatting. It serves as a
comprehensive and visually appealing representation of the analysis
results and findings.
The TSV files, on the other hand, provide structured data in a tabular
format, specifically capturing the filtered variants for each genetic
model. These files can be useful for further analysis, data integration,
or importing the variant information into other software or databases.
After the report and TSV files have been generated, they will be stored
in the Reports section of the Geneyx Analysis platform. This allows for
easy access and retrieval of the generated reports at a later time. You
can refer to the Reports section to view, download, or share the
generated PDF and TSV files as needed.
It’s important to note that the specific file formats and storage
locations may vary depending on the configuration of Geneyx Analysis or
any customization made by the platform administrators. It is recommended
to refer to the platform’s documentation or consult their support team
for detailed information on the output file formats and storage
locations within the platform.
3.0 Geneyx Manual
3.1 Accessing Geneyx
When you access Geneyx Analysis through the URL
https://analysis.geneyx.com/, you may come across a series of dialogs
before you can open or create a project. Let’s take a look at these
dialogs:
Login or Signup: If you’re a registered user, you’ll need to log in using your credentials. If you’re new to Geneyx, you can sign up for an account to gain access to the platform.
Welcome Dialog: Upon logging in, you’ll be greeted with a welcome dialog that provides an overview of Geneyx Analysis and its features.
Project Dashboard: Once you’ve selected or created a project, you’ll be directed to the project dashboard. This is the main interface where you can manage and access various components of your project, such as data sources, analyses, reports, and settings.
These dialogs serve as initial steps to ensure a smooth onboarding
process and to help users navigate the Geneyx Analysis platform
effectively. They provide essential information and options for users to
get started with their projects and utilize the platform’s
functionalities.
If you already have a Geneyx account with a valid license, enter your
email address and password in the provided fields. You also have the
option to check the “Remember Me” box, which will allow Geneyx to
automatically open for future access. Once you have entered your account
information, click on the “Log In” button to
proceed.
If you don’t have an existing account, click on the “Don’t have an
account? Sign up for free!” option, which will redirect you to the
registration form. Fill out the form with the required information and
make sure to read and accept the license agreement. Once all the
required fields are filled in and the license agreement is accepted, the
“Log In” button will become active.
You also have the option to check the “Remember Me” box, which will
enable automatic opening of Geneyx for future access. After filling in
your account information, click on the “Log In” button to proceed.
3.1.2 Navigating Geneyx Analysis
Geneyx Analysis offers a comprehensive navigation console located on the
left-hand side of the interface, providing easy access to various
sections and functionalities. Each section serves a specific purpose and
allows intuitive exploration of your data. Here are the details of each
field:
3.2. Dashboard: Provides an overview and summary of
relevant information. For more detailed information, refer to the
Dashboard section.
Analysis:
3.3. Analysis: Allows you to manage and view the analyses performed on your data. Further information can be found in the Analyses section.
3.4. VCF Samples: Provides access to the Variant Call Format (VCF) samples, allowing you to explore and analyze genetic variants. More details can be found in the VCF Samples section.
3.5. Reports: Enables the management and generation of reports based on your analysis results. Refer to the Reports section for additional information.
3.6. Subjects: Allows you to manage and organize subjects or individuals within your dataset. For more details, refer to the Subjects section.
Data Management:
3.7.1. Processing Tasks: Provides access to processing tasks, allowing you to monitor and manage various data processing operations. Additional information can be found in the Processing Tasks section.
3.7.2 Seq Samples: Allows management of sequencing samples, including importing, organizing, and analyzing genomic data. Refer to the Seq. Samples section for further details.
3.8. Batches (fastq) : Provides a way to group and manage related samples or analyses in batches. For more information, see the Batches section.
3.9. Variant Browser Offers a powerful tool for
exploring and investigating genetic variants. It allows filtering,
annotation, and visualization of variants. More details can be found in
the Variant Browser section.
3.10. Settings: Provides options for customizing your
Geneyx Analysis account, including preferences, permissions, and other
settings.
3.11 Accounts: Allows management of user accounts and
access privileges. Refer to the Accounts section for additional
information.
Geneyx
Analysis Navigation Window
The navigation console offers a centralized and intuitive way to access
all aspects of your data and analysis within the Geneyx Analysis
platform. Furthermore, if there are any specific questions, you can
always use the Help icon in the upper right corner, which will direct
you to a support specialist.
3.2 Dashboard
Clicking the Dashboard icon on the left-hand side will display the
most recent analyses and VCF samples in the account, and hovering over
the Dashboard icon will provide the ability to Start New
Analysis or go into Account features.
Dashboard View
3.2.1 Dashboard View
In the Dashboard view of Geneyx, you’ll find two main sections: Analyses
and VCF Samples. These sections serve different purposes and provide
important information for managing and analyzing genetic data. Here’s a
breakdown of each section and the details they display:
Analyses:
Analyses are preconfigured workflows that incorporate specific filter
logic and genetic models based on the utilized protocol. When an
analysis is initiated, it goes through an annotation process, which
involves adding relevant annotations to the genetic variants. Once the
annotation process is complete, you can analyze the data further by
clicking the “Analyze” option.
Analyses are displayed in a table format, sorted with the most recent
analysis at the top. The columns displayed in the table include:
MODIFIED: This column indicates the date of the most recent modification to the analysis.
ANALYSIS: The analysis ID is displayed as a hyperlink. Clicking on it will take you to the 2.5. Analysis Detailsdirectory for that specific case.
#SAMPLES: This column shows the number of samples present in the analysis.
STATUS: It represents the current status of the case, which can be modified in the Analysis Details window or during the analysis process.
TAT: Number of days waiting for analysis out of ‘required TAT’ as defined at protocol level
ASSIGNMENT: If multiple users have access to the account, sample assignments can be performed with administrative capabilities. 3.2.7 Account options provide more details on this feature.
SUBJECT: The subject identification is displayed in this column.
ANALYZE: Clicking on this option will direct you to the variant interpretation interface specific to that case.
Analyses dialog
VCF Samples
The VCF Samples section in Geneyx displays a table format view of the
uploaded VCF (Variant Call Format) files or samples. This section
provides an overview of the available samples for analysis and
annotation. If a VCF sample is present in the account, you can click on
“New Analysis” to select a protocol and create a new analysis
specifically for that sample.
The columns in the VCF Samples section include:
MODIFIED: This column displays the date of the last modification performed on the sample, indicating when any changes or updates were made.
VCF SAMPLE: The VCF sample identification is displayed in this column. It helps to identify and differentiate between multiple samples within the account.
SUBJECT: The SUBJECT column represents a unique identifier assigned to each VCF sample. It helps in tracking and organizing samples based on their specific subjects or individuals.
GENOME BUILD: This column indicates the reference genome assembly that was used for the VCF sample analysis. It specifies the genome version against which the genetic variants were compared.
VERSION SET: The VERSION SET column displays the version of Geneyx Analysis that was used to annotate the data of the VCF sample. It helps to keep track of the software version used for analysis.
STATUS: The STATUS column provides information about the progress of the annotation step for the sample. It notifies whether the annotation process for the sample has been completed or is still ongoing.
PERMISSION: The PERMISSION column indicates the designation of sample accessibility. It reflects the permissions assigned to each sample, determining who can access and work with the sample data. The permissions can be updated in the VCF Sample dialog, allowing you to control sample access and sharing.
VCF Samples dialog
By navigating through the Analyses and VCF Samples sections in the
Dashboard, you can access and manage your genetic data effectively,
initiate analyses, track modifications, and delve into the variant
interpretation interface for further analysis.
3.2.2 Start New Analysis
A new analysis can be initiated by clicking on the Start New
Analysis icon, in the upper right corner of the Dashboard view.
The following steps will guide the process of importing variants into a
protocol.
Start New Analysis
3.2.3 Protocol Selection
After selecting Start New Analysis, the user will be required to
select the protocol. A protocol is a template for analysis which
includes predefined filters and annotations. Default protocols are
provided and categorized into Germline, Somatic, and Health Screening
and if you want to customize or modify a protocol, this can be done in
the 3.10.9 Protocols section in Settings.
Default protocols include:
Germline:
Single Sample Analysis: Analyzes affected germline samples without associated maternal or paternal VCF files.
Trio Analysis: Utilizes inheritance models to analyze rare Mendelian variants using mother, father, and proband samples.
Single Sample SNV/SV: Analyzes affected germline samples for single nucleotide polymorphisms and structural variations.
CNV Only Analysis: Detects large structural variations in a specific file format that includes information such as chr, start, end, effect, score, and copy number.
Mitochondria Analysis: Specifically used for the analysis of mitochondrial variants.
Somatic:
Tumor Only Analysis: Analyzes tumor-only samples without a paired normal/germline sample.
Tumor Normal Analysis: Commonly used to analyze somatic samples together with their germline counterparts.
Cancer Screening: Detects the presence of somatic variants not isolated from tumor tissue.
Health Screening:
Carrier Screening: Analyzes variants that could be passed as an inherited genetic condition to offspring.
Protocol selection
3.2.4 Subject Information
In the Subject dialog, you can provide general details about the subject
of the analysis. These details are important for reference throughout
the analysis process and are automatically included in the clinical
report.
Here are the fields that can be entered for new subjects:
SUBJECT ID (required): This is a unique identifier for the subject. It is commonly the patient identifier used in the Electronic Medical Record (EMR) or Laboratory Information Management System (LIMS), but it can be any identifier you choose.
NAME (optional): You can enter the name of the subject if available.
DATE OF BIRTH (optional): Enter the date of birth of the subject if known.
GENDER (optional): Specify the gender of the subject as Male or Female. If the information is unknown or unspecified, you can leave this field blank.
CONSENT – PERSONAL DATA (optional): Indicate whether the subject has provided consent for the use of their personal data by selecting Yes or No. If the consent status is unknown or unspecified, you can leave this field blank.
CONSENT – CLINICAL DATA (optional): Specify whether the subject has provided consent for the use of their clinical data by selecting Yes or No. If the consent status is unknown or unspecified, you can leave this field blank.
CONSANGUINITY (optional): This field is commonly used to indicate parental consanguinity in cases of rare disease analysis. You can provide relevant information if applicable.
ETHNICITY (optional): Enter the ethnicity of the subject if known.
PATERNAL ANCESTRY (optional): Specify the ancestry or origin of the subject’s father.
MATERNAL ANCESTRY (optional): Specify the ancestry or origin of the subject’s mother.
FAMILY HISTORY (optional): You can add any notes or details regarding the family history of the subject in the provided text box.
Subject Information
Users can also select an existing subject if one has already been
created. This option requires previous samples to have been imported
into the account. In this case, once Subject ID is selected, there
will be a dropdown displaying the most recent submissions. As the name
is entered, it will populate with the available choices. Once selected,
all associated information will be applied.
Existing Subject
By entering this information for the subject, you ensure that it is
captured and utilized throughout the analysis process. This information
is valuable for interpreting the genetic data and generating a
comprehensive clinical report.
3.2.5 Importing Samples
In the Samples section, you have the option to either upload a new
sample by clicking on “Select Sample” next to the relevant sample or
select an existing pre-loaded sample for the analysis using the “All
Sample” option. Clicking on “Browse” will enable navigation to a
local directory to select a given VCF file. Once a new VCF file is
uploaded into the system, you can define additional parameters for the
sample. These parameters include:
Serial Number: You can specify a unique identifier for this analysis. By default, a value is automatically generated, but you have the flexibility to replace it with your own identifier, such as the ID of the case in your system.
RELATION: This field allows you to specify the relation of the sample in reference to the proband (e.g., mother, father, sibling, etc.) if applicable.
Sequencing Machine: You can indicate the machine used to automate the DNA sequencing process for this sample.
Enrichment Kit & Filtering: This parameter defines the target capture regions for the sample. You can also implement your own BED files if needed. See 3.10.1 Enrichment Kits & Smart Filtering in the Settings dialog.
Genome Build: Specify the reference genome assembly build that was used for alignment and mapping of the sample.
Taken Date: Enter the date when the sample was obtained or collected.
Sequence Date: Provide the date when the sample was sequenced.
Received Date: Enter the date when the sequencing data was received.
Sequencing Target: This field represents the Next-Generation Sequencing (NGS) data type associated with the sample.
Sample Source: Specify the method by which DNA was collected for this sample (e.g., blood, tissue, saliva, etc.).
Notes: You can include any additional comments or relevant information about the sample in this field.
Methylation file URL: If applicable, you can provide the URL of the methylation file associated with the sample. URL file must be present on http or https and enable public access to the file.
files must be bgzipped with corresponding index file (tbi)
xxx.bed.gz and xxx.bed.gz.tbi
BAM file URL: If applicable, you can provide the URL of the BAM file associated with the sample.
files must be bgzipped with corresponding index file (tbi)
eg. xxx.bam.gz and xxx.bam.gz.tbi
URL file must be present on http or https and enable public access to the file.
Exclude sample from local allele frequency: This option allows you to exclude the sample from local allele frequency calculations if desired.
VCF Sample upload
Alternatively, if you want to use a VCF sample that is already present
in the account, you can select the Search option. This will display
a drop down menu for you to select from.
Upload of existing VCF sample
If you wish to modify the associated information, you can click on the
edit icon next to the sample name.
Edit option for existing sample.
By entering these parameters, you can provide detailed information about
the sample, which will contribute to the accuracy and comprehensiveness
of the analysis.
3.2.6 Clinical Information
In the Clinical Information section, you can provide additional context
about the analysis and subject, which will assist in the interpretation
of the results. Here is the information that can be entered:
Serial Number (required): Specify a unique identifier for this analysis. A default value is automatically generated, but you can replace it with your own identifier, such as the ID of the case in your system.
Name (optional): Enter the name of the protocol used for the analysis. This helps identify the specific analysis protocol being applied.
Owner (Doctor) (optional): Enter the name of the referring clinician or the owner of the analysis. This provides information about the responsible healthcare professional.
Description (optional): You can provide free text to describe the clinical aspects of the case. This allows you to include relevant details or observations about the patient’s condition.
Disease frequency (required): Specify the expected frequency of the disease occurring in the population. This information is used in the filtering process to reduce the number of variants observed above the defined threshold. It helps prioritize variants based on their relevance to the disease.
Department(optional): Enter the name of the department or unit associated with the analysis. This provides organizational context for the analysis.
Phenotypes (optional): Enter any relevant biological terms (phenotypes) that will be used to score and rank the candidate variants. You can enter multiple phenotypes using the Advanced option. To ensure proper recognition of compound terms, such as “Developmental Delay,” enclose the term in quotation marks (e.g., “Developmental Delay”). After entering the phenotypes, you can select “Keywords,” which will display highly specific terms and indicate how many genes are associated with each condition.
Clicking on HPO will also provide the ability to enter in a list of HPO terms.
Entering clinical information
By entering the clinical information accurately, you enhance the
analysis process and improve the interpretation of the genetic variants.
Once completed you will be directed to the 2.5. Analysis Details page
for this case.
3.2.7 Account
Hovering over the Dashboard icon will provide a tool tip option for
Accounts, this will direct you to the Usage Dashboard. The
Usage Dashboard in Geneyx Analysis offers comprehensive and visually
appealing reports on the activities carried out within your Geneyx
Analysis account. It provides valuable insights into various workflows
conducted, enabling effective management of your genetic analysis
processes. Here are the key features and benefits of the Usage
Dashboard:
Detailed and Colorful Reports: The dashboard presents detailed reports in a visually appealing manner, making it easy to understand and analyze the data. The use of colors enhances the visual experience and enables quick identification of trends and patterns.
VCF Sample Import by Sequencing Target: The dashboard provides information on VCF samples imported based on sequencing targets. This feature allows you to track and monitor the distribution of samples across different sequencing targets, giving you insights into the utilization of specific genetic analysis workflows.
Distribution of Samples Processed using Secondary Pipelines: This offers insights into the distribution of samples processed using secondary pipelines. This feature helps you understand how samples are being processed through different secondary analysis pipelines, providing a deeper understanding of your analysis workflow efficiency.
Total Number of Analyses Created based on Protocols: The dashboard displays the total number of analyses created, categorized by different protocols. This information gives you a clear overview of the usage of specific analysis protocols, allowing you to evaluate their popularity and assess their effectiveness.
Longitudinal Analysis and Peak Seasons: This provides a longitudinal view of the data, enabling you to identify peak seasons or periods of high analysis activity. This insight helps in resource planning, ensuring optimal allocation of computational resources and staff during busy periods.
Management Tool: The Usage Dashboard serves as a powerful management tool, providing valuable data for decision-making, resource optimization, and process improvement. It allows you to track the usage of the Geneyx Analysis platform, monitor the performance of different workflows, and make informed decisions based on data-driven insights.
Usage Dashboard
By leveraging the Usage Dashboard, you can effectively manage and
optimize your genetic analysis processes, leading to improved
efficiency, resource allocation, and decision-making within your Geneyx
Analysis account.
3.2.7.1 Account Management
In the upper right corner of the Usage Dashboard, you will find the
Account Management icon. This feature is accessible to account
administrators, who are designated by Geneyx support during the creation
of the account. Account administrators have exclusive privileges and can
perform various administrative tasks. Here’s how the Account Management
feature works:
Accessing Account Management: By clicking on the Account Management icon, account administrators can access the administrative functions of the Geneyx Analysis account. This option is specifically available to administrators, ensuring proper control and management of the account.
Inviting Administrators and Users: With administration credentials, account administrators can invite other administrators and general users to the account by clicking on the Invite icon.
It is important to note that in order to invite someone to the account,
they must first be invited as a general user.
Once the user is added to the account as a general user, administrator
privileges can then be granted to provide them with additional
administrative capabilities.
Inviting someone as an administrator without first inviting them as a
user will not give them access to the account.
By inviting additional administrators, you can delegate administrative responsibilities and distribute the workload effectively. Inviting general users allows collaboration and participation in the analysis processes.Administering User Roles and Groups: Within the Account Management feature, administrators have the ability to assign roles and permissions to different users. This ensures that each user has appropriate access levels and restrictions based on their responsibilities and requirements. User roles can be customized to control access to sensitive data and critical functions.
Managing User Accounts: Account administrators can also manage user accounts within the Account Management section. This includes actions such as creating new user accounts, modifying user information, and disabling or removing user accounts as needed. Effective management of user accounts ensures the security and integrity of the Geneyx Analysis platform.
Monitoring and Oversight: The Account Management feature provides administrators with a comprehensive view of the user activity and account usage. It allows administrators to monitor user interactions, track analysis processes, and ensure compliance with organizational policies and guidelines.
By utilizing the Account Management feature, account administrators can
maintain control over the Geneyx Analysis account, invite and manage
users effectively, and ensure the smooth functioning of the platform.
This feature empowers administrators with the necessary tools to
administer and govern the account according to their organization’s
requirements.
Account Management
3.2.8 Roles
When you click on the Roles option in the Account Management
dashboard, a dialog will appear where you can create specific
permissions that can be assigned to users within the account. Here are
the available roles and their corresponding permissions:
Analyst: Users with the Analyst role have the ability to work with samples assigned to them within protocols. They can apply these samples to protocols and access analyses that are specifically assigned to them.
Assign Group: This role allows a user to assign samples to specific groups within the organization. They have the authority to organize samples into different groups based on organizational requirements or criteria.
Assign Users: Users with the Assign Users role can assign samples to specific users within the organization. This role enables them to delegate samples to specific individuals based on their responsibilities or expertise.
Data Management Operate: This role grants access to all samples in the account, but users with this role are restricted from modifying analyses and generating reports. They can perform data management tasks such as organizing and manipulating samples, but they do not have the ability to make changes to analyses or generate reports.
Data Management View: Users with the Data Management View role have access to all samples in the account, but they can only view the data in a viewer mode. They do not have editing privileges and are limited to viewing and reviewing the samples and associated information.
Viewer: The Viewer role provides access to all samples in the account but restricts users to viewing only. They cannot make any modifications to the samples or perform any actions beyond viewing. This role is suitable for users who need to review or analyze the data without the need for editing or modifying it.
These predefined roles and their associated permissions provide
flexibility in assigning appropriate access levels to different users
within the organization. By assigning specific roles to users,
administrators can ensure that each user has the necessary privileges
and restrictions to perform their designated tasks effectively while
maintaining data integrity and security.
Here is a table view that details functionalities of each:
Permissions available
3.2.9 Assign VCF Samples
To delegate samples to individual users or groups, you can utilize the
VCF level and Analysis Details window. Here’s a step-by-step guide for
delegating at the VCF level:
Navigate to the VCF Sample section in the console, located on the left-hand side.
Identify and select the sample you wish to delegate by clicking on it.
Scroll down to the Permissions section.
Click on the permissions, which will prompt a new dialog to appear.
In the dialog, you can assign the sample to a specific user or group.
Once assigned, only the designated user or group will have access to
that particular sample.
VCF Sample delegation
By delegating samples in this manner, you can ensure that access to
sensitive data is restricted to authorized individuals or specific
groups within your organization.
3.2.10 Analysis Delegation
To assign an analysis to an individual user, follow these steps:
Click on “Analyses” in the navigation console to access the list of analyses.
Select the analysis you want to assign by clicking on it.
You will be directed to the Analysis Details window, which provides detailed information about the analysis.
In the upper right corner of the Analysis Details window, you will find an option to assign the analysis.
Click on the assignment option, and a dialog box will appear.
From the dialog box, you can select the individual user to whom you want to assign the analysis.
After making the selection, save the changes.
By assigning the analysis to a specific user, you can ensure that it is
directed to the appropriate individual within a group. This can help
with task management, accountability, and efficient workflow
distribution.
Analysis Delegation
3.2.11 Usage
The Usage Dashboard includes an icon in the upper right corner that
allows users to access the usage statistics of their account. Clicking
on this icon will display the usage statistics, and on the right side,
there is a filter option that enables users to set a specific period of
time for the statistics. Once the filter is applied, users can
differentiate between the following categories:
Analyses: This category represents the number of analyses generated by applying a VCF to a protocol.
VCF Samples: It includes the count of VCF samples that were uploaded into the account. However, VCF samples derived from internal fastq processing are excluded from this count.
Seq Samples: This category displays the number of subjects created from the fastq processing pipeline.
Processing Time: It reflects the time taken for a fastq file to complete processing and generate a VCF file.
Annotation Pipeline: This category indicates the time taken to annotate a VCF file.
The table in the Usage Dashboard presents the following fields:
Date: Represents the date when the data was applied or entered into the system.
Category: Indicates the specific category to which the data belongs.
Details: Clicking on this field allows users to navigate to the specific details associated with the data.
Processing Time: If applicable, this field shows the processing time consumed for the respective data file.
Storage: Reflects the storage space consumed by the data file.
Source: Represents the source from which the sample was derived.
Target: Specifies the sequencing target of the sample.
VCF Sample: Displays the name or identifier of the VCF sample that was applied to the specific category.
User: Indicates the user responsible for entering the data file into the system.
Usage Dashboard
If the user selects “Summary” in the upper right corner of the Usage
Dashboard, it will display a comprehensive summary report for the
selected timeline. This summary report offers an easy-to-read overview
of sample usage, processing activities, time consumption, and storage
allocation within the specified period.
The summary report provides valuable insights into the following
aspects:
Sample Usage: It presents an overview of the number of samples utilized within the specified timeframe, giving users an understanding of the volume of data processed.
Processes: This section highlights the different processes performed during the specified period, such as analysis generation, VCF sample uploads, and subjects created from the fastq processing pipeline. It provides a summary of the key activities carried out within the account.
Time Consumption: The summary report includes information on the time consumed by specific processes, such as processing time for fastq files and annotation pipeline time for VCF files. This helps users identify areas where more time is being utilized and optimize their workflow accordingly.
Storage Allocation: It displays the storage space allocated or consumed by the data files within the specified timeline. This information allows users to track their storage usage and manage their resources effectively.
By selecting the “Summary” option, users can easily access a concise and
informative report that provides a snapshot of their sample usage,
processing activities, time utilization, and storage allocation. This
feature enables users to quickly assess and evaluate the overall account
performance and resource
utilization.
By providing these detailed usage statistics, the Usage Dashboard
enables users to track and analyze their account activity, resource
utilization, and processing times across different categories and data
types.
3.3 Analysis
The analysis console serves as a comprehensive repository for all the
data generated within the account, including samples, analyses, and
reports. It provides users with unrestricted access to their data,
ensuring that they can retrieve and review it whenever needed, as long
as their account has an active license. The analysis console comprises
the following key fields:
Analyses: This section displays a list of all the analyses conducted
within the account. It provides an overview of the analysis projects,
allowing users to access and review specific analysis details.
VCF Samples: The VCF samples field presents a collection of the uploaded
VCF (Variant Call Format) samples in the account. These samples contain
genetic variant information that can be utilized for further analysis
and interpretation.
Reports: In this section, users can find a compilation of generated
reports. These reports may include clinical summaries, variant
annotations, and other relevant information derived from the analysis
process. Accessing reports enables users to review and share valuable
findings.
Subjects: The subjects field encompasses the details of individual
subjects involved in the analyses. It includes information such as
subject IDs, names, demographics, and clinical data. Users can reference
this information to maintain context and track the subjects throughout
the analysis workflow.
Analysis console
By organizing and presenting data within these fields, the analysis
console provides a centralized location for users to manage and retrieve
their genetic analysis data efficiently and without any limitations
imposed on data volume.
3.3.1 Analyses
This section provides a comprehensive list of all analyses conducted
within the account, offering an overview of analysis projects for easy
access and review of specific analysis details. The table is sorted by
default in descending order of the most recent modification, ensuring
that the latest analyses are displayed at the top. The following fields
are available:
NAME: Displays the name of the analysis along with the designated panel or focus. Clicking on the name of the analysis will bring you to the 2.5. Analysis Details window.
PROBAND SAMPLE: Indicates the name of the affected VCF samples that serve as the primary subject of the analysis. Clicking on the hyperlink will bring you to the 3.4.1 VCF Details detail page.
SUBJECT ID: Represents a unique identifier assigned to each VCF sample. Typically, this identifier corresponds to the patient’s identification in EMR or LIMS systems, enabling seamless integration. Clicking on the subject will bring you to the 3.6. Subjects page.
PROTOCOL: Specifies the protocol utilized for the analysis, outlining the predefined filters and annotations employed.
PHENOTYPES: If applicable, this field captures the phenotypic terms associated with the analysis, providing additional context or criteria for variant analysis.
STATUS: Reflects the current status of the analysis, providing information on its progress or completion.MODIFIED: Indicates the date of the last modifications made to the analysis, enabling users to track the timeline of updates and changes.
Analyses window
These fields collectively offer a comprehensive overview of the analyses
conducted within the account, facilitating efficient analysis management
and tracking of key information.
3.3.2 Analysis Details
The Analysis Details page serves as an informative hub for the
specific analysis. This page provides an audit trail of the metrics and
details associated with the analysis. The following information is
available:
Created By and Creation Time: It displays the user who created the analysis and the timestamp indicating when it was created.
Description: If provided during import, the description field captures any additional information or notes related to the case.
Protocol: This field indicates the specific protocol that was selected for the analysis. The protocol determines the set of guidelines and criteria used for variant analysis.
Genetic Models: It lists the genetic models applied to the sample, which are integrated at the protocol level. These models define the inheritance patterns considered during the analysis.
Default Genetic Model: The default genetic model is the primary model displayed when opening the variant analysis interface for the analysis. It serves as the starting point for variant interpretation.
Disease Frequency (%): This field captures the expected frequency of the disease in the population. It is used as a filter to prioritize rare variants present in the sample.
Phenotype: It shows the phenotype information entered during the import process, providing context and additional clinical details about the case.
Version Set: The Version Set reflects the specific version of the Geneyx Analysis software used for the analysis.
Local Frequency Date: This field indicates the date of the in-house allele frequency database used for the analysis, which informs the interpretation of variant frequencies.
#SNV Samples in Local Frequency: It displays the number of single nucleotide variant (SNV) samples present in the in-house allele frequency database at the time of analysis.
#SV Samples in Local Frequency: This field shows the number of structural variant (SV) samples present in the in-house allele frequency database at the time of analysis.
Case ACMG Date: It represents the date of the ACMG (American College of Medical Genetics and Genomics) guidelines annotation that was used in the analysis.
Last Modified: This section indicates the user who made the last modification to the analysis and the timestamp of the modification.
Analysis Details page
Associated samples: Associated samples defined in the protocol will be present here. There is also the option to include additional associated samples at this stage using the ‘+Add button’
To add additional samples, cick the ‘Edit Associated Samples’ button
Then ‘+Add’ to include additional samples. When all samples have been added, click ‘Done Adding Additional Samples’ and then ‘Analyse’
Gene panels: Gene panels defined in the protocol will be present here. There is also the option to include gene panels at this stage using the ‘+Add button’
The Analysis Details page provides a comprehensive overview of the
analysis, including its origin, applied metrics, and important
timestamps. This information helps users track the progress and
modifications made to the analysis, ensuring transparency and
accountability throughout the workflow as well as an ‘Analysis notes’
feature to allow users to leave timestamped, user-tagged comments at the
analysis level
3.3.3 Creating a Focused Workflow
In the Analysis Details window, located to the left of the Analyze icon,
you will find a target icon that allows you to create a focused analysis
specifically for the uploaded VCF sample. By selecting the target icon,
you can utilize existing filtering options to create a new analysis
based on specific criteria.
One option is to select an existing enrichment kit, which provides
predefined target regions for analysis. This allows you to focus your
analysis on specific genomic regions of interest that are covered by the
chosen enrichment kit.
Creating a Focused workflow
Alternatively, if you prefer to define a new target region, you have the
flexibility to do so. This means you can customize the analysis by
specifying your own set of genomic regions to be included as the target.
By defining a new target region, you can tailor the analysis to your
specific research or clinical requirements.
When creating a New Target Region, the user will be prompted to enter
the following information:
Name: Provide a name for the enrichment kit you are creating. This helps in identifying and referencing the kit later.
Description: Add detailed information about the enrichment kit, such as its purpose, target regions, or any other relevant details.
Include Ref/Ref: Selecting this option will capture homozygous reference calls in the analysis. It ensures that variants with homozygous reference genotypes are included in the analysis results.
Include all modifiers: Enabling this option will capture all regulatory elements outside of coding regions. It ensures that regulatory elements, such as enhancers or promoters, are considered during the analysis.
Genes/Genomic Regions: Enter a list of genes or genomic regions that you want to include in the target region. This allows you to focus the analysis specifically on these genes or regions of interest.
CADD Scores: CADD scores are part of the smart filtering capabilities, which can be used to prefilter variants based on a CADD threshold. CADD scores are computed for single nucleotide variants and insertions/deletions in the human genome. You can set a specific PHRED-scaled score threshold to filter variants based on their deleteriousness. Variants with scores above the threshold will be included in the analysis.
HOM Allele Count: This prefiltering option allows you to filter out variants that have a homozygous allele count in the gnomAD database greater than a defined value. It helps in excluding variants that are commonly observed in the population.
HET Allele Count: Similar to the HOM Allele Count, this option allows you to filter variants based on their heterozygous allele count in the gnomAD database. Variants with a count greater than the defined value will be excluded from the analysis.
HEMI Allele Count: This prefiltering option is specifically for variants on the X chromosome. It filters variants based on their hemizygous allele count in the gnomAD database. Variants with a count greater than the defined value will be filtered out.
By providing these options during the creation of a New Target Region,
users can customize the analysis to focus on specific genomic regions,
apply smart filtering based on CADD scores, and set prefiltering
criteria to refine the variant selection process.
Creating a New Target Region
The ability to create a focused analysis using existing filtering or
defining a new target region provides users with flexibility and control
over the analysis process. This feature allows for more precise and
customized analyses based on the specific genomic regions or enrichment
kits of interest.
3.3.4 Creating a Copy of Analysis
On the Analysis Details page, located to the left of the Analyze
icon, you will find a “Create Copy” option. Selecting this option
allows you to create a copy of the analysis with the most recent version
of the software, including updated annotations and new features, if
available.
By choosing to create a copy, you can ensure that your analysis benefits
from the latest advancements in the software and incorporates the most
up-to-date annotations. This process maintains all the associated data
from the original analysis while updating the annotations to reflect the
most recent version.
Create Copy of Analysis
Creating a copy of the analysis with updated annotations and features
enables you to leverage the improved functionality and information
provided by the latest software version. It allows you to stay current
with advancements and enhancements in the analysis process, ensuring
that you have access to the most accurate and comprehensive results
3.3.5 Analysis History
On the Analysis Details page, you can find an “Analysis History” option
located to the left of the Analyze icon. By selecting this option, you
gain access to an audit trail that showcases the chronological sequence
of steps performed during the analysis.
The Analysis History feature is valuable from a management perspective
as it provides insights into the actions taken throughout the analysis
process. It allows you to track the progress of the analysis, review the
specific steps that were executed, and understand the order in which
they occurred.
By examining the Analysis History, you can gain a comprehensive overview
of the analysis workflow, ensuring transparency and accountability. This
feature facilitates effective project management by allowing you to
monitor the analysis’s progression, identify potential issues or
bottlenecks, and assess the overall efficiency of the process.
Analysis History
The Analysis History serves as a useful tool for tracking and
documenting the analysis’s evolution, enabling effective collaboration,
communication, and decision-making among team members involved in the
project.
3.3.6 Recalculate Local Frequency
On the Analysis Details page, you will find an option called
“Recalculate Local Frequency” located to the left of the Analyze icon.
By selecting this option, you can initiate the recalculation of the
local allele frequency within the analysis.
The local allele frequency refers to the frequency of specific genetic
variants observed within the analyzed samples. By recalculating the
local allele frequency, you ensure that the analysis incorporates the
most up-to-date information regarding variant frequencies.
When you choose to recalculate the local allele frequency, the analysis
will consider the most recent samples that have been applied to the
analysis. This ensures that the frequency calculations are based on the
latest available data.
It’s important to note that it may take up to 24 hours for the local
allele frequencies to be updated. This timeframe allows sufficient time
for the system to process and incorporate the new sample data into the
frequency calculations.
By utilizing the “Recalculate Local Frequency” option, you can ensure
that the analysis reflects the most current and accurate information
regarding variant frequencies, providing you with more reliable and
relevant results.
3.4 VCF Samples
The VCF Samples field in the account showcases the collection of
uploaded Variant Call Format (VCF) samples, which contain valuable
genetic variant information for further analysis and interpretation. The
field includes the following details for each VCF sample:
ID: Represents the name of the VCF file, serving as a unique identifier. Clicking the hyperlink will navigate to the 3.4.1 VCF Details window.
SUBJECT: Displays the unique identifier assigned to the VCF sample, typically corresponding to the patient’s identification or a specific reference in the EMR or LIMS system. Clicking the hyperlink will navigate to the 3.6.1 Subject Details window.
GENOME BUILD: Specifies the reference genome assembly utilized for alignment and mapping during the sequencing process.
SEQUENCING TARGET: Indicates the scope of the sequencing, such as Whole Genome, Exome, Gene Panel, Targeted Region, or Clinical Exome.
SAMPLE SOURCE: Describes the source from which the sample was derived, encompassing categories such as Germline, Tumor Biopsy, Blood, Buccal, Mitochondria, Saliva, Fetal, Other, and Parents.
TAKEN DATE: Represents the date when the sample was received or obtained for analysis.
SEQUENCE DATE: Reflects the date when the sample underwent the sequencing process.
EXCLUDE SAMPLE FROM LOCAL ALLELE FREQUENCY: Provides an option to exclude the sample from local allele frequency calculations.
CREATED BY: Specifies the user who imported the VCF file into the system.
CREATED: Indicates the date when the VCF file was initially uploaded.
MODIFIED BY: Displays the user who made modifications or updates to the VCF file.
MODIFIED: Represents the date of the last modification or update made to the VCF file.
To modify the details of a specific VCF file, an edit option is
available on the right side of each VCF file entry. Clicking on this
option will open a new interface where the appropriate modifications can
be made to the VCF sample details. Please see
3.4.1 VCF Details for available fields.
VCF Samples window
This comprehensive set of fields and the edit functionality provide
users with control and flexibility in managing and updating the details
of their VCF samples within the system.
3.4.1 VCF Details
The VCF details page can be accessed either in the Dashboard or in the
VCF Samples section by clicking on a specific VCF sample ID. This will
direct you to a directory containing various information and files
related to the sample, including clinical records, associated data
files, applied analyses, annotation history, annotation information, and
QC data (if available).
VCF Details
Here are the details of each field:
Details: Clicking on the edit option next to the VCF Sample name
provides the ability to update VCF level information. The fields
include:
Serial Number: You can specify a unique identifier for this analysis. By default, a value is automatically generated, but you have the flexibility to replace it with your own identifier, such as the ID of the case in your system.
RELATION: This field allows you to specify the relation of the sample in reference to the proband (e.g., mother, father, sibling, etc.) if applicable.
Sequencing Machine: You can indicate the machine used to automate the DNA sequencing process for this sample.
Enrichment Kit & Filtering: This parameter defines the target capture regions for the sample. You can also implement your own BED files if needed. See 3.10.1 Enrichment Kits & Smart Filtering in the Settings dialog.
Genome Build: Specify the reference genome assembly build that was used for alignment and mapping of the sample.
Taken Date: Enter the date when the sample was obtained or collected.
Sequence Date: Provide the date when the sample was sequenced.
Received Date: Enter the date when the sequencing data was received.
Sequencing Target: This field represents the Next-Generation Sequencing (NGS) data type associated with the sample.
Sample Source: Specify the method by which DNA was collected for this sample (e.g., blood, tissue, saliva, etc.).
Notes: You can include any additional comments or relevant information about the sample in this field.
Methylation file URL: If applicable, you can provide the URL of the methylation file associated with the sample. URL file must be present on http or https and enable public access to the file.
files must be bgzipped with corresponding index file (tbi)
xxx.bed.gz and xxx.bed.gz.tbi
BAM file URL: If applicable, you can provide the URL of the BAM file associated with the sample. URL file must be present on http or https and enable public access to the file.
files must be bgzipped with corresponding index file (tbi)
Exclude sample from local allele frequency: This option allows you to exclude the sample from local allele frequency calculations if desired.
File(s): This section displays all the files associated with the sample.
Please note that BAM and BAI files are available only for samples that
have undergone the secondary pipeline initiated in Geneyx Analysis.
Permissions: This feature enables users with assigning privileges to
assign the sample to a specific group or user within the account. It
provides control over the access and ownership of the sample within the
system.
VCF Sample Details, File(s), and Permissions
Analyses: This section provides a comprehensive list of all the analyses
that have been performed on the VCF sample. By clicking on the
hyperlink, you can navigate to the Analysis Details window, where you
can explore the specific details of each analysis.
Annotation History: This feature presents a detailed history of the
annotation steps carried out on the sample using the designated
enrichment kit. The information includes:
Version Set: Indicates the version of the application that was utilized during the annotation process. Clicking on the link will provide access to all versions of the annotation sources used.
Pipeline: Specifies the data file that was employed in the annotation pipeline.
Status: Displays the current status of the annotation pipeline, indicating whether it is in progress or completed.
Timestamp: Indicates the date and time when the annotation process was initiated.
Lines: Represents the number of variant lines in the annotated VCF file.
Variants: Indicates the total number of variants captured in the annotation pipeline.
#AA Changes: Reflects the count of amino acid changes detected during the annotation process.
Median Depth: Provides the median depth value calculated from all the captured variants.
For each entry in the annotation history, there are several actions and
downloadable files available, including:
TSV: This option enables you to download the entire annotated VCF file in TSV format.
Art: By selecting this, you can download an artifact file associated with the annotation.
Log: This allows you to download the annotation log file, providing a detailed record of the annotation process.
Delete: Clicking on this option will delete the annotation history for the specific sample, if desired.
QC Data: This section provides valuable information regarding the
quality control (QC) metrics of the sample, specifically related to its
coverage. These metrics are derived from the BAM file and are applicable
when using the secondary pipeline in Geneyx. The fields in this section
include:
Passed: Indicates the total number of positions that were captured in the sample.
Fail: Represents the number of positions that did not meet the QC metrics of the secondary pipeline and failed the quality check.
Mapped: Indicates the number of positions that were successfully mapped during the alignment process.
Aligned: Reflects the number of positions that were aligned correctly.
Paired: Represents the number of positions where the reads were paired correctly.
BED Pos: This field provides the count of positions captured within the defined target region, based on the applied enrichment kit.
Mean: Specifies the mean coverage of the sample, which is the average depth of coverage across all positions.
%5X: Indicates the percentage of positions covered at a depth of 5 times or more.
%20X: Represents the percentage of positions covered at a depth of 20 times or more.
%50X: Indicates the percentage of positions covered at a depth of 50 times or more.
Analyses, Annotation History, and QC Data
These metrics provide insights into the coverage and quality of the
sequencing data for the sample, giving an indication of its reliability
and suitability for further analysis.
3.5 Reports
The Reports section presents a collection of generated reports that
encompass clinical summaries, variant annotations, and other pertinent
information derived from the analysis process. Users can access these
reports to review and share valuable findings. The fields included in
this section are as follows:
Created: Represents the date and time when the report was created,
providing a timestamp for reference.
Subject: Displays the unique identifier assigned to the samples or
subjects associated with the report. Clicking the hyperlink will
navigate to the 3.6.1 Subject Details window.
Analysis: Specifies the name or identifier of the analysis from which
the report was generated. Clicking this link will navigate to the 3.3.2
Analysis Details window.
Summary: Provides a concise summary of the analysis, encapsulating key
findings and relevant information.
Report: Presents the report itself in PDF format, enabling users to view
and review the comprehensive analysis report. Provided are the
downloadable reports in PDF, word, or JSON.
Data file: Offers an excel or TSV file containing the detailed findings
and data derived from the analysis, allowing for further exploration and
analysis if required.
Report dialog
The Report window also includes “Reports for Review” and “Reports
History” under the Report icon in the navigation window. The
comprehensive Report History consolidates all generated reports, while
Reports for Review strategically filters reports with a Sub-status
assignment excluding “Closed.” This intuitive organization streamlines
accessibility and retrieval of relevant reports.
By organizing and presenting these reports in a structured manner, users
can conveniently access and retrieve important insights from their
analyses. The availability of both PDF reports and Excel data files
ensures flexibility and ease of sharing information with colleagues or
other stakeholders.
3.6 Subjects
The Subjects field provides comprehensive information about individual
subjects involved in the analyses. This data includes subject IDs,
names, demographics, and clinical information, offering users the
ability to maintain context and track subjects throughout the analysis
workflow. The available fields in this section are as follows:
ID: Represents the unique identification assigned to the subject within the system. The hyperlink will navigate to the 3.6. Subjects window.
NAME: Displays the name of the patient or subject.
DATE OF BIRTH: Specifies the date of birth for the subject, providing essential demographic information.
CONSENT – PERSONAL DATA: Indicates the consent status regarding the usage of personal data for the subject.
CONSENT – CLINICAL DATA: Represents the consent status regarding the usage of clinical data for the subject.
CREATED BY: Identifies the user who created the subject entry in the system.
CREATED: Displays the date and time when the subject was initially entered into the system.
MODIFIED BY: Indicates the user who made modifications to the subject information.
MODIFIED: Represents the date and time when the subject information was last modified.
Additionally, each subject entry includes a delete option, allowing
users to remove subjects from the internal database if necessary.
Subjects window
By organizing and presenting subject details in this manner, users can
easily access and manage subject information throughout the analysis
process, ensuring effective tracking and reference.
3.6.1 Subject Details
The Subject details page can be accessed either in the Dashboard or in
the Subject window by clicking on a specific Subject ID. This will
direct you to a directory containing various information and files
related to the Subject, including clinical records, associated data
files, applied analyses, and sequencing samples.
Subject details window
In this dialog users can update Subject information, including:
SUBJECT ID (required): This is a unique identifier for the subject. It is commonly the patient identifier used in the Electronic Medical Record (EMR) or Laboratory Information Management System (LIMS), but it can be any identifier you choose.
NAME (optional): You can enter the name of the subject if available.
DATE OF BIRTH (optional): Enter the date of birth of the subject if known.
GENDER (optional): Specify the gender of the subject as Male or Female. If the information is unknown or unspecified, you can leave this field blank.
Note: If left blank, gender is assumed as female and variants on chrX
will be annotated as homozygous. If ‘Male’ then chrX variants will be
annotated as hemizygous.
CONSENT – PERSONAL DATA (optional): Indicate whether the subject has provided consent for the use of their personal data by selecting Yes or No. If the consent status is unknown or unspecified, you can leave this field blank.
CONSENT – CLINICAL DATA (optional): Specify whether the subject has provided consent for the use of their clinical data by selecting Yes or No. If the consent status is unknown or unspecified, you can leave this field blank.
CONSANGUINITY (optional): This field is commonly used to indicate parental consanguinity in cases of rare disease analysis. You can provide relevant information if applicable.
ETHNICITY (optional): Enter the ethnicity of the subject if known.
PATERNAL ANCESTRY (optional): Specify the ancestry or origin of the subject’s father.
MATERNAL ANCESTRY (optional): Specify the ancestry or origin of the subject’s mother.
FAMILY HISTORY (optional): You can add any notes or details regarding the family history of the subject in the provided text box.
Users also have the option to add or create a new clinical record for
the VCF sample. Any updates made to this field will be automatically
applied to the analysis and will be reflected in the generated report.
The fields for the clinical record include:
Record Date: This is the date when the clinical information was provided or documented.
Description: Users can provide a clinical summary or description of the subject’s condition or relevant information.
Phenotype Codes: This field enables users to enter a list of Human Phenotype Ontology (HPO) phenotypes associated with the subject. HPO phenotypes are standardized terms used to describe observable characteristics or symptoms related to genetic disorders or diseases.
Create a new Clinical Record
By including this information in the clinical record, users can enhance
the analysis and reporting process, ensuring that the relevant clinical
details and phenotypic information are captured and considered during
the interpretation of the genomic data.
Users are also empowered with the capability to access and manage
applied analyses for the subject, VCF samples, and sequencing samples.
This functionality allows users to make modifications or delete the
applied analyses as needed, providing greater control and flexibility in
the analysis process. For information related to Seq. Samples, please
see 3.7.2 Seq Samples section.
Subject derived Analyses, VCF Samples, and Seq Samples
3.6.2 VUS Monitor
Keeping track of variant classification changes can be a daunting task.
To simplify this process, we have implemented an auto-notification
feature that alerts you whenever variants undergo classification changes
according to ClinVar. This proactive notification provides comprehensive
details, including a direct link to the analysis, enabling you to
perform retrospective analyses with ease and precision.
If a variant has been assigned as a VUS in the ‘Annotate variant’ pop-up
in an analysis, it will automatically be integrated into the VUS
monitor. In the next update of the application, these newly added
variants will appear and a notification provided in the toolbar.
Notification of variant classification update.
Specifically, there will be an Update Classification section which will
show the updated classification according to ClinVar.
A variant can also be easily investigated by clicking the refresh icon
on the last column. This will update the analysis with the given variant
and include the latest release of the annotation sources.
3.7 Data Management
The Data Management interface serves as a powerful tool for implementing
secondary pipelines, specifically for generating VCF files from fastq
data. This functionality leverages the efficient Illumina DRAGEN
pipeline for accurate variant calling. Users can take advantage of this
feature by manually uploading fastq files or utilizing batch upload
capabilities through integration with a cloud infrastructure hosting the
fastq files. It’s important to note that access to this feature is
typically limited to licenses that support fastq upload and processing,
ensuring optimal performance and compatibility.
Total FASTQ file size per sample should not exceed 120Gb.
3.7.1 Processing Tasks
The Processing Tasks interface provides detailed information on the
conversion progress from fastq files to VCF files. To facilitate easier
navigation, this dialog offers the option to filter tasks based on date
and category. In the upper right corner of the interface, users can
select either “Upload Single Sample” or “Batch Upload” to initiate the
variant calling pipeline for the desired samples.
Furthermore, the interface also displays information about the data
storage created from the fastq files. This allows users to track and
manage the storage associated with the processed files.
Within the Processing Tasks interface, you’ll find the following
relevant fields:
DATE: The date when the processing task was initiated.
SAMPLE ID: The unique identifier assigned to the sample.
SUBJECT: The subject associated with the sample.
BATCH: The title of the batch workflow, if applicable.
STATUS: The current status of the pipeline, indicating whether the task is in progress or completed.
DETAILS: Provides additional information and specific details about the pipeline and its execution.
PIPELINE: Specifies the pipeline that was utilized for variant calling and subsequent annotation.
PROTOCOLS: Identifies the specific protocol that was employed for variant analysis.
GENOME BUILD: Indicates the reference genome assembly that was used during the variant calling process.
Processing Tasks
Overall, the Processing Tasks interface offers a user-friendly and
comprehensive view of the fastq to VCF conversion process, giving users
control over their data and facilitating efficient management of variant
calling tasks.
3.7.2 Seq Samples
The Seq. Samples interface presents a comprehensive list of all
sequencing samples that have been created, along with their associated
clinical information. The table includes the following columns:
ID: This column displays the unique identifier for each sequencing sample, known as the Sequencing Sample ID.
Sample: The Sample column provides various details about the sample, including:
Consent: Indicates whether consent has been provided for the sample.
Target: Specifies the sequencing target of the sample.
Source: Indicates the source of the sample.
Relation: Describes the relationship of the sample, if applicable.
Enrichment Kit: Specifies the enrichment kit that was used for sequencing.
Collection ID: This column displays the collection sample ID associated with the sequencing sample.
Files: Indicates the number of fastq files associated with the sample.
Taken: Specifies the date when the sample was taken.
Sequenced Date: Indicates the date when the sample was sequenced.
Received: Specifies the date when the sample was received.
Created: Indicates the date when the sequencing sample was initially created.
Modified: Specifies the date when modifications were last made to the sequencing sample.
Subject
ID: Represents the unique identification assigned to the subject within the system. The hyperlink will navigate to the 3.6. Subjects window.
NAME: Displays the name of the patient or subject.
DATE OF BIRTH: Specifies the date of birth for the subject, providing essential demographic information.
CONSENT – PERSONAL DATA: Indicates the consent status regarding the usage of personal data for the subject.
CONSENT – CLINICAL DATA: Represents the consent status regarding the usage of clinical data for the subject.
CREATED BY: Identifies the user who created the subject entry in the system.
CREATED: Displays the date and time when the subject was initially entered into the system.
MODIFIED BY: Indicates the user who made modifications to the subject information.
MODIFIED: Represents the date and time when the subject information was last modified.
Phenotypes: If phenotypic terms and HPO ontologies have been entered for the sequencing sample, this column will display the associated phenotypes. It provides valuable information about the observed characteristics or traits related to the sample, as well as the hierarchical structure of the Human Phenotype Ontology (HPO) terms used to describe those phenotypes.
QC Data
Passed: Indicates the total number of positions that were captured in the sample.
Fail: Represents the number of positions that did not meet the QC metrics of the secondary pipeline and failed the quality check.
Mapped: Indicates the number of positions that were successfully mapped during the alignment process.
Aligned: Reflects the number of positions that were aligned correctly.
Paired: Represents the number of positions where the reads were paired correctly.
BED Pos: This field provides the count of positions captured within the defined target region, based on the applied enrichment kit.
Mean: Specifies the mean coverage of the sample, which is the average depth of coverage across all positions.
%5X: Indicates the percentage of positions covered at a depth of 5 times or more.
%20X: Represents the percentage of positions covered at a depth of 20 times or more
%50X: Indicates the percentage of positions covered at a depth of 50 times or more.
Sequencing Sample dialog
3.7.3 Upload Single Sample (fastq)
This section provides a step-by-step guide on initiating a secondary
pipeline to convert fastq files to VCF format for a single sample. The
process is applicable to both local and remote fastq files. If your
fastq files are stored on a cloud infrastructure, you will need to
configure the Data Source in the Settings dialog before
proceeding.
Once the setup is complete, you can proceed with initiating the
pipeline. Click on the Upload Single Sample icon in the Seq.
Samples or through the Processing Tasks window. Follow the
instructions provided in the interface to select and upload the fastq
files. Make sure to provide accurate and relevant information, such as
consent status, sequencing target, sample source, and any other required
fields.
Upload Single Sample icon.
The next sections will cover the steps and information for each section.
3.7.4 Entering Subject Information (fastq)
When you click on the Upload Single Sample icon, you will be prompted to
provide the necessary Subject information. This will allow you to create
a New Subject or use an Existing Subject. The fields include:
Subject ID (required): This field is used to assign a unique identifier to the subject.
Name (optional): You can enter the name of the subject if available.
Date of Birth (optional): If known, you can specify the date of birth of the subject.
Gender (optional): You can select the gender of the subject as Male or Female. If the gender is unknown or unspecified, you can leave this field blank. It’s important to note that selecting the appropriate gender is required if you’re interested in generating a CNV (Copy Number Variation) panel of normals to call events in non-autosomal chromosomes.
Consent – Personal Data (optional): This field allows you to specify whether consent has been given for the use of personal data associated with the subject.
Consent – Clinical Data (optional): You can indicate whether consent has been given for the use of clinical data related to the subject.
Consanguinity (optional): This field is commonly used to specify parental consanguinity in cases of rare disease analysis.
Ethnicity (optional): You can enter the ethnicity of the subject if known.
Paternal Ancestry (optional): If available, you can provide information about the paternal ancestry of the subject.
Maternal Ancestry (optional): If known, you can provide details about the maternal ancestry of the subject.
Family History (optional): This field allows you to include any additional family history information relevant to the subject.
By providing the necessary Subject information in these fields, you can
ensure proper identification and contextual information for the uploaded
single sample.
Subject information
By providing the necessary Subject information in these fields, you can
ensure proper identification and contextual information for the uploaded
single sample.
3.7.5 Entering Seq Samples Details (fastq)
The Sequencing Sample Details section contains important fields that
provide essential information about the sample. Here are the details of
each field:
Serial Number: The Serial Number refers to the prefix of the fastq file names for Illumina FASTQ. In this example, the fastq files are named MQ21111B_005.R1.fq.gz and MQ21111B_005.R1.fq.gz. In this case, the Serial Number would be “MQ21111B.”
For MGI FASTQ, the Serial Number refers to the sample ID: In this
example, the FASTQ files are named V450272805_L01_D12345_1.fq.gz. In
this case, the serial Number would be “D12345”.
Use Consent: Indicate whether you give consent for your colleagues to use the sample.
Expected # fastq files: Specify the number of fastq files available for the sample.
Sequencing Target: Select the target of the sequencing for the sample. This includes options such as Whole Genome, Exome, Gene Panel, Target Region, or Clinical Exome.
Enrichment Kit and Smart Filtering: The enrichment kit defines the target capture regions for the sample. You can choose from pre-defined enrichment kits or add custom enrichment kits in the Settings dialog. Smart filtering allows for intelligent filtering of variants based on specific criteria.
Taken Date: Enter the date when the sample was taken or obtained.
Sequence Date: Specify the date when the sample was sequenced.
Received Date: Enter the date when the sample was received for processing.
Sequencing Machine: Indicate the sequencing machine that was used for
the sample. If no Sequencing machine is selected the default of
Illumina will be assumed. To specify for MGI, please use the dropdown
and select the relevant machine.
Sample Source: Select the method by which the DNA sample was collected. Options include Germline, Tumor Biopsy, Blood, Buccal, Mitochondria, Saliva, Fetal, Bone Marrow, Other, and Parents. * If you plan to run a somatic pipeline, you will need to select Tumor Biopsy.
Notes: Provide any additional notes or relevant information about the sample.
Exclude Sample from local allele frequency: Select this option if you want to exclude the sample from being considered in the internal local allele frequency calculations.
Sequencing Sample Details
These fields collectively provide crucial details about the sequencing
sample, ensuring accurate processing and analysis.
3.7.6 Selecting File Source (fastq)
In the File Source section, you have the flexibility to upload local
files or fetch them remotely. If your fastq files are stored on a cloud
infrastructure, you can configure the connection by accessing the Data
Sources feature within the Settings Dialog.
When you choose the “Fetch Remote” option, you will be prompted to
select the specific Data Source to be utilized. Various Data sources are
supported, including FTP, sFTP, Amazon S3, BaseSpace, One Drive, and
Google Cloud Storage. This allows you to seamlessly retrieve the fastq
files from your preferred remote location.
Selecting File Source
By offering both local file upload and remote file fetching
capabilities, the system caters to different scenarios and enables
convenient access to your fastq files, regardless of their location.
3.7.7 Selecting Seq Sample Files (fastq)
In the next dialog, you will have the option to browse to the local
directory containing the samples of interest, if applicable. It is
necessary to upload the samples before proceeding by selecting the
appropriate files. It is important to note that a stable and reliable
internet connection is crucial for successful data transfer.
Insufficient connectivity may lead to data corruption and result in the
failure of the secondary pipeline.
To ensure optimal results and minimize potential issues, it is
recommended to establish a direct connection between the cloud
infrastructure hosting the data files and Geneyx. This integration
eliminates the need to upload the files into Geneyx separately. By
directly accessing the cloud infrastructure, you can streamline the data
transfer process and enhance the efficiency of the secondary
pipeline.
Selecting Seq Sample Files
3.7.8 Selecting Seq Sample Pipeline (fastq)
In the final step, you will need to define the sequencing pipeline to be
initiated. This selection is based on the reference genome assembly and
can be categorized into the following options:
DRAGEN Exome – hg19: This pipeline is specifically designed for exome sequencing data and utilizes the DRAGEN software with the hg19 reference genome assembly.
DRAGEN Exome – hg38: Similar to the previous option, this pipeline is optimized for exome sequencing data but utilizes the hg38 reference genome assembly.
DRAGEN Whole Genome – hg19: This pipeline is tailored for whole genome sequencing data and employs the DRAGEN software along with the hg19 reference genome assembly.
DRAGEN Whole Genome – hg38: Similar to the previous option, this
pipeline is specifically designed for whole genome sequencing data but
utilizes the hg38 reference genome assembly.Users also can select what
version of DRAGEN is being used for the secondary pipeline. The baseline
DRAGEN version will remain v4.0.5 but will be deprecated with newer
releases. This option allows users to conform to their existing
workflows whilst enabling exploration of updated DRAGEN features. The
optional DRAGEN version (4.2.4) enables several additional features,
discussed below, including advanced callers and low pass whole genome
options. For upgrading to the newest Dragen caller, please reach out to
our support (support@geneyx.com).
Advanced DRAGEN Callers: When running whole genome workflows from the
secondary pipeline, there are new sequence-graph realignment settings
that run on the backend to improve calling for genes that have high
identity paralogs. This includes genes such as: GBA, SMN1, HBA,
LPA, RH, CYP2D6, CYP21A2, CYP2B6. For analyses that have whole
genome workflows with at least 30X coverage, output metrics for these
genes will be present in the Advanced Analysis link in the “Info”
section of the analysis. All outputs are taken from Illumina and each
output will reference the associated hyperlink.
If advanced caller processing is performed outside of Geneyx, the system
supports the direct integration of JSON files generated by these
external specialty callers, including DRAGEN TruSight Oncology 500 for
display for tumour mutation burden (TMB), microsatellite instability
(MSI) and genomic instability (GIS) For further details see
https://github.com/geneyx/geneyx.analysis.api/tree/main/scripts/DragenTruSightOncology500
DRAGEN Low pass whole genome sequencing: Low-pass whole genome
sequencing (low-pass WGS) is a genomic sequencing approach where the
entire genome is sequenced at a relatively low depth, typically less
than 5x coverage. Low pass whole genome sequencing is now supported in
the secondary pipeline of Geneyx. To implement, the sequencing target
will need to set to Whole Genome Low Pass.
Selecting primary pipeline to be initiated
By selecting the appropriate pipeline based on your sequencing data and
reference genome assembly, you can ensure accurate and efficient variant
calling and analysis.
* Sentieon is also available as a secondary pipeline option within
Geneyx Analysis. Sentieon’s exceptional software has earned widespread
recognition in the genomics community for its exceptional accuracy,
speed, and scalability. Trusted by researchers in diverse fields such as
cancer genomics, rare disease studies, population genetics, and
precision medicine, Sentieon is now seamlessly integrated as part of
your secondary pipeline options. This pipeline is available upon
request.
Once you click Next, you will have completed the secondary pipeline
process.
3.7.9 Sequencing Samples Details
The Sequencing Samples details provide a comprehensive overview of the
options selected for the secondary pipeline, as well as the current
progress and results. The following sections offer valuable information:
Seq. Sample: Data Manager and Admins can choose to execute secondary pipeline(s) using the
icon.
Processing Tasks: This section displays the current status of the secondary pipeline. If any changes have been made to the pipeline or if an error was encountered, the user has the option to reset the process or delete it using the X icon. Upon successful completion, a hyperlink to the VCF Sample will be generated, allowing easy access to the VCF Details page.
QC Data: This section showcases the quality control metrics generated from the secondary pipeline. These metrics provide insights into the reliability and accuracy of the sequencing data, ensuring high-quality results.
Files: In this section, the fastq files that were utilized during the secondary pipeline are listed. This information helps track the source and origin of the data, facilitating traceability and reproducibility.
VCF Samples: Here, the VCF files that were generated from the pipeline are displayed. Clicking on the hyperlink associated with each VCF Sample will direct you to the VCF Details page. This page offers comprehensive information about the variants detected in the sample and allows for further analysis and exploration.
By providing these detailed sections, the Sequencing Samples interface
enables users to track the progress of the secondary pipeline, assess
the quality of the data, and access the resulting VCF files for further
analysis and interpretation.
Sequencing Samples Details
Once the secondary pipeline is completed, clicking on the VCF hyperlink
will take you to the VCF Details page, where you can explore the
generated output. The output files included in this page typically
consist of a BAM file, a BAM.bai index file, SNV VCF file, and CNV/SV
VCF file.
*The analysis of repeat sequences is not optimal for PCR-based
next-generation sequencing (NGS) due to the high degree of homology
between repeat regions. PCR amplification of repeat regions can result
in non-specific amplification, which can lead to errors in sequence
assembly and analysis. Furthermore, the presence of repeat regions can
result in difficulties in mapping reads back to a reference genome,
making accurate alignment and variant calling challenging. Therefore, it
is recommended to exercise caution when performing the analysis of
repeat sequences or to use specialized protocols and tools that can
account for the challenges posed by these regions in PCR-based NGS.
The VCF files are automatically associated with the enrichment kit that
was specified during the import process. As a result, the metrics
obtained from the enrichment kit will be displayed in the Annotation
History section, providing valuable information about the variants
detected and annotated in the sample.
To further analyze the VCF file, you have two options:
Dashboard: You can apply the VCF file to an Analysis by either clicking on “New Analysis” in the Dashboard. This allows you to perform in-depth analysis and interpretation of the variants within the VCF file.
VCF Details Page: Alternatively, you can directly access the analysis options by clicking on the “New Analysis” icon next to the sample name within the VCF Details page. This streamlined approach enables you to initiate an analysis specific to the selected VCF sample, ensuring efficient exploration and processing of the genomic data.
By providing these options, the platform empowers users to leverage the
generated VCF files for downstream analysis, allowing for detailed
investigations into the genetic variants present in the sample.
VCF Sample Details
In the QC Data section, there is an option to output Custom Coverage and
QC Metrics. QC Metrics will provide a downloadable file. Information for
these metrics can be obtained with the following links:
When a user selects Custom Coverage, it will provide the option for Gene
List or Gene Panel coverage. Gene List allows the user to enter a list
of genes, whereas the Gene Panel requires a panel to be integrated into
the account. The outputs will provide coverage for the genes at 10 and
20X. More information related to this can be found here:
https://github.com/geneyx/geneyx.analysis.api/wiki/Get-Coverage-For-Gene-Panel.
3.8 Batches (fastq)
Geneyx Analysis offers support for batch upload of fastq files using a
TSV (tab-separated values) file. This convenient feature allows users to
describe a set of fastq files available in external storage, such as
FTP, sFTP, Basespace, S3, Google Cloud, or OneDrive. To initiate the
batch upload process, it is necessary to configure a Data Source in
the Settings dialog, enabling seamless integration with external
storage systems. Once a Data Source has been configured, you are now
ready to initiate a batch workflow.
By clicking on the “Batches” section, users can access a comprehensive
overview of all batch workflows that have been initiated in their Geneyx
Analysis account. The displayed fields provide essential information
about each batch workflow, including:
Name: The descriptive name assigned to the batch workflow.
File Name: The name of the uploaded text file that contains the batch information.
Status: The current status of the batch workflow, indicating whether it is running, completed, or encountered any errors.
Created By: The user who initiated the batch workflow.
Created: The date and time when the batch workflow was initiated.
Modified By: The user who last made modifications to the batch workflow.
Modified: The date and time of the last modification made to the batch workflow.
Additionally, users are granted the flexibility to manage their batch
workflows. They have the option to delete a batch workflow that is no
longer needed or edit the details of an existing workflow, ensuring
efficient workflow management and customization.
Batch details overview
With the batch upload functionality and comprehensive batch workflow
management options, Geneyx Analysis facilitates the streamlined
processing of large-scale sequencing data, empowering users to
efficiently analyze and interpret genomic information at scale.
3.8.1 Batch Upload (fastq)
Clicking on the Batch Upload icon will navigate to a window where you
can download a TSV template. The template needs to be modified with the
updated sample information. The fields include:
# p_sn – Patient serial number / id [Required]. This is the prefix of the sample present in the Data Source that hosts the fastq file. For example, if you have two fastq files named MQ21111B_005.R1.fq.gz and MQ21111B_005.R2.fq.gz, the serial number would be “MQ21111B”.
# p_name – Patient name [Optional]
# p_gender – Patient gender (M, F) [Optional]. This is important to specify if you are interested in calling events in non-autosomal chromosomes.
# p_dob – Patient date of birth (yyyy/mm/dd) [Optional]
# p_consentPersonal – Patient consent for personal data (true, false, 0, 1) [Optional]
# p_consentClinical – Patient Consent for clinical data (true, false, 0, 1) [Optional]
# p_generallyHealthy – Patient indication to be generally healthy [Optional]
# c_date – Clinical Record date (yyyy/mm/dd) [Optional]
# c_description – Clinical Record description [Optional]
# c_phenotypes – Clinical Record list of HPO phenotype codes or names colon separated [Optional]
# ss_sn – Sequenced sample serial number/id [Required]. This should be the same as the Patient Sample Serial number.
# ss_seqMachine – Sequenced sample – sequencing machine name
[Optional] If no sequencing machine is specified in the
BatchImportTemplate, the default of Illumina will be assumed. To
specify for MGI, please specify the relevant machine in the #
ss_seqMachine field.
# ss_kit – Sequenced sample – enrichment kit [Optional]. List of enrichment kits are available in the Enrichment Kit and Smart filtering dialog in Settings.
# ss_takenDate – Sequenced sample – taken date (yyyy/mm/dd) [Optional]
# ss_seqDate – Sequenced sample – sequencing date (yyyy/mm/dd) [Optional]
# ss_receiveDate – Sequenced sample – Received date (yyyy/mm/dd) [Optional]
# ss_excludeLAF – Flags whether this sample should be excluded from the in-house allele frequencies [Optional]
# pipeline – The pipeline name to run – otherwise use the seq target default pipeline defined in account (skip or – to skip pipeline. overwrite or o to overwrite existing task). [Optional]
# build – the pipeline genome build to run – valid values are (hg19, hg38) [Required]
# protocols – The protocols to run (defined by the serial-numbers and comma separated) [Optional]. This will apply the output files directly into analyses and greatly improve sample automation workflows.
# genePanel – The gene panel by name to add for each created analysis [Optional]
To streamline the batch upload process, each sample in the batch
workflow should be displayed on a per-row basis, presenting the updated
information for the required fields. This allows users to conveniently
review and modify the details of each sample as needed. Once the
necessary modifications are made, the next step is to select the
“Browse” button, which enables users to navigate to the updated TSV
(tab-separated values) file that contains the batch information.
Upon selecting the TSV file, users can proceed by clicking the “Process
Batch” button. This action initiates the secondary pipeline for all the
samples defined within the batch workflow. The secondary pipeline
performs the necessary steps, such as variant calling and VCF file
generation, for each sample in the batch. By initiating the process,
users can efficiently process a large number of samples simultaneously,
saving time and effort.
Batch Upload Dialog
The Batch functionality in Geneyx Analysis empowers users to effectively
manage and execute batch workflows, automating the analysis of multiple
samples and enabling efficient processing of large-scale sequencing data
3.9 Variant Browser
The Variant Browser in Geneyx Analysis provides a powerful platform for
querying and exploring SNVs (Single Nucleotide Variants) and CNV/SVs
(Copy Number Variants/Structural Variants) across all samples within an
account. This feature allows users to apply unique filtering logic and
perform comprehensive analyses on a large-scale dataset.
With the Variant Browser, users can flexibly define filters based on
various criteria such as genomic coordinates, variant type, allele
frequency, functional impact, and more. This enables targeted
investigations and in-depth exploration of specific variants or genomic
regions of interest.
The Variant Browser is designed to handle large-scale datasets, allowing
users to query and analyze up to 200,000 samples at a time. This
capacity ensures that users can efficiently navigate and explore
extensive genomic data, gaining valuable insights into the genetic
variations present in their samples.
By leveraging the capabilities of the Variant Browser, researchers and
clinicians can conduct sophisticated variant analysis, identify
potential disease-causing variants, and unravel the genetic basis of
complex disorders.
Variant Browser
To run the Variant Browser workflow in Geneyx Analysis, follow these
steps:
Click on the “New” button located in the upper right corner of the Variant Browser interface. This will open the Query Builder dialog.
In the Query Builder dialog, you can specify the genome build for your analysis. Choose the appropriate genome build from the available options.
Configure the specific filters based on your analysis requirements. You can define filters based on various criteria such as genomic coordinates, variant type, allele frequency, functional impact, and more. This allows you to narrow down the variants of interest.
Once you have configured the filter logic, click “OK” to save the filter settings.
After setting up the filters, click on the “Run” button located in the upper right corner of the Variant Browser interface. This will initiate the query across all internal samples in your account.
The results of the query will be displayed in a table format, showing the variants that match your specified filters. You can explore the results and perform further analysis.
If desired, you can export the query results, edit them, or save them for downstream applications. This allows you to integrate the results into your research or clinical workflows.
By following these steps, you can effectively utilize the Variant
Browser workflow to query and analyze variants across all internal
samples, enabling you to gain valuable insights into the genomic
variations within your dataset.
Query Builder Dialog
3.10 Settings
The Settings dialog in Geneyx Analysis provides users with the ability
to customize and configure various aspects of their account. It allows
for both general user-initiated modifications and certain administrative
capabilities.
By accessing the Settings dialog, users can take advantage of these
customizable features and tailor their Geneyx Analysis account to suit
their specific requirements, allowing for a more personalized and
efficient genomic analysis experience.
Settings Dialog
3.10.1 Enrichment Kits & Smart Filtering
Enrichment kits play a crucial role in the pre-sequencing DNA
preparation process by selectively amplifying or capturing specific
regions of interest from the genome. Instead of sequencing the entire
genome, only the targeted regions are enriched and sequenced, resulting
in a more cost-effective and focused analysis.
In Geneyx Analysis, default enrichment kits are provided, but users also
have the flexibility to edit and customize them according to their
specific research or clinical requirements. The edit icon allows users
to modify the target capture regions, add or remove specific genomic
regions, or create entirely custom enrichment kits. This customization
ensures that the sequencing and analysis focus on the specific regions
of interest relevant to the study or investigation.
Smart filtering is another powerful feature in Geneyx Analysis that
empowers users to define frequency and CADD score thresholds for
filtering out irrelevant variants before the VCF annotation step.
Frequency thresholds enable users to filter variants based on their
frequency in specific populations or databases, helping to prioritize
rare or low-frequency variants that might be more relevant to the
analysis. CADD score thresholds allow users to filter variants based on
their predicted pathogenicity, focusing on variants with higher scores
that are more likely to have functional significance. In addition to
existing features, you can also apply SMART filtering based on Splice-AI
and Read Depth.
By leveraging smart filtering, users can efficiently narrow down the
variants of interest, reducing the analysis workload and improving the
accuracy of downstream interpretation. This feature helps researchers
and clinicians focus on the most relevant variants and prioritize their
investigation efforts, saving time and resources in the analysis
process.
Enrichment Kits
Geneyx provides a Default-Exons enrichment kit, which will be applied by
default unless a protocol is configured with a different enrichment kit.
Users can specify which enrichment kit they wish to use as default in
this dialog.
* Default – Exons only’ kit – selecting this enrichment kit applies a
.bed file curated from UCSC and ClinVar. This bed file is updated with
each release of Geneyx Analysis. Please contact support@geneyx.com for
the most current versions.
When selecting “New” in the upper right corner of the Enrichment Kit &
Smart Filtering interface, users can add new enrichment kits to the list
of selectable options. The following details need to be entered for
creating a new enrichment kit:
NAME: Provide a name for the enrichment kit.
DESCRIPTION: Enter a description or additional details about the enrichment kit.
BED FILE PADDING: Specify the number of bases to be added as padding to each region in the sequencing BED file. This padding helps ensure complete coverage of the target regions during sequencing.
DEFAULT: Enrichment kit assigned as ‘Default’ in the settings, will be applied when no ‘Default Enrichment kit and Smart filtering’ is applied to a protocol. Users can specify which enrichment kit they wish to use as default.
Under the Smart Filtering section, users can configure additional
pre-filtering options for efficient variant annotation. These options
include:
Exons +/-50bp BED file: This option captures all exonic regions with an additional 50 base pairs of padding on each side of the exon.
This uses the ‘Default-Exons only’ .bed
Filtering :Users can define which variants will be captured when the enrichment kit is applied. This includes options like including homozygous reference variants (ref/ref) and capturing all modifiers according to the specified rules.
Merge BED with ClinVAR BED: This option combines the existing BED file with the ClinVar BED file, capturing variants that are present in both.
Always include ClinVar P/LP variants: Include ClinVar Pathogenic/Likely Pathogenic variants regardless of CADD, splice-AI, gnomAD, and ColorsDB defined thresholds
VCF FILTER FIELD VALUES: If the VCF has a FILTER value, such as PASS or ., they can be specified to automate the annotation of the VCF file.
DP (>= OR N/A): Enables prefiltering based on read depth.
GQ (>= OR N/A): Enables prefiltering based on genotype quality.
CADD SCORE (≥ OR N/A): Pre-filter variants based on the specified CADD score threshold. Variants with a CADD score equal to or higher than the threshold will be included.
SPLICE AI: (>= OR N/A): Enables prefiltering based on SpliceAI.
HOM ALLELE COUNT (≤ OR N/A): Pre-filter variants based on the homozygous allele count from gnomAD exomes/genomes data. Variants with a homozygous allele count less than or equal to the specified threshold will be included.
HET ALLELE COUNT (≤ OR N/A): Pre-filter variants based on the heterozygous allele count from gnomAD exomes/genomes data. Variants with a heterozygous allele count less than or equal to the specified threshold will be included.
HEMI ALLELE COUNT (≤ OR N/A): Pre-filter variants based on the hemizygous allele count from gnomAD exomes/genomes data. Variants with a hemizygous allele count less than or equal to the specified threshold will be included.
CoLoRSdb AF (% <= or N/A): Pre-filter variants according to CoLoRSdb AF (%)
Once the user clicks Save, it will navigate to the Enrichment kit
details, where you can add sequencing and genotyping BED files. Once you
click Add, you will need to input the following:
TYPE: This includes Sequencing and Genotype. Genotype is used for a BED file that contains individual genetic variant locations. Sequencing is used for BED files that encompass regions of DNA to capture.
GENOME BUILD: hg19/GRCh37 or hg38
BED FILE: This includes ability to Select Existing, Select Shared, and Upload a new file. To upload a new file, select the Browse icon to locate the BED file. Once selected, click Save, and it will convert the BED file into a new enrichment kit that can be used. Note: Mitochondrial coordinates should be entered as chrMT 1 16569 (NC_012920) whether hg19 or GRCh38 – this is because chrM 1 16571 (NC_001807) is obsolete.
By customizing these options, users can tailor the enrichment kit and
smart filtering settings to their specific analysis requirements,
ensuring that only relevant variants are captured and processed during
the analysis pipeline.
Enrichment Kit Details
3.10.2 Gene Panels
To incorporate a specific subset of genes into the analysis workflow,
the Gene Panels feature can be utilized. To create a new gene list,
simply click on the “New” icon located in the top right corner of the
interface. If you wish to modify an existing gene list, select the
“Edit” icon next to the respective panel. Modification of gene panels is
restricted to users with Admin permissions.
When creating or editing a gene list, it is important to enter the gene
names in accordance with the standards set by the HUGO Gene Nomenclature
Committee (HGNC). This ensures accurate recognition and processing of
the gene names. In case a gene is not recognized correctly, a
notification will be displayed indicating the incorrect gene entry.
Adjacent to the “New” icon, you will find the “PanelApp” icon, which
provides access to PanelApp. PanelApp offers recommended gene panels
that can be explored and used as references for creating custom gene
lists.
By leveraging the Gene Panels feature, users can conveniently integrate
specific genes of interest into their analysis, allowing for more
targeted and focused investigations within the selected gene subset.
Gene Panels
Gene panels can be incorporated into the analysis workflow either at the
protocol level or during the analysis stage. When a gene panel is added
at the protocol level, it is automatically applied to capture variants
exclusively within the specified genes.
By including a gene panel at the protocol level, the analysis pipeline
is tailored to focus solely on the genes defined in the panel. This
ensures that the subsequent variant calling and annotation processes
exclusively consider variants within the designated gene set.
Adding gene panels at the protocol level offers the advantage of
streamlining the analysis workflow by predefining the targeted genes.
This approach allows for more efficient and accurate identification of
relevant variants, facilitating in-depth investigations within the
specific gene subset.
Gene Panel option at the protocol level
Gene panels can be disabled from use (for example if a newer version of
a panel is required), by selecting the Disable icon in the edit panel.
Disable gene panel
3.10.3 Variant Maps
Variant maps are utilized to define target capture regions within an
analysis, either through a BED file or specific genomic positions. They
enable the inclusion of specific variants of interest during the
analysis process. To create a new variant map, click on the New icon,
which will open a dialog where the following information needs to be
entered:
NAME: Provide a name for the variant map.
DESCRIPTION: Add a description that provides details about the variant map.
LOCATIONS: Enter the locations in a BED file format, or you can directly paste the contents from a BED file. This specifies the genomic regions to be targeted.
VARIANTS: Enter the variants in a tab-delimited format, with each line reserved for one variant. This allows for the inclusion of specific variant positions of interest.
RS IDS: Users can enter a list of rsIDs that they wish to capture in a comma separated format.
Creating a Variant Map
By defining variant maps, you can customize the analysis to focus on
specific genomic regions or variants, enabling more precise
investigations and interpretations within those designated regions.
3.10.4 Allele Frequency Backlog
The allele frequency backlog feature enables users to incorporate a
database of variant- and associated allele frequencies into their
internal catalog. By utilizing this feature, users can enhance the
representation of variants specific to different populations within
their Geneyx account.
The allele frequency backlog feature helps to improve the accuracy and
relevance of variant analysis by considering population-specific allele
frequencies. By incorporating this information into the internal
catalog, users can obtain more reliable and context-specific
interpretations of variants within their samples.
This feature contributes to a more comprehensive understanding of
genetic variation, particularly with regards to population-specific
variations and their impact on variant interpretation and disease
association.
Allele Frequency Backlog
When entering data into the Allele Frequency backlog, you can follow a
specific format to ensure accurate and structured information. The
format for entering data are provided as templates, and as follows:
Chromosome: Specify the chromosome number or identifier where the variant is located.
Position: Provide the genomic position of the variant on the specified chromosome.
Reference Allele: Enter the reference allele for the variant at the given position.
Alternate Allele: Specify the alternate allele(s) for the variant at the given position.
Population: Indicate the population or ethnic group associated with the allele frequency data.
Allele Frequency: Enter the frequency of the alternate allele in the specified population. This should be provided as a decimal value ranging from 0 to 1.
Each line represents a single variant with its corresponding allele
frequency information. You can add multiple variants by entering them on
separate lines, following the same format.
By accurately entering the data in this format, you can effectively
incorporate population-specific allele frequencies into your internal
catalog, enriching the analysis and interpretation of variants within
different populations.
Entering data into the Allele Frequency Backlog
3.10.5 Sequencing Machines
The Sequencing Machine console refers to the instrument used during the
primary analysis of the sequencing data. If the sequencing machine used
for your samples is not available by default in the system, you have the
option to create a new entry.
To create a new sequencing machine entry, click on the New icon next to
the Sequencing Machine field. This will open a dialog where you can
provide the necessary information for the new sequencing machine:
Sequencing company: Select sequencing company from the drop down
If no option is selected the default of Illumina will be used
To enable batch upload of MGI FASTQ the ‘MGI’ option must be selected.
Model: Enter the name or identifier of the sequencing machine.
Description (optional): Provide a description or additional details about the sequencing machine.
By entering the relevant details, you can add a new sequencing machine
entry to accurately reflect the instrument used for your sequencing
samples. This ensures that the information is properly recorded and can
be referenced in the future for analysis and interpretation
purposes.
Sequencing Machines
3.10.6 Data Sources
Data Sources in the context of variant calling refers to the storage
locations where the fastq files required for the secondary pipeline can
be accessed. These storage devices include:
Amazon S3: Amazon Simple Storage Service is a web service offered by Amazon Web Services (AWS) for storing and retrieving data. It provides scalable, reliable, and secure object storage in the cloud.
FTP: File Transfer Protocol is a standard network protocol used for transferring files between a client and a server on a computer network. It allows users to upload and download files from an FTP server.
Secure File Transfer Protocol (SFTP): SFTP is a secure version of FTP that uses encryption to protect the confidentiality and integrity of data during file transfer. It provides a more secure alternative to FTP.
BaseSpace: BaseSpace is a web-based platform provided by Illumina, a leading genomics company. It allows biologists and informaticians to store, analyze, and share genetic data generated from Illumina sequencing systems. Please note when using Basespace, the Illumina SampleSheet used to generate FASTQ should adhere to the following: either Sample_ID must be the same as Sample_Name, or only Sample_ID should be used. If Sample_ID does not = Sample_Name, files will not be detected.
OneDrive: OneDrive is a file hosting and synchronization service provided by Microsoft. It allows users to store files and access them from various devices. It also integrates with Microsoft Office applications.
Google Cloud Storage: Google Cloud Storage is a cloud-based object storage service provided by Google Cloud Platform. It offers scalable and durable storage for a wide range of data types and is accessible via RESTful APIs.
These data sources provide options for accessing and retrieving the
fastq files required for variant calling and subsequent analysis. By
configuring the appropriate data source in the Settings dialog, users
can seamlessly integrate their data from these storage devices into the
secondary pipeline.
To create a new Data Source for seamless communication with Geneyx
Analysis, follow these steps:
Go to the navigation pane on the left side of the Geneyx Analysis platform.
Select “Settings” from the navigation options. This will open the Settings menu.
Within the Settings menu, locate and select “Data Sources.” This will take you to the Data Sources configuration page.
On the Data Sources page, you can view and manage existing Data Sources. To create a new Data Source, click on the “New” button or the “+ Add Data Source” option.
A dialog or form will appear, allowing you to configure the details of the new Data Source.
Enter the required information for the new Data Source, including the cloud storage provider (e.g., Amazon S3, Google Cloud Storage), authentication credentials, and any additional settings specific to the chosen provider.
If you’re setting up a cloud-based Data Source, you may need to provide access keys or authentication tokens to establish the connection between Geneyx Analysis and the cloud storage service.
Once you have entered all the necessary information, click “Save” or “Create” to create the new Data Source.
The newly created Data Source will now be available for selection when configuring file sources or workflows within Geneyx Analysis.
*Please specify the specific directory in which the fastq files are
stored.
Data Source dialog
By setting up a Data Source with the cloud storage provider, you can
establish a seamless communication channel between Geneyx Analysis and
the designated cloud infrastructure where your fastq files are stored.
This is particularly advantageous for handling large fastq files
efficiently. Additionally, you also have the option to upload fastq
files locally if needed.
3.10.7 Case Sub Statuses
Case Sub Statuses are a way to further categorize and track the progress
or status of individual cases within a laboratory or workflow. They
allow for customization based on the specific requirements and needs of
the lab.
To configure Case Sub Statuses in Geneyx Analysis, follow these steps:
Access the Geneyx Analysis platform and navigate to the Settings section.
Within the Settings menu, locate and select “Case Sub Statuses” or a similar option related to case statuses or workflows.
On the Case Sub Statuses configuration page, you will see a list of existing sub statuses, if any.
To create a new Case Sub Status, click on the “New” button or an equivalent option to add a new sub status.
Provide a descriptive name for the new sub status. This should reflect the specific stage, progress, or status that you want to track for cases.
Save or apply the changes to create the new Case Sub Status.
Repeat the process if you need to create multiple sub statuses to cover different stages or statuses in your workflow.
Case Sub States can be applied in the Analysis Details window or in an
analysis by clicking “Open” and changing to the specific status of the
case.
Changing the status of a case
When a case status has been changed to “Closed”, the analysis will
convert into a read-only mode.
In this mode, variants can no longer be selected or deselected, ensuring
results remain frozen and protected from accidental modification. While
reports can no longer be generated, all other analysis functionalities
remain accessible for review. Only users with administrator permissions
can unlock the analysis for further editing.
By customizing Case Sub Statuses, you can create a more granular and
tailored system for tracking and managing cases within your lab. This
allows you to categorize cases based on specific milestones, progress,
or custom stages that align with your workflow and requirements.
3.10.8 ACMG Settings
The ACMG Settings in Geneyx Analysis allows users to customize specific
thresholds that are used in the ACMG (American College of Medical
Genetics and Genomics) criteria for variant analysis. These thresholds
help determine the pathogenicity or benign status of variants based on
various prediction algorithms and scores. Here are the thresholds that
can be modified:
REVEL Pathogenic Cutoff: This threshold determines the cutoff score for the REVEL (Rare Exome Variant Ensemble Learner) algorithm. Variants with a score greater than the specified cutoff are considered evidence of pathogenicity.
REVEL Benign Cutoff: This threshold sets the suggested cutoff for the REVEL algorithm to identify benign variants. Variants with scores below this cutoff are more likely to be classified as benign.
SIFT Cutoff: The SIFT algorithm predicts the deleteriousness of amino acid substitutions. The SIFT cutoff defines the threshold for normalized probabilities. Positions with probabilities below the cutoff are predicted to be deleterious, while those equal to or above the cutoff are predicted to be tolerated.
PhyloP Cutoff: PhyloP scores indicate whether sites are under purifying selection. A PhyloP score cutoff of less than 0 is typically used to filter out sites under selection, effectively removing them from the analysis.
GERP Cutoff: GERP scores also indicate the impact of mutations affected by purifying selection. A GERP score cutoff of less than 2 is generally used to filter out sites under selection.
dbNSFP dbscSNV Cutoff: dbNSFP dbscSNV provides a dichotomous effect score for variants. The specified threshold value determines the cutoff for considering variants as having dichotomous effects.
MetaSVM Cutoff: MetaSVM is a variant prediction tool. The MetaSVM cutoff defines the threshold value used to classify variants based on their predicted pathogenicity.
ACMG Settings
By modifying these thresholds in the ACMG Settings, users can fine-tune
the variant analysis criteria to align with their specific requirements
and interpretations of variant pathogenicity.
3.10.9 Protocols
The Protocols dialog in Geneyx Analysis provides an overview of
workflows that can be utilized for variant analysis. A protocol defines
the filtering logic and genetic models used to identify clinically
relevant variants in a streamlined manner. It serves as a standardized
template for workflows, incorporating specific criteria such as genetic
models, report structure, allele frequency thresholds, and associated
phenotypes and diseases. Here are the details included in the Protocols
dialog:
NAME: This field displays the name of the protocol, which helps identify and differentiate between different protocols.
TYPE: The protocol is categorized into specific types, allowing users to group and organize them based on their purpose or application.
DESCRIPTION: This section provides additional details and information about the protocol, helping users understand its purpose, objectives, or any specific considerations.
PHENOTYPES: The protocol may be associated with specific phenotypes, indicating the clinical characteristics or traits that are relevant for the variant analysis performed using this protocol.
DISEASE: This field lists the associated diseases or conditions that are the focus of the variant analysis conducted using the protocol.
TEMPLATE: The protocol may be based on a specific template, which serves as the foundation for the analysis workflow. The template provides a standardized structure and format for reporting results and interpreting variants.
CREATED BY: This field identifies the user who initially created the protocol, providing information about its origin.
CREATED: The date or timestamp when the protocol was originally created is displayed, enabling users to track the timeline of protocol development.
MODIFIED BY: If modifications have been made to the protocol, this field indicates the user responsible for those modifications.
MODIFIED: The date or timestamp of the most recent modifications made to the protocol is listed, allowing users to track any updates or changes to the protocol.
Protocol Dialog
By utilizing protocols, users can streamline their variant analysis
workflows, ensure consistency, and apply standardized criteria and
interpretations to identify clinically relevant variants.
To create a new protocol in Geneyx Analysis, follow these steps:
1. Click on the New icon in the top right corner of the protocol
interface.
2. Fill in the following fields to define the protocol:
NAME: Provide a name for the protocol.
TYPE: Select the focus or category of the protocol from options like Misc, Somatic, Germline, Health Screening, Panel Design, etc.
ORDINAL: Specify the position of the protocol in a series.
SERIAL NUMBER: Enter a unique identifier for the protocol.
Analysis Mode: Available options include Individual, Parental, and Familial. Individual: Used when the index is a singleton sample. Parental: Includes both mother and father, typically used for carrier screening. Familial: Does not designate an index and encompasses all variants across all samples within the family.
Required TAT: The turnaround time (TAT) for analysis, specified in days. Once the analysis is initiated, a timer will start and continue running until the case is closed.
AUTOMATIC ANALYSIS: Enable this option to apply pre-defined filter logic to the analysis and See appendix A for filters available.AUTO ANNOTATE: If a variant is reported on in one case and is seen in another, this feature ensures consistent relevance, classification, and comments.
DESCRIPTION: Provide user-applied details or additional information about the protocol.
PHENOTYPES: Enter associated phenotypic terms, if applicable.
ALWAYS SHOW PAT/LP: Enable this option to filter logic and always display pathogenic and likely pathogenic variants according to ClinVar.
ALWAYS SHOW PAT/LP IN-HOUSE: This allows a user to always display Likely Pathogenic (LP) and Pathogenic (P) variants based on In-house classifications. Enabling this option ensures that these variants are captured and presented to you, even if they fail to meet specific filtering criteria. This feature helps uncover potentially clinically significant variants that might have been missed otherwise.
DEFAULT GENETIC MODEL: Set the default genetic model that is displayed first when opening the variant analysis interface.
DEFAULT ENRICHMENT KIT & SMART FILTERING: Optionally select a custom enrichment kit for pre-sequencing DNA preparation.
Note: The option to hide columns has been moved from Protocol
settings to Genetic model settings ( 3.12.4 Genetic Models) from version
6.2.
Fill in additional fields as needed, such as:
DISEASE: Specify the disease related to the protocol.
DISEASE FREQUENCY (%): Enter the expected percentage of individuals presenting the symptoms in the general population. This represents the prevalence of the disorder and is not related to allele frequency. The default value is 0.1%, indicating that the disorder is found in less than 1 out of 1000 individuals.
CALCULATE LOCAL FREQUENCY: Enable this option to calculate the local frequency of the variants across internal samples within the protocol.
REPORT CONFIGURATIONS: Choose between default or custom configurations. Report Configurations can be accessed in the Settings dialog.
REPORT TEMPLATE: Select a report template from the available default or custom choices. This is available in the Settings window.
DEFAULT SORT BY COLUMN: Define the column by which the data will be displayed first when opening the variant analysis interface.
SORT DIRECTION: Specify the default data presentation order, either in descending or ascending manner.
PROBAND SAMPLE SOURCE: Define the location from which the DNA was collected for the proband. Options include Germline, Tumor Biopsy, Blood, Buccal, Mitochondria, Saliva, Fetal, Bone Marrow, Other..
VARIANT MAP: Select a target capture region defined by a BED file.
HIDDEN COLUMNS: Specify which annotation sources should be hidden in the variant analysis by default. This includes:
LICENSE: Provide license details for paid protocols.
DEFAULT SORT BY COLUMN: Enable the capability to sort variants in the variant analysis interface by an annotation field.
SORT DIRECTION: Choose whether the default column sorting should be in ascending or descending order.
SINGLE-GENETIC MODEL VARIANT SELECTION: Setting this option to ‘Yes’ will alert the user that a variant has already had a user annotation made in another genetic model.
SINGLE-GENETIC MODEL VARIANT SELECTION: The user is alertred to the
prior user annotation of the vatiant.
Once all the desired fields have been assigned for the new protocol,
click the Save icon. This will store the protocol with your Geneyx
Analysis account, allowing you to use it for all downstream variant
analysis and repeated workflows.
Protocol Configuration
In addition to the fields mentioned earlier, you can further customize
your protocol by incorporating Filters, Gene Panels, and Associated
samples. These options provide more flexibility and specificity to your
variant analysis. Here’s an overview of these customization features:
Filters: Filters allow you to define specific criteria to narrow down and focus your variant analysis within the protocol. By configuring filters, you can apply specific rules and conditions to select variants that meet your criteria. You can refer to the Filter section for detailed instructions on setting up filters and their application within the protocol.
Gene Panels: Gene Panels provide a way to target and analyze a select subset of genes within your protocol. You can integrate a gene list into your workflow using the Gene Panels option. To create a new gene list, click on the New icon in the top right corner. If you need to edit an existing gene list, select the edit icon next to each panel. Ensure that all genes are entered according to the HUGO Gene Nomenclature Committee (HGNC) guidelines. If any gene is not recognized correctly, you will receive a notification.
Associated samples: This feature allows you to associate specific samples with your protocol. By linking samples to the protocol, you can easily access and analyze relevant data within the context of the protocol. Associated samples can provide valuable insights and streamline the analysis process.
Remember that configuring filters within a protocol applies those
specific filters only within that particular protocol. It allows you to
tailor the analysis criteria and conditions according to the
requirements of the protocol.
Additional protocol customizations
By creating and utilizing protocols, you can streamline your variant
analysis processes, maintain consistency, and apply standardized
criteria and reporting templates across different analyses and cases.
3.10.10 Filters
The Filters dialog provides the capability to customize the logic used
within each protocol. Please note that modifying filter logic requires
Administrative Privileges. Users can modify existing filters or create
new ones by selecting the New icon in the upper right corner. The
following dialog will be displayed, with the fields explained below:
Name: Enter a name for the filter.
Description: Provide a description that explains the purpose or function of the filter.
Category: Choose the category under which the filter should be organized when accessing it in the analysis.
Query: Within each section, you can apply custom filter logic by defining a Query. The Query field allows you to input a Boolean JavaScript expression that will be evaluated for each row of VCF (Variant Call Format) data. You can refer to the available filters and examples in the following link: https://github.com/geneyx/geneyx.analysis.api/wiki/FilterQuery.
Commonly used functions include “||” for OR logic, “&&” for AND logic, “!==”” for NOT EQUALS, “==” for EQUALS, and “>” or “<” for LESS THAN or GREATER THAN comparisons. If you have more complex filters that you would like to implement, you can reach out to support@geneyx.com for assistance.
Query Data: The Query Data field can be used to capture multiple values. For example, you can capture relevant genes by entering the appropriate query.
Creating a New Filter
By utilizing the flexibility of the Query and Data fields, you can
create custom filter logic that precisely matches your variant analysis
requirements.
3.10.11 Genetic Model Management
Geneyx Analysis now provides advanced customization capabilities for
genetic models at the account level, offering enhanced flexibility
across analysis and reporting workflows. This update streamlines model
governance, improves cross-analysis consistency, and significantly
enhances customization for laboratories with specialized workflows.
Genetic Model Management and Column Settings
Key updates and features include:
Renaming Genetic Models: Users can directly rename existing genetic models from a dedicated settings panel.
Custom Model Creation: The system supports the creation of new custom genetic models for both Single Nucleotide Variants (SNVs) and Copy Number Variations/Structural Variants (CNVs/SVs), with the capacity to create up to four custom models for each type.
Model Reordering: Genetic models can be reordered to control their display sequence and execution priority within each analysis.
Column settings: Columns can be hidden or removed per genetic model. (Note: The feature to hide columns was previously in the Protocol settings)
Reporting Section Assignment: Genetic models can now be directly assigned to specific sections within patient reports. This enables tailored clinical reporting for various categories, such as Primary Findings, Carrier Screening, and Pharmacogenomics (PGx).
All existing relationships between genetic models and other system
components, including protocols, filters, gene panels, variant maps, and
auto-execution filters, remain fully intact and functional.
Please note that genetic model settings will apply to all protocols to
which the genetic model is applied.
3.10.12 Secondary Pipelines
The Secondary Pipelines provide a list of available options for variant
calling in Geneyx Analysis. These options include:
DRAGEN Exome: This pipeline enables ultra-rapid analysis of NGS (Next-Generation Sequencing) data for whole exomes. It is designed to efficiently analyze genetic variants in the exome regions of the genome.
DRAGEN Exome (Clinical Exome): This pipeline enables ultra-rapid analysis of NGS data for a limited number of genes, specifically focused on genes associated with clinical exome sequencing. It allows for targeted analysis of clinically relevant genes.
DRAGEN Exome (Panel): This pipeline enables ultra-rapid analysis of NGS data for a targeted gene panel. It is designed for analyzing specific gene panels, allowing for focused variant calling in the selected genes of interest.
DRAGEN Exome (Target Region): This pipeline enables ultra-rapid analysis of NGS data for targeted regions of interest. It allows for efficient analysis of specific genomic regions, providing a more focused variant calling approach.
DRAGEN Whole Genome: This pipeline enables ultra-rapid analysis of NGS data for whole genomes. It is designed to analyze genetic variants across the entire genome, providing comprehensive variant calling results.
Sentieon Exome: This pipeline enables ultra-rapid analysis of NGS (Next-Generation Sequencing) data for whole exomes. It is designed to efficiently analyze genetic variants in the exome regions of the genome. Note: It is currently restricted to SNV only.
Sentieon Genome: This pipeline enables ultra-rapid analysis of NGS (Next-Generation Sequencing) data for whole genomes. It is designed to efficiently analyze genetic variants in the exome regions of the genome. Note: It is currently restricted to SNV only.
It is important to note that for calling of Copy Number Variations
(CNVs) and Structural Variants (SVs), a panel of normal (PON) is
required for exome and panels (PON not required for WGS). The PON
creates a baseline coverage profile against which your sample of
interest can be compared. The PON should ideally consist of 50 samples
that have undergone the same library preparation methods, although they
do not necessarily have to come from the same sequencing run. To
incorporate a PON to your account please contact support@geneyx.com.
VCF files containing pre-annotated or marked variants can be applied to
the secondary pipeline to create a unique field in the output VCF file.
These marked variants will enable the user to have a customized
annotation in the variant interface. This is available on request.
Secondary Pipelines
By utilizing the appropriate Secondary Pipeline based on your specific
analysis needs, you can achieve efficient and accurate variant calling
results in Geneyx Analysis.
3.10.13 Report Configurations
The report modification feature in Geneyx Analysis allows users to
customize the appearance and content of their reports according to their
specific needs. To create a new report, simply click on the New icon
located in the upper right corner of the interface. This will open a
pop-up window where you can configure the report settings. Here’s an
overview of the options available:
Report Name: Enter a name for the report in the provided text box.
Company/Lab Logo: Upload the logo image of your company or laboratory to be included in the report. This helps in branding and customization.
Color Scheme: Choose the desired color scheme for the report to match your branding or personal preference.
Header: Write the header of the report, which typically includes information such as the name of the company or lab, address, contact details, and any other relevant information.
Introduction Text: Provide an introduction text that explains the type of test being conducted or the intended audience of the report. This section can help set the context and provide necessary background information.
Methods: Describe the methods applied during the analysis. This section can include details about the sequencing techniques, variant calling methods, and any other relevant information about the analysis workflow.
Limitations and References, in the Report Configuration. These fields come pre-populated with default information, providing a structured foundation for report content. This ensures a more detailed and standardized approach to report generation, enhancing the overall quality and accuracy of the information presented.
Disclaimer: Include a disclaimer section where you can add any necessary disclaimers, legal statements, or other relevant information regarding the report and its limitations.
Signature: Decide whether you want to include a signature section in the report, where authorized individuals can sign off on the findings or conclusions.
CIViC Evidence: Choose whether you want to include CIViC evidence in the report. CIViC is a database of clinically relevant genetic variants and their associated evidence. Adding CIViC evidence can provide additional information and context for the reported variants.
Report Configurations
To apply the report template configuration to a protocol, you can follow
these instructions:
Start by navigating to the Settings menu and open the Protocol dialog.
Select the specific protocol you want to modify and click on the edit icon.
Within the protocol settings, you’ll have the option to update various parameters, including the selection of a pre-configured report.
Choose the report template that you have previously configured and saved.
Once you have selected the desired report, save the changes to the protocol settings.
From now on, every time you generate a new report using the updated
protocol, it will automatically incorporate the configurations from the
selected report template. This ensures consistent and efficient report
generation.
Please note that for older analyses that were performed before the
protocol update, you will need to manually apply the changes. To do
this, simply access the navigation window and click on the Reset option.
This will ensure that the updated protocol configurations are applied to
the older analyses, enabling the generation of reports with the new
settings.
By utilizing this process, you can easily apply and generate reports
using specific configurations tailored to your requirements.
Report Configuration
Together, customizing these elements, users can create reports that
align with their branding, contain all the necessary information, and
meet their specific requirements. Once the report settings are
configured, they can be saved and applied to generate customized reports
for their analyses in Geneyx Analysis.
3.10.14 Genome Browser
Geneyx supports the ability to customize the default tracks in IGV
through the settings dialog. Once configured, the selected tracks will
display automatically when IGV is opened. In IGV, SNV and SV/CNV
variants from the proband and associated samples derived from the VCF
are displayed as a new track. This will allow the exact variant to be
investigated and determine if it is shared among samples in the
analysis.
Genome Browser in the Settings interface enables user to customize the
tracks in IGV.
3.10.15 Variant Tags
Variant tags are integrated when annotating variants and the values can
be configured in the Settings dialog. These will appear in the variant
and gene interpretations in the in-house allele frequency database.
Variant Tags can be configured in the Settings interface and used when
annotating variants.
3.10.16 Gene Descriptions
Within the Setting dialog, a powerful Gene Descriptions feature has been
integrated, allowing the meticulous curation of individual or bulk gene
interpretations. This curated information reflects in the analysis when
clicking on the Gene information column, and can be transferred to the
final report, enhancing the depth and specificity of genetic insights.
Gene Descriptions can be curated and automatically incorporated into
analyses and reports.
3.10.17 Variant Knowledgebase
Geneyx’s classified variants are organized within a knowledgebase,
uniquely tied to each account. This feature provides the capability to
export and import this valuable data. An exemplary use case involves
lifting over annotated variants to an alternate genome assembly,
leveraging the transformed data as an allele frequency backlog for
augmented evidence across different assemblies.
This interface provides five different tabs: Batch Actions, hg19 SNV,
hg19 CNV/SVs, hg38 SNV, hg38 CNV/SVs.
Variant Knowledgebase Interface
Batch Actions: The Batch Actions interface enables users to import and
export variant knowledgebases. In the upper right corner, there is an
Import and Export Option. The import option will enable users to import
external data into Geneyx, which will then be displayed in the In-House
Data column for all analyses.
Clicking on Import, will provide the following options:
Select Variant Knowledgebase Import File: This allows users to browse to a local file that will be imported into Geneyx. The tool tip of this also provides users with a template to use as a reference. Additional templates can be found here: https://github.com/geneyx/geneyx.analysis.api/tree/main/templates/Variant%20KB
Import Name: Name of the imported file.
Genome Build: The reference genome assembly of the input file.
Variant Type: Is this a SNV or CNV/SVs file.
Allow New Tags: If a tag is associated with a variant that is not present in the Variant Tag option, this will apply the new tags to the database.
Overwrite Variants if exists: If a variant already exists in the knowledgebase, selecting this option will overwrite the information with what is being imported.
Hg19 SNV: If variants have been imported in the Batch Action function
with hg19 reference genome and SNV as the variant type, they will be
displayed here. This dialog gives the user the ability to modify the
existing information or import new variants by selecting New in the
upper right corner.
Hg19 SNV interface
Hg19 CNV/SVs: If variants have been imported in the Batch Action
function with hg19 reference genome and CNV/SVs as the variant type,
they will be displayed here. This dialog gives the user the ability to
modify the existing information or import new variants by selecting New
in the upper right corner.
Hg19 CNV/SVs Interface
Hg38 SNV: If variants have been imported in the Batch Action function
with hg38 reference genome and SNV as the variant type, they will be
displayed here. This dialog gives the user the ability to modify the
existing information or import new variants by selecting New in the
upper right corner.
Hg38 SNV interface
Hg38 CNV/SVs: If variants have been imported in the Batch Action
function with hg38 reference genome and CNV/SVs as the variant type,
they will be displayed here. This dialog gives the user the ability to
modify the existing information or import new variants by selecting New
in the upper right corner.
Hg38 CNV/SVs Interface
When variants have been imported into the Variant Knowledgebase, the
information will display in the IN HOUSE column of analyses.
IN HOUSE column colored green, will display information from variant
knowledgebase.
If you click on the interpretation on the Variant or Gene level, the
associated information will be displayed, including the Source column.
Information imported using the Variant Knowledgebase are represented
with a notebook icon, and have a tool tip indicting the knowledgebase it
was derived from.
In House data showing information pulled from Variant Knowledgebase.
3.11 Accounts
If you require the capability to grant permission to other accounts or
would like to request access to someone else’s data, Geneyx Analysis
supports multiple account configurations to facilitate such
collaborations. For more detailed information and assistance with
setting up this feature, please reach out to support@geneyx.com. The
Geneyx support team will be able to provide you with further guidance
and address any specific requirements or questions you may have
regarding account permissions and data access.
3.12 Variant Analysis
Once an analysis has been created and the annotation process completed,
users can perform variant analysis by navigating to the Dashboard window
and clicking on Analyze next to the subject of interest.
The analysis screen will display all variant information provided by the
VCF file as well as the annotations and genetics models that are
populated based on the specified protocol.
3.12.1 Summary Information
Summary information regarding the sample is present in the upper right
corner of the analysis page. The information that can be accessed
includes:
Subject: Provides general details about the subject. Please note that if the calculation of variant ratios on chromosome X that are not in pseudo autosomal regions differs from the assigned gender, the user will receive a notification regarding the discrepancy.
Clinical: User can review the clinical details of the analysis and keywords selected for scoring and ranking genes.
VCF Samples: Here the user can see which samples were used in the analysis, the total number of variants for those samples, and Identity by Descent.
Identity by Descent: Underlies genetically mediated similarities among individuals.
HET in both parents: reflects the total number (Count) of heterozygous variants in both parents and the number of homozygous variants in the proband (should be around 25% for a real trio).
HOM-ALT in Mother, HOM-Ref in Father reflects the alternate alleles that are potentially transmitted from the mother to the proband, missing in the father and called as heterozygous in the proband (should be around 100% for a real trio).
HOM-ALT in Father, HOM-Ref in Mother reflects the alternate alleles that are potentially transmitted from the father to the proband, missing in the mother and called as heterozygous in the proband (should be around 100% for a real trio).
Uniparental Disomy (UPD) Analysis: The UPD analysis section provides valuable information about uniparental disomy, including the chromosome number, significance (p-value), and category of inheritance. Here’s how you can interpret and explore this analysis:
Chromosome: This indicates the specific chromosome number under investigation.
Significance (p-value): The p-value represents the statistical significance of the observed UPD. As a general guideline, a p-value less than 1e-40 is considered significant.
Category of Inheritance: This categorizes the type of inheritance associated with the UPD, such as maternal hetero- or isodisomy, paternal hetero- or isodisomy, or biparental inheritance.
It’s important to note that in some cases, there may be two values of significance associated with the UPD analysis. In such situations, it is recommended to select the greater value and perform a manual inspection for further investigation.
To explore additional insights and details regarding the UPD analysis, you can click on the “Details” option. This will provide more in-depth information and help you understand the specific characteristics and implications of the identified UPD.
Understanding the results of the UPD analysis is crucial for identifying potential genomic aberrations and gaining insights into the inheritance patterns of certain chromosomal regions. It can assist in the evaluation of genetic disorders and guide further investigations or medical decision-making.
By utilizing the UPD analysis feature, Geneyx Analysis empowers users to explore and interpret the significance of uniparental disomy, facilitating comprehensive variant analysis and aiding in the understanding of genetic conditions.
VCF Summary Information
3.12.2 Variant Analysis
The table interface in Geneyx Analysis offers a user-friendly and
interactive experience for variant analysis. Here are some key features
and customization options available within the table interface:
Tabs for Genetic Models: The table interface is organized into tabs, with each tab dedicated to a specific genetic model used for the analysis. This allows users to focus on the variants relevant to a particular model and mode of inheritance.
Interactive Table: Within each tab, the table presents the variants in rows, and each column represents a specific attribute of the variant. This tabular representation enables users to efficiently navigate and analyze the data.
Collapse/Expand Categories: The attributes within the table are grouped into five categories. Users have the option to collapse or expand these categories using the ‘>’ icon located to the right of each category title. This feature helps to declutter the view and focus on the desired information.
Customizable Column Width: Users can customize the width of individual columns within the table. This flexibility allows for better visibility and arrangement of the data. If needed, the settings icon provides an option to restore the default column width.
Variant Analysis Interface
By offering a well-organized and customizable table interface, Geneyx
Analysis empowers users to efficiently explore and analyze variants
across different genetic models. The interactive nature of the table
interface enhances usability and allows for a more tailored and
comprehensive variant analysis experience.
3.12.3 Filters and Tools
The Filters dialog, located on the left side of the interface, provides
a summary of the applied filters for the current tab. This pane gives
users an overview of the filters that have been applied to refine the
variant analysis.
The variant dashboard in Geneyx Analysis provides the flexibility to
apply filters to every column in the table. To apply a filter, simply
click on the filter icon located next to the column header. This action
will open a dialog where you can define the desired threshold or
criteria for filtering.
By setting a threshold, the variant table will dynamically update to
display only the variants that meet the specified criteria. This
interactive filtering capability allows users to focus on specific
subsets of variants based on their desired thresholds or criteria.
Furthermore, each applied column filter will be reflected in the filter
dashboard on the left side of the interface. This dashboard provides a
summary of all the applied filters, allowing users to easily track and
manage their filtering logic.
In addition to the filters derived from the columns in the table, users
can apply further filtering using predefined Gene Panels, Gene Lists,
and Variant Maps. To access these additional filters, click on the
Settings icon located on the left side of the window. Once the Settings
icon is selected, a new window will appear, providing options to create
and manage Gene Panels and Variant Maps. These custom filters can be
created based on specific criteria or genetic regions of interest.
Once created, these Gene Panels and Variant Maps are stored internally
within the Geneyx Analysis system and can be applied to all analysis
cases. This allows users to consistently apply the same predefined
filters across different analyses, saving time and ensuring consistency
in the filtering logic.
Filter logic pane
In addition to the basic column filters, Geneyx Analysis offers the
flexibility to create customizable filters using complex logic. These
advanced filters can be created and configured in the Filters section
within the Settings menu.
With the advanced filters, users have the ability to define intricate
criteria and conditions based on multiple attributes and parameters.
This allows for more refined and specific filtering of variants based on
research or clinical requirements. Filter presets can be created in the
Settings dialog and easily applied by searching for the name in the
Filter window. Furthermore, if the filters need to be reset to the
original filters, the user can select the gear icon on the right side.
This will revert the filters to the original presets.
Resetting the filter to original presets.
However, if you encounter any difficulties or challenges in creating the
desired criteria using the advanced filters, the Geneyx support team is
readily available to assist you. You can reach out to
support@geneyx.com for guidance and support in creating the custom
filters according to your specific needs.
The support team will be able to provide you with expert assistance and
help you configure the filters to effectively capture the variants of
interest, ensuring a seamless and efficient analysis experience.
Together, the ability to apply column filters empowers users to explore
and analyze the variants based on specific attributes or criteria of
interest. It facilitates the customization of the analysis to suit
individual research or clinical requirements, providing a comprehensive
and efficient variant filtering experience.
Another useful tool of the Geneyx application is the magnification icon
in the upper right corner. This feature enables the user to enter a gene
symbol and see if any SNV or CNV/SV events are present and is
independent of any filters that are applied. It serves as a holistic
approach to view all variants of a given gene and associated coverage.
Magnification icon enables user to search for all variants (SNV/CNV) and
coverage using gene input.
3.12.4 Genetic Models
Genetic models in Geneyx Analysis play a crucial role in identifying
relevant variants based on the mode of inheritance. To navigate between
different genetic models, simply click on the tab corresponding to the
model of interest. Each genetic model focuses on specific inheritance
patterns and provides a targeted analysis approach. Genetic models can
also be applied or removed from protocol, to do this you will need to go
to Protocols in the Settings dialog.
Genetic models can also incorporate specific filter logic tailored to
each model. These filters can be configured in the Filter dialog
accessible through the Settings menu. By customizing the filters within
each genetic model, you can refine the analysis results to meet your
specific requirements and focus on variants that are most relevant to
the selected mode of inheritance.
Here are some of the genetic models available in Geneyx Analysis:
Fast Track: This model applies filters designed to identify the most relevant variant(s) based on the clinical phenotype, if available, or the pathogenicity of the variants.
Recessive HOM: This model focuses on displaying clinically relevant variants associated with recessive inheritance, specifically those where the proband is homozygous.
Recessive Compound HET: This model highlights clinically relevant variants associated with recessive inheritance, specifically those where the proband is heterozygous.
Dominant HET: This model showcases clinically relevant variants associated with dominant inheritance, particularly those where the proband is heterozygous.
Mitochondrial: This model specifically displays clinically relevant variants associated with mitochondrial DNA.
CNV: This model is designed to identify clinically relevant copy number variants.
Incidental findings: This model focuses on displaying variants associated with ACMG incidental findings genes.
Genetic Models based on mode of inheritance
By selecting the appropriate genetic model and applying the associated
filters, you can effectively narrow down and analyze the variants that
are most relevant to the mode of inheritance or specific clinical
considerations. This enhances the efficiency and accuracy of variant
interpretation and aids in making informed decisions during the analysis
process.
3.12.5 Annotations
Within each genetic inheritance model, Geneyx Analysis provides
annotations and algorithms that aid in identifying candidate variants
highly associated with the patient’s phenotype. These annotations can be
applied to the filter logic using the filter icon next to each column.
By clicking on any of the ‘Relevance’, ‘Pathogenic’, ‘Notes’ or ‘Tags’
columns, the following annotations are available for completion:
Once a variant in annotated, the annotation information will be stored
for the variant, which will be available for all downstream analyses and
can be included in the final report. These columns are:
Relevance: This annotation allows you to set the relevance of a variant, determining whether it should be included in the final report.
Pathogenic: With this annotation, you can indicate the pathogenicity of a variant, determining its inclusion in the report.
Validation: Does the annotation require validation
Notes: You can add insights and notes about a variant, including its potential relevance to the diagnosis or any additional information deemed important.
Tags: Variant tags can be included in the annotation and are created in the Settings dialog.
Variant Interpretation: This can be used to provide context regarding the variant and will be stored for all downstream encounters of the variant.
Variant Level Recommendations: This can be used for recommendations in the annotation process.
References: Users can add PMID (up to 15) in the format PMID:XXXXXX or XXXXXX
Gene Description: Populated from the Gene Description section within the Settings interface.
Phenotype & Evidence: This enables users to select information from public annotation sources.
ACMG criteria applied to the variant will also be visible in this window.
Note: once pathogenicity or relevance are assigned to a variant the variant will be blocked for further editing. To allow continued editing return to the genetic model where pathogenicity/relevance were assigned and remove criteria.
The variant dashboard in Geneyx Analysis also consists of several public
annotations, each providing valuable information. Here are the column
details:
LOCATION: This column displays the genetic location of the variant based on the selected genetic assembly (GRCh37/38). When a variant is selected for visualization by clicking on the location, IGV will display in a new window. This will allow for visualization to occur on a different monitor and can be easily referenced using the URL. Clicking on the location will direct you to the adjusted view in the IGV genome browser. Alternatively, you can select the UCSC icon in the top right to view the variant in the UCSC genome browser.
GENE: This column displays the gene associated with the specific variant. The model of inheritance (AR = Autosomal Recessive, AD = Autosomal Dominant, XL = X linked, DG = Digenic) is indicated. Icons displayed next to genes indicate if a pseudogene (P), ACMG secondary finding gene (SF) or addition information (I), where tooltips will display the added information.
Gene information tips
– Clicking on the gene name will provide access to clinical
information, pathways, and drugs associated with the gene. This includes
OMIM/ClinVar, Expression, Pathways, and Coverage. Within this section,
you will find the following details:
GDI (Human Gene Damage Index): GDI measures the accumulated mutational damage of each human gene in the healthy human population. It helps identify genes with a higher likelihood of being disease-causing based on the mutational data from the 1000 Genomes Project database.
GnomAD Missense Z-score: This score measures the constraint or intolerance of a gene or transcript to missense variations based on the gnomAD dataset. Positive Z-scores indicate more constraint, while negative scores indicate less constraint.
The following metrics, derived from gnomAD (Genome Aggregation Database), are utilized within Geneyx Analysis to assess gene constraint and tolerance to variation. These metrics are now available for both gnomAD v2.1.1 and gnomAD v4.0.1 datasets:
pLI (Probability of Loss-of-function Intolerance): pLI scores measure how constrained or intolerant a gene or transcript is to loss-of-function variations. Higher pLI scores indicate intolerance to loss-of-function variants.
Loss-of-function: This section provides information on the gene’s tolerance to protein truncating variations, such as nonsense, frameshift, splicing, and deletion variations.
O/E (Observed/Expected): The O/E constraint score represents the ratio of observed to expected variants in a gene. It provides a measurement of a gene’s tolerance to a specific class of variation.
#HOM LOF: gnomAD homozygous LoF count
# HET LOF: gnomAD heterozygous LoF count
OMIM/ClinVar: This section displays gene-phenotype associations according to OMIM (Online Mendelian Inheritance in Man) and summarizes pathogenic/likely pathogenic single nucleotide variants (SNVs) and copy number variants (CNVs) in the gene based on ClinVar. A newly added Edit History column in the OMIM gene table displays the most recent update date for each entry (e.g., “05/03/2022”), providing users with context when reviewing genes.
Clicking on the “+” icon for a given phenotype will display the clinical synopsis according to OMIM. The Clinical Synopsis section in the Online Mendelian Inheritance in Man (OMIM) database offers concise gene-specific summaries of clinical features and symptoms associated with genetic disorder/syndrome, along with information about the molecular basis of the associated disease.
Phenotypes Associated with Gene: This section presents the known conditions associated with the gene based on OMIM phenotypes, along with the associated inheritance model and relevant documentation.
Diseases Associated with in ClinGen: ClinGen (Clinical Genome Resource) defines the clinical relevance of genes and variants, and the ClinGen Variant Curation Expert Panels (VCEPs) play a crucial role in this process by providing expert evaluation of genetic variants to determine their clinical significance. This information is now available in the Gene information for the disease if available.
Digenic Inheritance: Digenic inheritance refers to a mode of genetic inheritance in which two genes contribute jointly to the phenotype, meaning that variations in both genes are necessary to produce a specific trait or disease. Unlike classical Mendelian inheritance, where a single gene mutation is typically responsible for a disorder, digenic inheritance involves the interaction between mutations in two different genes. If a gene is associated with a phenotype with digenic inheritance, the second gene will be displayed in the Genes tab of the application with associated phenotype.
DECIPHER: Disease associated with the gene from the DECIHER resource
Orphanet: Disease associated with the gene from the Orphanet resource
SFARI: Phenotypes and disorders associated with the gene from SFARI resource
Variants: Provides a structured view of ClinVar variant classifications for the selected gene, including detailed tables on pathogenic SNVs, CNVs by condition/effect, and an overall pathogenicity summary.
Coverage: If applicable, this column displays the coverage of the gene, providing information on the sequencing depth for that specific gene.
Gene Description: This will pull in interpretations made on the gene level, which is available in the Settings dialog in gene descriptions.
This category provides information regarding
variant nomenclature. The fields include:
REF – The allele present in the reference genome.
ALT – The alternative call for the sample.
AA – The amino acid change, if applicable.
HGVS – The representation of the variant in HGVS (Human Genome Variation Society) nomenclature including full HGVS-compliant annotations for frameshift variants (e.g., p.Leu1204ValfsTer153) for clarity.
ZYG – The zygosity of the variant, indicating whether it is heterozygous, homozygous, or hemizygous.
A1, A2: Allele 1/Allele 2 (in case phasing blocks are present in the VCF file) – These provide insight into variant phasing by resolving haplotypes without maternal and paternal sequencing information.
Phasing block: The phasing block in which this variant is found. Phasing information will be available if present in the VCF file.
REFSEQ – The relevant transcript ID(s) that pertain to the effect on the protein. This track contains RefSeq Gene transcripts annotated by the NCBI Homo sapiens Annotation Release. The canonical transcript is now selected according to Matched Annotation from the NCBI and EMBL-EBI (MANE). If a MANE transcript is not available, the canonical transcript will be selected based on shortest transcript ID.
Exon #: The impacted exon of the gene.
Splice Distal: Distance of the splice site to the canonical splice site location.
CODON – The codon change caused by the variant.
DBSNP – Known identifier (dbSNP RSID). This track displays common single nucleotide polymorphisms (SNPs) from dbSNP build 151. dbSNP is a database that catalogs short variations in nucleotide sequences from various organisms. It includes single nucleotide variations, short nucleotide insertions and deletions, short tandem repeats, and microsatellites. Common SNPs are those with a minor allele frequency of at least 1% in at least one 1000Genomes population, contributed by two or more founders. Some entries in dbSNP may have additional information associated with them, such as disease associations, genotype information, and allele origin.
DBSNP version – The earliest dbSNP version that included the variant.
ACMG:
The American College of Medical Genetics and Genomics (ACMG) has
developed guidelines for interpreting sequencing variants. These
guidelines are used to classify variants based on various types of
evidence, including population data, computational data, functional
data, and segregation data.
Dom – This option aggregates data from public databases using the ACMG Guidelines for variants associated with dominant inheritance.
Rec – This option aggregates data from public databases using the ACMG Guidelines for variants associated with recessive inheritance.
By selecting one of these options, you will be directed to the full ACMG
classification panel. On the left side, you will find different
categories of the ACMG guidelines, and the middle dialog displays the
specific ACMG criteria. Applicable criteria will be solid, while empty
criteria indicate insufficient evidence for application. Crossed-out
criteria indicate that they are not applicable or that the variant
evidence contradicts the criteria.
Population frequency information is from gnomAD v2 for GNE and from
gnomAD v4.1 for GNG.
Hovering over an individual criterion will provide additional details,
and clicking on a cell allows for manual override of the
autoclassification. The supporting evidence for each criterion is
available on the right-hand side of the panel.
Associated Samples:
The “Associated samples” section provides information about the zygosity (heterozygous or homozygous) of the variant in parents, if parental samples were submitted for analysis. This allows for the assessment of variant inheritance patterns and can be useful in determining the potential relevance of the variant in relation to the patient’s phenotype.
Variant Calling Q&R (Quality and Read Information):
This category provides information extracted from the VCF (Variant Call
Format) file, which includes details about the quality and
characteristics of the variant calls.
Q&R: This field classifies the variant calls based on their quality. Variants are categorized as Low quality if the coverage is less than 10x and the genotype quality (GQ) is less than 15. Variants are categorized as Medium quality if the coverage is greater than 10x or the GQ is greater than 15. Variants are categorized as High quality if the coverage is greater than 20x and the GQ is greater than 50.
Depth: This indicates the total read depth at the variant position, which represents the number of reads covering that particular genomic region.
DP2: This field provides the read counts for the reference allele (Ref) and the alternative allele (Alt), respectively. It gives insights into the distribution of reads supporting each allele.
%ALT: The percentage of reads that display the alternative allele among all the reads at the variant position. This provides an estimation of the allelic frequency based on the sequencing data.
GQ: The variant calling quality score (GQ) represents the confidence in the genotype call for that specific variant. A higher GQ value indicates a higher confidence in the called genotype.
FILTER: This field indicates the quality filter status assigned by the variant caller algorithm. It helps determine if the variant passes specific quality control criteria or if it should be flagged as potentially unreliable.
PL: Phred-scaled genotype likelihoods (HOM REF, HET, HOM ALT) represent the log-scaled likelihoods of each possible genotype. These values provide additional information about the likelihood of each genotype given the observed sequencing data.
AMP SCORE: The amplification score represents the coverage at the variant position relative to the median coverage across the genome. It provides insights into the level of amplification or depletion of reads at the variant location compared to the average coverage.
These variant calling and read information metrics help assess the
quality, depth, genotype likelihoods, and other relevant characteristics
of the variant calls in the analysis.
Clinical Evidence:
This category provides information about the clinical evidence
associated with the variants.
PHENO: The PHENO score represents the association between the selected phenotypes and the gene. It is calculated using the Geneyx Knowledgebase and Geneyx Phenotyper (Pheneyx), which integrates major clinical data sources (OMIM, ClinVar, OrphaNet, and HPO) and utilizes advanced matching capabilities to establish direct and indirect associations between genes and biomedical information.
This feature includes: matched and unmatched phenotypes to the gene, diseases and variants related and the publication year of displayed publications. Now, users can easily discern the chronological context of phenotypic data, adding a valuable layer of context to the prioritization process.
MATCHED PHENOTYPES: This field indicates the number of matching terms between the selected phenotypes and the gene. It helps identify the degree of phenotypic overlap between the gene and the observed patient characteristics.
CLINVAR: The CLINVAR database, maintained by NCBI, provides information about the phenotypes and supporting evidence for variants listed in the dbSNP database. It offers valuable insights into the clinical significance and implications of specific variants. For more detailed information about the ClinVar database, please refer to the “About ClinVar” section.
REVIEW STATUS: One of the key features of ClinVar is its review status for each variant, which indicates the level of evidence and consensus supporting the clinical significance of the variant. The review status of ClinVar submissions is a separate column that can also be filtered on.
CLN-AA: The CLN-AA field represents the clinical significance of matching ClinVar entries based on amino acid changes. Clicking on the hyperlink provides access to clinical relevance and summaries for pathogenic entries, offering valuable insights into the clinical interpretation of the variant.
OMIM: OMIM (Online Mendelian Inheritance in Man) is a comprehensive and authoritative compendium of human genes and genetic phenotypes. It is freely available and regularly updated. The OMIM track displays the names of associated HPO (Human Phenotype Ontology) phenotypes, providing a valuable resource for understanding the genetic basis of various phenotypes.
OMIM Inheritance: The OMIM Inheritance track displays information about the inheritance patterns associated with the gene and its corresponding phenotypes. The inheritance patterns include Autosomal Dominant (AD), Autosomal Recessive (AR), Multifactorial (Mu), Somatic Mutation (SMu), Digenic Dominant (DD), Digenic Recessive (DR), Isolated Cases (IC), Mitochondrial (Mi), Pseudoautosomal Dominant (PADom), Pseudoautosomal Recessive (PARec), Somatic Mosaicism (SomMos), X-Linked (XL), X-Linked Dominant (XLD), X-Linked Recessive (XLR), and Y-Linked (YL).
CLINGEN HAPLO: ClinGen haplosensitivity and hyperlink (in CNV/SV genetic model only)
CLINGEN TRILPO: ClinGen triplosensitivity and hyperlink (in CNV/SV genetic model only)
CIVIC: CIViC is an open-access, community-driven web resource for the clinical interpretation of variants in cancer. It aims to facilitate precision medicine by providing an educational platform for disseminating knowledge and fostering active discussions on the clinical significance of cancer genome alterations. The information includes the CIViC Evidence ID (EID), Disease, and Drugs associated with the variant.
EL: The Evidence Level field denotes the experimental method from which the evidence statement is derived. It ranges from inferential associations from a single experiment to trusted associations that routinely inform clinical action.
ET: The Evidence Type field describes the predictive, prognostic, or diagnostic association between an evidence statement and a variant, providing insights into the nature of the evidence.
TR: The Trust Rating field provides a rating out of 5 stars, indicating the level of trustworthiness or reliability assigned to the evidence.
PUBS: The PUBS field represents the number of relevant publications available in the MasterMind database, offering additional scientific literature supporting the variant.
LitVar is an integrated annotation source that will enhance phenotypic prioritization through improved literature mining and can be used as an independent annotation source. LitVar allows the search and retrieval of variant specific information from relevant studies in the literature, with related concept (e.g., diseases) annotations. By normalizing variant names, LitVar returns the same results regardless of which name of a variant (e.g. BRCA1 p.P871L or c.2612C>T) is used in the query.
GoogleSearch, Geneyx supports the ability to easily search for the variant in google with a useful hyperlink output. This can then be used for other downstream purposes, and aims to enhance searching functionalities.
This clinical evidence section offers valuable insights into the
phenotypic associations, clinical significance, inheritance patterns,
and supporting literature associated with the variants under
consideration.
IN HOUSE:
This section provides information about the frequency of the variant
observed internally across samples.
V: The V field contains in-house annotations for this variant. It includes associated Date, Relevance, Pathogenicity, Notes, and Analyses, providing additional context and details about the variant. This will display the most recent classification (by pathogenicity) in the local allele frequency database. This classification can be applied as a filter and will enhance the understanding of variants and genes previously evaluated.
G: The G field includes the Date and Time of observed variants, amino acid (AA) changes, Zygosity Effect, Relevance, Pathogenicity, Notes, and Analyses. This information helps in understanding the specific characteristics and implications of the variant.
AF (%): The AF (%), or Allele Frequency, represents the frequency of the variant allele in all the samples of the current account. It is expressed as a percentage, providing an indication of how common the variant is within the sample population.
#Local Het: The #Local Het field indicates the count of samples within this account that have this heterozygous variant. It provides insights into the number of individuals carrying one copy of the variant allele in the local sample population.
#Local Hom: The #Local Hom field represents the count of samples within this account that have this homozygous variant. It helps identify the number of individuals with two copies of the variant allele in the local sample population.
#Local Hem: The #Local Hem field displays the count of samples within this account that have this hemizygous variant. It indicates the number of individuals with a single copy of the variant allele in the local sample population, typically applicable to variants on sex chromosomes.
The IN HOUSE section provides information about the variant’s frequency
and characteristics within the internal sample population, aiding in the
assessment and interpretation of the variant’s significance.
Matched CNV/SVs:
Enhancing the SNV genetic models, Geneyx incorporates CNV/SV visibility
within the analysis. This column displays the total number of CNVs
overlapping the given SNV position. Additionally, a dedicated column
highlights CNV events predicted to be damaging, encompassing events that
overlap exons and exhibit an allele frequency below 5%. To further
simplify your analysis, you can now search for CNVs using the magnifying
glass tool, providing efficient access to the desired information.
Effect & Prediction:
This section provides relevant information about the effect of the
variant, including functional and conservation annotations.
EFFECT: The EFFECT field describes the impact of the variant on the protein, including splicing effects. It provides various options, such as downstream, exon, frameshift, missense, synonymous, start gained, stop change, and more. This information helps understand the specific alteration caused by the variant.
SEV: The SEV field indicates the severity of genetic variants based on their coding or splicing effects. Variants are categorized into high, medium, or low severity. High severity variants include loss of function (LOF), splicing, and coding variants with specific in-silico prediction scores. Medium severity variants have less stringent criteria than high severity variants. Low severity variants meet specific criteria, such as being in the splice site region or a LOF variant with low scores for CADD or Splice-AI.
CADD (PHRED): The CADD (PHRED) field provides the PHRED-scaled scores (ranging from 1 to 99) for Combined Annotation Dependent Depletion (CADD) scores. These scores indicate the deleteriousness of single nucleotide variants, indels, and insertion/deletions in the human genome. Higher scores suggest a higher likelihood of deleterious effects.
CADD (RAW): The CADD (RAW) field provides the raw CADD scores, also known as “C-scores,” which come directly from the Support Vector Machine model. These scores have relative meaning, with higher values indicating a higher likelihood of deleterious effects.
PLI: The PLI field represents the gnomAD probability of being a LoF (Loss of Function) intolerant variant. The pLI score measures how constrained or intolerant a gene or transcript is to protein-truncating variation. Higher pLI scores indicate greater intolerance to truncating variants.
REVEL: The REVEL field displays the Rare Exome Variant Ensemble Learner (REVEL) score, which predicts the pathogenicity of missense variants. It combines scores from multiple individual tools to assess the likelihood of a variant being disease-causing. Scores range from 0 to 1, with higher scores indicating a higher likelihood of pathogenicity.
SPLICE-AI: The SPLICE-AI field provides information about the SpliceAI score, which predicts splice junctions and identifies noncoding genetic variants that cause cryptic splicing. The score has different cutoffs indicating high recall, recommended, and high precision. Precomputed scores for whole genomes are available to improve annotations for non-coding variants in addition to hyperlink to SpliceAI look-up (https://spliceailookup.broadinstitute.org/) for masked scores and to change max distance
AlphaMissense: AlphaMissense is a predictive model that generates pathogenicity scores for single nucleotide missense variants in human protein-coding genes. The predictions cover a vast number of possible variants (71 million) across approximately 19,000 genes. The pathogenicity scores provided by AlphaMissense range from 0 to 1 and can be interpreted as the predicted probability of a variant being clinically pathogenic. Higher scores suggest a higher likelihood of the variant contributing to a clinical condition.
PHYLOP: The PHYLOP field represents the PhyloP conservation score, which measures the evolutionary conservation at individual alignment sites based on multiple alignments of mammalian genomes. Higher scores indicate greater conservation.
GERP_RS: The GERP_RS field displays the Genomic Evolutionary Rate Profiling Rejected Substitutions (GERP NR) scores. These scores estimate the evolutionary constraint of specific positions, with higher scores indicating greater conservation.
LRT PRED, MUTTASTER, POLYPHEN2 HDIV, POLYPHEN2 HVAR, SIFT: These fields provide predictions from various tools, such as Likelihood Ratio Test (LRT), MutationTaster, PolyPhen2 (HDIV and HVAR), and SIFT. These tools assess the deleteriousness or tolerability of variants based on specific criteria.
UNIPROT: The UNIPROT field shows the accession of the protein in the Universal Protein Resource (UniProt), which is a comprehensive database of protein sequence and annotation data.
ADA SCORE, RF SCORE: The ADA SCORE and RF SCORE fields present prediction scores based on ada-boost and random forests, respectively. These scores estimate the probability of a splicing variant affecting splicing, with suggested cutoffs for binary predictions.
Frequency:
This section provides information about the frequency of the variant
across different population databases.
MAX AF (%): The MAX AF field displays the maximum observed allele frequency in population databases such as 1000 Genomes, , and gnomAD. It represents the highest frequency at which the variant has been observed in these databases.
GNEV4 AF (%): gnomAD v4 exomes allele frequency (for GRCh38 samples only)
GNGV4 AF (%) gnomAD v4 genomes allele frequency (for GRCh38 samples only)
COLORSDB AF (%): CoLoRSdb allele frequency
#HOM: The #HOM field represents the count of samples in gnomAD where the variant is homozygous. It indicates the number of individuals in the gnomAD database who have two copies of the allele.
#HET: The #HET field represents the count of samples in gnomAD where the variant is heterozygous. It indicates the number of individuals in the gnomAD database who have one copy of the allele.
#HEMI: The #HEMI field represents the count of samples in gnomAD where the variant is hemizygous. It indicates the number of individuals in the gnomAD database who have a single copy of the allele on the X chromosome.
GnomAD Flag: The GnomAD Flag field provides information about quality control filters applied to the variant in GnomAD exomes and genomes. These filters ensure the reliability and accuracy of the variant data.
GNOMAD-v4 Hom: The count of samples with this homozygous variant in gnomAD v4 (for GRCh38 samples only)
GNOMAD-v4 Het: The count of samples with this heterozygous variant in gnomAD v4 (for GRCh38 samples only)
COLORSDB HOM: The count of samples with this homozygous variant in CoLoRSdb
COLORSDB HET: The count of samples with this homozygous variant in CoLoRSdb
1K GENOME AF %: The 1K GENOME AF field displays the allele frequency of the variant in the 1000 Genomes Project database. It represents the frequency at which the variant has been observed in the 1000 Genomes population.
ESP African AF %: The ESP African AF field represents the allele frequency of the variant in the Exome Sequencing Project (ESP) African American population. It indicates the frequency at which the variant has been observed in this specific population.
ESP European AF %: The ESP European AF field represents the allele frequency of the variant in the ESP European population. It indicates the frequency at which the variant has been observed in this specific population.
DBSNP MAF %: The DBSNP MAF field displays the minor allele frequency of the variant in the dbSNP database. It represents the frequency at which the variant has been observed in the general population according to dbSNP.
HAN AF %: The HAN AF field represents the allele frequency of the variant in the CONVERGE project for the Han Chinese population. It indicates the frequency at which the variant has been observed in this specific population.
GNE AF: The GNE AF field displays the gnomAD Exomes allele frequency. It represents the frequency at which the variant has been observed in the gnomAD Exomes population.
GNG AF: The GNG AF field displays the gnomAD Genomes allele frequency. It represents the frequency at which the variant has been observed in the gnomAD Genomes population.
GNE controls AF (%): The GNE controls AF field represents the allele frequency of the variant in gnomAD exomes controls. It indicates the frequency at which the variant has been observed in the gnomAD exomes control group.
GNG controls AF (%): The GNG controls AF field represents the allele frequency of the variant in gnomAD genomes controls. It indicates the frequency at which the variant has been observed in the gnomAD genomes control group.
GNE AMR: gnomAD exomes AMR allele frequency
GNE AFR: gnomAD exomes AFR allele frequency
GNG AFR: gnomAD genomes AFR allele frequency
GNE ASJ: gnomAD exomes ASJ allele frequency
GNG ASJ: gnomAD genomes ASJ allele frequency
GNE EAS: gnomAD exomes EAS allele frequency
GNG EAS: gnomAD genomes EAS allele frequency
GNE FIN: gnomAD exomes FIN allele frequency
GNG FIN: gnomAD genomes FIN allele frequency
GNE NFE: gnomAD exomes NFE allele frequency
GNG NFE: gnomAD genomes NFE allele frequency
GNE OTH: gnomAD exomes OTH allele frequency
GNE SAS: gnomAD exomes SAS allele frequency
Mitochondrial:
This section provides information specific to mitochondrial variants.
Heteroplasmy: Heteroplasmy is a measure of the presence of different mitochondrial DNA variants within an individual. It is calculated by dividing the depth of informative reads supporting the reference allele by the total depth of informative reads at that particular location. Heteroplasmy indicates the degree of variability or mixture of mitochondrial DNA variants within an individual.
MitoMap: MitoMap is a valuable resource that contains published data on human mitochondrial DNA variation. The variant tables in MitoMap provide information such as frequencies of mitochondrial DNA variants, which are derived from 48,882 full-length human mitochondrial DNA sequences. MitoMap is a reliable source for exploring and understanding mitochondrial DNA variation. For more detailed information about MitoMap, please refer to the “About MITOMAP” section.
CNV/SV Annotations:
The CNV/SV (Copy Number Variant/Structural Variant) annotations provide
detailed information about structural variations and copy number
variants. These annotations offer insights into the analysis of
CNVs/SVs, including various annotation sources and visualization
capabilities. Additionally, they allow for a closer examination of
individual genes affected by the event.
Location: This indicates the genetic location of the structural variation. Selecting the location provides further details in the IGV genome browser.
Genomic and Genetic Data:
This category provides information about the size and characteristics of
the structural variation.
REF: Reference bases if available
Length (LEN): The length of the CNV/SV in base pairs.
%RMSK: The percentage of DNA sequences spanning interspersed repeats and low complexity DNA sequences.
Cytogenetic Region: The cytogenetic region associated with the variant. CNVs and SVs are annotated with frequency databases according to the matching effect. CNV/SV frequency databases include DGV, DGV Gold, gnomAD SV, and genomes.
Start: The starting position of the structural variation.
End: The ending position of the structural variation.
Zygosity: Zygosity of CNV/SVs
Junction reads: Number of reads spanning the event junction
Spanning reads: Number of reads aligning near the event junction
Associated samples: The number of overlapping CNV/SVs in associated samples (if applicable).
Variant Calling Q&R:
Score: Quality score of the event.
ROH Score: Score for runs of homozygosity events. The score increases with every additional homozygous variant (0.025) and decreases with a large penalty (1-0.025) for every heterozygous SNV.
CAMO(%): Dark and camouflage genes refer to those that are difficult to pinpoint due to various reasons, such as their location in genomic regions with complex structural variations, repetitive sequences, or areas prone to technical challenges in sequencing. Annotating Dark and Camouflage genes helps to improve the accuracy and comprehensiveness of CNV analysis. As such Dark (low sequencing depth) and Camouflage (ambiguous alignment) are now new columns integrated into the CNV/SV tab and displayed in IGV.
DARK(%):Dark and camouflage genes refer to those that are difficult to pinpoint due to various reasons, such as their location in genomic regions with complex structural variations, repetitive sequences, or areas prone to technical challenges in sequencing. Annotating Dark and Camouflage genes helps to improve the accuracy and comprehensiveness of CNV analysis. As such Dark (low sequencing depth) and Camouflage (ambiguous alignment) are now new columns integrated into the CNV/SV tab and displayed in IGV.
FFPM: Fusion fragments per million total reads
Effect & Prediction:
This section presents information related to the effect and predicted
consequences of the variant.
Effect: Describes the effect of the variant, including Breakend (BND), Deletion (DEL), Duplication (DUP), Insertion (INS), Runs of Homozygosity(ROH), and Inversion (INV).
ROH refers to a stretch of consecutive homozygous genetic markers on both chromosomes, which can be indicative of offspring from consanguineous marriages and increase the risk of autosomal recessive disorders.
Copy Number: The number of copies of the CNV/SV event.
Zygosity (ZYG): The genotype call in the proband sample, indicating if it is heterozygous (HET), homozygous (HOM), or hemizygous (HEMI).
A1/A2: This reflects the allele phasing if available in the VCF file.
Genes: The number of genes affected by the event.
Max Exons (M.EXONS): The maximum number of exons affected within genes.
Enhancers (ENHS): The number of enhancers impacted by the event.
Repeat: If the variant is a repeat variant, information about the repeat will be presented, including the number in the wild-type sequence, the repeat sequence itself, and the number in each allele.
Gene fusion: Pair of genes involved in fusion event
If associated samples are applied to the analysis, such as Trio
workflows, the interface will now display events (CNV/SV) that directly
match the proband. This will enable detection of de novo or transmitted
events. Geneyx also displays the Reference and Alternate alleles, if
present, as well as read depth and variant allele frequency.
Clinical Evidence:
This category displays clinical evidence associated with the variant
based on annotation sources.
Pheno (Max): The score of the association between selected phenotypes and the gene. This score utilizes the Geneyx Knowledgebase and Geneyx Phenotyper (Pheneyx) to consolidate clinical data sources such as OMIM, ClinVar, OrphaNet, and HPO.ClinVar: Clinical significance of matching ClinVar entries.
OMIM Inheritance Statistics: A summarized inheritance pattern provided by OMIM for impacted genes.
In House:
#LV: The total number of local variants out of all samples present in the account.
#Local Het: The count of samples with this heterozygous variant in the account.
#Local Hom: The count of samples with this homozygous variant in the account.
#Local Hemi: The count of samples with this hemizygous variant in the account.
Matched SNVs:
#SNV: The total number of SNVs found within the genomic location of the CNV/SV within the sample.
#D-SNVs: The total number of potentially deleterious SNVs (VUS, LP, and Pathogenic) found within the genomic location of the event within the sample.
Frequency:
Frequency information regarding the structural variation in population
databases.
Max AF (%): The maximum allele frequency for this event (>80% overlap) based on DGV, DGV-Gold, gnomAD-SV, and gnomAD genomes.
DGV AF (%): The highest overlapping variant allele frequency for this event in DGV.
DGV-Gold AF (%): The highest overlapping variant allele frequency for this event in DGV Gold.
GDS AF (%): The highest overlapping variant allele frequency for this event in gnomAD.
Genes Affected by CNV:
Once a CNV (Copy Number Variant) of interest has been identified, it is
possible to investigate the individual gene level to gain more insights.
When a CNV is selected, the bottom window will update to display the
genes affected by the CNV.
The information at the gene level includes:
GENE: Displays the impacted gene. Clicking on the gene will provide access to clinical information, pathways, and drugs associated with this gene.
OVERLAP: Indicates if the event is a full or partial overlap of the gene.
REFSEQ – Transcript according to which gene is affected
PHENO: The score of the association between the selected phenotypes and the gene. This scoring algorithm utilizes the Geneyx Knowledgebase and Geneyx Phenotyper (Pheneyx), which consolidates major clinical data sources such as OMIM, ClinVar, OrphaNet, and HPO. It employs advanced matching capabilities with direct and indirect associations between genes and biomedical information.
MATCHED PHENOTYPES: The number of matching terms and the relevant list of phenotype terms associated with the gene.
IMPACTED EXON(S): Exon(s) affected by the event according to MANE transcript
#CODING EXONS: number of coding exons affected within this gene by the event.
#NON_CODING EXONS: number of non-coding exons affected within this gene by the event.
ENH SCORE: Enhancer confidence score.
ENH-GENE SCORE: Gene-Enhancer associated score.CLINVAR: Clinical significance of matching ClinVar entries. Clicking on it provides relevance and summary for pathogenic entries.
OMIM (GENE): The gene phenotype as defined by OMIM.
OMIM INHERITANCE: The disease inheritance pattern defined by OMIM.
#GNF: Number of features in gnomAD SV that fully overlap this gene/enhancer.
#GNP: The number of features in gnomAD SV that partially overlap the gene/enhancer.
#D-HOM SNVs: Total number of homozygous deleterious SNVs (Single Nucleotide Variants) found within the genomic location of the CNV and the gene.
#D-HET SNVs: Total number of heterozygous deleterious SNVs found within the genomic location of the CNV and the gene.
The CNV event can also be visualized using the IGV genome browser, which
can be expanded using the icon in the upper right corner. Alternatively,
if preferred, this event can be plotted in the UCSC browser using the
hyperlink in the upper right corner.
a CNV/SV ACMG Guidelines
The CNV/SV tab supports the technical standards for reporting of
constitutional copy number variants according to the joint consensus
recommendations of ACMG and ClinGen, reference article here
https://pubmed.ncbi.nlm.nih.gov/31690835/ . ACMG guidelines are
calculated automatically for all deletion and duplication events and
there is an interface to modify criteria using internal evidence. The
interface reflects a similar approach as the ClinGen CNV calculator with
full transparency of activated or inactivated criteria, as well as audit
trails for all actions implemented.
ACMG Guidelines for CNV/SV
When you click on the ACMG criteria, a detailed view will appear,
showcasing the logic used to score the event. This view will present the
five different categories defined by the ACMG guidelines, each
automatically populated with the relevant criteria. Users will have the
ability to manually modify each criterion to tailor the assessment as
needed.
Legend for ACMG automated criteria
The options include:
Solid with border- Scored Criteria
Solid without border- Criteria with evidence
Border Only – Not an automated criterion
Slash Box- Unmet criteria (inactivated automatically based on evidence)
When you hover over a specific criterion, a tooltip will appear,
providing detailed information about the evidence used for scoring. This
tooltip will display a concise summary of the supporting data, such as
relevant studies, experimental results, computational predictions, or
clinical observations that contributed to the assessment of the
criterion. This feature ensures that users have immediate access to the
underlying evidence, allowing for a better understanding of how each
criterion was evaluated and scored.
Each criterion will have a tool
tip feature that shows what evidence was used to score it.
When you click on a specific criterion, a detailed information panel
will open, providing comprehensive insights into the focus and relevance
of that criterion. This panel will include a summary of the criterion’s
purpose and its role in the overall scoring process. Additionally, users
will have interactive options within this panel:
Activate or Inactivate: Users can toggle the criterion on or off, allowing them to include or exclude it from the final scoring. This flexibility is crucial for adapting the evaluation to specific contexts or new information.
Adjust the Score: Users will have the ability to manually adjust the score associated with the criterion. This feature allows for fine-tuning the assessment based on expert judgment or additional evidence that may not have been captured by the automated scoring system.
These functionalities provide users with a robust and flexible tool to
customize the evaluation process, ensuring that the scoring reflects the
most accurate and relevant information available.
ACMG criterion options
Once the criteria have been saved, the interpretation will be securely
stored for all future encounters of that specific variant. This feature
ensures that the detailed assessment, including any manual adjustments
and contextual notes, is readily available for reference in subsequent
analyses. By storing this information, users can maintain consistency
and accuracy in variant interpretation over time.
This capability is particularly valuable for the interpretation of Copy
Number Variants (CNVs) and Structural Variants (SVs), such as deletions
and duplications. By having a refined and reusable interpretation
framework, users can apply a more precise and efficient approach to
evaluating these variants. The stored interpretations will align with
the best practice workflows established by the American College of
Medical Genetics and Genomics (ACMG) and the Clinical Genome Resource
(ClinGen).
Through this feature, the system supports a streamlined and standardized
process, allowing users to leverage past evaluations and ensure that
each variant is interpreted according to the highest standards of
clinical genetics. This not only enhances the quality and reliability of
variant interpretation but also saves time and resources by reducing
redundant efforts in re-evaluating previously encountered variants.
3.12.5b Repeat Expansion Analysis
MRepeat expansion refers to a specific type of genetic mutation where a
segment of DNA, typically consisting of a short sequence of nucleotides,
is repeated multiple times within a gene. These repeated sequences are
also known as “tandem repeats” or “microsatellites.” When these repeat
sequences expand beyond a certain threshold, it can lead to various
genetic disorders and diseases. Geneyx supports the ability to import
repeats from different file formats, including DRAGEN, ONT, and PacBio.
Some well-known genetic disorders associated with repeat expansions
include:
Huntington’s Disease: This is a neurodegenerative disorder caused by an expanded CAG repeat in the HTT gene. The expanded repeat leads to the production of a mutant protein that damages nerve cells in the brain.
Fragile X Syndrome: Fragile X syndrome is caused by an expansion of CGG repeats in the FMR1 gene. It is the most common inherited cause of intellectual disability and autism.
Myotonic Dystrophy: Myotonic dystrophy is characterized by an expansion of CTG repeats in the DMPK gene. This expansion results in muscle wasting, myotonia, and a range of other symptoms.
Spinocerebellar Ataxias (SCAs): Several SCAs are caused by repeat expansions in different genes. These disorders lead to problems with coordination and balance due to damage in the cerebellum.
Frontotemporal Dementia (FTD): Some forms of FTD are associated with repeat expansions in specific genes, such as the C9orf72 gene.
The severity of these genetic diseases often depends on the length of
the repeat expansion. Longer expansions tend to be associated with more
severe and earlier onset of symptoms. The underlying molecular
mechanisms by which these repeat expansions cause disease are complex
and can vary from one disorder to another.
Repeat expansion analysis is integrated into the CNV/SV genetic model
and we have introduced a color-coding system, which adds a layer of
visual information to the analysis results. Repeats are now displayed in
one of four distinct colors, each of which represents a specific
clinical implication:
Red: A red color signifies that at least one of the repeat alleles found in the individual is within the pathogenic range. This color indicates a significant deviation from the wildtype, potentially warranting further medical attention.
Orange: Repeats are colored orange when at least one repeat is within the premutation range, but no repeat falls within the pathogenic range. This color serves as a warning sign that a repeat is in a range of clinical interest.
Green: A green color indicates that both repeats are outside of the mutation or premutation range. This is generally considered a normal result with no immediate clinical concern.
No Color: In cases where there isn’t enough clinical data available to establish a clear range threshold, the repeat is left uncolored. This signals that further research may be needed to determine the clinical implications.
Repeat Expansion analysis showing full mutation in TCF4 gene,
associated with Fuchs endothelial corneal dystrophy.
Repeat expansions are also split into two columns. The first will show
the observed copies on allele one and allele two, and the second column
will show the repeat unit. Hovering over the repeat will show the wild
type allele copy.
Enhanced Filtering Options
In addition to the color-coding system, Geneyx implements filtering
options to empower users in customizing their analysis. Users can filter
repeats based on their color or potential clinical impact. This feature
allows for more focused and efficient analysis, making it easier to
identify relevant mutations.
Repeats filtering options.
The ranges and clinical implications attributed to each color have been
curated and are derived from various authoritative sources. These
sources include:
GeneReviews: A comprehensive resource for genetic disorders, providing detailed information on the clinical characteristics and management of various conditions.
MedlinePlus: A trusted source of health information from the National Library of Medicine, offering in-depth information on diseases, conditions, and wellness.
OMIM (Online Mendelian Inheritance in Man): A catalog of human genes and genetic disorders, offering detailed information about genetic conditions.
Medical Literature: Our data curation process also involves reviewing medical literature and research studies to ensure that valuable information is incorporated.
Additional In-House Data Presentation
We’ve also improved the presentation of in-house data within the
analysis dashboard. Users will now find the total number of samples, as
well as the count of homozygotes, heterozygotes, and hemizygotes for the
repeat with lower frequency within their in-house database. Hovering
over the total number of samples value reveals the distribution of both
repeat alleles within user’s in house repository.
In-house distribution of the repeats
Clicking on the total sample count opens a pop-up window displaying a
table and histogram:
Table: Offers additional information for the less frequent allele reported in the in-house data, such as samples where it has been identified, associated analyses, phenotypes, and more.
Histogram: Visualizes the frequency of all repeats found in the reported gene.
Repeat counts distribution in TCF4 gene. Sample repeat alleles
indicated in orange.
3.12.6 Annotating Variants
The variant table provides three annotation fields to the left of each
variant row: Relevance, Pathogenic, and Notes. These fields can be
defined for each candidate variant and are available for reporting
purposes. Here is an overview of how these annotation fields work:
Clicking on Relevance opens a window where the user can categorize the
relevance of the variant and include additional information to be
rendered in the report. The window includes the following options:
Inheritance Model: Select between recessive or dominant inheritance.
Relevance: Choose from options like High (1), Medium (2), Low (3), or Not Relevant to indicate the relevance of the variant.
Pathogenicity: Select the appropriate pathogenicity category such as Pathogenic, Likely Pathogenic, Uncertain Significance, Likely Benign, Benign, or Risk Factor.
Validation: Specify whether the variant requires validation, has positive validation, or has negative validation.
Notes: Add free-text notes to provide additional details or comments.
Phenotype and Evidence: This allows the user to select relevant
disorders from OMIM (Online Mendelian Inheritance in Man) and associated
evidence to be included in the report. By selecting the checkbox next to
each entry, the information will be incorporated into the report.
ClinVar: This section presents information from ClinVar about the variant, including a hyperlinked accession number, the type of the variant, the associated conditions, and the clinical significance of the variant.
OMIM: This displays known conditions associated with the disease according to OMIM phenotypes. It provides the associated inheritance model and a hyperlink to relevant documentation.
Evidence: This section shows diseases, gene summaries, publications, and related variant and analysis terms for the given gene.
Diseases related to gene: An advanced search that identifies all publications present in databases such as OrphaNet, OMIM, UniProt, Decipher, and PubMed, containing conditions matching the entered phenotype and the given gene. Gene summaries are provided, with highlighted phenotypic terms if matched.
Publications related to gene: This displays publications in ClinVar and Entrez that are related to the given gene.
Variants related to gene: Shows all variants in ClinVar and OMIM that are present in the gene.
Analysis terms related to gene: Displays terms that are frequently, moderately, or rarely related to the entered phenotype.
Once information has been entered in the Relevance interface, the fields
will be populated in the variant table, providing relevant information
for each variant.
3.12.7 Reporting Variants
Once the relevance and associated information has been set, user can
review the selections across all of the genetic models using the
‘Selected Variants’ feature.
Geneyx Analysis includes a convenient Selected Variant view that
consolidates all the selected variants from different genetic models in
one place.
To access this feature, simply click on the open book icon located in
the upper right corner. It will open the Selected Variant view, which
pulls information from the Annotate Variant window.
Select Variant View
The Selected Variant view provides a comprehensive overview of all the
variants that have been selected, allowing for easy navigation and
analysis of the chosen variants across different genetic models. It
streamlines the process of reviewing and studying selected variants,
enhancing the efficiency of variant analysis and interpretation.
After selecting the clinically relevant variants and applying the
relevance scores and related notes, the next step is to review the
findings using the ‘Report Preview’ button located in the top right
corner. There is also a Report Configuration option, which empowers
users to tailor the report layout with case-specific modifications,
ensuring a dynamic and personalized presentation. Importantly,
alterations made within this feature are exclusive to the specific
report, avoiding unintended changes in other reports implementing the
template configuration.
Report icon displayed in upper right corner
By clicking the green Report Preview icon, a visual summary of all the
findings, along with the general information, will be generated. In this
view, users have the flexibility to choose which evidence sections
should be included in the final report.
The Report Preview provides a mock report that gives an overview of the
findings and allows for a comprehensive review before generating the
final report. To generate the report, simply click the Save button. This
will produce a complete report in PDF format, containing all the
relevant information, supporting evidence, including publications, and
more. Additionally, an Excel file will be generated, displaying the full
variant tables that were analyzed.
The generated report in PDF format provides a comprehensive and
organized presentation of the findings, making it easier to communicate
and share the results with colleagues, healthcare professionals, or
other stakeholders involved in the analysis process. If a customized
report is required, please contact support@geneyx.com.
3.13 Batch VCF Upload (& Joint VCF Files)
In addition to the option of uploading individual samples, Geneyx
Analysis provides users with the convenience of uploading multiple
samples in batch using the VCF Uploader. The VCF Uploader utilizes an
Excel spreadsheet with a list of samples and key fields, along with an
executable file that is configured to the user’s account. To utilize
this feature, users need their API ID and API Key, which can be obtained
by contacting support@geneyx.com. Additionally, this feature is only
supported for Windows OS, if you are using a MAC please contact support.
Here are the steps to use the VCF Uploader:
Open the “Tgex.VCFUploader.exe.config” file and modify it by adding your account credentials (API ID and API Key). If you don’t have this information, please reach out to support@geneyx.com for assistance.
Navigate to the “Resources” folder within the unzipped files. This folder contains an Excel document that needs to be updated with the files you want to import. Place this updated Excel file in the same directory where your VCF files are located.
Open the “Tgex.VCFUploader.exe” file.
In the bottom left corner of the application, click on “Select File” and choose the updated Excel file that contains the sample information.
Optionally, you can define a protocol that will be applied to all the samples in the batch. If you don’t need a protocol, you can leave this field blank.
Click on the “Import” button. This will initiate the upload process for all the samples listed in the Excel file to your Geneyx Analysis account.
VCF Uploader
Additionally, there is a video recording available, which can be
downloaded here,
https://github.com/geneyx/geneyx.analysis.api/tree/main/apps/VCF%20Uploader
that demonstrates how to use this. By utilizing the VCF Uploader, users
can streamline the process of uploading multiple samples in batch,
saving time and effort.
3.14 APIs
The API feature in Geneyx Analysis provides the ability to automate
sample workflows by integrating with Laboratory Information Management
Systems (LIMS) and Electronic Health Record (EHR) systems. It allows for
seamless data exchange between Geneyx and these systems, facilitating
the transfer of patient metadata and relevant information.
To utilize the API feature, you can access the scripts and details of
each field in the Geneyx Analysis API repository on GitHub:
https://github.com/geneyx/geneyx.analysis.api. This repository
contains all the necessary scripts and documentation to guide you
through the integration process.
Additionally, a collection of scripts is available for download from the
same repository. These scripts can be used with the popular API
development platform called Postman, which can be obtained from
https://www.postman.com/. Once you have downloaded and opened the
Postman application, you can import the Geneyx Analysis API collection
by selecting “Import” next to “My Workspace.” Choose the “Upload Files”
option and select the Geneyx Analysis API
collection.postman_collection.json file that you downloaded earlier.
Importing the collection will create a set of scripts with pre-defined
fields that can be used to push or extract information from your Geneyx
account. To use these scripts, you will need to obtain your API User ID
and API User Key. For this information, you can contact
support@geneyx.com, and they will assist you in obtaining the
necessary credentials. All available APIs and associated descriptions
can be found here: https://github.com/geneyx/geneyx.analysis.api/wiki.
When working with the scripts in Postman, you can update the fields by
navigating to the “Body” section. The field structure should be in JSON
format, allowing you to customize the data according to your specific
requirements.
By leveraging the Geneyx Analysis API feature and using the provided
scripts and tools, you can streamline your workflows and automate the
exchange of information between Geneyx and your LIMS or EHR systems.
3.15 Python Scripts/Command Line
Geneyx Analysis offers the capability to utilize Python scripts to
perform various functions within the application. These scripts can be
accessed from the Geneyx Analysis API repository on GitHub:
https://github.com/geneyx/geneyx.analysis.api/tree/main/scripts. Let’s
explore the available scripts:
**ga.config.yml**:
This configuration file should be placed in the same directory as the other Python scripts. It serves as a reference for the account configuration and specifies the server URL, `apiUserId`, and `apiUserKey`. You can obtain the `apiUserId` and `apiUserKey` by contacting support@geneyx.com.
**ga_CreateCase.py**:
This script allows users to create an analysis by utilizing existing VCF files in their account. The `CreateCase_Data.json` file, located in the templates directory, contains the data fields that can be used with this Python script. Descriptions of the available fields can be found in the following link: https://github.com/geneyx/geneyx.analysis.api/wiki/Create-Case.
This command creates a case in the account using the information provided in the JSON file.
**ga_addClinicalRecord.py**:
This Python script enables users to add a clinical record to an existing subject. The `clinicalRecord.json` file can be modified with the specific information and executed with the script. For the phenotypic codes, only “HP:” terms are accepted. Detailed field descriptions can be found here: https://github.com/geneyx/geneyx.analysis.api/wiki/AddClinicalRecord.
This Python script allows users to upload a VCF sample into a new or existing subject along with associated patient information. The `sample.json` file in the template directory should be modified with the relevant data. Field descriptions can be found here: https://github.com/geneyx/geneyx.analysis.api/wiki/Create-Sample.
To run this script, the following fields need to be specified: `–snvVCF`, `–svVCF`, and `–cnvVCF`.
By leveraging these Python scripts, users can perform specific actions
within Geneyx Analysis, such as creating cases, adding clinical records,
creating patients, and uploading VCF samples. The provided links offer
additional details on the required fields and their descriptions,
enabling users to customize and execute the scripts effectively.
3.15.1 CNV/SV/Repeat Unification Script
Streamlining the consolidation of multiple VCFs for Copy Number
Variations (CNVs), Structural Variations (SV), and repeats, a
comprehensive video tutorial has been created. This tutorial guides
users on efficiently merging these diverse files, enhancing the
understanding of script usage for proper sample uploads. The scripts can
be found here:
https://github.com/geneyx/geneyx.analysis.api/tree/main/scripts/UnifyVcf.
Tutorials can be found here: https://geneyxuk.com/?s=unification.
This script not only unifies the files but also modifies them when
necessary, since each provider provides slightly different vcf files and
the final file has to include specific fields for Geneyx application to
be able to read it properly. Hence, it’s important to run the script
that matches the pipeline with which the vcf files were created.
The options are:
DragenUnifyVcf.py – for DRAGEN pipeline ONTUnifyVcf.py – for Oxford
Nanopore sequences and pipeline PacBioUnifyVcf.py – for PacBio sequences
and pipeline
-o The path to the unified vcf file, including its name. The script
compresses the output file, so its name should end with “.vcf” and not
“.gz” -s The path to the structural variants (sv) vcf. This file can be
either gzipped or unzipped, but it must be a vcf file. -c The path to
the Copy Number Variants (CNV) vcf. This file can be either gzipped or
unzipped, but it must be a vcf file. -r The path to the tandem repeats
variants vcf. This file can be either gzipped or unzipped, but it must
be a vcf file. -b
Relevant only for PacBio. A bed file used to filter the repeats vcf
file. If the PacBioUnifyVcf.py script is called without this parameter,
the repeats vcf file (if given) will not be unified. When called with
this parameter, the PacBioUnifyVcf.py script creates a filtered repeats
vcf file and unifies it (rather than the full repeats file) with the
other vcf files. This bed file can be downloaded from PacBio’s github:
https://github.com/PacificBiosciences/trgt/tree/main/repeats (named
pathogenic_repeats.hg38.bed or the like) Please notice that when running
with this parameter you have to run on linux and have bedtools installed
3.15.2 Microarray to VCF Converter
For customers that have microarray data, such as those from Affymetrix,
Geneyx provides a script to convert the file into a compatible format
with Geneyx. The script is available here,
https://github.com/geneyx/geneyx.analysis.api/tree/main/apps/microarray.
The converted files can then be loaded into the CNV/SV genetic model.
Users can also select this as a sequencing target during data upload.
Pharmacogenomics (PGx) is a rapidly advancing field within precision
medicine, focusing on the relationship between genetic variations and
drug metabolism. The adoption of PGx in clinical and diagnostic settings
is gaining momentum as knowledgebases and genetic insights continue to
improve. At Geneyx, we recognize the significance of this field and are
committed to offering a comprehensive PGx workflow to our users.
The Geneyx PGx workflow provides thorough interpretations for 13 CPIC
level A/B genes that influence the metabolism of 68 commonly prescribed
drugs. This information is curated from multiple reliable annotation
sources, including the Clinical Pharmacogenomic Implementation
Consortium (CPIC), the Food and Drug Administration (FDA), and the
Pharmacogenomics Knowledge Base (PharmGKB). Each reported gene is
accompanied by useful hyperlinks to these databases, allowing easy
access to additional information, all of which is automatically
integrated into the report.
The Geneyx PGx report delivers detailed patient and drug information in
a user-friendly and customizable format. Genes with low quality or
missing SNPs can be conveniently excluded from the report. Moreover,
final modifications to the report can be made using the report editor,
and electronic signature sign-off ensures compliance and accountability.
The report presents the gene name, genotype, and the impact of genetic
variations on drug metabolism. Additionally, gene descriptions elucidate
the gene’s function, its association with drug metabolism, and the
effect of genetic changes on therapeutic efficacy.
By leveraging key databases and unique genetic profiles, the Geneyx
Pharmacogenomics workflow equips users with powerful insights to
mitigate potential adverse drug reactions and metabolic responses. If
you are affiliated with a clinical or diagnostic company and are
interested in providing PGx reports to your patients, we invite you to
schedule a meeting with us to experience the Geneyx PGx workflow in real
time.
The PGx protocol requires have a PGx license to use. To obtain a
license, please contact support@geneyx.com.
3.16.1 Instruction for using PGx
The PGx protocol in Geneyx Analysis requires capturing SNVs that are
typically located outside of coding regions and are homozygous reference
alleles. To ensure accurate results, the protocol necessitates
genotyping enrichment.
The PGx protocol is associated with an enrichment kit that contains PGx
BED files for both the hg19 and hg38 reference genome builds. When
starting from fastq files, please select the PGx enrichment during the
fastq upload process. This enrichment kit captures all exonic regions,
along with 50 base pairs into intronic regions, and merges them with the
ClinVar bed file. Additionally, a PGx BED file is used for genotyping,
enabling the capture of variants even if they are homozygous reference
at a specific position. This is particularly important for capturing
star alleles and facilitating subsequent reporting.
If you are starting from VCF files, please consider the following
factors. Ensure that the secondary pipeline used to call the VCF files
includes the regions defined in the PGx BED files. If these files are
not available, please contact support@geneyx.com for assistance. Your
development or IT team will need to implement the PGx BED files
accordingly. Alternatively, you can consider calling a gvcf output. For
guidance on using gvcf, please reach out to support@geneyx.com.
Once you have the enriched VCF ready, you can import it into the PGx
protocol. Please note that this protocol is specific to PGx analysis and
does not include other genetic models. The default filtering is already
applied, eliminating the need for modifying the displayed variants. To
view the star alleles, simply click on the star icon located in the
upper right corner. This option will display all the genes that will be
included in the report. If necessary, you can exclude genes that lack
coverage, exhibit low quality, or have missing SNPs.
PGx Star Allele icon
After making the required modifications, you can generate the report.
The report template has already been created but can be further
customized to match your lab’s design and specific requirements. We
encourage you to test this feature and reach out to us if you have any
questions or concerns.
PGx Result Preview
We are committed to providing a comprehensive and user-friendly PGx
workflow, and we appreciate your interest in our system. Should you
require further assistance or have additional inquiries, please do not
hesitate to contact us. We are here to support you.
3.17 Long-Read Sequencing
Long read sequencing, a powerful technology in genomics, offers several
advantages over traditional short read sequencing by generating longer
contiguous sequences of DNA. This ability to produce long reads, often
spanning thousands to millions of base pairs, provides a more
comprehensive view of complex regions of the genome. It enhances the
detection of structural variants, repetitive sequences, and phased
haplotypes, which are crucial for understanding genetic diversity and
disease mechanisms. Long read sequencing technologies, such as those
developed by Pacific Biosciences (PacBio) and Oxford Nanopore
Technologies, have become invaluable tools for researchers aiming to
achieve a more complete and accurate assembly of genomes.
Geneyx, a leading provider of genomic data analysis solutions, offers
robust support for the analysis of long read sequencing data. The
platform is equipped with advanced algorithms and tools specifically
designed to handle the unique characteristics of long read data. Geneyx
facilitates the accurate detection and annotation of variants, including
single nucleotide polymorphisms (SNPs), insertions, deletions, and
complex structural variants. The platform’s intuitive interface and
comprehensive analysis capabilities enable researchers to interpret long
read sequencing data efficiently, providing insights into genetic
disorders, personalized medicine, and evolutionary studies. By
leveraging Geneyx’s powerful analysis tools, researchers can maximize
the potential of long read sequencing to uncover novel genetic
information and advance genomic research.
3.17.1 Import
Geneyx supports the analysis of long-read sequencing data starting from VCF files. The platform accommodates various file formats, including SNV, CNV, SV, and repeat files. To ensure proper merging and annotation, SV, CNV, and repeat files must be processed through a unification script, available [here](https://github.com/geneyx/geneyx.analysis.api/tree/main/scripts/UnifyVcf).
Once the data is in the correct format, it can be uploaded into the
dedicated “Long-Read” protocols. To activate these protocols, please
contact support@geneyx.com. The upload process allows for the import
of SNV and associated files. If the unification script was used, the
resulting file can be imported into either the CNV or SV model, ensuring
comprehensive analysis and accurate annotation.
Long Read Protocols
After the files finish uploading, a new dialog will appear, allowing for
the entry of key fields associated with the files. By default, this
protocol utilizes a “Long Read Sequencing” enrichment kit. This kit
defines the target capture regions based on all transcripts in the human
genome, covering end-to-end, including all intronic regions and upstream
and downstream regions. Integrated SMART filtering helps remove
background noise specific to long-read samples. Detailed information on
the SMART filtering process can be viewed
here.
Import Dialog
The other key fields for import are BAM FILE URL and METHYLATION FILE
URL. If these files are hosted on a cloud infrastructure, the links for
these files can be placed in these fields. In some cases, the
configuration of the cloud infrastructure is needed. Please see section
3.17.2, AWS Configuration, for an example. Alternatively, local IGV can
be used, further details on configuration can be found in the section
Desktop IGV Integration.
Once the information is entered, click next and enter in any associated
phenotypes for the patient. This will help identify candidate variants
through a prioritization tool.
3.17.2 Analyzing Long-Read Data
There are unique features available for analyzing long-read data in
Geneyx Analysis. If the uploaded VCF files contain allele phasing, this
information will be pulled into the Genetic and Genomic categories.
Clicking the expansion icon of this column will display this
information. Allele phasing is a process in genetics that involves
determining which alleles at different loci belong to the same
chromosome. This information helps to understand whether certain genetic
variants are inherited together or independently. It is crucial for
studying genetic linkage, haplotype analysis, and understanding the
inheritance patterns of alleles within populations. Advances in
sequencing technologies and computational methods have significantly
improved our ability to accurately phase alleles, thereby enhancing our
understanding of genetic variation and its implications in health and
disease.
Allele Phasing in the Genomic and Genetic Data category
Another feature that enhances the analysis of long-read data is the
capability to visualize overlapping SNV and CNV events. When CNV/SV
files are imported alongside SNV files, Geneyx includes a “MATCHED
CNV/SVs” column. This feature is particularly useful as overlapping
events may suggest the presence of potential recessive compound
heterozygous variants involving both CNVs and SNVs. Understanding these
overlaps provides valuable insights into complex genetic interactions
and their potential impact on phenotypic expression.
Matched CNV/SVs column highlighted in yellow
Haplotype visualization in the BAM files involves accessing a genetic
variant’s specific location. By clicking on the variant, users can open
Integrative Genomics Viewer (IGV), which displays detailed coverage and
pileup information from the BAM file. Additionally, IGV allows users to
examine the associated haplotypes, providing a comprehensive view of how
genetic variations are distributed and potentially linked across the
genome. This capability supports detailed genomic analysis, aiding in
the identification and interpretation of haplotype structures and their
implications in genetic research and diagnostics.
Haplotypes present in BAM files
Methylation visualization within Geneyx is facilitated by linking a
specific URL during data import. This feature is particularly beneficial
for researchers and clinicians alike. By associating samples with a
case, users can compare regions that exhibit differential methylation
patterns. This comparative analysis provides insights into epigenetic
modifications across the genome, aiding in the exploration of potential
biomarkers or regulatory mechanisms underlying various diseases or
biological processes. The ability to visualize and analyze methylation
data enhances the depth and scope of genomic investigations, offering
valuable insights into gene expression regulation and its impact on
health and disease.
Methylation Plots in IGV
Repeat expansions, analyzed from long-read data, provide crucial
insights into genomic instability and disease mechanisms. This
capability is detailed in section 3.12.4a of the Geneyx platform,
focusing on Repeat Expansion Analysis. By leveraging long-read
sequencing technologies, researchers can accurately characterize and
quantify repetitive DNA sequences that are challenging to resolve with
traditional short-read methods. This analysis is particularly relevant
in studying disorders associated with unstable repeat expansions, such
as Huntington’s disease and various types of ataxias. Understanding the
dynamics of repeat expansions at a molecular level enhances our
comprehension of genetic diseases, paving the way for improved
diagnostic accuracy and potential therapeutic interventions.
3.19 Providing BAM & methylation tracks for samples not processed by Geneyx
There are cases that the customer executed the secondary pipeline
locally and started working with it from VCF. In those cases, if the
customer wants to visualize the BAM/Methylation tracks, then what needs
to be done
3.19.1 General
First of all, all BAM/Methylation files must be bgzipped with the
corresponding index file (tbi) in the following naming conventions:
Methylation: xxx.bed.gz and xxx.bed.gz.tbi
BAM: xxx.bam and xxx.bam.bai
3.19.2 File accessibility
Geneyx is using the JavaScript version of IGV; hence, the files must be
accessible via the user’s web browser. If all users are
connected to the same LAN, then having local LAN access to files is
enough.
Where to store the files?
There are several options for where and how to provide access to the
files
3.19.2.1 Option 1: Local Server
When files are stored locally and need to be accessed only locally (from
the same LAN), this approach can be used:
Files are installed locally on the file system of the local LAN
Create a web server that serves those files over HTTP or HTTPS – that web server can be a local server with a domain name or just an IP address
The server CORS settings will allow the Geneyx domain to access those files
The URLs set to geneyx are the local server’s file URLs
In case the server is over HTTP only, it requires allowing untrusted content at each client browser accessing those files
3.19.2.2 Option2: Cloud
The files could be stored at AWS S3 or Azure blob storage.
Those files could be directly or indirectly served from storage,
depending on the security and IT setup of the current organization
Containers and Buckets are public
This is the easiest solution, in this case
Set CORS permission from the Geneyx domain
Set direct URLs (over https)
Buckets are public, but are allowed access from a specific IP address
or IP range
Containers and Buckets are Private – Use Access Tokens
This scenario requires a web server that acts as a proxy. It gets the
object request and redirects to the target storage URL by adding a
time-limited access token (service of AWS or AZURE) for a specific IAM
user.
Create a web server that serves those files over HTTPS – that web server can be a local server with a domain name or just an IP address
The server CORS settings may allow the Geneyx domain to access those files
The URLs set to geneyx are the local server’s file URLs
Buckets are private: CDN (CloudFront with Secure Headers Policy) + AWS
WAF
This setup ensures that your S3 contents remain private, are accessed
securely via CloudFront, and only allow access from specified IPs with
AWS WAF protection.
Step 1: Ensure AWS S3 Buckets Are Private
Navigate to the AWS S3 Console.
Select the S3 bucket you want to secure.
Under the Permissions tab, confirm the following:
Block Public Access settings should be enabled.
No Bucket Policy allows public access.
No Access Control List (ACLs) grants public read access.
Save the settings to enforce privacy.
Step 2: Configure CloudFront Distribution with Secure Headers