
As our understanding of human health grows, we often see that similar biological dysfunction underlies the co-occurrence of various complex diseases. It remains difficult to determine if there are common genetic mechanisms contributing to clinically distinct conditions or if expression of both conditions relates to other shared risk factors. For example, in some situations, genetic variation may increase risk for one condition, and expression of this condition then increases risk for another disease. Identifying potentially pleiotropic genes is crucial for advancing the development of more effective treatment options, especially in instances where current therapies are insufficient. Genome-wide association studies (GWAS) provide cross-trait associations but do not provide the full functionality of how dysfunction in genes being tagged by GWAS hits are contributing to two or more distinct phenotypes. Fortunately, as other types of available data continue to grow exponentially (e.g., RNA-seq, mass spectrometry, mouse knock-out phenotype associations), these can be leveraged to help process GWAS results into meaningful information. The aim of this protocol is to provide clear instructions for using various databases and available software tools to identify key pleiotropic genes contributing to two distinct phenotypes of interest. The protocol uses information from various publicly available databases, including GWAS Catalog, Functional Mapping and Annotation (FUMA), Drosophila RNAi Screening Center Integrative Ortholog Prediction Tool (DIOPT), International Mouse Phenotype Consortium (IMPC), STRINGdb, Pharos, and Cytoscape for network visualization. This pipeline, with code written in R and RStudio software, helps the user identify and generate hypotheses about shared genetic mechanisms contributing to their selected phenotypes of interest as well as prioritize genes of interest to functionally follow up in model systems that are more likely to be clinically relevant. © 2025 Wiley Periodicals LLC. Basic Protocol: Pleiotropic gene prioritization pipeline for studies in model systems.
Balanced translocation carriers experience elevated reproductive risks, including pregnancy loss and children with anomalies due to generating chromosomally unbalanced gametes. While understanding the likelihood of producing unbalanced conceptuses is critical for individuals to make reproductive decisions, risk estimates are difficult to obtain as most balanced translocations are unique. To improve reproductive risk estimates, Drs. Trunca and Mendell created models based on a logistic regression analysis of a dataset of over 6000 individuals from over 1000 translocation families. While risk assessments using these models have been offered as a free service for years, this protocol aims to create a sustainable model for genetics professionals to obtain risk estimates for their patients directly. This protocol guides the user through collecting clinical information, using a risk-generating calculator based on the models, and interpreting the calculator outputs. This version of the protocol has been updated from the initial publication to introduce an additional resource to the community that is intended to improve accessibility and reduce error in calculating risks. In addition to the previously offered custom Java program, a web-based calculator that calculates age-adjusted miscarriage risks is now available. A practice tutorial is provided for both versions of the calculator to ensure competency in interpretation prior to use. © 2025 Wiley Periodicals LLC. Basic Protocol 1 : Estimation of reproductive risks for balanced translocation carriers Basic Protocol 2 : Practical examples of typical patient encounters with instructive interpretations
Omics biomarkers play a pivotal role in personalized medicine by providing molecular-level insights into the etiology of diseases, guiding precise diagnostics, and facilitating targeted therapeutic interventions. Recent advancements in omics technologies have resulted in an increasing abundance of multimodal omics data, providing unprecedented opportunities for identifying novel omics biomarkers for human diseases. Mendelian randomization (MR) is a practically useful causal inference method that uses genetic variants as instrumental variables (IVs) to infer causal relationships between omics biomarkers and complex traits/diseases by removing hidden confounding bias. In this article, we first present current challenges in performing MR analysis with omics data, and then describe four MR methods for analyzing multi-omics data including epigenomics, transcriptomics, proteomics, and metabolomics data, all executable within the R software environment.
DNA copy number variants (CNVs) are routinely evaluated as part of clinical diagnosis in both the prenatal and postnatal genetic settings. Current guidelines for interpreting the potential clinical significance of these CNVs, typically identified by chromosomal microarray, focus entirely on genes localized within the CNV region. However, recent work has suggested that some CNVs can actually produce clinical impacts by influencing transcription of genes outside the CNV region. These alterations of transcription appear to occur by disrupting the composition of DNA topologically associated domains (TADs), which strongly influence contacts between gene promoters and their associated enhancers. Here we present a set of detailed protocols for the use of the free software tool ClinTAD ( https://www.clintad.com ). This decision-support software allows for prediction as to whether a given CNV may potentially disrupt a TAD boundary, and offers phenotype matching to genes near, but not within the CNV region, whose expression could be influenced by altered TAD architecture and that have phenotypic impacts related to that reported in a given patient. Our protocols here provide specific examples of how to implement these tools. In addition, the software has the capability to impact genomic research by evaluating multiple cases in parallel. We propose that this decision-support tool can benefit and improve genetic diagnosis. © 2020 Wiley Periodicals LLC. Basic Protocol 1 : Evaluating a single case using ClinTAD Basic Protocol 2 : Evaluating a single case with multiple variants using ClinTAD Basic Protocol 3 : Evaluating multiple cases using ClinTAD Basic Protocol 4 : Creating tracks with custom data
Genetic research often utilizes or generates information that is potentially sensitive to individuals, families, or communities. For these reasons, genetic research may warrant additional scrutiny from investigators and governmental regulators, compared to other types of biomedical research. The informed consent process should address the range of social and psychological issues that may arise in genetic research. This article addresses a number of these issues, including recruitment of participants, disclosure of results, psychological impact of results, insurance and employment discrimination, community engagement, consent for tissue banking, and intellectual property issues. Points of consideration are offered to assist in the development of protocols and consent processes in light of contemporary debates on a number of these issues. © 2020 Wiley Periodicals LLC.
Our understanding of genetic disease(s) has increased exponentially since the completion of human genome sequencing and the development of numerous techniques to detect genetic variants. These techniques have not only allowed us to diagnose genetic disease, but in so doing, also provide increased understanding of the pathogenesis of these diseases to aid in developing appropriate therapeutic options. Additionally, the advent of next-generation or massively parallel sequencing (NGS/MPS) is increasingly being used in the clinical setting, as it can detect a number of abnormalities from point mutations to chromosomal rearrangements as well as aberrations within the transcriptome. In this article, we will discuss the use of multiple techniques that are used in genetic diagnosis. © 2020 by John Wiley & Sons, Inc.
Novel cytogenetic tools are increasingly based on genome sequencing for detecting chromosomal abnormalities. Different sequence-based techniques optimized for diagnosis of structural variants can be useful for narrowing down the localization of breakpoints of chromosomal abnormalities, but do not offer nucleotide resolution of breakpoints for proper interpretation of gene disruption. This protocol presents the characterization of structural variants at nucleotide resolution using Sanger sequencing after low-pass large-insert genome sequencing or other long-molecule methods. © 2020 Wiley Periodicals LLC. Basic Protocol 1: Primer design for junction amplification at translocations and inversions Basic Protocol 2: Amplification of derivative chromosomes using a long-range polymerase Alternate Protocol: Amplification of derivative chromosomes using a hot-start polymerase Basic Protocol 3: Preparation of DNA for Sanger sequencing Basic Protocol 4: Interpretation and reporting of breakpoints based on Sanger sequencing.
The AD Knowledge Portal (adknowledgeportal.org) is a public data repository that shares data and other resources generated by multiple collaborative research programs focused on aging, dementia, and Alzheimer's disease (AD). In this article, we highlight how to use the Portal to discover and download genomic variant and transcriptomic data from the same individuals. First, we show how to use the web interface to browse and search for data of interest using relevant file annotations. We demonstrate how to learn more about the context surrounding the data, including diagnostic criteria and methodological details about sample preparation and data analysis. We present two primary ways to download data-using a web interface, and using a programmatic method that provides access using the command line. Finally, we show how to merge separate sources of metadata into a comprehensive file that contains factors and covariates necessary in downstream analyses. © 2020 The Authors. Basic Protocol 1: Find and download files associated with a selected study Basic Protocol 2: Download files in bulk using the command line client Basic Protocol 3: Working with file annotations and metadata.
Profiling genetic variants-including single nucleotide variants, small insertions and deletions, copy number variations, and structural variations (SVs)-from both healthy individuals and individuals with disease is a key component of genetic and biomedical research. SVs are large-scale changes in the genome and involve breakage and rejoining of DNA fragments. They may affect thousands to millions of nucleotides and can lead to loss, gain, and reshuffling of genes and regulatory elements. SVs are known to impact gene expression and potentially result in altered phenotypes and diseases. Therefore, identifying SVs from the human genomes is particularly important. In this review, I describe advantages and disadvantages of the available high-throughput assays for the discovery of SVs, which are the most challenging genetic alterations to detect. A practical guide is offered to suggest the most suitable strategies for discovering different types of SVs including common germline, rare, somatic, and complex variants. I also discuss factors to be considered, such as cost and performance, for different strategies when designing experiments. Last, I present several approaches to identify potential SV artifacts caused by samples, experimental procedures, and computational analysis. © 2020 Wiley Periodicals LLC.
Transposable element (TE) mobilization is a significant source of genomic variation and has been associated with various human diseases. The exponential growth of population-scale whole-genome sequencing and rapid innovations in long-read sequencing technologies provide unprecedented opportunities to study TE insertions and their functional impact in human health and disease. Identifying TE insertions, however, is challenging due to the repetitive nature of the TE sequences. Here, we review computational approaches to detecting and genotyping TE insertions using short- and long-read sequencing and discuss the strengths and weaknesses of different approaches. © 2020 Wiley Periodicals LLC.
ATAC-seq, the assay for transposase-accessible chromatin using sequencing, is a quick and efficient approach to investigating the chromatin accessibility landscape. Investigating chromatin accessibility has broad utility for answering many biological questions, such as mapping nucleosomes, identifying transcription factor binding sites, and measuring differential activity of DNA regulatory elements. Because the ATAC-seq protocol is both simple and relatively inexpensive, there has been a rapid increase in the availability of chromatin accessibility data. Furthermore, advances in ATAC-seq protocols are rapidly extending its breadth to additional experimental conditions, cell types, and species. Accompanying the increase in data, there has also been an explosion of new tools and analytical approaches for analyzing it. Here, we explain the fundamentals of ATAC-seq data processing, summarize common analysis approaches, and review computational tools to provide recommendations for different research questions. This primer provides a starting point and a reference for analysis of ATAC-seq data. © 2020 Wiley Periodicals LLC.
In neurodegeneration studies, researchers are faced with problems such as limited material availability and late disease manifestation. Cell models provide the opportunity to investigate molecular mechanisms of pathogenesis. Moreover, genome editing technologies enable generation of isogenic cell models of hereditary diseases. Our protocol outlines an approach for introducing an expanded CAG repeat tract into the first exon of the HTT gene, the Huntington's disease causing mutation. The protocol allows modeling the disease at various severity levels by introducing different numbers of CAG repeats. Furthermore, the protocol can be applicable for modeling other diseases caused by trinucleotide repeat expansion. It is important to note there are many difficulties with cloning repeated sequences and amplification of GC-rich regions. Here, we also propose troubleshooting options, which overcome these problems. The protocol is based on CRISPR/Cas9-mediated homologous recombination with a uniquely designed donor plasmid harboring an expanded CAG tract flanked with long homology arms. © 2020 Wiley Periodicals LLC. Basic Protocol 1: Design and assembling donor and CRISPR/Cas9-expressing plasmids Basic Protocol 2: Transfection of cells with plasmids and sorting GFP-positive cells Basic Protocol 3: PCR screening single-cell clones and validation of the mutant HTT expression.
Clinical interpretation of DNA sequence variants is a critical step in reporting clinical genetic testing results. Application of next-generation sequencing technology in molecular genetic testing has facilitated diagnoses of genetic disorders in clinical practice. However, the large number of DNA sequence variants detected in clinical specimens, many of which have never been seen before, make clinical interpretation challenging. Recommendations by the American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP) have been widely adopted by clinical laboratories around the world to guide clinical interpretation of sequence variants. The ClinGen Sequence Variant Interpretation Working Group and various disease-specific variant curation expert panels have also developed specifications for the ACMG/AMP recommendations. Despite these efforts to standardize variant interpretation in clinical practice, different laboratories may subjectively use professional judgment to determine which criteria are applicable when classifying a variant. In addition, clinicians and researchers who are not familiar with the variant interpretation process may have difficulty understanding clinical genetic reports and communicating the clinical significance of genetic testing results. Here we provide a step-by-step protocol for clinical interpretation of sequence variants, including practical examples. By following this protocol, clinical laboratory geneticists can interpret the clinical significance of sequence variants according to the ACMG/AMP recommendations and ClinGen framework. Furthermore, this article will help clinicians and researchers to understand variant classification in clinical genetic testing reports and evaluate the quality of the reports. © 2020 by John Wiley & Sons, Inc. Basic Protocol: Interpreting the clinical significance of sequence variants Support Protocol: Reevaluating the clinical significance of sequence variants.
As genome sequencing methodologies have become more sensitive in detecting low-frequency rare-variant events, the link between post-zygotic mutagenesis and somatic mosaicism in the etiology of several human genetic conditions other than cancers has become more clear. Given that current clinical-genomics diagnostic methods have limited detection sensitivity for mosaic events, a copy-number variant (CNV) deletion inherited from a parent with low-level (<10%) mosaicism can be erroneously interpreted in the proband to represent a de novo germline event. Here, we describe three sensitive, precise, and cost-efficient methods that can quantitatively assess the potential degree of parental somatic mosaicism levels for CNV deletions: droplet digital PCR (ddPCR), PCR amplicon-based next-generation sequencing (NGS), and quantitative PCR. ddPCR using the EvaGreen fluorescent dye protocol can specifically quantify the deleted or non-deleted alleles by analyzing the number of droplets positive for a fluorescent signal for each event. PCR amplicon-based NGS assesses the allele frequencies of a heterozygous single-nucleotide polymorphism within a deletion region. The difference in number of reads between the two genotypes indicates the level of somatic mosaicism for the CNV deletion. Quantitative PCR can be applied where the relative quantity of the deletion junction-specific product represents the level of mosaicism. Clinical implementation of these quantitative variant-detection methods enables potentially more accurate assessment of disease recurrence risk in family-based genetic counseling, allowing couples to engage in more informed family planning. © 2020 by John Wiley & Sons, Inc. Basic Protocol: Droplet digital PCR (ddPCR) Alternate Protocol 1: PCR amplicon-based next-generation sequencing Alternate Protocol 2: Quantitative real-time PCR (qPCR).
In order to comply with regulations set by established local, state, and federal agencies and other regulatory organizations, such as the College of American Pathologists and the International Organization for Standardization, a clinical laboratory needs to develop procedures for the processes of validating laboratory-developed tests (LDTs) and establishing performance specifications for these assays prior to use in clinical testing. This is applicable to all fluorescence in situ hybridization (FISH) assays. Even Food and Drug Administration-approved FISH assays must undergo some form of verification before implementation in the clinical laboratory. The process of validating an assay as an LDT must include a plan, a procedure, and a report. The validation studies described here include metaphase and interphase FISH methodology for identification of the LSI EGR1/D5S23, D5S721 dual-color probe, which labels distinct biomarkers consistent with myeloid hematologic disorders, including myelodysplasias and acute myeloid leukemia. © 2020 by John Wiley & Sons, Inc. Basic Protocol 1: Validation plan for fluorescence in situ hybridization (FISH) probes for chromosome 5 monosomy and deletion Support Protocol: Normal cut-off calculation Basic Protocol 2: Validation procedure for FISH probes for chromosome 5 monosomy and deletion Basic Protocol 3: Validation report for FISH probes for chromosome 5 monosomy and deletion.
Genetic research often utilizes or generates information that is potentially sensitive to individuals, families, or communities. For these reasons, genetic research may warrant additional scrutiny from investigators and governmental regulators, compared to other types of biomedical research. The informed consent process should address the range of social and psychological issues that may arise in genetic research. This paper addresses a number of these issues, including recruitment of participants, disclosure of results, psychological impact of results, insurance and employment discrimination, community engagement, consent for tissue banking, and intellectual property issues. Points of consideration are offered to assist in the development of protocols and consent processes in light of contemporary debates on a number of these issues.
Standard clinical interpretation of DNA copy number variants (CNVs) identified by cytogenomic microarray involves examining protein-coding genes within the region and comparison to other CNVs. Emerging basic research suggests that CNVs can also exert a pathogenic effect through disruption of DNA structural elements such as topologically associated domains (TADs). To begin to integrate these discoveries with current practice, we developed ClinTAD, a free browser-based tool to assist with interpretation of CNVs in the context of TADs (www.clintad.com). We used ClinTAD to examine 209 randomly selected single-nucleotide polymorphism microarray cases with a total of 236 CNVs. We compared 118 CNVs classified as variants of uncertain clinical significance (VUS), where additional insight into pathogenicity of these CNVs would be of greatest utility, to 118 CNVs classified as benign. We found that a higher proportion of VUS had at least two genes in a nearby TAD related to a phenotype seen in the patient based on Human Phenotype Ontology (HPO) annotation. We present example cases demonstrating scenarios where ClinTAD may either increase or decrease clinical suspicion of pathogenicity for VUS, depending on disruption of TAD boundaries and HPO phenotype match. ClinTAD is an easy-to-use tool, based on emerging research in chromatin architecture, that can help inform CNV interpretation.
Genotype imputation infers missing genotypes in silico using haplotype information from reference samples with genotypes from denser genotyping arrays or sequencing. This approach can confer a number of improvements on genome-wide association studies: it can improve statistical power to detect associations by reducing the number of missing genotypes; it can simplify data harmonization for meta-analyses by improving overlap of genomic variants between differently-genotyped sample sets; and it can increase the overall number and density of genomic variants available for association testing. This article reviews the general concepts behind imputation, describes imputation approaches and methods for various types of genotype data, including family-based data, and identifies web-based resources that can be used in different steps of the imputation process. For practical application, it provides a step-by-step guide to implementation of a two-step imputation process consisting of phasing of the study genotypes and the imputation of reference panel genotypes into the study haplotypes. In addition, this review describes recently developed haplotype reference panel resources and online imputation servers that are capable of remotely and securely implementing an imputation workflow on uploaded genotype array data. © 2019 by John Wiley & Sons, Inc.
With the advent of Next Generation Sequencing (NGS) technologies, whole genome and whole exome DNA sequencing has become affordable for routine genetic studies. Coupled with improved genotyping arrays and genotype imputation methodologies, it is increasingly feasible to obtain rare genetic variant information in large datasets. Such datasets allow researchers to gain a more complete understanding of the genetic architecture of complex traits caused by rare variants. State-of-the-art statistical methods for the statistical genetics analysis of sequence-based association, including efficient algorithms for association analysis in biobank-scale datasets, gene-association tests, meta-analysis, fine mapping methods that integrate functional genomic dataset, and phenome-wide association studies (PheWAS), are reviewed here. These methods are expected to be highly useful for next generation statistical genetics analysis in the era of precision medicine. © 2019 by John Wiley & Sons, Inc.