BACKGROUND:Next-generation sequencing has enabled precision therapeutic approaches that have improved the lives of children with rare diseases. Congenital diarrhea and enteropathies (CODEs) are associated with high morbidity and mortality. Although treatment of these disorders is largely supportive, emerging targeted therapies based on genetic diagnoses include specific diets, pharmacologic treatments, and surgical interventions. METHODS:We analyzed the exomes or genomes of infants with suspected monogenic congenital diarrheal disorders. Using cell and zebrafish models, we tested the effects of variants in newly implicated genes. RESULTS:In our case series of 129 infant probands with suspected monogenic congenital diarrheal disorders, we identified causal variants, including a new founder NEUROG3 variant, in 62 infants (48%). Using cell and zebrafish models, we also uncovered and functionally characterized three novel genes associated with CODEs: GRWD1, MYO1A, and MON1A. CONCLUSIONS:We have characterized the broad genetic architecture of CODE disorders in a large case series of patients and identified three novel genes associated with CODEs. (Funded by the National Institutes of Health and others.).
Introduction: Genomic variants that lead to MET proto-oncogenem receptor tyrosine kinase (MET) exon 14 skipping represent a potential targetable molecular abnormality in NSCLC. Consequently, reliable molecular diagnostic approaches that detect these variants are vital for patient care. Methods: We screened tumor samples from patients with NSCLC for MET exon 14 skipping by using two distinct approaches: a DNA-based next-generation sequencing assay that uses an amplicon-mediated target enrichment and an RNA-based next-generation sequencing assay that uses anchored multiplex polymerase chain reaction for target enrichment. Results: The DNA-based approach detected MET exon 14 skipping variants in 11 of 856 NSCLC samples (1.3%). The RNA-based approach detected MET exon 14 skipping in 17 of 404 samples (4.2%), which was a statistically significant increase compared with the DNA-based assay. Among 286 samples tested by both assays, RNA-based testing detected 10 positives, six of which were not detected by the DNA-based assay. Examination of primer binding sites in the DNA-based assay in comparison with published MET exon 14 skipping variants revealed genomic deletion involving primer binding sequences as the likely cause of false negatives. Two samples positive via the DNA-based approach were uninformative via the RNA-based approach due to poor-quality RNA. Conclusions: By circumventing an inherent limitation of DNA-based amplicon-mediated testing, RNA-based analysis detected a higher proportion of MET exon 14 skipping cases. However, RNA-based analysis was highly reliant on RNA quality, which can be suboptimal in some clinical samples. (C) 2019 International Association for the Study of Lung Cancer. Published by Elsevier Inc. All rights reserved.
The epithelial cell adhesion molecule gene (EPCAM, previously known as TACSTD1 or TROP1) encodes a membrane-bound protein that is localized to the basolateral membrane of epithelial cells and is overexpressed in some tumors. Biallelic mutations in EPCAM cause congenital tufting enteropathy (CTE), which is a rare chronic diarrheal disorder presenting in infancy. Monoallelic deletions of the 3 ' end of EPCAM that silence the downstream gene, MSH2, cause a form of Lynch syndrome, which is a cancer predisposition syndrome associated with loss of DNA mismatch repair. Here, we report 13 novel EPCAM mutations from 17 CTE patients from two separate centers, review EPCAM mutations associated with CTE and Lynch syndrome, and structurally model pathogenic missense mutations. Statistical analyses indicate that the c.499dupC (previously reported as c.498insC) frameshift mutation was associated with more severe treatment regimens and greater mortality in CTE, whereas the c.556-14A>G and c.491+1G>A splice site mutations were not correlated with treatments or outcomes significantly different than random simulation. These findings suggest that genotype-phenotype correlations may be useful in contributing to management decisions of CTE patients. Depending on the type and nature of EPCAM mutation, one of two unrelated diseases may occur, CTE or Lynch syndrome.
Next-generation sequencing (NGS) diagnostic assays increasingly are becoming the standard of care in oncology practice. As the scale of an NGS laboratory grows, management of these assays requires organizing large amounts of information, including patient data, laboratory processes, genomic data, as well as variant interpretation and reporting. Although several Laboratory Information Systems and/or Laboratory Information Management Systems are commercially available, they may not meet all of the needs of a given Laboratory, in addition to being frequently cost-prohibitive. Herein, we present the System for Informatics in the Molecular Pathology Laboratory (SIMPL), a free and open-source Laboratory Information System/Laboratory Information Management System for academic and nonprofit molecular pathology NGS laboratories, developed at the Genomic and Molecular Pathology Division at the University of Chicago Medicine. SIMPL was designed as a modular end-to-end information system to handle all stages of the NGS laboratory workload from test order to reporting. We describe the features of SIMPL, its clinical validation at University of Chicago Medicine, and its installation and testing within a different academic center laboratory (University of Colorado), and we propose a platform for future community co-development and interlaboratory data sharing.
s(RU) and right lower (RL) distal duodenum and duodenum bulb (6 o'clock) were significantly higher compared to controls.8/9 active CeD had Marsh 3 classification on histology.There were no adverse events.Conclusion: Duodenal mucosal impedance is altered in patients with active CeD compared to those with inactive CeD or normal subjects.We hypothesize this might be due to flattening of the duodenal mucosa which might increase mucosal impedance.Further studies using modifications to the MI catheter to increase sensitivity for detecting changes in the columnar epithelium might provide additional insight into an alternative diagnostic modality in patients with CeD pre-histology.Table 1: Duodenal mucosal impedance measurements (in Ohms) stratified by location in active CeD, inactive CeD, and controls.The values represent median measurements with ranges in parentheses showing the interquartile range.*Kruskal-Wallis test Sa2003
Next-generation sequencing (NGS) diagnostic assays increasingly are becoming the standard of care in oncology practice. As the scale of an NGS laboratory grows, management of these assays requires organizing large amounts of information, including patient data, laboratory processes, genomic data, as well as variant interpretation and reporting. Although several Laboratory Information Systems and/or Laboratory Information Management Systems are commercially available, they may not meet all of the needs of a given laboratory, in addition to being frequently cost-prohibitive. Herein, we present the System for Informatics in the Molecular Pathology Laboratory (SIMPL), a free and open-source Laboratory Information System/Laboratory Information Management System for academic and nonprofit molecular pathology NGS laboratories, developed at the Genomic and Molecular Pathology Division at the University of Chicago Medicine. SIMPL was designed as a modular end-to-end information system to handle all stages of the NGS laboratory workload from test order to reporting. We describe the features of SIMPL, its clinical validation at University of Chicago Medicine, and its installation and testing within a different academic center laboratory (University of Colorado), and we propose a platform for future community co-development and interlaboratory data sharing. Next-generation sequencing (NGS) diagnostic assays increasingly are becoming the standard of care in oncology practice. As the scale of an NGS laboratory grows, management of these assays requires organizing large amounts of information, including patient data, laboratory processes, genomic data, as well as variant interpretation and reporting. Although several Laboratory Information Systems and/or Laboratory Information Management Systems are commercially available, they may not meet all of the needs of a given laboratory, in addition to being frequently cost-prohibitive. Herein, we present the System for Informatics in the Molecular Pathology Laboratory (SIMPL), a free and open-source Laboratory Information System/Laboratory Information Management System for academic and nonprofit molecular pathology NGS laboratories, developed at the Genomic and Molecular Pathology Division at the University of Chicago Medicine. SIMPL was designed as a modular end-to-end information system to handle all stages of the NGS laboratory workload from test order to reporting. We describe the features of SIMPL, its clinical validation at University of Chicago Medicine, and its installation and testing within a different academic center laboratory (University of Colorado), and we propose a platform for future community co-development and interlaboratory data sharing. Over the past few years, laboratories increasingly have adopted next-generation sequencing (NGS) technologies for molecular diagnostics in clinical oncology because of the expanding diversity of diagnostic, prognostic, and therapeutic genomic markers that require assessment in the context of various malignancies.1Kamps R. Brandão R.D. van den Bosch B.J. Paulussen A.D.C. Xanthoulea S. Blok M.J. Romano A. Next-generation sequencing in oncology: genetic diagnosis, risk prediction and cancer classification.Int J Mol Sci. 2017; 18 (pii: E308)Crossref PubMed Scopus (260) Google Scholar Onboarding NGS technologies into the laboratory and keeping up with the intense pace of change in oncology diagnostics via continuous test evolution can be immensely challenging. The most commonly addressed NGS-associated obstacles relate to the complexity of the underlying molecular biology applications and the scale and processing of the primary sequencing data to uncover meaningful tumor-related anomalies.2Kadri S. Long B.C. Mujacic I. Zhen C.J. Wurst M.N. Sharma S. McDonald N. Niu N. Benhamed S. Tuteja J.H. Seiwert T.Y. White K.P. McNerney M.E. Fitzpatrick C. Wang Y.L. Furtado L.V. Segal J.P. Clinical validation of a next-generation sequencing genomic oncology panel via cross-platform benchmarking against established amplicon sequencing assays.J Mol Diagn. 2017; 19: 43-56Abstract Full Text Full Text PDF PubMed Scopus (77) Google Scholar, 3Cheng D.T. Mitchell T.N. Zehir A. Shah R.H. Benayed R. Syed A. Chandramohan R. Liu Z.Y. Won H.H. Scott S.N. Brannon A.R. O'Reilly C. Sadowska J. Casanova J. Yannes A. Hechtman J.F. Yao J. Song W. Ross D.S. Oultache A. Dogan S. Borsu L. Hameed M. Nafa K. Arcila M.E. Ladanyi M. Berger M.F. Memorial sloan kettering-integrated mutation profiling of actionable cancer targets (MSK-IMPACT): a hybridization capture-based next-generation sequencing clinical assay for solid tumor molecular oncology.J Mol Diagn. 2015; 17: 251-264Abstract Full Text Full Text PDF PubMed Scopus (1179) Google Scholar, 4Pritchard C.C. Salipante S.J. Koehler K. Smith C. Scroggins S. Wood B. Wu D. Lee M.K. Dintzis S. Adey A. Validation and implementation of targeted capture and sequencing for the detection of actionable mutation, copy number variation, and gene rearrangement in clinical cancer specimens.J Mol Diagn. 2014; 16: 56-67Abstract Full Text Full Text PDF PubMed Scopus (198) Google Scholar However, a less-appreciated problem is the general organization of the laboratory and the management of laboratory data and information flows, which can become urgent and compelling as the laboratory scale grows with few or no straightforward solutions. Proper management of NGS diagnostic assays requires the organization of large amounts of information about patients, specimens, laboratory processes, and process status, as well as storage and management of genetic variants, interpretations, and reports. Often, laboratories use spreadsheets and e-mails to organize these data, but these methods can be insecure and are inefficient as volumes inevitably increase. Laboratory Information Systems (LISs) and/or Laboratory Information Management Systems (LIMSs) are not new to the molecular pathology laboratory, but the need for specialized systems is greatly heightened by the complexity of oncology NGS sample management, library preparation, sequencing, and data interpretation, compared with more traditional molecular pathology assays and workflows.5Roy S. Durso M.B. Wald A. Nikiforov Y.E. Nikiforova M.N. SeqReporter: automating next-generation sequencing result interpretation and reporting workflow in a clinical laboratory.J Mol Diagn. 2014; 16: 11-22Abstract Full Text Full Text PDF PubMed Scopus (23) Google Scholar, 6Aronson S.J. Clark E.H. Babb L.J. Baxter S. Farwell L.M. Funke B.H. Hernandez A.L. Joshi V.A. Lyon E. Parthum A.R. Russell F.J. Varugheese M. Venman T.C. Rehm H.L. The GeneInsight Suite: a platform to support laboratory and provider use of DNA-based genetic testing.Hum Mutat. 2011; 32: 532-536Crossref PubMed Scopus (65) Google Scholar, 7Sharma M.K. Phillips J. Agarwal S. Wiggins W.S. Shrivastava S. Koul S.B. Bhattacharjee M. Houchins C.D. Kalakota R.R. George B. Meyer R.R. Spencer D.H. Lockwood C.M. Nguyen T.T. Duncavage E.J. Al-Kateb H. Cottrell C.E. Godala S. Lokineni R. Sawant S.M. Chatti V. Surampudi S. Sunkishala R.R. Darbha R. Macharla S. Milbrandt J.D. Virgin H.W. Mitra R.D. Head R.D. Kulkarni S. Bredemeyer A. Pfeifer J.D. Seibert K. Nagarajan R. Clinical genomicist workstation.AMIA Jt Summits Transl Sci Proc. 2013; 2013: 156-157PubMed Google Scholar Some of the most challenging areas are as follows. i) Specimens: molecular oncology specimen workflow processes are difficult in general, requiring review of perhaps multiple specimens, block selections, management of recuts, and assessment of adequacy and tumor purity. NGS analysis may compound these difficulties because of potentially more stringent specimen requirements compared with single-gene tests. ii) Workflow tracking: compared with PCR-based molecular pathology assays, NGS laboratory workflows may be highly variable (eg, amplicon versus hybrid capture) and may require a large number of steps over multiple days, potentially with more than one technologist participating in the preparation. NGS also has the unique feature of library pooling before sequencing, based on planned complementarity of sample-specific barcode sequences. Thus, sequencing batches typically include multiple sample libraries, which may be a problematic piece of logic to manage for many traditional molecular laboratory information systems. iii) Bioinformatics processing: every laboratory that performs clinical NGS uses either commercially available or custom data processing pipelines, which may vary significantly, raising the issue of whether and to what degree this aspect of the laboratory should or could be integrated into an information management system. As pipelines are updated, it also is critical to track the pipeline version that was used to process each specimen. iv) Interpretation and reporting: for an NGS clinical laboratory to function properly, it is essential to have a support platform for reviewing and interpreting final NGS data and creating reports. Historical variants and interpretations need to be archived and should be searchable to allow for easy review of new cases, and there is a need to assemble all relevant case information into a final document for reporting. v) Overall case management: to prevent confusion and minimize turnaround time, the status of each of these steps needs to be continuously tracked such that laboratory staff can quickly determine which samples require which processing step. As laboratory volume increases, the difficulty of maintaining awareness of the status of every specimen and analyzed data set in the laboratory grows, and because of the complexity of NGS workflows, it ultimately can become unmanageable without an effective status tracking system. Team sign-out organization also can be problematic, and a mechanism for clear assignment of responsibility for case review and completion can be extremely beneficial. As the available options were investigated, it was found that the available commercial LIMS/LIS options did not meet all of our requirements and also frequently were extremely cost-prohibitive. As a result, a modular end-to-end information system was generated to handle this workload to cover all stages from test order to reporting. During the development process, workflow tracking and data management issues common across molecular pathology laboratories were focused on avoiding implementation of logic unique to our laboratory whenever possible, in the interest of creating a system with the greatest potential to support the continued evolution in more than one laboratory. Here, we present the System for Informatics in the Molecular Pathology Laboratory (SIMPL), a free and open-source LIS/LIMS system for nonprofit molecular pathology NGS laboratories, developed at the Genomic and Molecular Pathology Division at the University of Chicago Medicine (UCM-GMP). We also describe its features, clinical validation of the system at UCM-GMP, its installation, testing within a different academic center laboratory, and propose options for possible future community co-development and interlaboratory data sharing. The authors should be contacted to obtain a copy of the SIMPL codebase. SIMPL is a web-based LIS implemented largely in Django (Django, https://www.djangoproject.com, last accessed December 19, 2017), programmed in Python, because of its straightforward architecture and approachable database design. It was developed as a result of UCM-GMP's collaboration with the University of Chicago Center for Research Informatics, and takes advantage of powerful and secure infrastructure that was already available. In UCM-GMP's configuration, SIMPL runs on virtual machines within a large secure computing cluster maintained by the Center for Research Informatics, following the organization's Information Technology security policies, which are based on the NIST 800-53 Cybersecurity Framework (https://www.nist.gov/cyberframework, last accessed November 1, 2017). The main software runs on a web server connected to a database server running MySQL software version 5.6.36 (Oracle Corporation, Redwood City, CA) (Supplemental Figure S1). SIMPL incorporates Secure Sockets Layer encryption and allows for Lightweight Directory Access Protocol (LDAP) authentication for user login. This allows straightforward connection to existing hospital user verification systems for security and password management. A part-time employee is responsible for maintaining security and functional updates to SIMPL. SIMPL is designed to help manage molecular pathology information management across the entirety of the laboratory testing process, including pre-analytic, analytic, and postanalytic phases. This includes recording patient information and associated NGS test orders, specimen tracking processes, DNA/RNA extraction, library preparation, and sequencing batches, as well as storage of variants (and other result types) from the assay performed. Interpretations can be added to each genomic result, and previous interpretations can be searched, copied, or modified to assist with ongoing analysis. The system has the ability to autogenerate editable reports for each patient test including patient and specimen details, results, interpretations, and general information about the test. In addition to the clinical module, SIMPL also is designed to support some functionality for research samples via a research module (see Research Module), because NGS clinical laboratories often participate simultaneously in clinical care and translational scientific projects. Figure 1 shows a schematic representation of the three main modules of the system, with the first module handling patient, order, and specimen information (patient and order tracking); the second module handling laboratory process batch information (laboratory process tracking); and the last module handling interpretation and reporting (genomics and reporting). Supplemental Figure S2 shows the dashboard of the SIMPL web interface, which is the screen seen by the user after logging in. This screen shows a snapshot of all of the samples currently in process by the laboratory and is dynamically updated. The plus sign separates the clinical and research samples. SIMPL was beta-tested over a period of 2 years, during which each module of the system was incrementally developed, tested, and improved using mock data in a test environment on a development server. This system is now clinically live at the Molecular Pathology Laboratory at UCM-GMP. In SIMPL, limited protected health information is stored for each patient in the system including the full name, medical record number, date of birth, and sex. Each patient can be assigned to one or multiple categories, each with a three-letter prefix that decides the internal unique identifier for this patient in its category. At UCM-GMP, patients receive the designation CGL (for Clinical Genomics Laboratory) or other customized prefixes for particular research projects, determined within the research module described below. A patient thus may have multiple linked identifiers. Every time a new patient is added to SIMPL in a specific category, the system automatically increments and assigns the next available number in the category. Figure 1 shows the various status gates that each test order may proceed through in dashed boxes. At any given point, a user logged into the system can access an order and check the status of the order. The "Order Status" section in Figure 2 shows the progression of an example test order. Users also can ask for reports detailing which cases are awaiting particular steps in the process. SIMPL allows each subject to receive multiple test orders, either on the same or separate specimens. Duplicate orders can be detected based on the associated CoPath (Cerner, Kansas City, MO) IDs. Recorded order information includes the date, requesting physician, hospital, test, and diagnosis (International Classification of Diseases, 10th revision code from the order). The system is designed to accept orders from multiple hospitals, and thus stores basic information about the hospitals and the physicians. After a test order is generated for a specific patient, the test order number is automatically incremented in the system, starting with T1 for the first order. The status of the order is set to "New Orders" at this time (Figure 1). The order entry process is currently manual in SIMPL, performed by accessioning staff. In a future update, an interface between SIMPL and other hospital information systems may be included. The test orders are linked to specific specimen processes belonging to each subject (patient) and each specimen can be tracked in SIMPL (on the "Specimen Screen"). Multiple specimens from the same patient can be linked back to the patient, and the specimen process screen in SIMPL can be used to see the previous specimens tested in a dropdown menu. Logic has been introduced to the web server such that depending on the type of specimen source (formalin-fixed, paraffin-embedded; peripheral blood; cytology smear; bone marrow), certain fields related to the specimen process are mandatory. As an example, the collection date is mandatory for the blood and bone marrow samples, whereas the tumor cell percentage is mandatory only for cytology smears and formalin-fixed, paraffin-embedded specimens (a separate free-text descriptive field is available for blood and bone marrow specimens to describe associated hematopathology findings). Recut request and receipt dates as well as overall reviews of specimen adequacy can be recorded. SIMPL also is equipped to handle special cases in which instead of the specimen the laboratory directly receives DNA/RNA or the specimen type is unknown. OncoTree specimen and diagnosis classifications are built-in and can be recorded for each specimen to facilitate retrospective data mining (OncoTree, http://oncotree.mskcc.org, last accessed December 19, 2017). The specimen process in SIMPL allows for recording of all information associated with this process, before the decision is made to either deliver the specimen to the laboratory extraction ("In Lab") or to fail the specimen. Specimens also may be failed from the laboratory if, for example, the DNA yield is suboptimal. In such cases, either a new specimen process may be generated if another potential specimen exists ("Rescreen Orders"), or the entire test may be canceled ("Cancelled" status in SIMPL) (Figure 1). SIMPL performs batch level management of specimens as they move through the laboratory for extraction, library preparation, and sequencing (Figure 1). Specimen processes that pass adequacy checks can be included in batches for nucleic acid extraction, with the requirement for DNA versus RNA based on the recorded features of the laboratory assay. After extraction, test status is updated to "DNA extracted" or "RNA extracted." The extracted samples then are available for batches of library preparation ("Library Prep"), which then are available for batches of sequencing runs ("Sequenced"). For each batch, the laboratory technician (or operator) creating the batch is logged into the system. If a test is ordered on a specimen that previously was extracted/tested, the system automatically alerts the user of this situation. At UCM-GMP, the technician then checks whether adequate material is available and can decide to either make a new extraction of the specimen or use the previously extracted material. Once the samples are sequenced and data are available, all bioinformatics pipelines are run on secure high-performance computing clusters hosted at the University of Chicago. Currently, the pipeline processing and data management in SIMPL are kept separate, but may be linked in the future. The sequencing data for each test order is processed according to the latest version of the clinically validated bioinformatics pipeline specific for the test on a high-performance computing cluster, and the pipeline versions are logged in SIMPL when the data are uploaded along with other run-specific metadata. The "Lab Assay" section in Figure 2 shows the run-specific information of an example test order in SIMPL. The "+ Add Variants" button shown in Figure 2 can be used to perform variant uploads through the web interface. The genomics module of SIMPL is used after the assay results are available ("Analysis completed"). The data can be uploaded individually for each sample using the web interface or using an application program interface for batch uploads. The system can handle both variant and nonvariant results, which include copy number, fusion, and other structural rearrangements reported by UCM NGS clinical assays. Each variant is stored as a combination of chromosome, position, reference, and mutation, and sample-specific variants store the pipeline version generating that variant as well as depth information at that genomic position and the variant allele frequency. Each variant is also linked to its annotation, specific to the annotation software. Variant calls are annotated and converted to Human Genome Variation Society nomenclature using Alamut Batch software version 1.4.4 (Interactive Biosoftware, Rouen, France), which also pulls from publicly available databases such as COSMIC (http://www.sanger.ac.uk/cosmic, last accessed December 19, 2017),8Forbes S.A. Bindal N. Bamford S. Cole C. Kok C.Y. Beare D. Jia M. Shepherd R. Leung K. Menzies A. Teague J.W. Campbell P.J. Stratton M.R. Futreal P.A. COSMIC: mining complete cancer genomes in the catalogue of somatic mutations in cancer.Nucleic Acids Res. 2011; 39: D945-D950Crossref PubMed Scopus (1812) Google Scholar NCBI dbSNP,9Sherry S.T. Ward M.H. Kholodov M. Baker J. Phan L. Smigielski E.M. Sirotkin K. dbSNP: the NCBI database of genetic variation.Nucleic Acids Res. 2001; 29: 308-311Crossref PubMed Scopus (4834) Google Scholar Scale Invariable Feature Transformation (SIFT) algorithm,10Kumar P. Henikoff S. Ng P.C. Predicting the effects of coding non-synonymous variants on protein function using the SIFT algorithm.Nat Protoc. 2009; 4: 1073-1081Crossref PubMed Scopus (5005) Google Scholar and so forth. Older assays at UCM-GMP were annotated using ANNOVAR,11Wang K. Li M. Hakonarson H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data.Nucleic Acids Res. 2010; 38: e164Crossref PubMed Scopus (7866) Google Scholar and SIMPL can store these annotations as well, but the database can be easily modified to adapt to a center's annotation system. Although research samples are stored in SIMPL, genomic results are not stored for research samples because these are not interpreted in the system and thus do not influence the database statistics used for result interpretation. However, users have the ability to upload results for research samples as well. After the bioinformatics pipelines are completed, results are uploaded to SIMPL as described above, and the order status is changed to "Analysis Completed." At this time, the cases are available for assignment by the pathologists through a case assignment window (Supplemental Figure S3). The case reviewers and the pathologists review the primary data, make interpretations of the detected variants, and assemble and sign out a final report detailing the case findings along with comments or recommendations that may be appropriate within the context of each patient's disease process. The case can be assigned to one reviewer to draft a report and one molecular pathologist to finalize and sign out the report. Once a case is assigned (Supplemental Figure S3), the case will be available for the assigned users and the assignment status and related comments are tracked by SIMPL. When the primary reviewer completes the case and finalizes their report, an e-mail is generated automatically and sent to the case pathologist, and the case will appear in the work list of the case pathologist. When the case is officially signed out in SIMPL, its status will change from "Analysis Completed" to "Reported," and it will drop from the queue of the case pathologist. The overall workflow for case review includes the following: i) variant review and creation/modification of interpretations, ii) creation/modification of any necessary nonvariant interpretations, iii) report generation and review, and iv) case completion/sign-out (Figure 3). The variant review window is implemented as a scrollable window, with all column headings allowing sorting or filtering using arrow buttons, selectors, or blank fields (Figure 3A). For example, to remove common inherited variants one might filter out variants present at appreciable frequency (eg, 1%) using the Max1000 (1000 Genomes Project Max allele frequency field) field by entering "0.01" in the "To" field. Variants then may continue to be reviewed as per assay-specific guidelines. Any variants deemed worthy of reporting may have interpretations assigned to them using the edit button. The interpretation window prepopulates with the Human Genome Variation Society nomenclature from the annotation, which can be edited by the pathologist, and provides a box to add interpretive text and a drop-down choice of pathogenic rating (Figure 3B). If the same variant has been seen in the laboratory before, the previous interpretation will autopopulate to the bottom of the window for review, and there is an advanced search button that allows for searching of previous interpretations by gene, diagnosis, pathogenic rating, and so forth. Interpretations can be finalized using an "Is final" button. The use of a database to store this information increases the scope for data mining projects, and value for data sharing in larger genomic data-sharing initiatives such as the GEnetics of Nephropathy—an International Effort (GENIE) consortium.12AACR Project GENIE ConsortiumAACR Project GENIE: powering precision medicine through an international consortium.Cancer Discov. 2017; 7: 818-831Crossref PubMed Scopus (748) Google Scholar An example is shown in Figure 4, which shows the top 25% genes with pathogenic mutations in all cases run on UCM-OncoPlus,2Kadri S. Long B.C. Mujacic I. Zhen C.J. Wurst M.N. Sharma S. McDonald N. Niu N. Benhamed S. Tuteja J.H. Seiwert T.Y. White K.P. McNerney M.E. Fitzpatrick C. Wang Y.L. Furtado L.V. Segal J.P. Clinical validation of a next-generation sequencing genomic oncology panel via cross-platform benchmarking against established amplicon sequencing assays.J Mol Diagn. 2017; 19: 43-56Abstract Full Text Full Text PDF PubMed Scopus (77) Google Scholar split by hematologic (Figure 4A) and solid tumor (Figure 4B) cases. The knowledge base stores the more recently updated annotation for each variant along with every interpretation that has ever been reported for the variant. The knowledge base can be queried by authorized individuals using advanced search pages generated on the website, based on pathogenic level, OncoTree classifications, interpretive text, coding change, and so forth. Nonvariant interpretations are any interpretation of an identified genomic anomaly as a result of the test that is not a variant. Potential nonvariant findings include copy number abnormalities, gene fusions, rearrangements, and so forth. These interpretations contain the type of anomaly and a free text box in which to enter the proper nomenclature for the finding, along with a clinical interpretation and a pathogenic rating. Reports can be generated automatically in SIMPL by clicking "Generate Report," which becomes available after the results are uploaded and until a report is "Finalized." SIMPL uses a predefined report template specific to the clinical test and populates it with pertinent case information including diagnosis, specimen information, and all saved interpretations from the SIMPL database into a single text-based document, which is available for review and editing. Only laboratory directors have the privileges to edit the report templates for each test (see the User Roles and Groups section below). UCM-GMP uses only text-based reporting because of the nature of the hospital information systems, but modifying the SIMPL code to instead produce a formatted PDF report would be quite straightforward. SIMPL is a web application that was built on top of the Django Web Framework (version 1.11.9), which was developed in Python. The front end (client side) was written using HTML 5, CSS 3.0, and JavaScript. The major user interface was developed using Bootstrap 3.3 (https://getbootstrap.com, last accessed June 8, 2018) and jQuery 2.2 (http://www.cs.ubc.ca/labs/spl/projects/jquery, last accessed June 8, 2018). The back end (server side) was scripted using Django version 1.11 on Python 3.5. MySQL version 5.7 was selected as the relational database management system to store all user-generated data. Users could be authenticated either through one or more institutional LDAP servers or through local user accounts stored in SIMPL. Elasticsearch version 2.3 (Elasticsearch BV, Mountain View, CA) was used as the search platform for generalized full-text searches. All of the software used in SIMPL is open-source. Hardware requirements are modest and the system runs in Windows/IIS (Microsoft, Redmond, WA) and Linux/Apache (Apache Software Foundation, Forest Hill, MD) or Nginx (https://www.nginx.com, last accessed June 8, 2018) webserver environments. Supplemental Figure S1 shows the diagram of the system architecture and software environment of SIMPL. The SIMPL database and code are structured so that all user interactions with the system are logged, allowing for retrospective evaluation of all SIMPL user activity since the implementation of the system. This feature is extremely helpful for helping troubleshoot laboratory problems or errors. The SIMPL database is backed up nightly to redundant tape drive systems housed within the Center f
Disturbed mitochondrial fusion and fission have been linked to various neurodegenerative disorders. In siblings from two unrelated families who died soon after birth with a profound neurodevelopmental disorder characterized by pontocerebellar hypoplasia and apnoea, we discovered a missense mutation and an exonic deletion in the SLC25A46 gene encoding a mitochondrial protein recently implicated in optic atrophy spectrum disorder. We performed functional studies that confirmed the mitochondrial localization and pro-fission properties of SLC25A46. Knockdown of slc24a46 expression in zebrafish embryos caused brain malformation, spinal motor neuron loss, and poor motility. At the cellular level, we observed abnormally elongated mitochondria, which was rescued by co-injection of the wild-type but not the mutant slc25a46 mRNA. Conversely, overexpression of the wild-type protein led to mitochondrial fragmentation and disruption of the mitochondrial network. In contrast to mutations causing non-lethal optic atrophy, missense mutations causing lethal congenital pontocerebellar hypoplasia markedly destabilize the protein. Indeed, the clinical severity appears inversely correlated with the relative stability of the mutant protein. This genotype-phenotype correlation underscores the importance of SLC25A46 and fine tuning of mitochondrial fission and fusion in pontocerebellar hypoplasia and central neurodevelopment in addition to optic and peripheral neuropathy across the life span.
Background . Both prospective and retrospective studies have indicated that ex- trapulmonary dissemination of coccidioidomycosis occurs at an elevated rate among otherwise healthy African Americans and Filipinos, suggesting genetic predisposing factors may be present. The growing fi eld of immunogenetics seeks to catalog and un- derstand genetic differences so that diagnosis and treatment of infection can be opti-mized. Methods . With approval from the Of fi ce for Protection of Research subjects, pe- ripheral blood was obtained from a group of 20 ethnically diverse adults with extrapulmonary dissemination of coccidioidomycosis . DNA was extracted from peripheral bloodmononuclearcellsandsubmitted forwholeexomesequence analysisbytheUni-versity of California Los Angeles Core Microarray Laboratory on an Illumina HiSeq 2500 instrument. We fi ltered the resulting datafor variants in genes known or suspect-ed to be in pathways relating to responses to fungal infections. Results . We identi fi ed a mean of 246,708 variants from the human reference ge-nome per sample, ofwhicha mean of6336per sample
ADAM metallopeptidase domain 17 (ADAM17) is responsible for processing large numbers of proteins. Recently, 1 family involving 2 patients with a homozygous mutation in ADAM17 were described, presenting with skin lesions and diarrhea. In this report, we describe a second family confirming the existence of this syndrome. The proband presented with severe diarrhea, skin rash, and recurrent sepsis, eventually leading to her death at the age of 10 months. We performed exome sequencing and detailed pathological and immunological investigations. We identified a novel homozygous frameshift mutation in ADAM17 (NM_003183.4:c.308dupA) leading to a premature stop codon. CD4+ and CD8+ T-cell stimulation assays showed severely diminished tumor necrosis factor–α and interleukin-2 production. Skin biopsies indicated a focal neutrophilic infiltrate and spongiotic dermatitis. Interestingly, the patient developed unexplained systolic hypertension and nonspecific hepatitis with apoptosis. This report provides evidence for an important role of ADAM17 in human immunological response and underscores its multiorgan involvement.
High-throughput DNA sequencing has become a mainstay for the discovery of genomic variants that may cause disease or affect phenotype. A next-generation sequencing pipeline typically identifies thousands of variants in each sample. A particular challenge is the annotation of each variant in a way that is useful to downstream consumers of the data, such as clinical sequencing centers or researchers. These users may require that all data storage and analysis remain on secure local servers to protect patient confidentiality or intellectual property, may have unique and changing needs to draw on a variety of annotation data sets and may prefer not to rely on closed-source applications beyond their control. Here we describe scalable methods for using the plugin capability of the Ensembl Variant Effect Predictor to enrich its basic set of variant annotations with additional data on genes, function, conservation, expression, diseases, pathways and protein structure, and describe an extensible framework for easily adding additional custom data sets.
IMPORTANCE:Clinical exome sequencing (CES) is rapidly becoming a common molecular diagnostic test for individuals with rare genetic disorders.OBJECTIVE:To report on initial clinical indications for CES referrals and molecular diagnostic rates for different indications and for different test types.DESIGN, SETTING, AND PARTICIPANTS:Clinical exome sequencing was performed on 814 consecutive patients with undiagnosed, suspected genetic conditions at the University of California, Los Angeles, Clinical Genomics Center between January 2012 and August 2014. Clinical exome sequencing was conducted as trio-CES (both parents and their affected child sequenced simultaneously) to effectively detect de novo and compound heterozygous variants or as proband-CES (only the affected individual sequenced) when parental samples were not available.MAIN OUTCOMES AND MEASURES:Clinical indications for CES requests, molecular diagnostic rates of CES overall and for phenotypic subgroups, and differences in molecular diagnostic rates between trio-CES and proband-CES.RESULTS:Of the 814 cases, the overall molecular diagnosis rate was 26% (213 of 814; 95% CI, 23%-29%). The molecular diagnosis rate for trio-CES was 31% (127 of 410 cases; 95% CI, 27%-36%) and 22% (74 of 338 cases; 95% CI, 18%-27%) for proband-CES. In cases of developmental delay in children (<5 years, n = 138), the molecular diagnosis rate was 41% (45 of 109; 95% CI, 32%-51%) for trio-CES cases and 9% (2 of 23, 95% CI, 1%-28%) for proband-CES cases. The significantly higher diagnostic yield (P value = .002; odds ratio, 7.4 [95% CI, 1.6-33.1]) of trio-CES was due to the identification of de novo and compound heterozygous variants.CONCLUSIONS AND RELEVANCE:In this sample of patients with undiagnosed, suspected genetic conditions, trio-CES was associated with higher molecular diagnostic yield than proband-CES or traditional molecular diagnostic methods. Additional studies designed to validate these findings and to explore the effect of this approach on clinical and economic outcomes are warranted.
Four siblings presented with congenital diarrhea and various endocrinopathies. Exome sequencing and homozygosity mapping identified five regions, comprising 337 protein-coding genes that were shared by three affected siblings. Exome sequencing identified a novel homozygous N309K mutation in the proprotein convertase subtilisin/kexin type 1 (PCSK1) gene, encoding the neuroendocrine convertase 1 precursor (PC1/3) which was recently reported as a cause of Congenital Diarrhea Disorder (CDD). The PCSK1 mutation affected the oxyanion hole transition state-stabilizing amino acid within the active site, which is critical for appropriate proprotein maturation and enzyme activity. Unexpectedly, the N309K mutant protein exhibited normal, though slowed, prodomain removal and was secreted from both HEK293 and Neuro2A cells. However, the secreted enzyme showed no catalytic activity, and was not processed into the 66 kDa form. We conclude that the N309K enzyme is able to cleave its own propeptide but is catalytically inert against in trans substrates, and that this variant accounts for the enteric and systemic endocrinopathies seen in this large consanguineous kindred.
Bipolar disorder is a common, complex, and severe psychiatric disorder with cyclical disturbances of mood and a high suicide rate. Here, we describe a family with four siblings, three affected females and one unaffected male. The disease course was characterized by early-onset bipolar disorder and co-morbid anxiety spectrum disorders that followed the onset of bipolar disorder. Genetic risk factors were suggested by the early onset of the disease, the severe disease course, including multiple suicide attempts, and lack of adverse prenatal or early life events. In particular, drug and alcohol abuse did not contribute to the disease onset. Exome sequencing identified very rare, heterozygous, and likely protein-damaging variants in eight brain-expressed genes: IQUB, JMJD1C, GADD45A, GOLGB1, PLSCR5, VRK2, MESDC2, and FGGY. The variants were shared among all three affected family members but absent in the unaffected sibling and in more than 200 controls. The genes encode proteins with significant regulatory roles in the ERK/MAPK and CREB-regulated intracellular signaling pathways. These pathways are central to neuronal and synaptic plasticity, cognition, affect regulation and response to chronic stress. In addition, proteins in these pathways are the target of commonly used mood stabilizing drugs, such as tricyclic antidepressants, lithium and valproic acid. The combination of multiple rare, damaging mutations in these central pathways could lead to reduced resilience and increased vulnerability to stressful life events. Our results support a new model for psychiatric disorders, in which multiple rare, damaging mutations in genes functionally related to a common signaling pathway contribute to the manifestation of bipolar disorder.
BACKGROUND & AIMS Proprotein convertase 1/3 (PC1/3) deficiency, an autosomal-recessive disorder caused by rare mutations in the proprotein convertase subtilisin/kexin type 1 (PCSK1) gene, has been associated with obesity, severe malabsorptive diarrhea, and certain endocrine abnormalities. Common variants in PCSK1 also have been associated with obesity in heterozygotes in several population-based studies. PC1/3 is an endoprotease that processes many prohormones expressed in endocrine and neuronal cells. We investigated clinical and molecular features of PC1/3 deficiency. METHODS We studied the clinical features of 13 children with PC1/3 deficiency and performed sequence analysis of PCSK1. We measured enzymatic activity of recombinant PC1/3 proteins. RESULTS We identified a pattern of endocrinopathies that develop in an age-dependent manner. Eight of the mutations had severe biochemical consequences in vitro. Neonates had severe malabsorptive diarrhea and failure to thrive, required prolonged parenteral nutrition support, and had high mortality. Additional endocrine abnormalities developed as the disease progressed, including diabetes insipidus, growth hormone deficiency, primary hypogonadism, adrenal insufficiency, and hypothyroidism. We identified growth hormone deficiency, central diabetes insipidus, and male hypogonadism as new features of PCSK1 insufficiency. Interestingly, despite early growth abnormalities, moderate obesity, associated with severe polyphagia, generally appears. CONCLUSIONS In a study of 13 children with PC1/3 deficiency caused by disruption of PCSK1, failure of enteroendocrine cells to produce functional hormones resulted in generalized malabsorption. These findings indicate that PC1/3 is involved in the processing of one or more enteric hormones that are required for nutrient absorption.
Background Common single nucleotide polymorphisms (SNPs) in proprotein convertase subtilisin/kexin type 1 with modest effects on PC1/3 in vitro have been associated with obesity in five genome-wide association studies and with diabetes in one genome-wide association study. We here present a novel SNP and compare its biosynthesis, secretion and catalytic activity to wild-type enzyme and to SNPs that have been linked to obesity. Methodology/Principal Findings A novel PC1/3 variant introducing an Arg to Gln amino acid substitution at residue 80 (within the secondary cleavage site of the prodomain) (rs1799904) was studied. This novel variant was selected for analysis from the 1000 Genomes sequencing project based on its predicted deleterious effect on enzyme function and its comparatively more frequent allele frequency. The actual existence of the R80Q (rs1799904) variant was verified by Sanger sequencing. The effects of this novel variant on the biosynthesis, secretion, and catalytic activity were determined; the previously-described obesity risk SNPs N221D (rs6232), Q665E/S690T (rs6234/rs6235), and the Q665E and S690T SNPs (analyzed separately) were included for comparative purposes. The novel R80Q (rs1799904) variant described in this study resulted in significantly detrimental effects on both the maturation and in vitro catalytic activity of PC1/3. Conclusion/Significance Our findings that this novel R80Q (rs1799904) variant both exhibits adverse effects on PC1/3 activity and is prevalent in the population suggests that further biochemical and genetic analysis to assess its contribution to the risk of metabolic disease within the general population is warranted.
High throughput, massively parallel DNA sequencing provides a powerful technology to study the human genome and to identify variations in DNA that cause disease. Sequencing the protein coding region of the genome (`whole-exome sequencing') is a cost effective method to search the part of the genome that is most likely to harbor disease related mutations.We developed software methods to process sequencing data and to annotate variants with data on genes, function, conservation, expression, diseases, pathways, and protein structure. We applied whole-exome sequencing to search for the molecular basis of disease in three projects: 1) a cohort of patients with congenital diarrheal disorders (CDDs); 2) a cohort of patients with congenital chronic intestinal pseudo-obstruction (CIPO) or the related disease, megacystis-microcolon-intestinal hypoperistalsis syndrome (MMIH); and 3) four siblings with infantile pontocerebellar hypoplasia and spinal motor neuron degeneration.We sequenced 45 probands from diverse ethnic backgrounds who were diagnosed with a variety of CDDs of probable, but unknown genetic cause. Patients had been diagnosed with generalized malabsorptive diarrhea, selective nutrient malabsorption, secretory diarrhea, and infantile IBD. We found homozygous or compound heterozygous mutations, 25 of them novel, in genes known to be associated with CDDs in 27 cases (60%). The genes implicated were ADAM17, DGAT1, EPCAM, IL10RA, MALT1, MYO5B, NEUROG3, PCSK1, SI, SKIV2L, SLC26A3, and SLC5A.With whole-exome sequencing in a cohort of 20 patients with congenital CIPO or MMIH, we identified a subset of 10 cases with potentially damaging de-novo dominant acting mutations at highly conserved loci in the ACTG2 gene, encoding actin, gamma-enteric smooth muscle precursor, a protein essential to the functioning of muscle cells in the intestinal wall.By exome sequencing, we discovered rare recessive mutations in EXOSC3 (encoding exosome component 3) that were responsible for pontocerebellar hypoplasia and spinal motor neuron degeneration in the four probands, and identified identical and additional novel mutations in a large percentage of other children with the same disorder.In conclusion, we demonstrated that whole-exome sequencing is an effective approach for the identification of casual mutations in that may escape detection with standard practice involving a complex diagnostic workup and targeted gene sequencing.
The intestinal mucus layer plays a key role in the maintenance of host-microbiota homeostasis.The production of goblet cells, which secrete mucus, is modulated by microbiota, but neither the species nor the mechanisms involved in this process are still unknown.We studied how two prominent commensal bacteria may influence the mucus production by goblet cells and the profile of mucin glycosylation in gnotobiotic rats.We have chosen Bacteroides thetaiotaomicron, which is characterized by its high mucus-polysaccharides degrading potential, and Faecalibacterium prausnitzii, which is a sensor of intestinal health.Germ free rats (GF) were orally inoculated with B. thetaiotaomicron either alone or with a mix of B. thetaiotaomicron and F. prausnitzii leading respectively to mono-associated (Bt-rats) and di-associated rats (Bt+Fp-rats).A panel of goblet cells markers was analyzed by histological staining, immunohistochemistry, quantitative PCR and Western blot in colon epithelium.The mucin O-glycosylation was determined by MALDI TOF mass spectrometry.In Bt-rats, the goblet cells number and the expression of mucus-related genes (muc2, muc4, klf4, c1galt1 and b4galt4 mRNAs) were increased compared to GF ones.KLF4 protein, a transcription factor involved in goblet cell terminal differentiation, was also increased in Bt-rats, whereas a decrease in Chromogranin A protein, a marker of enteroendocrine cells was observed.We propose that B. thetaiotaomicron provokes an imbalance inside the secretory lineage by favoring mucus production at the expense of enteroendocrine cells.When B. thetaiotaomicron was associated to F. prausnitzii, the effects on goblet cells were reduced/ decreased/diminished. Indeed, the number of goblet cells per crypt and the amount of KLF4 protein were lower in Bt+Fp-rats than in Bt-rats.We then analyzed the mucus quality by studying the profile of mucin O-glycosylation.In Bt-rats, a decrease in the production of sulfated (4.5% of total oligosaccharides instead of 12.9%) and neutral (40.1% instead of 52.8%) oligosaccharides was observed and was correlated to an increased proportion of sialylated O-glycans carrying NeuAc (24.2% instead of 18.9%) or NeuGc (31.2% instead of 15.4%) residues compared to GF rats.Thus, B. thetaiotaomicron impacts the composition of mucin O-glycans, with a decrease in sulfated and neutral oligosaccharides in favor of sialylated ones.Furthermore, glycosylation of mucins from Bt+Fp-rats resembled to those of GF rats.As previously observed for goblet cells, F. prausnitzii seemed to decrease the effect of B. thetaiotaomicron on mucus.Using a novel gnotobiotic model, which is the first described with F. prausnitzii, we showed how the balance between B. thetaiotaomicron and F. prausnitzii plays a key role in protecting epithelium via their respective effects on mucus.
Objectives: Pontocerebellar hypoplasia with spinal muscular atrophy, also known as PCH1, is a group of autosomal recessive disorders characterized by generalized muscle weakness and global developmental delay commonly resulting in early death. Gene defects had been discovered only in single patients until the recent identification of EXOSC3 mutations in several families with relatively mild course of PCH1. We aim to genetically stratify subjects in a large and well-defined cohort to define the clinical spectrum and genotype–phenotype correlation. Methods: We documented clinical, neuroimaging, and morphologic data of 37 subjects from 27 families with PCH1. EXOSC3 gene sequencing was performed in 27 unrelated index patients of mixed ethnicity. Results: Biallelic mutations in EXOSC3 were detected in 10 of 27 families (37%). The most common mutation among all ethnic groups was c.395A>C, p.D132A, responsible for 11 (55%) of the 20 mutated alleles and ancestral in origin. The mutation-positive subjects typically presented with normal pregnancy, normal birth measurements, and relative preservation of brainstem and cortical structures. Psychomotor retardation was profound in all patients but lifespan was variable, with 3 subjects surviving beyond the late teens. Abnormal oculomotor function was commonly observed in patients surviving beyond the first year. Major clinical features previously reported in PCH1, including intrauterine abnormalities, postnatal hypoventilation and feeding difficulties, joint contractures, and neonatal death, were rarely observed in mutation-positive infants but were typical among the mutation-negative subjects. Conclusion: EXOSC3 mutations account for 30%–40% of patients with PCH1 with variability in survival and clinical severity that is correlated with the genotype.