The novel HLA-DQB1*06:469 allele differs from HLA-DQB1*06:01:01:01 by one nucleotide substitution in codon 187 in exon 3.
The novel HLA-DPB1*03:01:27 allele differs from HLA-DPB1*03:01:01 by one nucleotide substitution in codon 199 exon 4.
e22524 Background: The increase in large panel testing has facilitated the detection of numerous events, including those in genes associated with inherited cancer. Further, large panel sequencing may uncover variants in these genes in non-inherited cancers. Hence the frequency of pathogenic/likely pathogenic (P/LP) events in some common germline genes was studied across cancers in the Indian population. Methods: A statistical model was used to derive a method to label variants as likely germline. Variants in hereditary breast and ovarian cancer (HBOC) and mismatch repair (MMR) genes from 1553 cases sequenced at our referral laboratory on the TruSight Oncology 500 panel (TSO500) were analyzed for their variant allele frequency (VAF) distribution. A 3 component Gaussian mixture model (intended to separate somatic, germline heterozygous, homozygous variants) was fit to this VAF distribution and a cutoff was determined such that the area under the curve of the somatic component above the cutoff was below 0.01. This gave a threshold value of 40% VAF for determining likely germline variants (LGVs). Data from 1713 TSO500 cases was then analyzed to identify LGVs. Results: The dataset of 1713 cases comprised 20.67% NSCLC, 12.61% breast, 12.49% ovarian and 10.5% colorectal cancer and < 10% of other cancer types including 2.39% pancreatic and 2.04% prostate cancer. 23.8% of ovarian, 17.14% of prostate, 12.5% of breast and 12.2% of pancreatic cancer cases had variants in BRCA1/BRCA2. Among these, 18.69% of ovarian, 17.14% of prostate, 8.33% of breast and 7.32% of pancreatic cancers had LGVs. LGVs in BRCA1/2 were also found in 3.33% of colorectal and 1.13% of NSCLC. The BRCA1/BRCA2 LGV ratio was found to be 3.0 in ovarian, 1.0 in breast and 2.0 in pancreatic but 0.2 in prostate and colorectal and 0.33 in NSCLC cancer. Adding other HBOC genes (ATM, BRCA1/2, BRIP1, CHEK2, PALB2, RAD51C, RAD51D) increased the LGV frequency to 22.43% of ovarian, 22.86% of prostate, 11.11% of breast, 9.76% of pancreatic, 4.44% of colorectal and 2.26% of NSCLC. MMR gene LGVs were seen in 1.23% of all cancers mainly in 10.71% of endometrial and 4.44% of colorectal cancers. Unlike the HBOC genes, MMR gene variants were not observed across multiple cancers and did not show a predominantly germline distribution (somatic/LGV 1:1). In contrast, the background P/LP frequency calculated from gnomAD v4 South Asian data (excluding those with conflicting labels) was 0.74% for BRCA1/2, 1.82% for the extended HBOC list and 0.5% for the MMR genes. Conclusions: Somatic testing was used to identify LGVs and determine the frequencies of HBOC and MMR genes in cancers in the Indian population. Somatic testing also identified LGVs in multiple tissues at frequencies above the background P/LP frequency for this population, suggesting a possible role for these genes in other cancers. This information underscores the need for further study and may be useful to define a clinically relevant subset of patients.
AbstractNext-generation sequencing (NGS) technologies have transformed biomarker discovery, enabling the detection of disease-associated markers at the earliest stages of illness. In this study, we introduce a blood-based, non-invasive test for multi-cancer detection using cell-free DNA (cfDNA) methylation sequencing. The test employs a novel methylation scoring system derived from sequencing data and integrates machine learning to analyze a retrospective cohort of newly diagnosed cancer cases and controls recruited from multiple centers across India. To enhance robustness, the study includes a substantial proportion of controls with habitual tobacco and alcohol use, ensuring the test’s resilience against confounding factors. The test’s accuracy was further validated through synthetic data augmentation, demonstrating reliability under conditions of random signal perturbation. At an approximate specificity of 97%, the assay achieves sensitivities of 79.3% for Stage I, 78.4% for Stage II, 78.4% for Stage III, and 86.8% for Stage IV cancers in an independent validation cohort. Additionally, the test demonstrates Top 2 Tissue of Origin (TOO) accuracies of 78.3% for Stage I, 79.3% for Stage II, 82.8% for Stage III, and 69.7% for Stage IV cancers. This blood-based test holds considerable promise for early cancer detection, offering a precise test for cancer screening.
e15089 Background: Diagnosis of CRC is biased towards later stages in India (3.8% Stage I, 16.7% Stage II, 50.7% Stage III, 28.8% Stage IV), and five-year survival at < 40% is one of the lowest in the world. A blood-based non-invasive screening test for CRC using cell-free DNA (cfDNA) methylation sequencing is developed here from the blood of 212 controls and 67 treatment naive CRC patients from 21 sites across India and processed in Strand’s reference lab in Bangalore. Methods: Steps involved cfDNA extraction, NEB Enzymatic Methyl-Seq library preparation, Twist Human Methylome hybridisation capture, 2x150bp sequencing on NovaSeq 6000/X. Methylation + fragmentomic features were calculated for each target region. Samples were split randomly into a leave-in set of 170 controls + 53 cancers (I: 8, II 16, III: 23, IV: 6) and a leave-out set of 42 controls + 14 cancers, (I: 5, II: 4, III: 2, IV: 3) with 20 rounds of 4-fold cross-validation done on the leave-in set (random splits). Feature selection per fold was performed using the KS test without access to ¼ of the leave-in set and to the entire leave-out set. Gradient boosted trees with monotonic constraints reflecting the expected association of the scores with cancer were used to build “explainable” models. Test robustness was assessed using differentially methylated regions (DMRs) from other studies, and by assessing predictability using sample metadata alone. Results: At a 91% specificity level, the ensemble model had a median sensitivity of 62.5% for Stage I (95%CI 38%-86%), 87% for Stage II (95% CI 75%-95%), 87% for Stage III (95% CI 75%-95%) and 83.4% for Stage IV (95% CI 84%-100%) in cross-validation, and 60% for Stage I (95%CI 60%-80%), 100% for Stage II (95% CI 80%-100%), 100% for Stage III (95% CI 100%-100%) and 100% for Stage IV (95% CI 100%-100%) on the leave-out set. At a ~98% specificity level, the model had a median sensitivity of 37.5% for Stage I (95%CI 37%-50%), 69% for Stage II (95% CI 56%-75%), 69% for Stage III (95% CI 60%-84%) and 67% for Stage IV (95% CI 50%-84%) in cross-validation, and 40% for Stage I (95%CI 20%-60%), 75% for Stage II (95% CI 75%-100%), 100% for Stage III (95% CI 100%-100%) and 100% for Stage IV (95% CI 66%-100%) with a slight decrease in specificity to 95.2% on the leave-out set. Using DMRs derived from TCGA data and other publications, yielded a comparable (to our models) cross-validation area under the curve (AUC: 0.93-0.95). Cross-validation performance using only the sample metadata in table 1 and without access to the data was significantly poorer (AUC 0.75-0.81). Conclusions: cfDNA-based methylation profiles are consistent across studies and ethnicities, leading to robust and “explainable” CRC screening predictions. [Table: see text]
Background Dementia is a complex disorder. Genetic factors, both rare and common variants, may contribute to risk. Early onset, and family history, may point towards those harbouring rare variants of large effect. In addition to well-established risk variants, other Variants of Uncertain Significance (VUS) are commonly encountered during clinical exome sequencing. There should be a pipeline efficient to profile such variants and their interpretation across diverse populations and more complex phenotypes Profiling rare genetic risk variants for dementia in diverse populations may lead to a better understanding of the pathogenesis and identification of new therapeutic targets. Methods The DNA of patients with an ICD-10 diagnosis of dementia (N 30; F13; age at assessment 57±12 years; age at onset of dementia 54±13 years) identified from the Geriatric Clinic of National Institute of Mental Health and Neurosciences, Bengaluru, India, was archived in the Molecular genetics laboratory. Clinical exome sequencing was carried out for selected samples with strong family history or early onset, to identify possible causative variants of these archived samples. Standard quality control procedures were employed for exome data analysis. Variants were prioritised with a custom pipeline designed to identify variants relevant to neuropsychiatric disorders. To evaluate the possible impact of identified VUS, we applied the recently developed alpha-missense tool. Results Of the thirty patients, nineteen (63%) were diagnosed with Fronto temporal dementia (FTD), three (10%) with Alzheimer's Disease (AD), five (16.6%) with mixed dementia, (3%) one with Parkinson's disease (PD), one (3%) patient had dementia with PD. In two patients (one each with AD and FTD), had comorbid Bipolar disorder.In twenty patients (66%) we were able to detect rare exonic variants. Eight individuals were carriers of APO E4 alleles, including three homozygotes and six carried an additional VUS of potential interest.Pathogenic variants were also identified in four individuals which are known candidate genes of dementia. Among them, two unrelated patients had a known pathogenic variant in MAPT gene (chr 17:44087755C > T). The other two individuals carried mutations in GRN and PRNP genes. Additionally, eight individuals harboured other VUS in CCNF, TREM2, SORL1, ATP6AP2, EIF4GI, MATR3 genes which have been implicated in various dementias including AD, FTD/ Lewy body dementia. Upon further analysis with the alpha mis sense model, twelve out of twenty-one VUS could be re-classified into likely benign, nine had moderate to severe biological consequences. Discussion Significant exonic variants were found in twenty (66.6%) of our patients. APOE4 is an established high-risk allele for AD, and in its homozygous state, it is a major genetic form of AD. Mutation in MAPT gene causes abnormal tau protein aggregation, significantly implicated in FTD and AD. GRN gene mutations are causative of FTD, as haploinsufficiency of this gene typically leads to FTD. Alpha missense model enabled us to predict whether the variants were benign or pathogenic. To confirm this, genotype-phenotype correlations or investigation of biological impact would be necessary. Studying the biology and epidemiology of rare variants, across populations, may be quite useful.
The benefits of large-scale genetic studies for healthcare of the populations studied are well documented, but these genetic studies have traditionally ignored people from some parts of the world, such as South Asia. Here we describe whole genome sequence (WGS) data from 4806 individuals recruited from the healthcare delivery systems of Pakistan, India and Bangladesh, combined with WGS from 927 individuals from isolated South Asian populations. We characterize population structure in South Asia and describe a genotyping array (SARGAM) and imputation reference panel that are optimized for South Asian genomes. We find evidence for high rates of reproductive isolation, endogamy and consanguinity that vary across the subcontinent and that lead to levels of rare homozygotes that reach 100 times that seen in outbred populations. Founder effects increase the power to associate functional variants with disease processes and make South Asia a uniquely powerful place for population-scale genetic studies.
COVID-19 is a respiratory illness caused by a novel coronavirus called SARS-CoV-2. The viral spike (S) protein engages the human angiotensin-converting enzyme 2 (ACE2) receptor to invade host cells with ~10-15-fold higher affinity compared to SARS-CoV S-protein, making it highly infectious. Here, we assessed if ACE2 polymorphisms can alter host susceptibility to SARS-CoV-2 by affecting this interaction. We analyzed over 290,000 samples representing >400 population groups from public genomic datasets and identified multiple ACE2 protein-altering variants. Using reported structural data, we identified natural ACE2 variants that could potentially affect virus-host interaction and thereby alter host susceptibility. These include variants S19P, I21V, E23K, K26R, T27A, N64K, T92I, Q102P and H378R that were predicted to increase susceptibility, while variants K31R, N33I, H34R, E35K, E37K, D38V, Y50F, N51S, M62V, K68E, F72V, Y83H, G326E, G352V, D355N, Q388L and D509Y were predicted to be protective variants that show decreased binding to S-protein. Using biochemical assays, we confirmed that K31R and E37K had decreased affinity, and K26R and T92I variants showed increased affinity for S-protein when compared to wildtype ACE2. Consistent with this, soluble ACE2 K26R and T92I were more effective in blocking entry of S-protein pseudotyped virus suggesting that ACE2 variants can modulate susceptibility to SARS-CoV-2.
Abstract Background Duchenne muscular dystrophy (DMD) is an X‐linked recessive neuromuscular disorder characterised by progressive irreversible muscle weakness, primarily of the skeletal and the cardiac muscles. DMD is characterised by mutations in the dystrophin gene, resulting in the absence or sparse quantities of dystrophin protein. A precise and timely molecular detection of DMD mutations encourages interventions such as carrier genetic counselling and in undertaking therapeutic measures for the DMD patients. Results In this study, we developed a 2.1 Mb custom DMD gene panel that spans the entire DMD gene, including the exons and introns. The panel also includes the probes against 80 additional genes known to be mutated in other muscular dystrophies. This custom DMD gene panel was used to identify single nucleotide variants (SNVs) and large deletions with precise breakpoints in 77 samples that included 24 DMD patients and their matrilineage across four generations. We used this panel to evaluate the inheritance pattern of DMD mutations in maternal subjects representing 24 DMD patients. Conclusion Here we report our observations on the inheritance pattern of DMD gene mutations in matrilineage samples across four generations. Additionally, our data suggest that the DMD gene panel designed by us can be routinely used as a single genetic test to identify all DMD gene variants in DMD patients and the carrier mothers.
Formulating strategies for species conservation requires knowledge of evolutionary and genetic history. Tigers are among the most charismatic of endangered species and garner significant conservation attention. However, the evolutionary history and genomic variation of tigers remain poorly known. With 70% of the worlds wild tigers living in India, such knowledge is critical for tiger conservation. We re-sequenced 65 individual tiger genomes across their extant geographic range, representing most extant subspecies with a specific focus on tigers from India. As suggested by earlier studies, we found strong genetic differentiation between the putative tiger subspecies. Despite high total genomic diversity in India, individual tigers host longer runs of homozygosity, potentially suggesting recent inbreeding, possibly because of small and fragmented protected areas. Surprisingly, demographic models suggest recent divergence (within the last 10,000 years) between populations, and strong population bottlenecks. Amur tiger genomes revealed the strongest signals of selection mainly related to metabolic adaptation to cold, while Sumatran tigers show evidence of evolving under weak selection for genes involved in body size regulation. Depending on conservation objectives, our results support the isolation of Amur and Sumatran tigers, while geneflow between Malayan and South Asian tigers may be considered. Further, the impacts of ongoing connectivity loss on the health and persistence of tigers in India should be closely monitored.
Population-scale genetic studies can identify drug targets and allow disease risk to be predicted with resulting benefit for management of individual health risks and system-wide allocation of health care delivery resources. Although population-scale projects are underway in many parts of the world, genetic variation between population groups means that additional projects are warranted. South Asia has a population whose genetics is the least characterized of any of the world’s major populations. Here we describe GenomeAsia studies that characterize population structure in South Asia and that create tools for economical and accurate genotyping at population-scale. Prior work on population structure characterized isolated population groups, the relevance of which to large-scale studies of disease genetics is unclear. For our studies we used whole genome sequence information from 4,807 individuals recruited in the health care delivery systems of Pakistan, India and Bangladesh to ensure relevance to population-scale studies of disease genetics. We combined this with WGS data from 927 individuals from isolated South Asian population groups, and developed a custom SNP array (called SARGAM) that is optimized for future human genetic studies in South Asia. We find evidence for high rates of reproductive isolation, endogamy and consanguinity that vary across the subcontinent and that lead to levels of homozygosity that approach 100 times that seen in outbred populations. We describe founder effects that increase the power to associate functional variants with disease processes and that make South Asia a uniquely powerful place for population-scale genetic studies.
BACKGROUND:First Zika virus (ZIKV) positive case from North India was detected on routine surveillance of Dengue-Like Illness in an 85-year old female. Objective of the study was to conduct an investigation for epidemiological, clinical and genomic analysis of first ZIKV outbreak in Rajasthan, North India and enhance routine ZIKV surveillance.METHOD:Outbreak investigation was performed in 3 Km radius of the index case among patient contacts, febrile cases, and pregnant women. Routine surveillance was enhanced to include samples from various districts of Rajasthan. Presence of ZIKV in serum and urine samples was detected by real time PCR test and CDC trioplex kit. Few ZIKV positive samples were sequenced using the next-generation sequencing method for genomic analysis.RESULT:On outbreak investigation 153/2043 (7.48%) cases were found positive: 1/153 (0.65%) among contacts, 90/153 (58.8%) in fever cases, 62/153(40.5%) in pregnant females. In routine surveillance, 6/4722 (0.12%) serum samples were ZIKV positive.Majority of patients had mild signs and symptoms, no case of microcephaly and Guillain- Barre Syndrome was seen, 25 (40.3%) pregnant females delivered healthy babies, four (6.4%) reported abortion and three (4.8%) had intrauterine death, one (1.6%) child had colorectal malformation and died after few days of birth. ZIKV was found to belong to Asian lineage, mutation related to enhanced neuro-virulence and transmission in animal models was not found.CONCLUSION:ZIKV was endogenous to India belonging to Asian Lineage. Disease profile of the ZIKV was asymptomatic to mild. No major anomaly was observed in infants born to ZIKV positive mothers; however, long term follow up of these children is required. There is need to scale up surveillance in the virology lab network of India for early detection and control.SUMMARY LINE:Zika virus infection was endogenous due to Asian Lineage with mild disease, no case of microcephaly or Guillain- Barre Syndrome was seen but children need to be followed for anomalies and surveillance of ZIKV needs to be enhanced in the country.
The PRKAG2 syndrome is a rare autosomal dominant phenocopy of sarcomeric hypertrophic cardiomyopathy (HCM), characterized by ventricular pre-excitation, progressive conduction system disease and left ventricular hypertrophy. This study describes the phenotype, genotype and clinical outcomes of a South-Asian PRKAG2 cardiomyopathy cohort over a 7-year period. Clinical, electrocardiographic, echocardiographic, and cardiac MRI data from 22 individuals with PRKAG2 variants (68% men; mean age 39.5 ± 18.1 years), identified at our HCM centre were studied prospectively. At initial evaluation, all of the patients were in NYHA functional class I or II. The maximum left ventricular wall thickness was 22.9 ± 8.7 mm and left ventricular ejection fraction was 53.4 ± 6.6%. Left ventricular hypertrophy was present in 19 individuals (86%) at baseline. 17 patients had an WPW pattern (77%). After a mean follow-up period of 7 years, 2 patients had undergone accessory pathway ablation, 8 patients (36%) underwent permanent pacemaker implantation (atrio-ventricular blocks—5; sinus node disease—2), 3 patients developed atrial fibrillation, 11 patients (50%) developed progressive worsening in NYHA functional class, and 6 patients (27%) experienced sudden cardiac death or equivalent. PRKAG2 cardiomyopathy must be considered in patients with HCM and progressive conduction system disease.
Gallbladder cancer (GBC) is an aggressive gastrointestinal malignancy with no approved targeted therapy. Here, we analyze exomes ( n = 160), transcriptomes ( n = 115), and low pass whole genomes ( n = 146) from 167 gallbladder cancers (GBCs) from patients in Korea, India and Chile. In addition, we also sequence samples from 39 GBC high-risk patients and detect evidence of early cancer-related genomic lesions. Among the several significantly mutated genes not previously linked to GBC are ETS domain genes ELF3 and EHF , CTNNB1 , APC , NSD1 , KAT8 , STK11 and NFE2L2 . A majority of ELF3 alterations are frame-shift mutations that result in several cancer-specific neoantigens that activate T-cells indicating that they are cancer vaccine candidates. In addition, we identify recurrent alterations in KEAP1/NFE2L2 and WNT pathway in GBC. Taken together, these define multiple targetable therapeutic interventions opportunities for GBC treatment and management.
Deregulated HER2 is a target of many approved cancer drugs. We analyzed 111,176 patient tumors and identified recurrent mutations in HER2 transmembrane domain (TMD) and juxtamembrane domain (JMD) that include G660D, R678Q, E693K, and Q709L. Using a saturation mutagenesis screen and testing of patient-derived mutations we found several activating TMD and JMD mutations. Structural modeling and analysis showed that the TMD/JMD mutations function by improving the active dimer interface or stabilizing an activating conformation. Further, we found that HER2 G660D employed asymmetric kinase dimerization for activation and signaling. Importantly, anti-HER2 antibodies and small-molecule kinase inhibitors blocked the activity of TMD/JMD mutants. Consistent with this, a G660D germline mutant lung cancer patient showed remarkable clinical response to HER2 blockade.
Maturity-onset diabetes of the young (MODY) is an early-onset, autosomal dominant form of non-insulin dependent diabetes. Genetic diagnosis of MODY can transform patient management. Earlier data on the genetic predisposition to MODY have come primarily from familial studies in populations of European origin.
Bestinopathies are a spectrum of retinal disorders associated with mutations in BEST1 including autosomal recessive bestrophinopathy (ARB) and autosomal dominant Best vitelliform macular dystrophy (BVMD). We applied whole-exome sequencing on four unrelated Indian families comprising eight affected and twelve unaffected individuals. We identified five mutations in BEST1, including p.Tyr131Cys in family A, p.Arg150Pro in family B, p.Arg47His and p.Val216Ile in family C and p.Thr91Ile in family D. Among these, p.Tyr131Cys, p.Arg150Pro and p.Val216Ile have not been previously reported. Further, the inheritance pattern of BEST1 mutations in the families confirmed the diagnosis of ARB in probands in families A, B and C, while the inheritance of heterozygous BEST1 mutation in family D (p.Thr91Ile) was suggestive of BVMD. Interestingly, the ARB families A and B carry homozygous mutations while family C was a compound heterozygote with a mutation in an alternate BEST1 transcript isoform, highlighting a role for alternate BEST1 transcripts in bestrophinopathy. In the BVMD family D, the heterozygous BEST1 mutation found in the proband was also found in the asymptomatic parent, suggesting an incomplete penetrance and/or the presence of additional genetic modifiers. Our report expands the list of pathogenic BEST1 genotypes and the associated clinical diagnosis.