In susceptible patients, COVID-19 causes life-threatening disease driven by immune-mediated inflammatory lung injury. We have previously shown that multiple common host genetic variants are significantly associated with susceptibility to critical Covid-19, and in one case, we demonstrated that such variants can inform development of new, effective drug treatment. Here we report an association analysis of whole-genome sequences (WGS) from 11,423 cases from the GenOMICC study and 60,628 controls, together with meta-analyses with available genome-wide data. We identify a rare association signal at SLC50A1, primarily driven by a missense variant rs147850817 (1:155138217:G:T, Arg201Leu) that may interfere with transport function, and we identify four common association signals near ARF1, ZNF462, KLF13 and MVP genes. Finally, we build a WGS-derived polygenic risk score (PRS) for critical Covid-19, which offers only marginal improvement in risk estimation for the general population but may provide clinically-valuable discrimination for extreme susceptibility. ### Competing Interest Statement The authors have declared no competing interest. ### Clinical Protocols ### Funding Statement GenOMICC was funded by Sepsis Research (the Fiona Elizabeth Agnew Trust), the Intensive Care Society, a Wellcome Trust Senior Research Fellowship (J.K.Baillie, 223164/Z/21/Z), the Department of Health and Social Care (DHSC), Illumina, LifeArc, the Medical Research Council, UKRI, a BBSRC Institute Strategic Program Support Grant to the Roslin Institute (BBS/E/D/20002172, BBS/E/D/10002070 and BBS/E/D/30002275) and UKRI grants MC PC 20004, MC PC 19025, MC PC 1905, and MRNO2995X/1. ADB acknowledges funding from the Wellcome PhD training fellowship for clinicians (204979/Z/16/Z), the Edinburgh Clinical Academic Track (ECAT) programme. This research is supported in part by the Data and Connectivity National Core Study, led by Health Data Research UK in partnership with the Office for National Statistics and funded by UK Research and Innovation (grant ref MC PC 20029). This study owes a great deal to the National Institute for Healthcare Research Clinical Research Network (NIHR CRN) and the Chief Scientist's Office (Scotland), who facilitate recruitment into research studies in NHS hospitals, and to the global ISARIC and InFACT consortia. This work forms part of the translational research portfolio of the National Institute for Health and Care Research Barts Biomedical Research Centre. T.M. is supported by Cancer Research UK grant DRCRPG-May23/100002 to C. Siebold. Genomics England: This research was made possible through access to data in the National Genomic Research Library, which is managed by Genomics England Limited (a wholly owned company of the Department of Health and Social Care). The National Genomic Research Library (\url{https://www.genomicsengland.co.uk/research}) holds data provided by patients and collected by the NHS as part of their care and data collected as part of their participation in research. The National Genomic Research Library is funded by the National Institute for Health Research and NHS England. The Wellcome Trust, Cancer Research UK and the Medical Research Council have also funded research infrastructure. REACT: National Institute for Health and Care Research (NIHR) and UK Research and Innovation (UKRI) - REACT-Genomics England (REACT-GE) (MR/V030841/1) and REACT-Long COVID (REACT-LC) (COV-LT-0040). The REACT study was funded by the UK Department of Health and Social Care with supplemental funding from the Huo Family Foundation. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: GenOMICC was approved by the following research ethics committees: Scotland A Research Ethics Committee (15/SS/0110) and Coventry and Warwickshire Research Ethics Committee (England, Wales and Northern Ireland) (19/WM/0247). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All other data produced in the present study are available upon reasonable request to the authors
To identify host factors that affect Bovine Herpes Virus Type 1 (BoHV-1) infection we previously applied a genome wide CRISPR knockout screen targeting all bovine protein coding genes. By doing so we compiled a list of both pro-viral and anti-viral proteins involved in BoHV-1 replication. Here we provide further analysis of those that are potentially involved in viral entry into the host cell. We first generated single cell knockout clones deficient in some of the candidate genes for validation. We provide evidence that Polio Virus Receptor-related protein (PVRL2) serves as a receptor for BoHV-1, mediating more efficient entry than the previously identified Polio Virus Receptor (PVR). By knocking out two enzymes that catalyze HSPG chain elongation, HST2ST1 and GLCE, we further demonstrate the significance of HSPG in BoHV-1 entry. Another intriguing cluster of candidate genes, COG1, COG2 and COG4-7 encode six subunits of the Conserved Oligomeric Golgi (COG) complex. MDBK cells lacking COG6 produced fewer but bigger plaques compared to control cells, suggesting more efficient release of newly produced virions from these COG6 knockout cells, due to impaired HSPG biosynthesis. We further observed that viruses produced by the COG6 knockout cells consist of protein(s) with reduced N-glycosylation, potentially explaining their lower infectivity. To facilitate candidate validation, we also detailed a one-step multiplex CRISPR interference (CRISPRi) system, an orthogonal method to KO that enables quick and simultaneous deployment of three CRISPRs for efficient gene inactivation. Using CRISPR3i, we verified eight candidates that have been implicated in the synthesis of surface heparan sulfate proteoglycans (HSPGs). In summary, our experiments confirmed the two receptors PVR and PVRL2 for BoHV-1 entry into the host cell and other factors that affect this process, likely through the direct or indirect roles they play during HSPG synthesis and glycosylation of viral proteins.
Lumpy skin disease virus (LSDV) is a member of the capripoxvirus (CPPV) genus of the Poxviridae family. LSDV is a rapidly emerging, high-consequence pathogen of cattle, recently spreading from Africa and the Middle East into Europe and Asia. We have sequenced the whole genome of historical LSDV isolates from the Pirbright Institute virus archive, and field isolates from recent disease outbreaks in Sri Lanka, Mongolia, Nigeria and Ethiopia. These genome sequences were compared to published genomes and classified into different subgroups. Two subgroups contained vaccine or vaccine-like samples (“Neethling-like” clade 1.1 and “Kenya-like” subgroup, clade 1.2.2). One subgroup was associated with outbreaks of LSD in the Middle East/Europe (clade 1.2.1) and a previously unreported subgroup originated from cases of LSD in west and central Africa (clade 1.2.3). Isolates were also identified that contained a mix of genes from both wildtype and vaccine samples (vaccine-like recombinants, grouped in clade 2). Whole genome sequencing and analysis of LSDV strains isolated from different regions of Africa, Europe and Asia have provided new knowledge of the drivers of LSDV emergence, and will inform future disease control strategies.
Background Patients with cancer are at greater risk of dying from COVID-19 than many other patient groups. However, how this risk evolved during the pandemic remains unclear. We aimed to determine, on the basis of the UK national pandemic protocol, how factors influencing hospital mortality from COVID-19 could differentially affect patients undergoing cancer treatment. We also examined changes in hospital mortality and escalation of care in patients on cancer treatment during the first 2 years of the COVID-19 pandemic in the UK. Methods We conducted a prospective cohort study of patients aged older than 19 years and admitted to 306 health-care facilities in the UK with confirmed SARS-CoV-2 infection, who were enrolled in the International Severe Acute Respiratory and emerging Infections Consortium (ISARIC) WHO Clinical Characterisation Protocol (CCP) across the UK from April 23, 2020, to Feb 28, 2022; this analysis included all patients in the complete dataset when the study closed. The primary outcome was 30-day in-hospital mortality, comparing patients on cancer treatment and those without cancer. The study was approved by the South Central-Oxford C Research Ethics Committee in England (Ref: 13/SC/0149) and the Scotland A Research Ethics Committee (Ref 20/SS/0028), and is registered on the ISRCTN Registry (ISRCTN66726260). Findings 177 871 eligible adult patients either with no history of cancer (n=171 303) or on cancer treatment (n=6568) were enrolled; 93 205 (524%) were male, 84 418 (475%) were female, and in 248 (139%) sex or gender details were not specified or data were missing. Patients were followed up for a median of 13 (IQR 6-21) days. Of the 6568 patients receiving cancer treatment, 2080 (317%) died at 30 days, compared with 30 901 (180%) of 171 303 patients without cancer. Patients aged younger than 50 years on cancer treatment had the highest age-adjusted relative risk (hazard ratio [HR] 52 [95% CI 40-66], p<00001; vs 50-69 years 24 [22-26], p<00001; 70-79 years 18 [16-20], p<00001; and >80 years 15 [13-16], p<00001) but a lower absolute risk (51 [67%] of 763 patients <50 years died compared with 459 [302%] of 1522 patients aged >80 years). In -hospital mortality decreased for all patients during the pandemic but was higher for patients on cancer treatment than for those without cancer throughout the study period. Interpretation People with cancer have a higher risk of mortality from COVID-19 than those without cancer. Patients younger than 50 years with cancer treatment have the highest relative risk of death. Continued action is needed to mitigate the poor outcomes in patients with cancer, such as through optimising vaccination, long-acting passive immunisation, and early access to therapeutics. These findings underscore the importance of the ISARIC-WHO pandemic preparedness initiative.
The advances in gene editing bring unprecedented opportunities in high throughput functional genomics to animal research. Here we describe a genome wide CRISPR knockout library, btCRISPRko.v1, targeting all protein coding genes in the cattle genome. Using it, we conducted genome wide screens during Bovine Herpes Virus type 1 (BoHV-1) replication and compiled a list of pro-viral and anti-viral candidates. These candidates might influence multiple aspects of BoHV-1 biology such as viral entry, genome replication and transcription, viral protein trafficking and virion maturation in the cytoplasm. Some of the most intriguing examples are VPS51, VPS52 and VPS53 that code for subunits of two membrane tethering complexes, the endosome-associated recycling protein (EARP) complex and the Golgi-associated retrograde protein (GARP) complex. These complexes mediate endosomal recycling and retrograde trafficking to the trans Golgi Network (TGN). Simultaneous loss of both complexes in MDBKs resulted in greatly reduced production of infectious BoHV-1 virions. We also found that viruses released by these deficient cells severely lack VP8, the most abundant tegument protein of BoHV-1 that are crucial for its virulence. In combination with previous reports, our data suggest vital roles GARP and EARP play during viral protein packaging and capsid re-envelopment in the cytoplasm. It also contributes to evidence that both the TGN and the recycling endosomes are recruited in this process, mediated by these complexes. The btCRISPRko.v1 library generated here has been controlled for quality and shown to be effective in host gene discovery. We hope it will facilitate efforts in the study of other pathogens and various aspects of cell biology in cattle.
Critical illness in COVID-19 is an extreme and clinically homogeneous disease phenotype that we have previously shown 1 to be highly efficient for discovery of genetic associations 2 . Despite the advanced stage of illness at presentation, we have shown that host genetics in patients who are critically ill with COVID-19 can identify immunomodulatory therapies with strong beneficial effects in this group 3 . Here we analyse 24,202 cases of COVID-19 with critical illness comprising a combination of microarray genotype and whole-genome sequencing data from cases of critical illness in the international GenOMICC (11,440 cases) study, combined with other studies recruiting hospitalized patients with a strong focus on severe and critical disease: ISARIC4C (676 cases) and the SCOURGE consortium (5,934 cases). To put these results in the context of existing work, we conduct a meta-analysis of the new GenOMICC genome-wide association study (GWAS) results with previously published data. We find 49 genome-wide significant associations, of which 16 have not been reported previously. To investigate the therapeutic implications of these findings, we infer the structural consequences of protein-coding variants, and combine our GWAS results with gene expression data using a monocyte transcriptome-wide association study (TWAS) model, as well as gene and protein expression using Mendelian randomization. We identify potentially druggable targets in multiple systems, including inflammatory signalling ( JAK1 ), monocyte–macrophage activation and endothelial permeability ( PDE4A ), immunometabolism ( SLC2A5 and AK5 ), and host factors required for viral entry and replication ( TMPRSS2 and RAB2A ).
A common experimental output in biomedical science is a list of genes implicated in a given biological process or disease. The results of a group of studies answering the same, or similar, questions can be combined by meta-analysis to find a consensus or a more reliable answer. Ranking aggregation methods can be used to combine gene lists from various sources in meta-analyses. Evaluating a ranking aggregation method on a specific type of dataset before using it is required to support the reliability of the result since the property of a dataset can influence the performance of an algorithm. Evaluation of aggregation methods is usually based on a simulated database especially for the algorithms designed for gene lists because of the lack of a known truth for real data. However, simulated datasets tend to be too small compared to experimental data and neglect key features, including heterogeneity of quality, relevance and the inclusion of unranked lists. In this study, a group of existing methods and their variations which are suitable for meta-analysis of gene lists are compared using simulated and real data. Simulated data was used to explore the performance of the aggregation methods as a function of emulating the common scenarios of real genomics data, with various heterogeneity of quality, noise level, and a mix of unranked and ranked data using 20000 possible entities. In addition to the evaluation with simulated data, a comparison using real genomic data on the SARS-CoV-2 virus, cancer (NSCLC), and bacteria (macrophage apoptosis) was performed. We summarise our evaluation results in terms of a simple flowchart to select a ranking aggregation method for genomics data.
Pulmonary inflammation drives critical illness in Covid-19, [1][1];[2][2] creating a clinically homogeneous extreme phenotype, which we have previously shown to be highly efficient for discovery of genetic associations. [3][3];[4][4] Despite the advanced stage of illness, we have found that immunomodulatory therapies have strong beneficial effects in this group. [1][1];[5][5] Further genetic discoveries may identify additional therapeutic targets to modulate severe disease. [6][6] In this new data release from the GenOMICC (Genetics Of Mortality in Critical Care) study we include new microarray genotyping data from additional critically-ill cases in the UK and Brazil, together with cohorts of severe Covid-19 from the ISARIC4C [7][7] and SCOURGE [8][8] studies, and meta-analysis with previously-reported data. We find an additional 14 new genetic associations. Many are in potentially druggable targets, in inflammatory signalling (JAK1, PDE4A), monocyte-macrophage differentiation (CSF2), immunometabolism (SLC2A5, AK5), and host factors required for viral entry and replication (TMPRSS2, RAB2A). As with our previous work, these results provide tractable therapeutic targets for modulation of harmful host-mediated inflammation in Covid-19. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement GenOMICC was funded by Sepsis Research (the Fiona Elizabeth Agnew Trust), the Intensive Care Society, a Wellcome Trust Senior Research Fellowship (J.K.Baillie, 223164/Z/21/Z), the Department of Health and Social Care (DHSC), Illumina, LifeArc, the Medical Research Council, UKRI, a BBSRC Institute Program Support Grant to the Roslin Institute (BBS/E/D/20002172, BBS/E/D/10002070 and BBS/E/D/30002275) and UKRI grants MC\_PC\_20004, MC\_PC\_19025, MC\_PC\_1905, and MRNO2995X/1. This research is supported in part by the Data and Connectivity National Core Study, led by Health Data Research UK in partnership with the Office for National Statistics and funded by UK Research and Innovation (grant ref MC\_PC\_20029). We acknowledge NHS Digital, Public Health England and the Intensive Care National Audit and Research Centre who provided clinical data on the participants. This study owes a great deal to the National Institute for Healthcare Research Clinical Research Network (NIHR CRN) and the Chief Scientist Office (Scotland), who facilitate recruitment into research studies in NHS hospitals, and to the global ISARIC and InFACT consortia. GenOMICC genotype controls were obtained using UK Biobank Resource under project 788 funded by Roslin Institute Strategic Programme Grants from the BBSRC (BBS/E/D/10002070 and BBS/E/D/30002275) and Health Data Research UK (references HDR-9004 and HDR-9003) ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: We performed TWAS in the MetaXcan framework and the GTExv8 eQTL and sQTL MASHR-M models available for download in http://predictdb.org/. Publicly-available HGI data was downloaded from https://www.covid19hg.org/results/r6/. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable. Yes All data, including downloadable summary data and access applications for individual-level data, can be obtained through the GenOMICC gateway site https://genomicc.org/data. [1]: #ref-1 [2]: #ref-2 [3]: #ref-3 [4]: #ref-4 [5]: #ref-5 [6]: #ref-6 [7]: #ref-7 [8]: #ref-8
Critical COVID-19 is caused by immune-mediated inflammatory lung injury. Host genetic variation influences the development of illness requiring critical care1 or hospitalization2-4 after infection with SARS-CoV-2. The GenOMICC (Genetics of Mortality in Critical Care) study enables the comparison of genomes from individuals who are critically ill with those of population controls to find underlying disease mechanisms. Here we use whole-genome sequencing in 7,491 critically ill individuals compared with 48,400 controls to discover and replicate 23 independent variants that significantly predispose to critical COVID-19. We identify 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2). Using transcriptome-wide association and colocalization to infer the effect of gene expression on disease severity, we find evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease. Mendelian randomization provides evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets. Our results are broadly consistent with a multi-component model of COVID-19 pathophysiology, in which at least two distinct mechanisms can predispose to life-threatening disease: failure to control viral replication; or an enhanced tendency towards pulmonary inflammation and intravascular coagulation. We show that comparison between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
Lumpy skin disease virus (LSDV) is an emerging poxviral pathogen of cattle that is currently spreading throughout Asia. The disease situation is of high importance for farmers and policy makers in Asia. In October 2020, feral cattle in Hong Kong developed multifocal cutaneous nodules consistent with lumpy skin disease (LSD). Gross and histological pathology further supported the diagnosis and samples were sent to the OIE Reference Laboratory at The Pirbright Institute for confirmatory testing. LSDV was detected using quantitative polymerase chain reaction (qPCR) and additional molecular analyses. This is the first report of LSD in Hong Kong. Whole genome sequencing (WGS) of the strain LSDV/HongKong/2020 and phylogenetic analysis were carried out in order to identify connections to previous outbreaks of LSD, and better understand the drivers of LSDV emergence. Analysis of the 90 core poxvirus genes revealed LSDV/HongKong/2020 was a novel strain most closely related to the live-attenuated Neethling vaccine strains of LSDV and more distantly related to wildtype LSDV isolates from Africa, the Middle East and Europe. Analysis of the more variable regions located towards the termini of the poxvirus genome revealed genes in LSDV/HongKong/2020 with different patterns of grouping when compared to previously published wildtype and vaccine strains of LSDV. This work reveals that the LSD outbreak in Hong Kong in 2020 was caused by a different strain of LSDV than the LSD epidemic in the Middle East and Europe in 2015-2018. The use of WGS is highly recommended when investigating LSDV disease outbreaks.
AbstractCritical illness in COVID-19 is caused by inflammatory lung injury, mediated by the host immune system. We and others have shown that host genetic variation influences the development of illness requiring critical care1or hospitalisation2;3;4following SARS-Co-V2 infection. The GenOMICC (Genetics of Mortality in Critical Care) study recruits critically-ill cases and compares their genomes with population controls in order to find underlying disease mechanisms.Here, we use whole genome sequencing and statistical fine mapping in 7,491 critically-ill cases compared with 48,400 population controls to discover and replicate 22 independent variants that significantly predispose to life-threatening COVID-19. We identify 15 new independent associations with critical COVID-19, including variants within genes involved in interferon signalling (IL10RB, PLSCR1), leucocyte differentiation (BCL11A), and blood type antigen secretor status (FUT2). Using transcriptome-wide association and colocalisation to infer the effect of gene expression on disease severity, we find evidence implicating expression of multiple genes, including reduced expression of a membrane flippase (ATP11A), and increased mucin expression (MUC1), in critical disease.We show that comparison between critically-ill cases and population controls is highly efficient for genetic association analysis and enables detection of therapeutically-relevant mechanisms of disease. Therapeutic predictions arising from these findings require testing in clinical trials.
We produced a genome wide CRISPR knockout library, btCRISPRko.v1, targeting all protein coding genes in the cattle genome and used it to identify host genes important for Bovine Herpes Virus Type 1 (BHV-1) replication. By infecting library transduced MDBK cells with a GFP tagged BHV-1 virus and FACS sorting them based on their GFP intensity, we identified a list of pro-viral and anti-viral candidate host genes that might affect various aspects of the virus biology, such as cell entry, RNA transcription and viral protein trafficking. Among them were VPS51, VPS52 and VPS53 that encode for subunits of two membrane tethering complexes EARP and GARP. Simultaneous loss of both complexes in MDBKs resulted in a significant reduction in the production of infectious cell free BHV-1 virions, suggesting the vital roles they play during capsid re-envelopment with endocytosed membrane tubules prior to endosomal recycling mediated cellular egress. We also observed potential capsid retention and aggregation in the nuclei of these cells, indicating that they might also indirectly affect capsid egress from the nucleus. The btCRISPRko.v1 library generated here greatly expanded our capability in BHV-1 related host gene discovery; we hope it will facilitate efforts intended to study interactions between the host and other pathogens in cattle and also basic host cell biology.
In order to identify host factors that impact Bovine Herpes Virus Type 1 (BHV-1) infection we previously applied a genome wide CRISPR knockout screen with a library covering all bovine protein coding genes. We compiled a list of both pro-viral and anti-viral proteins involved in BHV-1 replication; here we provide further analysis of those that are potentially involved in viral entry into the host cell. These entry related factors include the cell surface proteins PVR and PVRL2, a group of enzymes directly or indirectly associated with the biosynthesis of Heparan Sulfate Proteoglycans (HSPG), and proteins that reside in the Golgi apparatus engaging in intra-Golgi trafficking. For the first time, we provide evidence that PVRL2 serves a receptor for BHV-1, mediating more efficient entry than the previously identified PVR. By knocking out two enzymes that catalyze HSPG chain elongation, HST2ST1 and GLCE, we demonstrated the significance of HSPG in BHV-1 entry. Another intriguing cluster of genes, COG1, COG2 and COG4-7 encodes for six subunits of the conserved oligomeric Golgi (COG) complex. MDBK cells lacking COG6 were less infectable by BHV-1 but release newly produced virions more efficiently as evidenced by fewer but bigger plaques compared to control cells, suggesting impaired HSPG biosynthesis. To facilitate candidate validation, we devised a one-step multiplex CRISPR interference (CRISPRi) system named CRISPR3i that enables quick and simultaneous deployment of three CRISPRs for efficient gene inactivation. Using CRISPR3i, we verified an additional 23 candidates, with many implicated in cellular entry.
The increasing body of literature describing the role of host factors in COVID-19 pathogenesis demonstrates the need to combine diverse, multi-omic data to evaluate and substantiate the most robust evidence and inform development of therapies. Here we present a dynamic ranking of host genes implicated in human betacoronavirus infection (SARS-CoV-2, SARS-CoV, MERS-CoV, seasonal coronaviruses). We conducted an extensive systematic review of experiments identifying potential host factors. Gene lists from diverse sources were integrated using Meta-Analysis by Information Content (MAIC). This previously described algorithm uses data-driven gene list weightings to produce a comprehensive ranked list of implicated host genes. From 32 datasets, the top ranked gene was PPIA, encoding cyclophilin A, a druggable target using cyclosporine. Other highly-ranked genes included proposed prognostic factors (CXCL10, CD4, CD3E) and investigational therapeutic targets (IL1A) for COVID-19. Gene rankings also inform the interpretation of COVID-19 GWAS results, implicating FYCO1 over other nearby genes in a disease-associated locus on chromosome 3. Researchers can search and review the gene rankings and the contribution of different experimental methods to gene rank at https://baillielab.net/maic/covid19 . As new data are published we will regularly update the list of genes as a resource to inform and prioritise future studies.
The major histocompatibility complex (MHC) region contains many genes that are key regulators of both innate and adaptive immunity including the polymorphic MHCI and MHCII genes. Consequently, the characterisation of the repertoire of MHC genes is critical to understanding the variation that determines the nature of immune responses. Our current knowledge of the bovine MHCI repertoire is limited with only the Holstein-Friesian breed having been studied in any depth. Traditional methods of MHCI genotyping are of low resolution and laborious and this has been a major impediment to a more comprehensive analysis of the MHCI repertoire of other cattle breeds. Next-generation sequencing (NGS) technologies have been used to enable high throughput and much higher resolution MHCI typing in a number of species. In this study we have developed a MiSeq platform approach and requisite bioinformatics pipeline to facilitate typing of bovine MHCI repertoires. The method was validated initially on a cohort of Holstein-Friesian animals and then demonstrated to enable characterisation of MHCI repertoires in African cattle breeds, for which there was limited or no available data. During the course of these studies we identified >140 novel classical MHCI genes and defined 62 novel MHCI haplotypes, dramatically expanding the known bovine MHCI repertoire.
Background A major step towards the success of chickens as a domesticated species was the separation between maternal care and reproduction. Artificial incubation replaced the natural maternal behaviour of incubation and, thus, in certain breeds, it became possible to breed chickens with persistent egg production and no incubation behaviour; a typical example is the White Leghorn strain. Conversely, some strains, such as the Silkie breed, are prized for their maternal behaviour and their willingness to incubate eggs. This is often colloquially known as broodiness. Results Using an F 2 linkage mapping approach and a cross between White Leghorn and Silkie chicken breeds, we have mapped, for the first time, genetic loci that affect maternal behaviour on chromosomes 1, 5, 8, 13, 18 and 19 and linkage group E22C19W28. Paradoxically, heterozygous and White Leghorn homozygous genotypes were associated with an increased incidence of incubation behaviour, which exceeded that of the Silkie homozygotes for most loci. In such cases, it is likely that the loci involved are associated with increased egg production. Increased egg production increases the probability of incubation behaviour occurring because egg laying must precede incubation. For the loci on chromosomes 8 and 1, alleles from the Silkie breed promote incubation behaviour and influence maternal behaviour (these explain 12 and 26 % of the phenotypic difference between the two founder breeds, respectively). Conclusions The over-dominant locus on chromosome 5 coincides with the strongest selective sweep reported in chickens and together with the loci on chromosomes 1 and 8, they include genes of the thyrotrophic axis. This suggests that thyroid hormones may play a critical role in the loss of incubation behaviour and the improved egg laying behaviour of the White Leghorn breed. Our findings support the view that loss of maternal incubation behaviour in the White Leghorn breed is the result of selection for fertility and egg laying persistency and against maternal incubation behaviour.
Computational tools are quickly becoming the main bottleneck to analyze large-scale genomic and genetic data. This big-data problem, affecting a wide range of fields, is becoming more acute with the fast increase of data available. To address it, we developed DISSECT, a new, easy to use, and freely available software able to exploit the parallel computer architectures of supercomputers to perform a wide range of genomic and epidemiologic analyses which currently can only be carried out on reduced sample sizes or in restricted conditions. We showcased our new tool by addressing the challenge of predicting phenotypes from genotype data in human populations using Mixed Linear Model analysis. We analyzed simulated traits from half a million individuals genotyped for 590,004 SNPs using the combined computational power of 8,400 processor cores. We found that prediction accuracies in excess of 80% of the theoretical maximum could be achieved with large numbers of training individuals.
Large-scale genetic and genomic data are increasingly available and the major bottleneck in their analysis is a lack of sufficiently scalable computational tools. To address this problem in the context of complex traits analysis, we present DISSECT. DISSECT is a new and freely available software that is able to exploit the distributed-memory parallel computational architectures of compute clusters, to perform a wide range of genomic and epidemiologic analyses, which currently can only be carried out on reduced sample sizes or under restricted conditions. We demonstrate the usefulness of our new tool by addressing the challenge of predicting phenotypes from genotype data in human populations using mixed-linear model analysis. We analyse simulated traits from 470,000 individuals genotyped for 590,004 SNPs in ∼4 h using the combined computational power of 8,400 processor cores. We find that prediction accuracies in excess of 80% of the theoretical maximum could be achieved with large sample sizes.