Reconstructing gene regulatory networks from large-scale heterogeneous data is a key challenge in biology. In multi-omics data analysis, networks based on pairwise statistical association measures remain popular, as they are easy to build and understand. In the presence of mixed-type (discrete and continuous) data, however, the choice of good association measures remains an important issue. It is proposed a novel approach based on the Gaussian copula, the parameters of which represent the links of the network. Novel properties of the model are obtained to guide the interpretation of the network. To estimate the copula parameters, a semiparametric pairwise likelihood for mixed data was calculated. An extensive simulation study showed that the proposed estimation procedure was able to accurately estimate the copula correlation matrix. The proposed methodology was also applied to a real ICGC dataset on breast cancer, and is implemented in a freely available R package heterocop.
BACKGROUND:Inferring partial correlation networks is essential in systems biology to uncover direct interactions between biological entities. Traditional Gaussian graphical models rely on the assumption of normally distributed data; this assumption is not satisfied when dealing with multi-omics datasets comprising heterogeneous data types such as continuous and discrete variables. RESULTS:We propose a novel likelihood-based approach for network inference using a Gaussian copula model with semiparametric pairwise-likelihood estimation of the latent correlation matrix. The inferred correlation structure is then inverted and regularized via the graphical lasso to recover latent partial correlations. Compared to a moment-based approach employing bridge functions, our method demonstrates significantly improved computational efficiency and estimation accuracy, particularly for discrete data with many categories and/or large values, such as count data. This result is important for biological applications, especially for the integration of RNA-seq count data. An application to a breast cancer data set from the International Cancer Genome Consortium (ICGC) successfully identified biologically relevant interactions. CONCLUSIONS:The proposed approach, based on the Gaussian copula and likelihood-based estimation, provides a novel, effective and computationally efficient mathematical framework for integrative multi-omics data analysis and network inference.
Reconstructing gene regulatory networks from large-scale heterogeneous data is a key challenge in biology. In multi-omics data analysis, networks based on pairwise statistical association measures remain popular, as they are easy to build and understand. In the presence of mixed-type (discrete and continuous) data, however, the choice of good association measures remains an important issue. We propose here a novel approach based on the Gaussian copula, the parameters of which represent the links of the network. Novel properties of the model are obtained to guide the interpretation of the network. To estimate the copula parameters, we calculated a semiparametric pairwise likelihood for mixed data. In an extensive simulation study, we showed that the proposed estimation procedure was able to accurately estimate the copula correlation matrix. The proposed methodology was also applied to a real ICGC dataset on breast cancer, and is implemented in a freely available R package heterocop.
The kinesin light chain 3 protein (KLC3) is the only member of the kinesin light chain protein family that was identified in post-meiotic mouse male germ cells. It plays a role in the formation of the sperm midpiece through its association with both spermatid mitochondria and outer dense fibers (ODF). Previous studies showed a significant correlation between its expression level and sperm motility and quantitative semen parameters in humans, while the overexpression of a KLC3-mutant protein unable to bind ODF also affected the same traits in mice. To further assess the role of KLC3 in fertility, we used CRISPR/Cas9 genome editing in mice and investigated the phenotypes induced by the invalidation of the gene or of a functional domain of the protein. Both approaches gave similar results, i.e. no detectable change in male or female fertility. Testis histology, litter size and sperm count were not altered. Apart from the line-dependent alterations of Klc3 mRNA levels, testicular transcriptome analysis did not reveal any other changes in the genes tested. Western analysis supported the absence of KLC3 in the gonads of males homozygous for the inactivating mutation and a strong decrease in expression in males homozygous for the allele lacking one out of the five tetratricopeptide repeats. Overall, these observations raise questions about the supposedly critical role of this kinesin in reproduction, at least in mice where its gene mutation or inactivation did not translate into fertility impairment.
Bull fertility is an important economic trait, and the use of subfertile semen for artificial insemination decreases the global efficiency of the breeding sector. Although the analysis of semen functional parameters can help to identify infertile bulls, no tools are currently available to enable precise predictions and prevent the commercialization of subfertile semen. Because male fertility is a multifactorial phenotype that is dependent on genetic, epigenetic, physiological and environmental factors, we hypothesized that an integrative analysis might help to refine our knowledge and understanding of bull fertility. We combined -omics data (genotypes, sperm DNA methylation at CpGs and sperm small non-coding RNAs) and semen parameters measured on a large cohort of 98 Montbéliarde bulls with contrasting fertility levels. Multiple Factor Analysis was conducted to study the links between the datasets and fertility. Four methodologies were then considered to identify the features linked to bull fertility variation: Logistic Lasso, Random Forest, Gradient Boosting and Neural Networks. Finally, the features selected by these methods were annotated in terms of genes, to conduct functional enrichment analyses. The less relevant features in -omics data were filtered out, and MFA was run on the remaining 12,006 features, including the 11 semen parameters and a balanced proportion of each type of-omics data. The results showed that unlike the semen parameters studied the-omics datasets were related to fertility. Biomarkers related to bull fertility were selected using the four methodologies mentioned above. The most contributory CpGs, SNPs and miRNAs targeted genes were all found to be involved in development. Interestingly, fragments derived from ribosomal RNAs were overrepresented among the selected features, suggesting roles in male fertility. These markers could be used in the future to identify subfertile bulls in order to increase the global efficiency of the breeding sector.
Lactation is an essential process for mammals. In sheep, the R96C mutation in suppressor of cytokine signaling 2 (SOCS2) protein is associated with greater milk production and increased mastitis sensitivity. To shed light on the involvement of R96C mutation in mammary gland development and lactation, we developed a mouse model carrying this mutation (SOCS2KI/KI). Mammary glands from virgin adult SOCS2KI/KI mice presented a branching defect and less epithelial tissue, which were not compensated for in later stages of mammary development. Mammary epithelial cell (MEC) subpopulations were modified, with mutated mice having three times as many basal cells, accompanied by a decrease in luminal cells. The SOCS2KI/KI mammary gland remained functional; however, MECs contained more lipid droplets versus fat globules, and milk lipid composition was modified. Moreover, the gene expression dynamic from virgin to pregnancy state resulted in the identification of about 3000 differentially expressed genes specific to SOCS2KI/KI or control mice. Our results show that SOCS2 is important for mammary gland development and milk production. In the long term, this finding raises the possibility of ensuring adequate milk production without compromising animal health and welfare.
Identifying differentially methylated cytosine-guanine dinucleotide (CpG) sites between benign and tumour samples can assist in understanding disease. However, differential analysis of bounded DNA methylation data often requires data transformation, reducing biological interpretability. To address this, a family of beta mixture models (BMMs) is proposed that (i) objectively infers methylation state thresholds and (ii) identifies differentially methylated CpG sites (DMCs) given untransformed, beta-valued methylation data. The BMMs achieve this through model-based clustering of CpG sites and by employing parameter constraints, facilitating application to different study settings. Inference proceeds via an expectation-maximisation algorithm, with an approximate maximization step providing tractability and computational feasibility. Performance of the BMMs is assessed through thorough simulation studies, and the BMMs are used for differential analyses of DNA methylation data from a prostate cancer study. Intuitive and biologically interpretable methylation state thresholds are inferred and DMCs are identified, including those related to genes such as GSTP1, RASSF1 and RARB, known for their role in prostate cancer development. Gene ontology analysis of the DMCs revealed significant enrichment in cancer-related pathways, demonstrating the utility of BMMs to reveal biologically relevant insights. An R package betaclust facilitates widespread use of BMMs.
BACKGROUND:Milk composition is complex and includes numerous components essential for offspring growth and development. In addition to the high abundance of miR-30b microRNA, milk produced by the transgenic mouse model of miR-30b-mammary deregulation displays a significantly altered fatty acid profile. Moreover, wild-type adopted pups fed miR-30b milk present an early growth defect. OBJECTIVE:This study aimed to investigate the consequences of miR-30b milk feeding on the duodenal development of wild-type neonates, a prime target of suckled milk, along with comprehensive milk phenotyping. METHODS:The duodenums of wild-type pups fed miR-30b milk were extensively characterized at postnatal day (PND)-5, PND-6, and PND-15 using histological, transcriptomic, proteomic, and duodenal permeability analyses and compared with those of pups fed wild-type milk. Milk of miR-30b foster dams collected at mid-lactation was extensively analyzed using proteomic, metabolomic, and lipidomic approaches and hormonal immunoassays. RESULTS:At PND-5, wild-type pups fed miR-30b milk showed maturation of their duodenum with 1.5-fold (P < 0.05) and 1.3-fold (P < 0.10) increased expression of Claudin-3 and Claudin-4, respectively, and changes in 8 duodenal proteins (P < 0.10), with an earlier reduction in paracellular and transcellular permeability (183 ng/mL fluorescein sulfonic acid [FSA] and 12 ng/mL horseradish peroxidase [HRP], respectively, compared with 5700 ng/mL FSA and 90 ng/mL HRP in wild-type; P < 0.001). Compared with wild-type milk, miR-30b milk displayed an increase in total lipid (219 g/L compared with 151 g/L; P < 0.05), ceramide (17.6 μM compared with 6.9 μM; P < 0.05), and sphingomyelin concentrations (163.7 μM compared with 76.3 μM; P < 0.05); overexpression of 9 proteins involved in the gut barrier (P < 0.1); and higher insulin and leptin concentrations (1.88 ng/mL and 2.04 ng/mL, respectively, compared with 0.79 ng/mL and 1.06 ng/mL; P < 0.01). CONCLUSIONS:miR-30b milk displays significant changes in bioactive components associated with neonatal duodenal integrity and maturation, which could be involved in the earlier intestinal closure phenotype of the wild-type pups associated with a lower growth rate.
A bstract Optimizing rabbit does preparation during early life to improve reproductive potential is a major challenge for breeders. Does selected for reproduction have specific nutritional needs, which may not be supplied with the common practice of feed restriction during rearing in commercial rabbit production. Nutrition during early life was already known to influence metabolism, reproduction and mammary gland development later in life, in particular during pregnancy. The aim of this study was to analyze the impact of four different feeding strategies in the early life of rabbit females (combination of high or moderate feed restriction from 5 to 9 weeks of age with restricted or ad libitum feeding regime from 9 to 12 weeks of constituting the pubertal period) on their growth, reproductive capacities and mammary development at mid-pregnancy. Unlike food intake, which remains regular, mean body weight gain was inversely proportional to the dietary restriction applied over the considered periods. The feeding strategies in place for the four groups had no effect on the reproductive parameters of the females at mid-pregnancy, as opposed to certain metabolic parameters such as cholesterolemia, that decreased with dietary intake at puberty (p≤0.05). Furthermore, restriction programs have impacted mammary tissular structures at mid-pregnancy. The expression of lipid metabolism enzymes (Fatty acid synthase N and Stearoyl co-A desaturase) is also increased in mammary epithelial tissue at mid-pregnancy by the dietary strategies implemented (p≤0.05). Moreover, milk gene expression, used as differentiation markers, indicates a better mammary epithelial development regarding further lactation, in the case of the less restrictive strategies during early life period, especially the higher feeding allowance. Our results highlight the importance of investigating feeding conditions of young female rabbits and nutrition in early life rearing, in order to provide specific recommendations for optimizing lactation and thus preventing neonatal mortality of the offspring.
Conflicting results regarding alterations to sperm DNA methylation in cases of spermatogenesis defects, male infertility and poor developmental outcomes have been reported in humans. Bulls used for artificial insemination represent a relevant model in this field, as the broad dissemination of bull semen considerably alleviates confounding factors and enables the precise assessment of male fertility. This study was therefore designed to assess the potential for sperm DNA methylation to predict bull fertility. A unique collection of 100 sperm samples was constituted by pooling 2–5 ejaculates per bull from 100 Montbéliarde bulls of comparable ages, assessed as fertile (n = 57) or subfertile (n = 43) based on non-return rates 56 days after insemination. The DNA methylation profiles of these samples were obtained using reduced representation bisulfite sequencing. After excluding putative sequence polymorphisms, 490 fertility-related differentially methylated cytosines (DMCs) were identified, most of which were hypermethylated in subfertile bulls. Interestingly, 46 genes targeted by DMCs are involved in embryonic and fetal development, sperm function and maturation, or have been related to fertility in genome-wide association studies; five of these were further analyzed by pyrosequencing. In order to evaluate the prognostic value of fertility-related DMCs, the sperm samples were split between training (n = 67) and testing (n = 33) sets. Using a Random Forest approach, a predictive model was built from the methylation values obtained on the training set. The predictive accuracy of this model was 72% on the testing set and 72% on individual ejaculates collected from an independent cohort of 20 bulls. This study, conducted on the largest set of bull sperm samples so far examined in epigenetic analyses, demonstrated that the sperm methylome is a valuable source of male fertility biomarkers. The next challenge is to combine these results with other data on the same sperm samples in order to improve the quality of the model and better understand the interplay between DNA methylation and other molecular features in the regulation of fertility. This research may have potential applications in human medicine, where infertility affects the interaction between a male and a female, thus making it difficult to isolate the male factor.
Research about mare's milk is mainly focused on quality and information about quantity is incomplete partly due to the lack of a consensus on the method of measuring milk yield. The live weight, body con-dition at foaling and age of mares are factors influencing milk yield. The influence of mare parity, how-ever, remains unclear. Over a period of 2 years (2018-2019), milk yield was evaluated on 65 mares (51 multiparous and 13 primiparous). Mares and foals were kept in a group at pasture. One method of milk yield measurement and one proxy method were applied; milking and weight-suckle-weight (WSW), respectively. The procedure was performed at five timepoints during the lactation period (3-30-60-90 and 180 days) without repetition. The relevance of WSW was addressed by studying the correlation between the two methods on 23 individuals. Factors influencing milk yield, through milking data, were studied on 57 individuals. Data was divided into two subsets. The first was an explanatory matrix con-taining the live weight of mares 24 h after parturition, parity, age, year of lactation and foal gender. The second was a response matrix containing data from milking at the five timepoints of the lactation. A correlation was found (RV = 0.41) between milking and WSW at day 3, however no correlation was found for other timepoints (RV < 0.15). The live weight of the mare 24 h after foaling, age and parity appeared to have a significant impact on milk production (P < 0.05). Thus, older or multiparous mares showed a higher milk yield than younger or primiparous mares. In addition, mares with a higher live weight after foaling produced more milk than those with a lower live weight. Overall, results can lead us to two main conclusions. First, the WSW method performed at five different timepoints of the lacta-tion, but without repeated measurements, is not an efficient way to estimate the milk yield of mares. Secondly, results concerning the live weight and age of mares were in accordance with previous studies. The influence of parity was also highlighted, confirming trends showed by other authors. Age and parity are closely related in our population, making it difficult to differentially assess their effects. Being able to identify the impact of both factors independently would benefit several sectors of the horse industry from sport to mare milk producers.(c) 2022 The Authors. Published by Elsevier B.V. on behalf of The Animal Consortium. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
The Shadoo and PrP prion protein family members are thought to be functionally related, but previous knockdown/knockout experiments in early mouse embryogenesis have provided seemingly contradictory results. In particular, Shadoo was found to be indispensable in the absence of PrP in knockdown analyses, but a double-knockout of the two had little phenotypic impact. We investigated this apparent discrepancy by comparing transcriptomes of WT, Prnp0/0 and Prnp0/0Sprn0/0 E6.5 mouse embryos following inoculation by Sprn- or Prnp-ShRNA lentiviral vectors. Our results suggest the possibility of genetic adaptation in Prnp0/0Sprn0/0 mice, thus providing a potential explanation for their previously observed resilience.
Background Integrating data from different sources is a recurring question in computational biology. Much effort has been devoted to the integration of data sets of the same type, typically multiple numerical data tables. However, data types are generally heterogeneous: it is a common place to gather data in the form of trees, networks or factorial maps, as these representations all have an appealing visual interpretation that helps to study grouping patterns and interactions between entities. The question we aim to answer in this paper is that of the integration of such representations. Results To this end, we provide a simple procedure to compare data with various types, in particular trees or networks, that relies essentially on two steps: the first step projects the representations into a common coordinate system; the second step then uses a multi-table integration approach to compare the projected data. We rely on efficient and well-known methodologies for each step: the projection step is achieved by retrieving a distance matrix for each representation form and then applying multidimensional scaling to provide a new set of coordinates from all the pairwise distances. The integration step is then achieved by applying a multiple factor analysis to the multiple tables of the new coordinates. This procedure provides tools to integrate and compare data available, for instance, as tree or network structures. Our approach is complementary to kernel methods, traditionally used to answer the same question. Conclusion Our approach is evaluated on simulation and used to analyze two real-world data sets: first, we compare several clusterings for different cell-types obtained from a transcriptomics single-cell data set in mouse embryos; second, we use our procedure to aggregate a multi-table data set from the TCGA breast cancer database, in order to compare several protein networks inferred for different breast cancer subtypes.
Avec pour thematique Ma mere, mon Placenta et Moi, ce colloque a pour objectif d’encourager la recherche sur le placenta et les pathologies obstetricales associees, de faciliter les collaborations internationales francophones, et les echanges scientifiques et methodologique dans notre domaine.