Understanding the cellular composition of complex tissues, such as tumors, is a key challenge in biology and medicine. A common approach, known as deconvolution, aims to estimate the cellular composition from bulk molecular measurements. With the growing availability of multiple types of molecular data, it is often assumed that combining data sources should improve deconvolution performance. Here, we present HADACA3, a community-driven benchmark designed to evaluate this assumption. We conducted a four-day collaborative competition followed by a large-scale computational benchmark, testing more than 250,000 analysis pipelines across nine datasets with matched DNA methylation (DNAm) and RNA profiles, representing a wide range of biological and experimental conditions. Our framework jointly evaluates the impact of preprocessing, feature selection, modeling, and integration strategies. We find that DNAm alone achieves the highest median performance across datasets, making it the most stable and reliable single-modality approach. However, multi-omics integration strategies can regularly achieve higher top performance in specific datasets and pipeline configurations. Among the tested strategies, late integration based on error-weighted averaging provides a strong and reliable baseline, while non-linear early integration methods, such as optimal transport, show promising results on real biological datasets. Overall, our results show that multi-omics integration does not systematically improve average performance over DNAm alone, but can improve best-case performance in specific settings. This highlights a trade-off between robustness and peak performance, and emphasizes the importance of aligning integration strategies with the statistical properties of the data. All data, code, and evaluation tools are publicly available to support reproducible research and future method development.
Cellular heterogeneity is a hallmark of biological tissues and plays a central role in disease progression, diagnosis, and prognosis. Yet, accurately characterizing this heterogeneity from bulk molecular profiles remains challenging because observed signals arise from mixtures of multiple cell populations. Cell deconvolution aim to recover the relative abundance of constituent cell types from such heterogeneous measurements, but most existing approaches implicitly rely on restrictive assumptions on residual errors, including independence, homoscedasticity, and normality. These assumptions are rarely satisfied in omics data, which are inherently bounded and overdispersed. In this work, we show that whole-genome cell-type specific DNA methylation profiles exhibit latent group structures that can substantially impair deconvolution accuracy when ignored. We therefore propose a mixture of non-negative Beta regression models estimated through an Expectation-Maximization algorithm for DNA methylation rates. Our framework naturally incorporates a feature selection mechanism through mixture component identification, making component selection a critical step of the inference procedure. We further propose a dedicated criterion for component selection and assess the performance of the approach through an extensive comparative study across several in vitro benchmark datasets. Our results demonstrate that deconvolution accuracy is highly sensitive to latent component structure and show that explicitly modeling this heterogeneity yields substantial improvements over standard whole-genome deconvolution strategies. Altogether, this work establishes mixture modeling of DNA methylation data as a powerful new direction for robust and accurate cell deconvolution.
BACKGROUND:Tumour heterogeneity significantly affects cancer progression and therapeutic response, yet quantifying it from bulk molecular data remains challenging. Deconvolution algorithms, which estimate cell type proportions in bulk samples, offer a potential solution. However, there is no consensus on the optimal algorithm for transcriptomic or methylomic data. RESULTS:Here, we present an unbiased evaluation framework for the first comprehensive comparison of deconvolution algorithms across both omic types, including reference-based and -free approaches. Our evaluation covers raw performance, stability, and computational efficiency under varying conditions, such as gene dependencies, missing or additional cell types and diverse sample compositions. We apply this framework across multiple benchmark datasets, including a novel multi-omics dataset generated specifically for this study. To ensure transparency and re-usability, we have designed a reproducible workflow using containerization and publicly available code. CONCLUSIONS:Our results highlight the strengths and limitations of various algorithms, and provides practical guidance for selecting the best method based on data type and analysis context. This benchmark sets a new standard for evaluating deconvolution methods and analysing tumour heterogeneity.
The study aimed to assess the extent to which protein aggregation, and even the modality of aggregation, can affect gastric digestion, down to the nature of the hydrolyzed peptide bonds. By controlling pH and ionic strength during heating, linear or spherical ovalbumin (OVA) aggregates were prepared, then digested with pepsin. Statistical analysis characterized the peptide bonds specifically hydrolyzed versus those not hydrolyzed for a given condition, based on a detailed description of all these bonds. Aggregation limits pepsin access to buried regions of native OVA, but some cleavage sites specific to aggregates reflect specific hydrolysis pathways due to the denaturation-aggregation process. Cleavage sites specific to linear aggregates indicate greater denaturation compared to spherical aggregates, consistent with theoretical models of heat-induced aggregation of OVA. Thus, the peptides released during the gastric phase may vary depending on the aggregation modality. Precisely tuned aggregation may therefore allow subtle control of the digestion process.
Dependence within a high-dimensional profile of explanatory variables affects estimation and prediction performance of regression models. However, the strong belief that dependence should not be ignored, based on our well-proven knowledge of low-dimensional regression modeling, is not necessarily true in high dimension. To investigate this point, we introduce a new class of prediction scores defined as linear combinations of a same random vector, including the naive prediction score obtained when ignoring dependence and the Ordinary Least Squares (OLS) prediction score that, on the contrary, fully accounts for dependence by a preliminary whitening of the explanatory variables. Interestingly, the former class also contains Ridge and Partial Least Squares prediction scores, that both offer intermediate ways of dealing with dependence. Through a theoretical comparative study, it is first shown how the best handling of dependence should depend on the interplay between the structure of conditional dependence across explanatory variables and the pattern of the association signal. We also derive the closed form expression of the prediction score with best prediction performance within the proposed class, leading to an adaptive handling of dependence. Finally, it is demonstrated through simulation studies and using benchmark datasets that this prediction score outperforms existing methods in various settings. Supplementary materials for this article are available online.
Makeup is a form of body art which has been used for more than 7,000 years and is present in the great majority of human cultures, often used to enhance facial attractiveness and to accentuate features that represent femininity. This study examines how cumulative levels of facial makeup influenced approach and avoidance tendencies and on facial muscle responses associated with emotional response obtained through facial electromyography (EMG) in a passive viewing task. Experiment 1 used the joystick variant of the approach-avoidance task, where 30 subjects categorized female faces by visual orientation (portrait/landscape) in seven cumulatively added makeup levels. In Experiment 2, facial EMG was recorded from 40 subjects in the passive viewing of the same images. The present study shows that makeup application modulates implicit responses and reveals two distinct implicit preferences, behavioral and affective, with a male behavioral preference for heavy eye cosmetics, a female behavioral preference for light makeup, and an overall affective preference in both men and women for makeup accentuating visual contrast in the eye and mouth regions. These results are consistent with the conception that perceptual cues underlying cosmetic enhancement are key determinants in aesthetic facial preferences.
The importance of poly-unsaturated fatty acids (PUFAs) in food is crucial for the animal and human development and health. As a complementary strategy to nutrition approaches, genetic selection has been suggested to improve fatty acids (FAs) composition in farmed fish. Gas chromatography (GC) is used as a reference method for the quantification of FAs; nevertheless, the high cost prevents large scale phenotyping as needed in breeding programs. Therefore, a calibration by means of Raman scattering spectrometry has been established in order to predict FA composition of visceral adipose tissue in rainbow trout Onchorhynchus mykiss. FA composition was analyzed by both GC and Raman micro-spectrometry techniques on 268 individuals fed with three different feeds, which have different FA compositions. Among the possible regression methods, the ridge regression method, was found to be efficient to establish calibration models from the GC and spectral data. The best cross-validated R2 values were obtained for total PUFAs, omega-6 (Ω-6) and omega-3 (Ω-3) PUFA (0.79, 0.83 and 0.66, respectively). For individual Ω-3 PUFAs, α-linolenic acid (ALA, C18:3), eicosapentaenoic acid (EPA, C20:5) and docosahexenoic acid (DHA, C22:6) were found to have the best R2 values (0.82, 0.76 and 0.81, respectively). This study demonstrates that Raman spectroscopy could be used to predict PUFAs with good correlation coefficients on adipocytes, for future on adipocytes physiology or for large scale and high throughput phenotyping in rainbow trout.
Genetic interaction is considered as one of the main heritable component of complex traits. With the emergence of genome-wide association studies (GWAS), a collection of statistical methods dedicated to the identification of interaction at the SNP level have been proposed. More recently, gene-based gene-gene interaction testing has emerged as an attractive alternative as they confer advantage in both statistical power and biological interpretation. Most of the gene-based interaction methods rely on a multidimensional modeling of the interaction, thus facing a lack of robustness against the huge space of interaction patterns. In this paper, we study a global testing approaches to address the issue of gene-based gene-gene interaction. Based on a logistic regression modeling framework, all SNP-SNP interaction tests are combined to produce a gene-level test for interaction. We propose an omnibus test that takes advantage of (1) the heterogeneity between existing global tests and (2) the complementarity between allele-based and genotype-based coding of SNPs. Through an extensive simulation study, it is demonstrated that the proposed omnibus test has the ability to detect with high power the most common interaction genetic models with one causal pair as well as more complex genetic models where more than one causal pair is involved. On the other hand, the flexibility of the proposed approach is shown to be robust and improves power compared to single global tests in replication studies. Furthermore, the application of our procedure to real datasets confirms the adaptability of our approach to replicate various gene-gene interactions.
Simultaneous tests of a huge number of hypotheses is a core issue in high flow experimental methods such as microarray for transcriptomic data. In the central debate about the type I error rate, Benjamini and Hochberg (1995) have proposed a procedure that is shown to control the now popular False Discovery Rate (FDR) under assumption of independence between the test statistics. These results have been extended to a larger class of dependency by Benjamini and Yekutieli (2001) and improvements have emerged in recent years, among which step-up procedures have shown desirable properties. The present paper focuses on the type II error rate. The proposed method improves the power by means of double-sampling test statistics integrating external information available both on the sample for which the outcomes are measured and also on additional items. The small sample distribution of the test statistics is provided and simulation studies are used to show the beneficial impact of introducing relevant covariates in the testing strategy. Finally, the present method is implemented in a situation where microarray data are used to select the genes that affect the degree of muscle destructuration in pigs. A phenotypic covariate is introduced in the analysis to improve the search for differentially expressed genes.
The specificity of pepsin, the major protease of gastric digestion, has been previously investigated, but only regarding the primary sequence of the protein substrates. The present study aimed to consider in addition physicochemical and structural characteristics, at the molecular and sub-molecular scales. For six different proteins submitted to in vitro gastric digestion, the peptide bonds cleaved were determined from the peptides released and identified by LC-MS/MS. An original statistical approach, based on propensity scores calculated for each amino acid residue on both sides of the peptide bonds, concluded that preferential cleavage occurred after Leu and Phe, and before Ile. Moreover, reliable statistical models developed for predicting peptide bond cleavage, highlighted the predominant role of the amino acid residues at the N-terminal side of the peptide bonds, up to the seventh position (P7 and P7'). The significant influence of hydrophobicity, charge and structural constraints around the peptide bonds was also evidenced.
In global testing, where a large number of pointwise test statistics are aggregated to simultaneously test for a collection of null hypotheses, the handling of dependence is a crucial issue. In various fields, more particularly in genetic epidemiology and functional data analysis, many testing methods for detecting an association signal between a response and explanatory variables have been proposed. Some aggregation procedures ignore dependence across pointwise test statistics whereas others introduce a model for decorrelation, with unclear conclusions on their relative performance. Indeed, the benefit that can be expected from decorrelation highly depends on the interplay between the structure of dependence across pointwise test statistics and the pattern of the association signal. Within a large class of test statistics covering a continuum of decorrelation approaches, an optimal procedure is introduced. This procedure is based on the maximization of an ad-hoc cumulant generating function-based distance between the null and nonnull distributions of a global test statistic, in order to adapt the aggregation of the pointwise statistics to the pattern of the association signal. A comparative study including simulations and applications to genetic association studies demonstrates that the ability of this test to detect a signal is more robust to the dependence structure than existing methods.
Abstract Background In response to major challenges regarding the supply and sustainability of marine ingredients in aquafeeds, the aquaculture industry has made a large-scale shift toward plant-based substitutions for fish oil and fish meal. But, this also led to lower levels of healthful n−3 long-chain polyunsaturated fatty acids (PUFAs)—especially eicosapentaenoic (EPA) and docosahexaenoic (DHA) acids—in flesh. One potential solution is to select fish with better abilities to retain or synthesise PUFAs, to increase the efficiency of aquaculture and promote the production of healthier fish products. To this end, we aimed i) to estimate the genetic variability in fatty acid (FA) composition in visceral fat quantified by Raman spectroscopy, with respect to both individual FAs and groups under a feeding regime with limited n-3 PUFAs; ii) to study the genetic and phenotypic correlations between FAs and processing yields- and fat-related traits; iii) to detect QTLs associated with FA composition and identify candidate genes; and iv) to assess the efficiency of genomic selection compared to pedigree-based BLUP selection. Results Proportions of the various FAs in fish were indirectly estimated using Raman scattering spectroscopy. Fish were genotyped using the 57 K SNP Axiom™ Trout Genotyping Array. Following quality control, the final analysis contained 29,652 SNPs from 1382 fish. Heritability estimates for traits ranged from 0.03 ± 0.03 (n-3 PUFAs) to 0.24 ± 0.05 (n-6 PUFAs), confirming the potential for genomic selection. n-3 PUFAs are positively correlated to a decrease in fat deposition in the fillet and in the viscera but negatively correlated to body weight. This highlights the potential interest to combine selection on FA and against fat deposition to improve nutritional merit of aquaculture products. Several QTLs were identified for FA composition, containing multiple candidate genes with indirect links to FA metabolism. In particular, one region on Omy1 was associated with n-6 PUFAs, monounsaturated FAs, linoleic acid, and EPA, while a region on Omy7 had effects on n-6 PUFAs, EPA, and linoleic acid. When we compared the effectiveness of breeding programmes based on genomic selection (using a reference population of 1000 individuals related to selection candidates) or on pedigree-based selection, we found that the former yielded increases in selection accuracy of 12 to 120% depending on the FA trait. Conclusion This study reveals the polygenic genetic architecture for FA composition in rainbow trout and confirms that genomic selection has potential to improve EPA and DHA proportions in aquaculture species.
Many factors could influence simultaneously soil spectra. We aimed to study the single effect of organic carbon and total iron in soil visible and short-wave near-infrared spectra and to quantify their contents. Two datasets of soil mixture samples were prepared by mixing, in various fractions, an organic carbon-rich material with a total iron-rich material and then with a total iron-poor material. For these two datasets, contents in organic carbon are quite similar but contents in total iron are significantly different. Results show that samples of the same dataset have the same overall spectral shape. Organic carbon has a decreasing effect that affects the whole spectral range without showing any specific absorption peaks. By contrast, total iron has specific absorption peaks. Spectra of the second dataset characterized by soil mixtures with higher total iron contents were more compact within the spectral bands 400-440 and 920-950 nm. Besides, continuum removal enables to exaggerate absorption peaks of wavelengths linked to total iron content. Partial Least Squares Regression (PLS R) models of both total organic carbon and total iron assign high coefficients to the wavelengths that are considered relevant and conversely low coefficients to those that are considered irrelevant. Both organic carbon content and total iron content were well predicted. For these models, coefficients of determination were superior to 0.9 and RMSE was closed to zero. The global models calibrated on all the samples demonstrated that PLS R was able to integrate sample heterogeneity.
The present document reproduces some data analyses and simulation studies presented in the submitted paper ‘A functional generalized F-test for signal detection with applications to Event-Related Potentials significance analysis’. In the manuscript, the whole dataset is used to illustrate the method, whereas in the following, the analyses are implemented on the data restricted to a region of interest made of three contiguous channels.
Motivated by the analysis of complex dependent functional data such as event-related brain potentials (ERP), this paper considers a time-varying coefficient multivariate regression model with fixed-time covariates for testing global hypotheses about population mean curves. Based on a reduced-rank modeling of the time correlation of the stochastic process of pointwise test statistics, a functional generalized F-test is proposed and its asymptotic null distribution is derived. Our analytical results show that the proposed test is more powerful than functional analysis of variance testing methods and competing signal detection procedures for dependent data. Simulation studies confirm such power gain for data with patterns of dependence similar to those observed in ERPs. The new testing procedure is illustrated with an analysis of the ERP data from a study of neural correlates of impulse control.
BACKGROUND:Because the cost of cereals is unstable and represents a large part of production charges for meat-type chicken, there is an urge to formulate alternative diets from more cost-effective feedstuff. We have recently shown that meat-type chicken source is prone to adapt to dietary starch substitution with fat and fiber. The aim of this study was to better understand the molecular mechanisms of this adaptation to changes in dietary energy sources through the fine characterization of transcriptomic changes occurring in three major metabolic tissues - liver, adipose tissue and muscle - as well as in circulating blood cells.RESULTS:We revealed the fine-tuned regulation of many hepatic genes encoding key enzymes driving glycogenesis and de novo fatty acid synthesis pathways and of some genes participating in oxidation. Among the genes expressed upon consumption of a high-fat, high-fiber diet, we highlighted CPT1A, which encodes a key enzyme in the regulation of fatty acid oxidation. Conversely, the repression of lipogenic genes by the high-fat diet was clearly associated with the down-regulation of SREBF1 transcripts but was not associated with the transcript regulation of MLXIPL and NR1H3, which are both transcription factors. This result suggests a pivotal role for SREBF1 in lipogenesis regulation in response to a decrease in dietary starch and an increase in dietary PUFA. Other prospective regulators of de novo hepatic lipogenesis were suggested, such as PPARD, JUN, TADA2A and KAT2B, the last two genes belonging to the lysine acetyl transferase (KAT) complex family regulating histone and non-histone protein acetylation. Hepatic glycogenic genes were also down-regulated in chickens fed a high-fat, high-fiber diet compared to those in chickens fed a starch-based diet. No significant dietary-associated variations in gene expression profiles was observed in the other studied tissues, suggesting that the liver mainly contributed to the adaptation of birds to changes in energy source and nutrients in their diets, at least at the transcriptional level. Moreover, we showed that PUFA deposition observed in the different tissues may not rely on transcriptional changes.CONCLUSION:We showed the major role of the liver, at the gene expression level, in the adaptive response of chicken to dietary starch substitution with fat and fiber.
We consider the problem of estimating the response probabilities in the context of weighting for unit nonresponse. The response probabilities may be estimated using either parametric or nonparametric methods. In practice, nonparametric methods are usually preferred because, unlike parametric methods, they protect against the misspeci cation of the nonresponse model. In this work, we conduct an extensive simulation study to compare methods for estimating the response probabilities in a nite population setting. In our study, we attempted to cover a wide range of (parametric and nonparametric) simple methods as well as aggregation methods like Bagging, Random Forests, Boosting. For each method, we assessed the performance of the propensity score estimator and the Hajek estimator in terms of relative bias and relative e ciency.