BackgroundDespite significantly improving outcomes in non-small cell lung cancer (NSCLC), neoadjuvant chemoimmunotherapy (NCIT) fails to achieve a major pathological response (MPR) in over 40% of patients. Consequently, the early identification of non-responders prior to treatment initiation remains a critical unmet clinical need.MethodsIn this study, we performed single-cell RNA sequencing (scRNA-seq) on pre-treatment NSCLC tissue samples and integrated data from two public databases to identify signaling pathways associated with poor treatment response. Key findings were subsequently validated using multiplex immunofluorescence (mIF), and the predictive value of identified molecules was finally assessed in our cohort and GEO datasets.ResultsAmong the 83 patients, 23/57 (40.35%) of these radiological responders failed to achieve MPR. Data analysis revealed activation of stress-related signaling pathways in cancer-associated fibroblasts (CAFs) and T cells from nMPR patients, with elevated expression of stress-related markers, including BAG3 and IFITM2. MIF confirmed that BAG3+IFITM2+ CAFs and BAG3+CD8+ T cells were spatially adjacent and significantly more abundant in nMPR patients. In our cohort and the two public databases, the BAG3+ CAF-T Cell Neighborhood was significantly more abundant in the nMPR group compared to the MPR group (p<0.05). In the MPR group, there was no significant difference in BAG3+ CAF-T Cell Neighborhood between the radiological PR and non-PR subgroups. The AUC values of BAG3+IFITM2+ CAFs, BAG3+CD8+ T cells, and the BAG3+ CAF-T Cell Neighborhood were 0.84 [95%CI: 0.746-0.931], 0.72 [95%CI: 0.603-0.835], and 0.87 [95%CI: 0.787-0.948], respectively. The OS and DFS of the BAG3+ CAF-T Cell Neighborhood high group are significantly decreased than that of the low group (p<0.05). The level of BAG3+ CAF-T Cell Neighborhood outperformed other two indicators in predicting non-response to NCIT. Consistent results were observed in GSE126044 and GSE135222.ConclusionThe BAG3+ CAF-T Cell Neighborhood may serve as a biomarker for predicting non-response to NCIT in NSCLC, with significant potential to inform clinical decision-making.
Since the pioneering work of sliced inverse regression, sufficient dimension reduction has been growing into a mature field in statistics and it has broad applications to regression diagnostics, data visualisation, image processing and machine learning. In this paper, we provide a review of several popular inverse regression methods, including sliced inverse regression (SIR) method and principal hessian directions (PHD) method. In addition, we adopt a conditional characteristic function approach and develop a new class of slicing-free methods, which are parallel to the classical SIR and PHD, and are named weighted inverse regression ensemble (WIRE) and weighted PHD (WPHD), respectively. Relationship with recently developed martingale difference divergence matrix is also revealed. Numerical studies and a real data example show that the proposed slicing-free alternatives have superior performance than SIR and PHD.
Martingale difference divergence measures the departure of conditional mean independence of two random vectors. Generalized martingale difference divergence and its correlation are developed based on symmetric Lévy measures to detect such an independence. Then the proposed generalized martingale difference correlation is utilized as a marginal utility to do high-dimensional variable screening. Both simulation results and real data illustrations show the promising performance of the developed indexes.
Variable screening as a fast and effective dimension reduction tool plays an important role in analyzing ultrahigh dimensional data. While a very large number of actual datasets contain both continuous and categorical variables, existing methods are mostly designed for continuous data. Partial sufficient variable screening, which aims to reduce the predictive set of primary interest without loss of regression information in the presence of some control variables, is proposed with theoretical guarantees. Specifically, for regression analyses involving mixed types of predictors, variable screening is approached under the notion of sufficiency by constraining the reduction of the continuous variables through the subpopulations identified by the categorical variables. The effectiveness of the proposed method is demonstrated through simulation studies encompassing a variety of regression and classification models, and an application in prognostic gene screening for diffuse large-B-cell lymphoma.
For univariate variable x ∈ R, the Lévy measures for the two real-valued continuous negative definite functions ln and lncosh(x) are shown in Böttcher et al. (2018). In this note, by utilizing the Fourier transform, we present the corresponding Lévy measures for ln and lncosh(|x|) with x ∈ Rp. These Lévy measures have significant potential applications across various domains.
Sliced inverse regression (SIR) has propelled sufficient dimension reduction (SDR) into a mature and versatile field with wide-ranging applications in statistics, including regression diagnostics, data visualisation, image processing and machine learning. However, traditional inverse regression techniques encounter challenges associated with sparsity arising from slicing operations. Weighted inverse regression ensemble (WIRE) presents a novel slicing-free approach to SDR. In this paper, we establish the asymptotic test theory to determine the dimension as estimated by WIRE. Moreover, we propose a permutation-based method for determining the order. Extensive numerical studies and real data analysis confirm the excellent performance of the proposed order determination method based on WIRE.
Sufficient dimension reduction reduces the dimension of a regression model without loss of information by replacing the original predictor with its lower-dimensional linear combinations. Partial (sufficient) dimension reduction arises when the predictors naturally fall into two sets X and W, and pursues a partial dimension reduction of X. Though partial dimension reduction is a very general problem, only very few research results are available when W is continuous. To the best of our knowledge, none can deal with the situation where the reduced lower-dimensional subspace of X varies with W. To address such issue, we in this paper propose a novel variable-dependent partial dimension reduction framework and adapt classical sufficient dimension reduction methods into this general paradigm. The asymptotic consistency of our method is investigated. Extensive numerical studies and real data analysis show that our variable dependent partial dimension reduction method has superior performance compared to the existing methods.
We propose two variable selection methods in multivariate linear regression with high-dimensional covariates. The first method uses a multiple correlation coefficient to fast reduce the dimension of the relevant predictors to a moderate or low level. The second method extends the univariate forward regression of Wang [(2009). Forward regression for ultra-high dimensional variable screening. Journal of the American Statistical Association, 104(488), 1512-1524. ] in a unified way such that the variable selection and model estimation can be obtained simultaneously. We establish the sure screening property for both methods. Simulation and real data applications are presented to show the finite sample performance of the proposed methods in comparison with some naive method.
The research is about a systematic investigation on the following issues. First, we construct different outcome regression-based estimators for conditional average treatment effect under, respectively, true, parametric, nonparametric and semiparametric dimension reduction structure. Second, according to the corresponding asymptotic variance functions when supposing the models are correctly specified, we answer the following questions: what is the asymptotic efficiency ranking about the four estimators in general? how is the efficiency related to the affiliation of the given covariates in the set of arguments of the regression functions? what do the roles of bandwidth and kernel function selections play for the estimation efficiency; and in which scenarios should the estimator under semiparametric dimension reduction regression structure be used in practice? Meanwhile, the results show that any outcome regression-based estimation should be asymptotically more efficient than any inverse probability weighting-based estimation. Several simulation studies are conducted to examine the finite sample performances of these estimators, and a real dataset is analyzed for illustration.
In this paper, we propose to use sufficient dimension reduction (SDR) in conjunction with nonparametric techniques to estimate the average treatment effect on the treated (ATT), a parameter of common interest in causal inference. The proposed method is applicable under a general low‐dimensional structure in the data, and avoids both the risk of model misspecification and the “curse of dimensionality,” for which it often outperforms the existing parametric and nonparametric methods. We develop the theoretical properties of the proposed method, including its asymptotic normality, its asymptotic super‐efficiency, and its equivalent form as an augmented inverse probability weighting estimator. We also consider the impact of SDR estimation in the asymptotic studies. These theoretical results are further illustrated by the simulation studies at the end.
Human leukocyte antigen G (HLA-G) is known as a novel immune checkpoint molecule in cancer; thus, HLA-G and its receptors might be targets for immune checkpoint blockade in cancer immunotherapy. The aim of this study was to systematically identify the roles of checkpoint HLA-G molecules across various types of cancer. ONCOMINE, GEPIA, CCLE, TRRUST, HAP, PrognoScan, Kaplan-Meier Plotter, cBioPortal, LinkedOmics, STRING, GeneMANIA, DAVID, TIMER, and CIBERSORT were utilized. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed. In this study, we comprehensively analysed the heterogeneous expression of HLA-G molecules in various types of cancer and focused on genetic alterations, coexpression patterns, gene interaction networks, HLA-G interactors, and the relationships between HLA-G and pathological stage, prognosis, and tumor-infiltrating immune cells. We first identified that the mRNA expression levels of HLA-G were significantly upregulated in both most tumor tissues and tumor cell lines on the basis of in-depth analysis of RNAseq data. The expression levels of HLA-G were positively associated with those of the other immune checkpoints PD-1 and CTLA-4. Abnormal expression of HLA-G was significantly correlated with the pathological stage of some but not all tumor types. There was a significant difference between the high and low HLA-G expression groups in terms of overall survival (OS) or disease-free survival (DFS). The results showed that HLA-G highly expressed have positive associations with tumor-infiltrating immune cells in the microenvironment in most types of tumors (P<0.05). Additionally, we identified the key transcription factor (TF) targets in the regulation of HLA-G expression, including HIVEP2, MYCN, CIITA, MYC, and IRF1. Multiple mutations (missense, truncating, etc.) and the methylation status of the HLA-G gene may explain the differential expression of HLA-G across different tumors. Functional enrichment analysis showed that HLA-G was primarily related to T cell activation, T cell regulation, and lymphocyte-mediated immunity. The data may provide novel insights for blockade of the HLA-G/ILT axis, which holds potential for the development of more effective antitumour treatments.
We thank all discussants for their insightful comments on ‘A selective overview of sparse sufficient dimension reduction’. We agree that sparse sufficient dimension reduction (sparse SDR) as a mode...
High-dimensional data analysis has been a challenging issue in statistics. Sufficient dimension reduction aims to reduce the dimension of the predictors by replacing the original predictors with a minimal set of their linear combinations without loss of information. However, the estimated linear combinations generally consist of all of the variables, making it difficult to interpret. To circumvent this difficulty, sparse sufficient dimension reduction methods were proposed to conduct model-free variable selection or screening within the framework of sufficient dimension reduction. We review the current literature of sparse sufficient dimension reduction and do some further investigation in this paper.
High-dimensional data analysis has been a challenging issue in statistics. Sufficient dimension reduction aims to reduce the dimension of the predictors by replacing the original predictors with a minimal set of their linear combinations without loss of information. However, the estimated linear combinations generally consist of all of the variables, making it difficult to interpret. To circumvent this difficulty, sparse sufficient dimension reduction methods were proposed to conduct model-free variable selection or screening within the framework of sufficient dimension reduction. We review the current literature of sparse sufficient dimension reduction and do some further investigation in this paper.
Sufficient dimension reduction aims for reduction of dimensionality of a regression without loss of information by replacing the original predictor with its lower-dimensional subspace. Partial (sufficient) dimension reduction arises when the predictors naturally fall into two sets, X and W, and we seek dimension reduction on X alone while considering all predictors in the regression analysis. Though partial dimension reduction is a very general problem, only very few research results are available when W is continuous. To the best of our knowledge, these methods generally perform poorly when X and W are related, furthermore, none can deal with the situation where the reduced lower-dimensional subspace of X varies dynamically with W. In this paper, We develop a novel dynamic partial dimension reduction method, which could handle the dynamic dimension reduction issue and also allows the dependency of X on W. The asymptotic consistency of our method is investigated. Extensive numerical studies and real data analysis show that our {\it Dynamic Partial Dimension Reduction} method has superior performance comparing to the existing methods.
Surface integrity of 3D medical data is crucial for surgery simulation or virtual diagnoses. However, undesirable holes often exist due to external damage on bodies or accessibility limitation on scanners. To bridge the gap, hole-filling for medical imaging is a popular research topic in recent years. Considering that medical image, e.g. CT or MRI, has the natural form of tensor, we recognize the problem of medical hole-filling as the extension of PCP problem from matrix case to tensor case. Since the new problem in tensor case is much more difficult than the matrix case, we design an efficient algorithm for the extension by relaxation technique. The most significant feature of our algorithm is that unlike traditional methods which follow a strictly local approach, our method fixes the hole by the global structure in the specific medical data. Another important difference to previous algorithm is that our algorithm is able to automatically separate the completed data from the hole in an implicit manner. Our experiments demonstrate that the proposed method can lead to satisfactory result.
A novel gender classification method based on frontal face images is presented.In the method,the global features are extracted by using an AdaBoost algorithm.The active appearance model(AAM) locates 83 landmarks,from which the local features are characterized.After the fusion of the local and global features,the mixed features are used to train support vector machine(SVM) classifiers.The method is evaluated by the recognition rates over a mixed face database containing over 14 700 images from 4 sources(AR,FERET,WWW and a database collected by the lab).Experimental results show that the hybrid method outperforms the unmixed appearance-or geometry-feature based methods and achieves a classification rate over 90%.
A novel method for single image super resolution without any training samples is presented in the paper. By sparse representation, the method attempts to recover at each pixel its best possible resolution increase based on the self similarity of the image patches across different scale and rotation transforms. The experiments indicate that the proposed method can produce robust and competitive results.
In this paper, we present the sexual dimorphism analysis in 3D human face and perform gender classification based on the result of sexual dimorphism analysis. Four types of features are extracted from a 3D human-face image. By using statistical methods, the existence of sexual dimorphism is demonstrated in 3D human face based on these features. The contributions of each feature to sexual dimorphism are quantified according to a novel criterion. The best gender classification rate is 94% by using SVMs and Matcher Weighting fusion method.This research adds to the knowledge of 3D faces in sexual dimorphism and affords a foundation that could be used to distinguish between male and female in 3D faces.
A novel gender classification method is presented which fuses information acquired from multiple facial regions for improving overall performance. It is able to compensate for facial expression even when training samples contain only neutral expression. We perform experimental investigation to evaluate the significance of different facial regions in the task of gender classification. Three most significant regions are used in our fusion-based method. The classification is performed by using support vector machines based on the features extracted using two-dimension principal component analysis. Experiments show that our fusion-based method is able to compensate for facial expressions and obtained the highest correct classification rate of 95.33%.