We propose a self-supervised pre-training framework for multivariate time-series classification that addresses the mismatch between fixed-window tokenization and the inherently variable temporal structure of real-world signals. Our framework combines Dynamic-Segment Masking (DSM) with a channel-independent Transformer encoder. DSM uses a recursive linear-fit validator to partition each sequence into content-adaptive, ϵ -linear segments and then randomly masks a proportion of segments. Besides, a lightweight, channel-independent Transformer encoder is trained to reconstruct the missing intervals, thereby learning temporal dependencies between observed and missing intervals. Despite having only 2.4 million parameters, our model achieves an average accuracy of 0.74 across 14 UEA benchmark datasets—exceeding a randomly initialized baseline by 7
Deep learning has advanced electromyography (EMG) based gesture recognition, yet existing models face robustness challenges. We propose a novel Transformer-based framework to enhance classification accuracy. Our architecture introduces two key innovations: a customized Patch Embedding module to adapt 1D time-series EMG signals for self-attention, and a novel Parametric Tanh Activation Function (DyT) that replaces conventional Layer Normalization to improve training stability and generalization. We evaluated our model on two public datasets representing distinct modalities: a sparse sEMG dataset (NinaPro DB2 Exercise B) and a high-density EMG dataset. The framework achieved high classification average test accuracies of 85.18
When sample sizes are small, it becomes challenging for an asymptotic test requiring diverging sample sizes to maintain an accurate Type I error rate. In this paper, we consider one-sample, two-sample and ANOVA tests for mean vectors when data are high-dimensional but sample sizes are very small. We establish asymptotic t-distributions of the proposed U-statistics, which only require data dimensionality to diverge but sample sizes to be fixed and no less than 3. The proposed tests maintain accurate Type I error rates for a wide range of sample sizes and data dimensionality. Moreover, the tests are nonparametric and can be applied to data which are normally distributed or heavy-tailed. Simulation studies confirm the theoretical results for the tests. We also apply the proposed tests to an fMRI dataset to demonstrate the practical implementation of the methods.
One important task in online data analysis is detecting network change, such as dissociation of communities or formation of new communities. Targeting on this type of application, we develop an online change-point detection procedure in the covariance structure of high-dimensional data. A new stopping rule is proposed to terminate the process as early as possible when a network change occurs. The stopping rule incorporates spatial and temporal dependence, and can be applied to non-Gaussian data. An explicit expression for the average run length (ARL) is derived, so that the level of threshold in the stopping rule can be easily obtained with no need to run time-consuming Monte Carlo simulations. We also establish an upper bound for the expected detection delay (EDD), the expression of which demonstrates the impact of data dependence and magnitude of change in the covariance structure. Simulation studies are provided to confirm accuracy of the theoretical results. The practical usefulness of the proposed procedure is illustrated by detecting brain's network change in a resting-state fMRI dataset.
We propose two procedures to detect a change in the mean of high-dimensional online data. One is based on a max-type U-statistic and another is based on a sum-type U-statistic. Theoretical properties of the two procedures are explored in the high dimensional setting. More precisely, we derive their average run lengths (ARLs) when there is no change point, and expected detection delays (EDDs) when there is a change point. Accuracy of the theoretical results is confirmed by simulation studies. The practical use of the proposed procedures is demonstrated by detecting an abrupt change in PM2.5 concentrations. The current study attempts to extend the results of the CUSUM and Shiryayev-Roberts procedures previously established in the univariate setting.
This article considers the problem of testing temporal homogeneity of p-dimensional population mean vectors from repeated measurements on n subjects over T times. To cope with the challenges brought about by high-dimensional longitudinal data, we propose methodology that takes into account not only the "large p, large T, and small n" situation but also the complex temporospatial dependence. We consider both the multivariate analysis of variance problem and the change point problem. The asymptotic distributions of the proposed test statistics are established under mild conditions. In the change point setting, when the null hypothesis of temporal homogeneity is rejected, we further propose a binary segmentation method and show that it is consistent with a rate that explicitly depends on p,T, and n. Simulation studies and an application to fMRI data are provided to demonstrate the performance and applicability of the proposed methods.
Extended from the spatially defined nearest neighbor index, the nearest neighbor index measures the levels of spatiotemporal clustering of a set of points, using only their spatial locations and the time associated with each point. The extended index is particularly suitable to use when there is no attribute information associated with geographic events except for locations and times when events occurred. In addition, it allows users to assess and visualize spatiotemporally distributed geographic events and to test if the events are more (or less) spatiotemporally clustered than would be expected by random chances. As a demonstration, this index was applied to crime and health data sets to demonstrate its usefulness. This article details the mathematical steps that formulate the extension. The calculation of the index values is fast and efficient with mathematical equations presented in the article.
High-dimensional time series are characterized by a large number of measurements and complex dependence, and often involve abrupt change points. We propose a new procedure to detect change points in the mean of high-dimensional time series data. The proposed procedure incorporates spatial and temporal dependence of data and is able to test and estimate the change point occurred on the boundary of time series. We study its asymptotic properties under mild conditions. Simulation studies demonstrate its robust performance through the comparison with other existing methods. Our procedure is applied to an fMRI dataset.
This paper considers testing the equality of two high dimensional means. Two approaches are utilized to formulate L-2-type tests for better power performance when the two high dimensional mean vectors differ only in sparsely populated coordinates and the differences are faint. One is to conduct thresholding to remove the nonsignal bearing dimensions for variance reduction of the test statistics. The other is to transform the data via the precision matrix for signal enhancement. It is shown that the thresholding and data transformation lead to attractive detection boundaries for the tests. Furthermore, we demonstrate explicitly the effects of precision matrix estimation on the detection boundary for the test with thresholding and data transformation. Extension to multi-sample ANOVA tests is also investigated. Numerical studies are performed to confirm the theoretical findings and demonstrate the practical implementations.
High throughput gene expression analysis using qPCR is commonly used to identify molecular markers of complex cellular processes. However, statistical analysis of multi-dimensional, temporal gene expression data is complicated by limited biological replicates and large number of measurements. Moreover, many available statistical tools for analysis of time series data assume that the data sequence is static and does not evolve over time. With this assumption, the parameters used to model the time series are fixed and thus, can be estimated by pooling data together. However, in many cases, dynamic processes of biological systems involve abrupt changes at unknown time points, making the assumption of stationary time series break down. We addressed this problem using a combination of statistical methods including hierarchical clustering, change point detection, and multiple testing. We applied this multi-step method to multi-dimensional, temporal gene expression data that resulted from our study of colony size-dependent neural cell differentiation of stem cells. The gene expression data were time series as the observations were recorded sequentially over time. Hierarchical clustering segregated the genes into three distinct clusters based on their temporal expression profiles; change point detection identified specific time points at which the entire dataset was divided into several homogenous subsets to allow a separate analysis of each subset; and multiple testing procedure identified the differentially expressed genes in each cluster within each subset of data. We established that our multi-step approach pinpoints specific sets of genes that underlie colony size-mediated neural differentiation of stem cells and demonstrated its advantages over conventional parametric and non-parametric tests that do not take into account temporal dynamics of the data. Importantly, our proposed approach is broadly applicable to any multivariate data sets of limited sample size from high throughput and high content screening such as in drug and biomarker discovery studies.
Tumor stroma is a major contributor to the biological aggressiveness of cancer cells. Cancer cells induce activation of normal fibroblasts to carcinoma-associated fibroblasts (CAFs), which promote survival, proliferation, metastasis, and drug resistance of cancer cells. A better understanding of these interactions could lead to new, targeted therapies for cancers with limited treatment options, such as triple negative breast cancer (TNBC). To overcome limitations of standard monolayer cell cultures and xenograft models that lack tumor complexity and/or human stroma, we have developed a high throughput tumor spheroid technology utilizing a polymeric aqueous two-phase system to conveniently model interactions of CAFs and TNBC cells and quantify effects on signaling and drug resistance of cancer cells. We focused on signaling by chemokine CXCL12, a hallmark molecule secreted by CAFs, and receptor CXCR4, a driver of tumor progression and metastasis in TNBC. Using three-dimensional stromal-TNBC cells cultures, we demonstrate that CXCL12 - CXCR4 signaling significantly increases growth of TNBC cells and drug resistance through activation of mitogen-activated protein kinase (MAPK) and phosphoinositide 3-kinase (PI3K) pathways. Despite resistance to standard chemotherapy, upregulation of MAPK and PI3K signaling sensitizes TNBC cells in co-culture spheroids to specific inhibitors of these kinase pathways. Furthermore, disrupting CXCL12 - CXCR4 signaling diminishes drug resistance of TNBC cells in co-culture spheroid models. This work illustrates the capability to identify mechanisms of drug resistance and overcome them using our engineered model of tumor-stromal interactions.
This paper aims to revive the classical Hotelling's $T^2$ test in the "large $p$, small $n$" paradigm. A Neighborhood-Assisted Hotelling's $T^2$ statistic is proposed to replace the inverse of sample covariance matrix in the classical Hotelling's $T^2$ statistic with a regularized covariance estimator. Utilizing a regression model, we establish its asymptotic normality under mild conditions. We show that the proposed test is able to match the performance of the population Hotelling's $T^2$ test with a known covariance under certain conditions, and thus possesses certain optimality. Moreover, the test has the ability to attain its best power possible by adjusting a neighborhood size to unknown structures of population mean and covariance matrix. An optimal neighborhood size selection procedure is proposed to maximize the power of the Neighborhood-Assisted $T^2$ test via maximizing the signal-to-noise ratio. Simulation experiments and case studies are given to demonstrate the empirical performance of the proposed test.
Microenvironmental factors have a major impact on differentiation of embryonic stem cells (ESCs). Here, a novel phenomenon that size of ESC colonies has a significant regulatory role on stromal cells induced differentiation of ESCs to neural cells is reported. Using a robotic cell microprinting technology, defined densities of ESCs are confined within aqueous nanodrops over a layer of supporting stromal cells immersed in a second, immiscible aqueous phase to generate ESC colonies of defined sizes. Temporal protein and gene expression studies demonstrate that larger ESC colonies generate disproportionally more neural cells and longer neurite processes. Unlike previous studies that attribute neural differentiation of ESCs solely to interactions with stromal cells, it is found that increased intercellular signaling of ESCs significantly enhances neural differentiation. This study offers an approach to generate neural cells with improved efficiency for potential use in translational research.
Many tests have been proposed to remedy the classical Hotelling's T^2 test in the "large p, small n" paradigm, but the existence of an optimal sum-of-squares type test has not been explored. This paper shows that under certain conditions, the population Hotelling's T^2 test with the known Σ^-1 attains the best power among all the L_2-norm based tests with the data transformation by Σ^η for η∈ (-∞, ∞). To extend the result to the case of unknown Σ^-1, we propose a Neighborhood-Assisted Hotelling's T^2 statistic obtained by replacing the inverse of sample covariance matrix in the classical Hotelling's T^2 statistic with a regularized covariance estimator. Utilizing a regression model, we establish its asymptotic normality under mild conditions. We show that the proposed test is able to match the performance of the population Hotelling's T^2 test under certain conditions, and thus possesses certain optimality. Moreover, it can adaptively attain the best power by empirically choosing a neighborhood size to maximize its signal-to-noise ratio. Simulation experiments and case studies are given to demonstrate the empirical performance of the proposed test.
The paper considers the problem of recovering the sparse different components between two high-dimensional means of column-wise dependent random vectors. We show that dependence can be utilized to lower the identification boundary for signal recovery. Moreover, an optimal convergence rate for the marginal false nondiscovery rate (mFNR) is established under dependence. The convergence rate is faster than the optimal rate without dependence. To recover the sparse signal bearing dimensions, we propose a Dependence-Assisted Thresholding and Excising (DATE) procedure, which is shown to be rate optimal for the mFNR with the marginal false discovery rate (mFDR) controlled at a pre-specified level. Extensions of the DATE to recover the differences in contrasts among multiple population means and differences between two covariance matrices are also provided. Simulation studies and case study are given to demonstrate the performance of the proposed signal identification procedure.
This paper considers the problem of testing temporal homogeneity of $p$-dimensional population mean vectors from the repeated measurements of $n$ subjects over $T$ times. To cope with the challenges brought by high-dimensional longitudinal data, we propose a test statistic that takes into account not only the "large $p$, large $T$ and small $n$" situation, but also the complex temporospatial dependence. The asymptotic distribution of the proposed test statistic is established under mild conditions. When the null hypothesis of temporal homogeneity is rejected, we further propose a binary segmentation method shown to be consistent for multiple change-point identification. Simulation studies and an application to fMRI data are provided to demonstrate the performance of the proposed methods.
We consider hypothesis testing problems for low-dimensional coefficients in a high dimensional additive hazard model. A variance reduced partial profiling estimator (VRPPE) is proposed and its asymptotic normality is established, which enables us to test the significance of each single coefficient when the data dimension is much larger than the sample size. Based on the p-values obtained from the proposed test statistics, we then apply a multiple testing procedure to identify significant coefficients and show that the false discovery rate can be controlled at the desired level. The proposed method is also extended to testing a low-dimensional sub-vector of coefficients. The finite sample performance of the proposed testing procedure is evaluated by simulation studies. We also apply it to two real data sets, with one focusing on testing low-dimensional coefficients and the other focusing on identifying significant coefficients through the proposed multiple testing procedure.
The paper considers the problem of identifying the sparse different components between two high dimensional means of column-wise dependent random vectors. We show that the dependence can be utilized to lower the identification boundary for signal recovery. Moreover, an optimal convergence rate for the marginal false non-discovery rate (mFNR) is established under the dependence. The convergence rate is faster than the optimal rate without dependence. To recover the sparse signal bearing dimensions, we propose a Dependence-Assisted Thresholding and Excising (DATE) procedure, which is shown to be rate optimal for the mFNR with the marginal false discovery rate (mFDR) controlled at a pre-specified level. Simulation studies and case study are given to demonstrate the performance of the proposed signal identification procedure.
Non-perturbative Hamiltonian light-front quantum field theory presents opportunities and challenges that bridge particle physics and nuclear physics. Fundamental theories, such as Quantum Chromodynmamics (QCD) and Quantum Electrodynamics (QED) offer the promise of great predictive power spanning phenomena on all scales from the microscopic to cosmic scales, but new tools that do not rely exclusively on perturbation theory are required to make connection from one scale to the next. We outline recent theoretical and computational progress to build these bridges and provide illustrative results for nuclear structure and quantum field theory. As our framework we choose light-front gauge and a basis function representation with two-dimensional harmonic oscillator basis for transverse modes that corresponds with eigensolutions of the soft-wall AdS/QCD model obtained from light-front holography.