.We propose a loss function class called convex p-Lipschitz loss functions, which includes hinge loss, pinball loss and least square loss and so on. For regularized kernel networks and bias corrected regularized kernel networks with general convex pLipschitz loss, we establish the error analysis frameworks by employing the leave one out technique [16]. Under a mild source condition which describes how the minimizer f* of the generalization error can be approximated by the hypothesis space HK, satisfying error bounds and learning rates are deduced. Moreover, our proofs also show that bias correction method can indeed decrease the learning error.
Single cell RNA sequencing (scRNA-seq), a powerful tool for studying the tumor microenvironment (TME), does not preserve/provide spatial information on tissue morphology and cellular interactions. To understand the crosstalk between diverse cellular components in proximity in the TME, we performed scRNA-seq coupled with spatial transcriptomic (ST) assay to profile 41,700 cells from three colorectal cancer (CRC) tumor-normal-blood pairs. Standalone scRNA-seq analyses revealed eight major cell populations, including B cells, T cells, Monocytes, NK cells, Epithelial cells, Fibroblasts, Mast cells, Endothelial cells. After the identification of malignant cells from epithelial cells, we observed seven subtypes of malignant cells that reflect heterogeneous status in tumor, including tumor_CAV1, tumor_ATF3_JUN | FOS, tumor_ZEB2, tumor_VIM, tumor_WSB1, tumor_LXN, and tumor_PGM1. By transferring the cellular annotations obtained by scRNA-seq to ST spots, we annotated four regions in a cryosection from CRC patients, including tumor, stroma, immune infiltration, and colon epithelium regions. Furthermore, we observed intensive intercellular interactions between stroma and tumor regions which were extremely proximal in the cryosection. In particular, one pair of ligands and receptors (C5AR1 and RPS19) was inferred to play key roles in the crosstalk of stroma and tumor regions. For the tumor region, a typical feature of TMSB4X-high expression was identified, which could be a potential marker of CRC. The stroma region was found to be characterized by VIM-high expression, suggesting it fostered a stromal niche in the TME. Collectively, single cell and spatial analysis in our study reveal the tumor heterogeneity and molecular interactions in CRC TME, which provides insights into the mechanisms underlying CRC progression and may contribute to the development of anticancer therapies targeting on non-tumor components, such as the extracellular matrix (ECM) in CRC. The typical genes we identified may facilitate to new molecular subtypes of CRC.
In this paper, we study the learning performance of regularized large-margin unified machines (LUMs) for classification problem. The hypothesis space is taken to be a reproducing kernel Hilbert space H-K, and the penalty term is denoted by the norm of the function in H-K. Since the LUM loss functions are differentiable and convex, so the data piling phenomena can be avoided when dealing with the high-dimension low-sample size data. The error analysis of this classification learning machine mainly lies upon the comparison theorem [3] which ensures that the excess classification error can be bounded by the excess generalization error. Under a mild source condition which shows that the minimizer f(V) of the generalization error can be approximated by the hypothesis space H-K, and by a leave one out variant technique proposed in [13], satisfying error bound and learning rate about the mean of excess classification error are deduced.
Principal component analysis (PCA) may be the most popular dimension reduction method. In this paper, the learning scheme of kernel PCA methods is established. Moreover, for the uncentered case, we introduce the error representation, and prove the comparison theorem that the learning error can be bounded by the excess generalization error. Under the condition that the positive eigenvalues of [Formula: see text] are all single, the satisfied error bound [Formula: see text] is deduced.
We study the use of kernel ridge regression (KRR) in the block-wise streaming data. The algorithm works in an online manner: when a new data block comes in, the algorithm computes a local estimator based on the incoming data block and updates the predictive model by weighted average of all local estimators. Assuming the block data sizes increase at a mild rate and the regularization parameters are selected adaptively according to the sample size of all available data at the time of updating the model, we prove the convergence of the average KRR estimator. The rate is optimal when the regression function can be well approximated by the reproducing kernel Hilbert space in the L-2 sense.
Loss function is the key element of a learning algorithm. Based on the regression learning algorithm with an offset, the coefficient-based regularization network with variance loss is proposed. The variance loss is different from the usual least quare loss, hinge loss and pinball loss, it induces a kind of samples cross empirical risk. Also, our coefficient-based regularization only relies on general kernel, i.e. the kernel is required to possess continuity, boundedness and satisfy some mild differentiability condition. These two characteristics bring essential difficulties to the theoretical analysis of this learning scheme. By the hypothesis space strategy and the error decomposition technique in [L. Shi, Learning theory estimates for coefficient-based regularized regression, Appl. Comput. Harmon. Anal. 34 (2013) 252–265], a capacity-dependent error analysis is completed, satisfactory error bound and learning rates are then derived under a very mild regularity condition on the regression function. Also, we find an effective way to deal with the learning problem with samples cross empirical risk.
Distributed machine learning systems have been receiving increasing attentions for their efficiency to process large scale data. Many distributed frameworks have been proposed for different machine learning tasks. In this paper, we study the distributed kernel regression via the divide and conquer approach. This approach has been proved asymptotically minimax optimal if the kernel is perfectly selected so that the true regression function lies in the associated reproducing kernel Hilbert space. However, this is usually, if not always, impractical because kernels that can only be selected via prior knowledge or a tuning process are hardly perfect. Instead it is more common that the kernel is good enough but imperfect in the sense that the true regression can be well approximated by but does not lie exactly in the kernel space. We show distributed kernel regression can still achieves capacity independent optimal rate in this case. To this end, we first establish a general framework that allows to analyze distributed regression with response weighted base algorithms by bounding the error of such algorithms on a single data set, provided that the error bounds has factored the impact of the unexplained variance of the response variable. Then we perform a leave one out analysis of the kernel ridge regression and bias corrected kernel ridge regression, which in combination with the aforementioned framework allows us to derive sharp error bounds and capacity independent optimal rates for the associated distributed kernel regression algorithms. As a byproduct of the thorough analysis, we also prove the kernel ridge regression can achieve rates faster than $N^{-1}$ (where $N$ is the sample size) in the noise free setting which, to our best knowledge, are first observed and novel in regression learning.
Recently, density-based clustering algorithms have been garnering considerable attention in the unsupervised learning field as they can identify arbitrary cluster shapes. However, classical density-based algorithms are non-backtracking such as density-based spatial clustering of applications with noise (DBSCAN) and its improved variants. Therefore, the allocation of samples cannot be optimised during the clustering process. An incorrect division can cause subsequent errors. Furthermore, when the density is unevenly distributed, the globally uniform parameters will inevitably lead to catastrophic error consequences. To address these limitations, this study proposes an adaptive density-based clustering algorithm using the shared k-nearest neighbours (SKNN) conflict game (DC-SKCG). A local density-based adaptive cut-off distance-setting method is designed for the DC-SKCG, and a conflict game method based on SKNN is proposed to optimise the clustering process. Moreover, an automatic fusion mechanism for redundant high-density core regions is presented to (1) reduce the parameter sensitivity of the algorithm and (2) improve the parameter tolerance capacity. A series of experiments are conducted using various datasets demonstrate that DC-SKCG offers higher accuracy and robustness in most cases than state-of-the-art methods. (c) 2021 Elsevier Inc. All rights reserved.
Distributed learning is an effective way to analyze big data. In distributed regression, a typical approach is to partition the sample set into m disjoint data subsets of equal size, and then applies the kernel ridge regression algorithm to each sample subset to derive a local estimator, then averages them to get the global estimator. This paper mainly considers distributed regression learning with dependent samples of regularized least squares withα – mixing inputs that is involved in pre-existing literature [15]. Error bound in the K – metric has been derived and a novel error division method has been used to prove the asymptotic convergence for this distributed regularization learning. Learning rate of this algorithm will be obtained under a standard regularity condition on the regression function and the polynomial decay strongly mixing condition. It is proved that distributed learning is applicable to not only the i. i. d. samples but also dependent samples.
针对复杂的动态场景中平稳重复运动计数准确率较低的问题,提出基于线性回归分析方法对特定重复动作进行计数估计,根据同一动作在运动频率和运动形态上的相似性,对动作执行时间和重复次数之间的相关关系进行建模分析,在自建的数据集中进行实验,对视频重复动作计数进行预测和评估.结果 表明:该方法不受复杂现实场景的干扰,对多样性的动作特征不敏感,且学习参数较少,简单快速;在自建的数据集中测试平均绝对错误率为12.94%,说明该方法对重复动作计数是有效的.
针对金属有机化学气相沉积(MOVCD)反应室中因感生电流的集肤效应而导致衬底温度分布不均匀,从而影响生长薄膜质量的问题,通过对电磁加热式MOCVD反应室建立数值仿真模型,分析电流强度和电流频率对磁场分布和焦耳热分布的影响,同时对不同材料组成的基座结构中的磁场分布和焦耳热分布进行研究.结果 表明:磁场分布和焦耳热分布不随电流强度的改变而改变,但磁场和焦耳热的数值与电流强度成正比;磁场分布、焦耳热分布和数值随着电流频率的变化而变化,数值与电流频率成正比,且随着电流频率的增大,趋肤效应越明显;改变组成基座的材料组成可以改变基座中焦耳热分布,从而提高衬底温度分布的均匀性.
In this paper, We focus on conditional quantile regression learning algorithms based on the pinball loss and lq-regularizer with 1≤q≤2. Our main goal is to study the consistency of this kind of regularized quantile regression learning. With concentration inequality and operator decomposition techniques, we obtained satisfied error bounds and convergence rates.
We study distributed learning with partial coefficients regularization scheme in a reproducing kernel Hilbert space (RKHS). The algorithm randomly partitions the sample set [Formula: see text] into [Formula: see text] disjoint sample subsets of equal size. In order to reduce the complexity of algorithms, we apply a partial coefficients regularization scheme to each sample subset to produce an output function, and average the individual output functions to get the final global estimator. The error bound in the [Formula: see text]-metric is deduced and the asymptotic convergence for this distributed learning with partial coefficients regularization is proved by the integral operator technique. Satisfactory learning rates are then derived under a standard regularity condition on the regression function, which reveals an interesting phenomenon that when [Formula: see text] and [Formula: see text] is small enough, this distributed learning has the same convergence rate with the algorithm processing the whole data in one single machine.
Canonical correlation analysis (CCA) is a useful tool in detecting the latent relationship between two sets of multivariate variables. In theoretical analysis of CCA, a regularization technique is utilized to investigate the consistency of its analysis. This letter addresses the consistency property of CCA from a least squares view. We construct a constrained empirical risk minimization framework of CCA and apply a two-stage randomized Kaczmarz method to solve it. In the first stage, we remove the noise, and in the second stage, we compute the canonical weight vectors. Rigorous theoretical consistency is addressed. The statistical consistency of this novel scenario is extended to the kernel version of it. Moreover, experiments on both synthetic and real-world data sets demonstrate the effectiveness and efficiency of the proposed algorithms.
脉络丛癌(choroid plexus carcinoma,CPC)是一种颅内少见病,临床上报道较少.本组7例患儿,其中男5例,女2例.年龄1~9岁,平均年龄4.7岁.7例均为单发.5例患儿有颅内压增高症状,头痛、呕吐为主,其中1例有抽搐;2例患者有单侧肢体肌力减退,活动障碍.影像学表现为所有病例术前均行MRI检查,术后随访患者均MRI或CT复查.MRI T1平扫提示肿瘤呈类圆形、团块状等信号,边缘略呈分叶状,周围有轻度水肿.增强后病变呈不均匀异常强化(图1).治疗方法为7例患儿全部行显微外科手术切除.本组病例术中均出血较多,全部予以输血治疗.术后留置硬膜外引流管1根2~3 d,每日记录脑脊液性状和引流量.
In this paper, we study the performance of kernel-based regression learning with non-iid sampling. The non-iid samples are drawn from different probability distributions with the same conditional distribution. A more general marginal distribution assumption is proposed. Under this assumption, the consistency of the regularization kernel network (RKN) and the coefficient regularization kernel network (CRKN) are proved. Satisfactory capacity independently error bounds and learning rates are derived by the techniques of integral operator.
We study the asymptotical properties of indefinite kernel network with lq-norm regularization. The framework under investigation is different from classical kernel learning. Positive semidefiniteness is not required by the kernel function. By a new step stone technique, without any interior cone condition for input space X and Lτ condition for the probability measure ρX, satisfied error bounds and learning rates are deduced.
This paper proposes a new conditional kernel CCA (canonical correlation analysis) algorithm and exploits statistical consistency of it via modified Tikhonov regularization scheme, which is a continuous study of [11]. A new measure which characterizes consistency of learning ability is discussed based on the notion of distance between feature subspaces. The consistency analysis is conducted under the assumptions of normalized cross-covariance operators, which is mild and can be constructed by means of mean square contingency. Meantime, the relationship between this new measure and previous consistency scheme is investigated. Furthermore, we study conditional kernel CCA in a more general scenario by means of the trace operator.
This paper considers the regularized learning schemes based on ℓ1-regularizer and the ε-insensitive pinball loss in a data dependent hypothesis space. The target is the error analysis for the conditional quantile regression learning. Except for continuity and boundedness, the kernel function is not necessary to satisfy any further regularity conditions. The data dependent nature of the algorithm leads to an extra error term called hypothesis error. By concentration inequality with ℓ2-empirical covering numbers and operator decomposition techniques, satisfied error bounds and convergence rates are explicitly derived.
This paper considers the regularized learning schemes based on l(1)-regularizer and the epsilon-insensitive pinball loss in a data dependent hypothesis space. The target is the error analysis for the conditional quantile regression learning. Except for continuity and boundedness, the kernel function is not necessary to satisfy any further regularity conditions. The data dependent nature of the algorithm leads to an extra error term called hypothesis error. By concentration inequality with l(2)-empirical covering numbers and operator decomposition techniques, satisfied error bounds and convergence rates are explicitly derived.