语义解析的目标是将自然语言表达映射为机器可理解的逻辑表达,该任务的关键挑战在于难以刻画自然语言中蕴含的组合语义。目前,结合深度神经网络模型的语义解析方法已经成为该领域的主流方法,该类方法通常采用编码器—解码器框架,通过设计树形结构的解码器或者在解码器中添加语法限制,从语法层面上提升逻辑表达生成的准确率。与现有的神经语义解析方法不同,该文从语义建模角度出发,以语义框架作为中间形式,通过自顶向下的生成方式,显式地建模自然语言表达中蕴含的层次化语义结构。模型先根据自然语言输入,自顶向下地生成语义框架,再将语义框架表示融入到逻辑表达的生成过程中。三个数据集上的实验结果表明,该文提出的模型能更准确地生成语义框架,并且在语义解析任务中取得更好的效果。
Channel interpolation is an essential technique for providing high-accuracy estimation of the channel state information for wireless systems design where the frequency-space structural correlations of multi-antenna channel are typically hidden in matrix or tensor forms. In this correspondence paper, a modified extreme learning machine (ELM) that can process tensorial data, or ELM model with tensorial inputs (TELM), is proposed to handle the channel interpolation task. The TELM inherits many good properties from ELMs. Based on the TELM, the Tucker decomposed extreme learning machine is proposed for further improving the performance. Furthermore, we establish a theoretical argument to measure the interpolation capability of the proposed learning machines. Experimental results verify that our proposed learning machines can achieve comparable mean squared error (MSE) performance against the traditional ELMs but with 15% shorter running time, and outperform the other methods for a 20% margin measured in MSE for channel interpolation.
As the process of identifying the modulation format of the received signal, automatic modulation classification (AMC) has various applications in spectrum monitoring and signal interception. In this paper, we propose a dictionary learning-based AMC framework, where a dictionary is trained using signals with known modulation formats and the modulation format of the target signal is determined by its sparse representation on the dictionary. We also design a dictionary learning algorithm called block coordinate descent dictionary learning (BCDL). Furthermore, we prove the convergence of BCDL and quantify its convergence speed in a closed form. Simulation results show that our proposed AMC scheme offers superior performance than the existing methods with low complexity.
为解决超限学习机复杂度较高的问题,提出了一种新型的超限学习机更新策略,称为序列超限学习机.避免了复杂度较高的逆矩阵运算,而且能够应用于嵌入式系统中.序列超限学习机比各种广泛应用的机器学习分类器具有更低的计算复杂度.基于实际数据集的仿真结果表明,序列超限学习机的分类精度比传统超限学习机和其他广泛应用的分类器更高,而且具有更短的训练时间.
In this letter, we propose a novel pilot assignment scheme for the pilot contamination problem in massive multiple-input multiple-output multi-cell networks. Based on the asymptotic signal-to-interference-plus-noise ratio (SINR), the proposed pilot assignment scheme adopts the harmonic SINR utility function to quantify the fairness of all users in the network. Specifically, we formulate the pilot assignment problem as a minimum-weight multi-index assignment problem. For a two-cell network, this problem can be solved by the Hungarian algorithm with a strongly polynomial complexity. For the general multi-cell networks with more than two cells, this problem is in general NP-hard and we propose an efficient algorithm to obtain a suboptimal solution. The numerical experiments show that the proposed algorithms outperform other conventional pilot assignment schemes.
Submodular optimization plays a significant role in combinatorial problems, since it captures the structure of the edge cuts in graphs, the coverage of sets, and so on. Many data mining and machine learning problems can be cast as submodular maximization problems with applications in recommendation systems and data diversification. In this paper, we focus on the problem of maximizing a monotone submodular function subject to a d-knapsack constraint, for which we propose a streaming algorithm that achieves a ((1/1 + 2d) - e)-approximation of the optimal value, while it only needs one single pass through the data set without storing all the data in the memory. In our experiments, we extensively evaluate the effectiveness of our proposed algorithm via two applications: news recommendation and scientific literature recommendation. It is observed that the proposed streaming algorithm achieves both execution speedup and memory saving by several orders of magnitude, compared with existing approaches.
Automatic modulation classification (AMC) is the process of identifying the modulation format of the received signal. It is generally a difficult task due to the limited knowledge of the signal. In this letter, we propose a data driven dictionary-learning-based AMC framework, where we first use the known training signals to train the dictionary set and then classify the unknown modulation format via certain sparse representations, for which we design a dictionary-learning-based algorithm called block coordinate descent dictionary learning. Simulation results show that the proposed method out performs other existing approaches, achieving higher accuracy with a less training time.
BACKGROUND:Genotype-phenotype association has been one of the long-standing problems in bioinformatics. Identifying both the marginal and epistatic effects among genetic markers, such as Single Nucleotide Polymorphisms (SNPs), has been extensively integrated in Genome-Wide Association Studies (GWAS) to help derive "causal" genetic risk factors and their interactions, which play critical roles in life and disease systems. Identifying "synergistic" interactions with respect to the outcome of interest can help accurate phenotypic prediction and understand the underlying mechanism of system behavior. Many statistical measures for estimating synergistic interactions have been proposed in the literature for such a purpose. However, except for empirical performance, there is still no theoretical analysis on the power and limitation of these synergistic interaction measures.RESULTS:In this paper, it is shown that the existing information-theoretic multivariate synergy depends on a small subset of the interaction parameters in the model, sometimes on only one interaction parameter. In addition, an adjusted version of multivariate synergy is proposed as a new measure to estimate the interactive effects, with experiments conducted over both simulated data sets and a real-world GWAS data set to show the effectiveness.CONCLUSIONS:We provide rigorous theoretical analysis and empirical evidence on why the information-theoretic multivariate synergy helps with identifying genetic risk factors via synergistic interactions. We further establish the rigorous sample complexity analysis on detecting interactive effects, confirmed by both simulated and real-world data sets.
As the market of electric vehicles is gaining popularity, large-scale commercialized or privately-operated charging stations are expected to play a key role as a technology enabler. In this paper, we study the problem of charging electric vehicles at stations with limited charging machines and power resources. The purpose of this study is to develop a novel profit maximization framework for station operation in both offline and online charging scenarios, under certain customer satisfaction constraints. The main goal is to maximize the profit obtained by the station owner and provide a satisfactory charging service to the customers. The framework includes not only the vehicle scheduling and charging power control, but also the managing of user satisfaction factors, which are defined as the percentages of finished charging targets. The profit maximization problem is proved to be NPcomplete in both scenarios (NP refers to "nondeterministic polynomial time"), for which two-stage charging strategies are proposed to obtain efficient suboptimal solutions. Competitive analysis is also provided to analyze the performance of the proposed online two-stage charging algorithm against the offline counterpart under non-congested and congested charging scenarios. Finally, the simulation results show that the proposed two-stage charging strategies achieve performance close to that with exhaustive search. Also, the proposed algorithms provide remarkable performance gains compared to the other conventional charging strategies with respect to not only the unified profit, but also other practical interests, such as the computational time, the user satisfaction factor, the power consumption, and the competitive ratio.
An important problem in the field of bioinformatics is to identify interactive effects among profiled variables for outcome prediction. In this paper, a logistic regression model with pairwise interactions among a set of binary covariates is considered. Modeling the structure of the interactions by a graph, our goal is to recover the interaction graph from independently identically distributed (i.i.d.) samples of the covariates and the outcome. When viewed as a feature selection problem, a simple quantity called influence is proposed as a measure of the marginal effects of the interaction terms on the outcome. For the case when the underlying interaction graph is known to be acyclic, it is shown that a simple algorithm that is based on a maximum-weight spanning tree with respect to the plug-in estimates of the influences not only has strong theoretical performance guarantees, but can also outperform generic feature selection algorithms for recovering the interaction graph from i.i.d. samples of the covariates and the outcome. Our results can also be extended to the model that includes both individual effects and pairwise interactions via the help of an auxiliary covariate.
For an acyclic directed network with multiple pairs of sources and sinks and a set of Menger's paths connecting each pair of source and sink, it is known that the number of mergings among these Menger's paths is closely related to net encoding complexity. In this paper, we focus oil networks with two pairs of sources and sinks and we derive bounds on and exact values of two functions relevant to encoding complexity for such networks.
In practice, signals may be interfered by hostile jamming or illegal transmission and it is a very challenging task to determine the modulation formats of mixed signals. To tackle this problem, we propose a three-step algorithm called PFS algorithm. In the first step, principal component analysis (PCA) is conducted to suppress the noise. In the second step, the mixed signals are separated via fast independent component analysis (FICA), which transforms the received signals into the components that are maximally independent of each other. In the third step, high-order cumulants (HOCs) and support vector machines (SVMs) are adopted to determine the modulation format of the signal. The numerical experiments show that the PFS algorithm has a superior performance compared to other existing methods.
In this paper, we focus our attention on the Rényi entropy rate of hidden Markov processes under certain positivity assumptions. The existence of the Rényi entropy rate for such processes is established. Furthermore, we show that, with some extra “fast-forgetting” assumptions, the Rényi entropy rate of the approximating Markov processes exponentially converges to that of the original hidden Markov process, as the Markov order goes to infinity.
High-throughput next-generation sequencing technologies are producing a flood of cheap genomic information, providing precision medicine with the opportunity to better understand the primary cause of complicated diseases like cancer. However, even current state-of-the-art approaches still have large gaps with data generation due to limited scalability, accuracy and computational efficiency. To explore how to efficiently and effectively synthesize genomic data into knowledge, we propose GATK-Spark, a balanced parallelization approach that implements an in-memory version of GATK using Apache Spark. First, we performed a rigorous analysis of current GATK optimization strategies. We identify that compute resource utilization, text-based data format and long time single-thread file cutting and mergence operations are three major scalable bottlenecks. Second, we share our experiences designing a new approach optimized for GATK with big-data computing frameworks Apache Spark - GATK-Spark, which reduces the original execution of 20 hours to 30 minutes with a speedup in excess of 37 at 256 CPU cores. This work will facilitate the understanding of genomics analytics pipeline and design of strategies for accelerating large scale genomic analysis applications.
Submodular maximization problems belong to the family of combinatorial optimization problems and enjoy wide applications. In this paper, we focus on the problem of maximizing a monotone submodular function subject to a d-knapsack constraint, for which we propose a streaming algorithm that achieves a (1/1+2d - ϵ)-approximation of the optimal value, while it only needs one single pass through the dataset without storing all the data in the memory. In our experiments, we extensively evaluate the effectiveness of our proposed algorithm via an application in scientific literature recommendation. It is observed that the proposed streaming algorithm achieves both execution speedup and memory saving by several orders of magnitude, compared with existing approaches.
An important problem in the field of bioinformaties is to identify interactive effects among profiled variables for outcome prediction. In this paper, a simple logistic regression model with pairwise interactions among a set of binary covariates is considered. Modeling the structure of the interactions by a graph, our goal is to recover the interaction graph from independently identically distributed (i.i.d.) samples of the covariates and the outcome. When viewed as a feature selection problem, a simple quantity called influence is proposed as a measure of the marginal effects of the interaction terms on the outcome. For the case when the underlying interaction graph is known to be acyclic, it is shown that a simple algorithm that is based on a maximum-weight spanning tree with respect to the plug-in estimates of the influences not only has strong theoretical performance guarantees, but can also significantly outperform generic feature selection algorithms for recovering the interaction graph from i.i.d. samples of the covariates and the outcome.
In this paper, a hub refers to a non-terminal vertex of degree at least three. We study the minimum number of hubs needed in a network to guarantee certain flow demand constraints imposed between multiple pairs of sources and sinks. We prove that under the constraints, regardless of the size of the network, such minimum number is always upper bounded and we derive tight upper bounds for some special parameters. In particular, for two pairs of sources and sinks, we present a novel path-searching algorithm, the analysis of which is instrumental for the derivations of the tight upper bounds. Our results are of both theoretical and practical interest: in theory, they can be viewed as generalizations of the classical Menger's theorem to a class of undirected graphs with multiple sources and sinks; in practice, our results, roughly speaking, suggest that for some given flow demand constraints, not "too many" hubs are needed in a network.
For an acyclic directed network with multiple pairs of sources and sinks and a group of edge-disjoint paths connecting each pair of source and sink, it is known that the number of mergings among different groups of edge-disjoint paths is closely related to network encoding complexity. Using this connection, we derive exact values of and bounds on two functions relevant to encoding complexity for such networks.
For an acyclic directed network with multiple pairs of sources and sinks and a set of Menger's paths connecting each pair of source and sink, it is well known that the number of mergings among these Menger's paths is closely related to network encoding complexity. In this paper, we focus on networks with two distinct sinks and we derive bounds on and exact values of two functions relevant to encoding complexity for such networks.