The past two decades have seen significant advances in learning-related technologies, with the attendant recognition (and expectation) of their impact on the higher education system. Many new teaching methods can now be employed and their efficacy and scalability studied. In parallel, the demand for Electrical and Computer Engineering (ECE) education continues to grow world-wide, as an increasing world population seeks educational opportunities. There is now a growing acknowledgement that technical education must be complemented with skills for professional success such as design, leadership, communication, understanding historical and contemporary social contexts, lifelong learning, creativity, entrepreneurship, and teamwork. It is also widely accepted that solving today’s major challenges requires a multidisciplinary approach. The time is ripe for large-scale experimentation and adoption of possibly revolutionary changes in ECE education.
We present a fast contour-integral eigensolver for finding selected or all of the eigenpairs of a non-Hermitian matrix based on a series of analytical and computational techniques, such as the analysis of filter functions, quick and reliable eigenvalue count via low-accuracy matrix approximations, and fast shifted factorization update. The quality of some quadrature rules for approximating a relevant contour integral is analyzed. We show that a filter function based on the trapezoidal rule has nearly optimal decay in the complex plane away from the unit circle (as the mapped contour) and is superior to the Gauss--Legendre rule. The eigensolver needs to count the eigenvalues inside a contour. We justify the feasibility of using low-accuracy matrix approximations for the quick and reliable count. Both deterministic and probabilistic studies are given. With high probabilities, the matrix approximations give counts very close to the exact one. Our eigensolver is built upon an accelerated FEAST algorithm. Both the eigenvalue count and the FEAST eigenvalue solution need to solve linear systems with multiple shifts and right-hand sides. For this purpose and also to conveniently control the approximation accuracy, we use a type of rank structured approximations and show how to update the factorization for varying shifts. The eigensolver may be used to find a large number of eigenvalues, where a search region is then partitioned into subregions. We give an optimal threshold for the number of eigenvalues inside each bottom level subregion so as to minimize the complexity which is $\mathcal{O}(rn^{2})+\mathcal{O}(r^{2}n)$ to find all the eigenpairs of an order-$n$ matrix with maximum off-diagonal rank or numerical rank $r$. Numerical tests demonstrate the efficiency and accuracy and confirm the benefit of our acceleration techniques.
Least Square Policy Iteration (LSPI) is a model-free Reinforcement Learning algorithm capable of dealing with continuous states and actions. Chebyshev polynomials are utilized as the approximator in LSPI while Kalman Filtering handles sampled, corrupted and delayed data. Since LSPI solves optimal problems, the algorithm needs to have an exploration phase in order to avoid local minima and to cope with non-stationary cost-to-go functions. The chapter investigates how often information between neighbors in cooperative Multi-Agent Systems (MAS) needs to be exchanged in order to meet a desired performance. It suggests that stabilizing upper bounds for intervals between two consecutive broadcasting instants of each individual agent, thereby giving rise to asynchronous communication. It analyses the optimal intermittent feedback problem for MASs. The chapter explains the optimal intermittent feedback …
We design a distributed-memory randomized structured multifrontal solver for large sparse matrices. Two layers of hierarchical tree parallelism are used. A sequence of innovative parallel methods are developed for randomized structured frontal matrix operations, structured update matrix computation, skinny extend-add operation, selected entry extraction from structured matrices, etc. Several strategies are proposed to reuse computations and reduce communications. Unlike an earlier parallel structured multifrontal method that still involves large dense intermediate matrices, our parallel solver performs the major operations in terms of skinny matrices and fully structured forms. It thus significantly enhances the efficiency and scalability. Systematic communication cost analysis shows that the numbers of words are reduced by factors of about $O(\sqrt{n}/r)$ in two dimensions and about $O(n^{2/3}/r)$ in three dimensions, where $n$ is the matrix size and $r$ is an off-diagonal numerical rank bound of the intermediate frontal matrices. The efficiency and parallel performance are demonstrated with the solution of some large discretized PDEs in two and three dimensions. Nice scalability and significant savings in the cost and memory can be observed from the weak and strong scaling tests, especially for some 3D problems discretized on unstructured meshes.
We present a superfast divide-and-conquer method for finding all the eigenvalues as well as all the eigenvectors (in a structured form) of a class of symmetric matrices with off-diagonal ranks or numerical ranks bounded by $r$, as well as the approximation accuracy of the eigenvalues due to off-diagonal compression. More specifically, the complexity is $O(r^{2}n\log n)+O(rn\log^{2}n)$, where $n$ is the order of the matrix. Such matrices are often encountered in practical computations with banded matrices, Toeplitz matrices (in Fourier space), and certain discretized problems. They can be represented or approximated by hierarchically semiseparable (HSS) matrices. We show how to preserve the HSS structure throughout the dividing process that involves recursive updates and how to quickly perform stable eigendecompositions of the structured forms. Various other numerical issues are discussed, such as computation reuse and deflation. The structure of the eigenvector matrix is also shown. We further analyze the structured perturbation, i.e., how compression of the off-diagonal blocks impacts the accuracy of the eigenvalues. They show that rank structured methods can serve as an effective and efficient tool for approximate eigenvalue solutions with controllable accuracy. The algorithm and analysis are very useful for finding the eigendecomposition of matrices arising from some important applications and can be modified to find SVDs of nonsymmetric matrices. The efficiency and accuracy are illustrated in terms of Toeplitz and discretized matrices. Our method requires significantly fewer operations than a recent structured eigensolver, by nearly an order of magnitude.
We propose a fast structured selected inversion method for extracting the diagonal blocks of the inverse of a sparse symmetric matrix A, using the multifrontal method and rank structures. When A arises from the discretization of some PDEs and has a low-rank property (the intermediate dense matrices in the factorization have small off-diagonal numerical ranks), structured approximations of the diagonal blocks and certain off-diagonal blocks of A(-1) (that are needed to find the diagonal blocks of A(-1)) can be quickly computed. A structured multifrontal LDL factorization is first computed for A with a forward traversal of an assembly tree, which yields a sequence of local data-sparse factors. The factors are used in a backward traversal of the tree for the structured inversion. The intermediate operations in the inversion are performed in hierarchically semiseparable or low-rank forms. With the assumptions of data sparsity and appropriate rank conditions, the theoretical structured inversion cost is proportional to the matrix size n times a low-degree polylogarithmic function of n after structured factorizations. The memory counts are similar. In comparison, existing direct selected inversion methods cost O(n(3/2)) flops in two dimensions and O(n(2)) flops in three dimensions for both the factorization and the inversion, with O(n(4/3)) memory in three dimensions. Additional formulas for efficient structured operations are also derived. Numerical tests on two-and three-dimensional discretized PDEs and more general sparse matrices are done to demonstrate the performance.
This letter presents a linear-complexity finite-element-based eigenvalue solver for efficient analysis of 3-D on-chip integrated circuits. In this solver, from the original 3-D quadratic eigenvalue problem governing the on-chip circuits, we formulate a new generalized eigenvalue problem to efficiently compute the eigenvalues and eigenvectors of physical interest. We also develop an efficient linear-complexity solution for the matrix equation involved in the solution of the generalized eigenvalue problem. Numerical results demonstrate the accuracy and efficiency of the proposed eigenvalue solver.
PURPOSE:The adoption of multichannel compressed sensing (CS) for clinical magnetic resonance imaging (MRI) hinges on the ability to accurately reconstruct images from an undersampled dataset in a reasonable time frame. When CS is combined with SENSE parallel imaging, reconstruction can be computationally intensive. As an alternative to iterative methods that repetitively evaluate a forward CS+SENSE model, we introduce a technique for the fast computation of a compact inverse model solution.METHODS:A recently proposed hierarchically semiseparable (HSS) solver is used to compactly represent the inverse of the CS+SENSE encoding matrix to a high level of accuracy. To investigate the computational efficiency of the proposed HSS-Inverse method, we compare reconstruction time with the current state-of-the-art. In vivo 3T brain data at multiple image contrasts, resolutions, acceleration factors, and number of receive channels were used for this comparison.RESULTS:The HSS-Inverse method allows for >6× speedup when compared to current state-of-the-art reconstruction methods with the same accuracy. Efficient computational scaling is demonstrated for CS+SENSE with respect to image size. The HSS-Inverse method is also shown to have minimal dependency on the number of parallel imaging channels/acceleration factor.CONCLUSIONS:The proposed HSS-Inverse method is highly efficient and should enable real-time CS reconstruction on standard MRI vendors' computational hardware.
We present some superfast ($O((m+n)\log^{2}(m+n))$ complexity) and stablestructured direct solvers for $m\times n$ Toeplitz least squares problems.Based on the displacement equation, a Toeplitz matrix $T$ is first transformedinto a Cauchy-like matrix $\mathcal{C}$, which can be shown to have smalloff-diagonal numerical ranks when the diagonal blocks are rectangular. Wegeneralize standard hierarchically semiseparable (HSS) matrix representationsto rectangular ones, and construct a rectangular HSS approximation to$\mathcal{C}$ in nearly linear complexity with randomized sampling and fastmultiplications of $\mathcal{C}$ with vectors. A new URV HSS factorization anda URV HSS solution are designed for the least squares solution. We alsopresent two structured normal equation methods. Systematic error and stabilityanalysis for our HSS methods is given, which is also useful for studying otherHSS and rank structured methods. We derive the growth factors and the backwarderror bounds in the HSS factorizations, and show that the stability resultsare generally much better than those in dense LU factorizations with partialpivoting. Such analysis has not been done before for HSS matrices. The solversare tested on various classical Toeplitz examples ranging fromwell-conditioned to highly ill-conditioned ones. Comparisons with some recentfast and superfast solvers are given. Our new methods are generally muchfaster, and give better (or at least comparable) accuracies, especially forill-conditioned problems.
This paper presents a new method to design the digital filters for correcting uncertain 2-1 cascaded sigma-delta (Sigma Delta) modulators. The main contribution of this paper consists of two parts. First, we develop a new filter design method, based on H-infinity loop shaping technique, to deal with a certain weighted matching condition with polytopic uncertainties in parameters. The feature of the proposed method is to show the filter order can be independent of the weighting function and determined beforehand. Therefore, in contrast to the conventional H. loop shaping design method, lower-order filters can be obtained by using the proposed method. The second contribution is the application of the proposed method. For uncertain cascaded Sigma Delta modulators, a low-order filter with the same order of the nominal filter is designed, which can efficiently reduce the H-infinity norm of the noise transfer function in the signal frequency band. Consequently, the signal-tonoise ratio (SNR) performance is improved. We compare the proposed method 'with other existing designs and establish its efficacy.
Relay channels aid in increasing the rate of communication possible from the source to the destination. However, in general, the capacity of a relay channel is still an open problem. In this work, we provide a lower bound for the capacity of the three-terminal relay channel with destination-source feedback in the presence of correlated noise using the recent results in [1] and employing linear processing at the relay node. We also extend our model to the one involving multiple relays in either series or parallel configuration. Simulation results show that significant improvements in the lower bound can be obtained via our formulation of the problem.
We propose a global placement algorithm that employs size scaling of circuit components to provide continuity during placement. In the context of mixed-size placement, size scaling is utilized to handle significant variations among the sizes of the components, thereby avoiding additional complexity that is often associated with multiple levels of smoothing. By using the optimal region approach to first determine an initial placement, the size scaling approach allows the global placement algorithm to converge to better placement solutions.
Stephen F Cauley, Yuanzhe Xi, Berkin Bilgic, Kawin Setsompop, Jianlin Xia, Elfar Adalsteinsson, V. Ragu Balakrishnan, and Lawrence L Wald A.A. Martinos Center for Biomedical Imaging, Dept. of Radiology, MGH, Charlestown, MA, United States, Department of Mathematics, Purdue University, West Lafayette, IN, United States, Department of Electrical Engineering and Computer Science, MIT, Cambridge, MA, United States, Harvard Medical School, Boston, MA, United States, School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, United States, Harvard-MIT Division of Health Sciences and Technology, Cambridge, Massachusetts, United States
We propose a new technique for the design of state feedback controller for piecewise-linear systems, such that, the closed-loop systems are well-posed and asymptotically stable. First, a new criterion for the avoidance of sliding motion on the boundaries is presented. Then, the piecewise affine controller is constructed in a way such that the resulting closed-loop system satisfies the proposed criterion and a piecewise quadratic Lyapunov function can be used to establish the asymptotical stability. In this way, the control design problem is formulated as a numerical optimization problem under Linear Matrix Inequality constraints. The results are illustrated by application to control an aerobatic helicopter and balance an inverted pendulum on a cart, respectively, which demonstrates the efficacy and advantage of the proposed approach.
The simulation of realistically sized devices under the Non-Equilibrium Greens Function (NEGF) formalism typically requires prohibitive amounts of memory and computation time. In order to meet the rising computational challenges associated with quantum-scale device simulation we offer a 2-D domain decomposition technique. This technique is applicable to a large class of atomistic and spatial simulation problems. Considering a decomposition along both the cross section and length of the device, the framework presented in this work ensures efficient distribution of both memory and computation based upon the underlying device structure. As an illustration we stably generate the density of states and transmission, under the NEGF formalism, for the atomistic-based simulation of square 5 nm cross section silicon nanowires consisting of over one million atomic orbitals.
Linear coding schemes for noiseless feedback channels have been well studied. These schemes have been shown to display favorable properties for a broad class of Gaussian noise processes. However, the design of linear codes for noisy feedback channels is much less developed. In this work, an analytical expression for the optimal linear code for the additive white Gaussian noise feedback channel is derived. Also, bounds are obtained on the post-processed signal-to-noise ratio for communication systems with similar noise statistics in the forward and the feedback channel. For the general case of arbitrary forward and feedback noisy channels, an iteratively optimized linear coding scheme is proposed.
Very large scale integration (VLSI) has been a central technology for the realization of modern-day systems. The number of components in a VLSI design may run in the billions and it continues to grow, while the advantages associated with these advances are realized in multiple fields. However, the increased complexity of the designs poses new challenges for today's electronic design automation. VLSI computer-aided design tools have to address the effects of technology scaling on interconnects. More specifically, interconnect delay has become a dominant factor in modern designs. During the physical design process, placement of the components on a chip and routing of the connections among them are performed, and the final routed wirelength is used as a metric for determining the performance of the design. By minimizing final routed wirelength, power dissipation and interconnect delay can be reduced. The purpose of this Dissertation is to examine the implications of nanometer-scale VLSI technology in the physical design process and to introduce effective approaches to facilitate the overall physical design flow. With this objective, a clustering algorithm, called SafeNet, is developed in the context of VLSI placement, in order to improve both scalability and performance of the placer. The clustering algorithm applies fine clustering of the hypergraphs, thereby preserving the connectivity of the original VLSI circuits and avoiding any significant modification. Moreover, a flat placement algorithm, called PlaceD, is developed as an approach to the wirelength-driven placement problem on application-specific integrated circuits. The algorithm formulates the placement problem as a non-linear constrained optimization problem and applies size scaling in order to minimize the total wire-length of the design, starting with an initial placement obtained using an optimal region-based approach. Simultaneously, the placement algorithm also satisfies various placement constraints. Furthermore, a two-level placement algorithm, called PlaceR, is proposed to address the routability-driven placement problem. This placement algorithm estimates the routed wirelength during global placement using wire density. After the routed wirelength inside multiple regions is estimated, the algorithm incorporates the routability information into an objective function for the global placement in order to guide the placement process. PlaceR also combines a method for the spreading of wire density with clustering, pin congestion control, and size scaling. In this way, the placer minimizes wire congestion and the final routed wirelength of the design. Finally, a post-processing algorithm for placement, called Allagi, is developed to further improve the placement quality. The algorithm formulates the placement problem as a linear program and reduces the total routed wirelength by identifying congested regions in the design. Circuit components that have been placed in the congested areas of the design are relocated to regions that are optimal in terms of minimizing wirelength. The path for the minimization of congestion and improvement in routability is a necessary step for achieving design optimization and enhancing performance. Empirical results show that the proposed methods efficiently reduce the wirelength and improve the routability of the designs. Each of these methods can be incorporated in the overall physical design flow to improve the placement solution and facilitate the operation of a router.
Hybrid automatic repeat request (ARQ) protocols have become common in many packet transmission systems due to their incorporation in various standards. Hybrid-ARQ combines the normal ARQ method with forward error correction (FEC) codes to increase reliability and throughput. In this paper, we look at improving upon this performance using feedback information from the destination, in particular, using a powerful FEC code in conjunction with a proposed linear feedback code for the Rayleigh block fading channels. The new hybrid-ARQ scheme is initially developed for full received packet feedback in a point-to-point link. It is then extended to various multiple-antenna scenarios [e.g., multiple-input single-output (MISO), multiple-input multiple-output (MIMO), etc.] with varying amounts of packet feedback information. Simulations illustrate gains in throughput.
In this paper, we present a new robust matching filter design method for uncertain 2-1 cascaded sigma-delta modulators. This method addresses a well known limitation of H-infinity loop shaping techniques that they yield filters of high order (equal to the sum of the plant order and the order of the weighting function), thus increasing the complexity of circuit implementation. In contrast, the new method yields filters whose order is equal to the plant order, independent of the weighting function. We compare the new method with other existing fixed-order designs, and establish its efficacy.