We study three convolutions of polynomials in the context of free probability theory. We prove that these convolutions can be written as the expected characteristic polynomials of sums and products of unitarily invariant random matrices. The symmetric additive and multiplicative convolutions were introduced by Walsh and Szegö in different contexts, and have been studied for a century. The asymmetric additive convolution, and the connection of all of them with random matrices, is new. By developing the analogy with free probability, we prove that these convolutions produce real rooted polynomials and provide strong bounds on the locations of the roots of these polynomials.
We define the rectangular additive convolution of polynomials with nonnegative real roots as a generalization of the asymmetric additive convolution introduced by Marcus, Spielman and Srivastava. We then prove a sliding bound on the largest root of this convolution. The main tool used in the analysis is a differential operator derived from the "rectangular Cauchy transform" introduced by Benaych-Georges. The proof is inductive, with the base case requiring a new nonasymptotic bound on the Cauchy transform of Gegenbauer polynomials which may be of independent interest.
We use techniques from finite free probability to analyze matrix processes related to eigenvalues, singular values, and generalized singular values of random matrices. The models we use are quite basic and the analysis consists entirely of expected characteristic polynomials. A number of our results match known results in random matrix theory, however our main result (regarding generalized singular values) seems to be more general than any of the standard random matrix processes (Hermite/Laguerre/Jacobi) in the field. To test this, we perform a series of simulations of this new process that, on the one hand, confirms that this process can exhibit behavior not seen in the standard random matrix processes, but on the other hand provides evidence that the true behavior is captured quite well by our techniques. This, coupled with the fact that we are able to compute the same statistics for this new model that we are for the standard models, suggests that further investigation could be both interesting and fruitful.
We prove that there exist bipartite, biregular Ramanujan graphs of every degree and every number of vertices provided that the cardinalities of the two sets of the bipartition divide each other. This generalizes a result of Marcus, Spielman, and Srivastava and, similar to theirs, the proof is based on the analysis of expected polynomials. The primary difference is the use of some new machinery involving rectangular convolutions, developed in a companion paper. We also prove the constructibility of such graphs in polynomial time in the number of vertices, extending a result of Cohen to this biregular case.
We use the method of interlacing families of polynomials to derive a simple proof of Bourgain and Tzafriri's Restricted Invertibility Principle, and then to sharpen the result in two ways. We show that the stable rank can be replaced by the Schatten 4-norm stable rank and that tighter bounds hold when the number of columns in the matrix under consideration does not greatly exceed its number of rows. Our bounds are derived from an analysis of the smallest zeros of Jacobi and associated Laguerre polynomials.
We prove two "master" convolution theorems for multivariate determinantal polynomials. The methods used include basic properties of what we call a "minor-orthogonal" ensemble as well as properties of the mixed discriminant of matrices. We also give applications, including a rederivation of a result of Barvinok on computing the permanent of a low rank matrix and a polynomial convolution corresponding to the unitarily invariant addition of generalized singular values.
Checklists and action plans are a proven mechanism for project-based collaboration. Synthesizing project-specific plans is challenging, as project managers must consider multiple sources of information, from structured surveys to semi-structured conversations with stakeholders. In a needfinding study with project managers, we identified challenges in creating action plans for teams. We built MixTAPE, a mixed-initiative system that addressed these challenges with three components: a semi-structured note-taking interface for capturing stakeholder conversations, a plan generator for automatically combining multi-source information into action plans, and classification models for assigning and prioritizing action items. We evaluated MixTAPE in an observational study of 32 website design projects. Compared to a previously unstructured process, MixTAPE generated 1.45X as many tasks that are more consistent, while reducing the plan creation time by 33.70%. Through interviews and surveys, we found that participants rate MixTAPE highly across several measures. Based on our findings, we discuss the implications and opportunities for mixed-initiative action plan creation.
Checklists and guidelines have played an increasingly important role in complex tasks ranging from the cockpit to the operating theater. Their role in creative tasks like design is less explored. In a needfinding study with expert web designers, we identified designers' challenges in adhering to a checklist of design guidelines. We built Critter, which addressed these challenges with three components: Dynamic Checklists that progressively disclose guideline complexity with a self-pruning hierarchical view, AutoQA to automate common quality assurance checks, and guideline-specific feedback provided by a reviewer to highlight mistakes as they appear. In an observational study, we found that the more engaged a designer was with Critter, the fewer mistakes they made in following design guidelines. Designers rated the AutoQA and contextual feedback experience highly, and provided feedback on the tradeoffs of the hierarchical Dynamic Checklists. We additionally found that a majority of designers rated the AutoQA experience as excellent and felt that it increased the quality of their work. Finally, we discuss broader implications for supporting complex creative tasks.
We prove that there exist bipartite Ramanujan graphs of every degree and every number of vertices. The proof is based on analyzing the expected characteristic polynomial of a union of random perfect matchings, and involves three ingredients: (1) a formula for the expected characteristic polynomial of the sum of a regular graph with a random permutation of another regular graph, (2) a proof that this expected polynomial is real rooted and that the family of polynomials considered in this sum is an interlacing family, and (3) strong bounds on the roots of the expected characteristic polynomial of a union of random perfect matchings, established using the framework of finite free convolutions introduced recently by the authors.
As people increasingly work different jobs, the responsibility of building long-term career satisfaction and stability increasingly falls more on the workers rather than individual employers. With technologies playing a central role in how people choose and access employment opportunities, it is necessary to understand how online technologies are shaping career development. We bring together leading human-computer interaction researchers, industry members, and community organizers, who have worked with systems and people across the socio-economic spectrum during the career development process. We will discuss research on the role of online technologies in career development, including but not limited to topics in crowd work, social media sites, and freelance work sites.
The solution to the Kadison–Singer conjecture used techniques that intersect a number of areas of mathematics. The goal of this Arbeitsgemeinschaft was to bring together people from each of these fields to support interactions between these areas. While the majority of the talks centered around topics in polynomial geometry, combinatorics, and real algebraic geometry, participants came from areas such as harmonic analysis, convex geometry, and frame theory.
These lecture notes are meant to accompany two lectures given at the CDM 2016 conference, about the Kadison-Singer Problem. They are meant to complement the survey by the same authors (along with Spielman) which appeared at the 2014 ICM. In the first part of this survey we will introduce the Kadison-Singer problem from two perspectives ($C^*$ algebras and spectral graph theory) and present some examples showing where the difficulties in solving it lie. In the second part we will develop the framework of interlacing families of polynomials, and show how it is used to solve the problem. None of the results are new, but we have added annotations and examples which we hope are of pedagogical value.
Crowdsourcing refers to solving large problems by involving human workers that solve component sub-problems or tasks. In data crowdsourcing, the problem involves data acquisition, management, and analysis. In this paper, we provide an overview of data crowdsourcing, giving examples of problems that the authors have tackled, and presenting the key design steps involved in implementing a crowdsourced solution. We also discuss some of the open challenges that remain to be solved.
We study rank 1 perturbations of matrices where the perturbation vectors are drawn uniformly from the unit ball associated with a general β ensemble. To do this, we use distributions on the unit ball in Rn derived from the Dirichlet distribution to mimic the behavior of a unit vector drawn uniformly from a general β regime. Our main tool is an identity that expresses certain functions of these random vectors in terms of the Jack symmetric functions. Using this, we extend several properties of random matrices that are well known in the β = {1, 2, 4} case to general β. We then study the effect of additive rank 1 perturbations drawn from these general β distributions and present some open problems.
A system and method for data classification are presented. A plurality of training tokens are identified by at least one server communicatively coupled to a network. Each training token includes a token retrieved from a content source and a classification of the token. For each training token in the plurality of training tokens, a plurality of n-gram sequences are identified, a plurality of features for the plurality of n-gram sequences are generated, and first training data is generated using the token retrieved from the content source, the plurality of features, and the classification of the token. A first classifier is trained with the first training data, and the first classifier is stored into a storage system in communication with the at least one server.
We show that certain determinantal functions of multiple matrices, when summed over the symmetries of the cube, decompose into functions of the original matrices. These are shown to be true in complete generality; that is, no properties of the underlying vector space will be used apart from normal ring properties, and therefore hold in any commutative ring. All proofs are elementary --- in fact, the majority are simply derivations.
Crowdsourcing and human computation enable organizations to accomplish tasks that are currently not possible for fully automated techniques to complete, or require more flexibility and scalability than traditional employment relationships can facilitate. In the area of data processing, companies have benefited from crowd workers on platforms such as Amazon's Mechanical Turk or Upwork to complete tasks as varied as content moderation, web content extraction, entity resolution, and video/audio/image processing. Several academic researchers from diverse areas, ranging from the social sciences to computer science, have embraced crowdsourcing as a research area, resulting in algorithms and systems that improve crowd work quality, latency, and cost. Despite the relative nascence of the field, the academic and the practitioner communities have largely operated independently of each other for the past decade, rarely exchanging techniques and experiences. Crowdsourced Data Management: Industry and Academic Perspectives aims to narrow the gap between academics and practitioners. On the academic side, it summarizes the state of the art in crowd-powered algorithms and system design tailored to large-scale data processing. On the industry side, it surveys 13 industry users - such as Google, Facebook, and Microsoft - and four marketplace providers of crowd work - such as CrowdFlower and Upwork - to identify how hundreds of engineers and tens of million dollars are invested in various crowdsourcing solutions. Crowdsourced Data Management: Industry and Academic Perspectives simultaneously introduces academics to real problems that practitioners encounter every day, and provides a survey of the state of the art for practitioners to incorporate into their designs. Through the surveys, it also highlights the fact that crowdpowered data processing is a large and growing field. Over the next decade, most technical organizations are likely to benefit in some way from crowd work, and this monograph can help guide the effective adoption of crowdsourcing across these organizations.
The databases community works hard on the scale, performance, and correctness of the storage and query processing systems that our users depend on. Researchers are therefore frustrated to see less principled, and often incorrect 2010s implementations of concepts that were introduced in the 1970s. The lens of usability can help us understand how certain systems see adoption, regardless of the soundness of their implementation. Usability deficiencies in best-of-class systems can explain the success of systems with poor transactional semantics or unideal query languages, and even the use of spreadsheets for large-scale data management.