In this work we demonstrate that Dark Matter (DM) evaporation severely hinders the effectiveness of exoplanets and Brown Dwarfs as sub -GeV DM probes. Moreover, we find useful analytic closed form approximations for DM capture rates for arbitrary astrophysical objects, valid in four distinct regions in the sigma - m(X )parameter space. As expected, in one of those regions the Dark Matter capture saturates to its geometric limit, i.e. the entire flux crossing an object. As a consequence of this region, which for many objects falls within the parameter space not excluded by direct detection experiments, we point out the existence of a DM parameter dependent critical temperature (T-crit ), above which astrophysical objects lose any sensitivity as Dark Matter probes. For instance, Jupiters at the Galactic Center have a T(crit )ranging from 700 K (for a 3M(J) Jupiter) to 950 K (for 14M(J)). This limitation is rarely (if ever) considered in the previous literature of indirect Dark Matter detection based on observable signatures of captured Dark Matter inside celestial bodies.
Graph representation learning (also called graph embeddings) is a popular technique for incorporating network structure into machine learning models. Unsupervised graph embedding methods aim to capture graph structure by learning a low-dimensional vector representation (the embedding) for each node. Despite the widespread use of these embeddings for a variety of downstream transductive machine learning tasks, there is little principled analysis of the effectiveness of this approach for common tasks. In this work, we provide an empirical and theoretical analysis for the performance of a class of embeddings on the common task of pairwise community labeling. This is a binary variant of the classic community detection problem, which seeks to build a classifier to determine whether a pair of vertices participate in a community. In line with our goal of foundational understanding, we focus on a popular class of unsupervised embedding techniques that learn low rank factorizations of a vertex proximity matrix (this class includes methods like GraRep, DeepWalk, node2vec, NetMF). We perform detailed empirical analysis for community labeling over a variety of real and synthetic graphs with ground truth. In all cases we studied, the models trained from embedding features perform poorly on community labeling. In constrast, a simple logistic model with classic graph structural features handily outperforms the embedding models. For a more principled understanding, we provide a theoretical analysis for the (in)effectiveness of these embeddings in capturing the community structure. We formally prove that popular low-dimensional factorization methods either cannot produce community structure, or can only produce “unstable" communities. These communities are inherently unstable under small perturbations.
Dark matter (DM) can be trapped by the gravitational field of any star, since collisions with nuclei in dense environments can slow down the DM particle below the escape velocity (v(est)) at the surface of the star. If captured, the DM particles can self-annihilate, and, therefore, provide a new source of energy for the star. We investigate this phenomenon for capture of DM particles by the first generation of stars [Population III (Pop III) stars], by using the multiscatter capture formalism. Pop III stars are particularly good DM captors, since they form in DM-rich environments, at the center of similar to 10(6) M-circle dot DM minihalos, at redshifts z similar to 15. Assuming a DM-proton scattering cross section (sigma) at the current deepest exclusion limits provided by the XENON1T experiment, we find that captured DM annihilations at the core of Pop III stars can lead, via the Eddington limit, to upper bounds in stellar masses that can be as low as a few M-circle dot if the ambient DM density (rho(X)) at the location of the Pop III star is sufficiently high. Conversely, when Pop III stars are identified, one can use their observed mass (M-*) to place bounds on rho(X)sigma. Using adiabatic contraction to estimate the ambient DM density in the environment surrounding Pop III stars, we place projected upper limits on sigma, for M-* in the 100 M-circle dot-1000 M-circle dot range, and find bounds that are competitive with, or deeper than, those provided by the most sensitive current direct detection experiments for both spin-independent and spin-dependent (SD) interactions, for a wide range of DM masses. Most intriguingly, we find that Pop III stars with mass M-* greater than or similar to 300 M-circle dot could be used to probe the SD proton-DM cross section below the "neutrino floor," i.e. the region of parameter space where DM direct detection experiments will soon become overwhelmed by neutrino backgrounds.
This work examines how to train fair classifiers in settings where training labels are corrupted with random noise, and where the error rates of corruption depend both on the label class and on the membership function for a protected subgroup. Heterogeneous label noise models systematic biases towards particular groups when generating annotations. We begin by presenting analytical results which show that naively imposing parity constraints on demographic disparity measures, without accounting for heterogeneous and group-dependent error rates, can decrease both the accuracy and the fairness of the resulting classifier. Our experiments demonstrate these issues arise in practice as well. We address these problems by performing empirical risk minimization with carefully defined surrogate loss functions and surrogate constraints that help avoid the pitfalls introduced by heterogeneous label noise. We provide both theoretical and empirical justifications for the efficacy of our methods. We view our results as an important example of how imposing fairness on biased data sets without proper care can do at least as much harm as it does good.
bands simultaneously was tested on simulated images. Based on our tests, we are confident that we can detect LSB galaxies down to a central surface brightness level of only 1.5 times the standard deviation from the mean pixel value in the image background. To assess the robustness of our method, the method was applied to a set of 18 B- and I-band images (covering 1.3 deg{sup 2} in total) of the Virgo Cluster to which Sabatini et al. previously applied a matched-filter dwarf LSB galaxy search algorithm. We have detected all 20 objects from the Sabatini et al. catalog which we could classify by eye as bona fide LSB galaxies. Our method has also detected four additional Virgo Cluster LSB galaxy candidates undetected by Sabatini et al. To further assess the completeness of the results of our method, both MARSIAA, SExtractor, and DetectLSB were applied to search for (1) mock Virgo LSB galaxies inserted into a set of deep Next Generation Virgo Survey (NGVS) gri-band subimages and (2) Virgo LSB galaxies identified by eye in a full set of NGVS square degree gri images. MARSIAA/DetectLSB recovered {approx}20% more mock LSB galaxies and {approx}40% more LSB galaxies identified by eye than SExtractor/DetectLSB. With a 90% fraction of false positives from an entirely unsupervised pipeline, a completeness of 90% is reached for sources with r{sub e} > 3'' at a mean surface brightness level of {mu}{sub g} = 27.7 mag arcsec{sup -2} and a central surface brightness of {mu}{sup 0}{sub g} = 26.7 mag arcsec{sup -2}. About 10% of the false positives are artifacts, the rest being background galaxies. We have found our proposed Markovian LSB galaxy detection method to be complementary to the application of matched filters and an optimized use of SExtractor, and to have the following advantages: it is scale free, can be applied simultaneously to several bands, and is well adapted for crowded regions on the sky.
We show that the mere observation of the first stars (Pop III stars) in the universe can be used to place tight constraints on the strength of the interaction between dark matter and regular, baryonic matter. We apply this technique to a candidate Pop III stellar complex discovered with the Hubble Space Telescope at z ∼ 7 and find bounds that are competitive with, or even stronger than, current direct detection experiments, such as XENON1T, for dark matter particles with mass (m_X) larger than about 100 GeV. We also show that the discovery of sufficiently massive Pop III stars could be used to bypass the main limitations of direct detection experiments: the neutrino background to which they will be soon sensitive.
Let $T$ be a binary search tree. We prove two results about the behavior of the Splay algorithm (Sleator and Tarjan 1985). Our first result is that inserting keys into an empty binary search tree via splaying in the order of either $T$'s preorder or $T$'s postorder takes linear time. Our proof uses the fact that preorders and postorders are pattern-avoiding: i.e. they contain no subsequences that are order-isomorphic to $(2,3,1)$ and $(3,1,2)$, respectively. Pattern-avoidance implies certain constraints on the manner in which items are inserted. We exploit this structure with a simple potential function that counts inserted nodes lying on access paths to uninserted nodes. Our methods can likely be extended to permutations that avoid more general patterns. Second, if $T'$ is any other binary search tree with the same keys as $T$ and $T$ is weight-balanced (Nievergelt and Reingold 1973), then splaying $T$'s preorder sequence or $T$'s postorder sequence starting from $T'$ takes linear time. To prove this, we demonstrate that preorders and postorders of balanced search trees do not contain many large "jumps" in symmetric order, and exploit this fact by using the dynamic finger theorem (Cole et al. 2000). Both of our results provide further evidence in favor of the elusive "dynamic optimality conjecture."
We introduce the zip tree,1 a form of randomized binary search tree that integrates previous ideas into one practical, performant, and pleasant-to-implement package. A zip tree is a binary search tree in which each node has a numeric rank and the tree is (max)-heap-ordered with respect to ranks, with rank ties broken in favor of smaller keys. Zip trees are essentially treaps [8], except that ranks are drawn from a geometric distribution instead of a uniform distribution, and we allow rank ties. These changes enable us to use fewer random bits per node. We perform insertions and deletions by unmerging and merging paths (unzipping and zipping) rather than by doing rotations, which avoids some pointer changes and improves efficiency. The methods of zipping and unzipping take inspiration from previous top-down approaches to insertion and deletion by Stephenson [10], Martínez and Roura [5], and Sprugnoli [9]. From a theoretical standpoint, this work provides two main results. First, zip trees require only O(log log n) bits (with high probability) to represent the largest rank in an n-node binary search tree; previous data structures require O(log n) bits for the largest rank. Second, zip trees are naturally isomorphic to skip lists [7], and simplify Dean and Jones’ mapping between skip lists
Consider the task of performing a sequence of searches in a binary search tree. After each search, an algorithm is allowed to arbitrarily restructure the tree, at a cost proportional to the amount of restructuring performed. The cost of an execution is the sum of the time spent searching and the time spent optimizing those searches with restructuring operations. This notion was introduced by Sleator and Tarjan in 1985 [27], along with an algorithm and a conjecture. The algorithm, Splay, is an elegant procedure for performing adjustments while moving searched items to the top of the tree. The conjecture, called dynamic optimality, is that the cost of splaying is always within a constant factor of the optimal algorithm for performing searches. The conjecture stands to this day.We offer the first systematic proposal for settling the dynamic optimality conjecture. At the heart of our methods is what we term a simulation embedding: a mapping from executions to lists of keys that induces a target algorithm to simulate the execution. We build a simulation embedding for Splay by inducing it to perform arbitrary subtree transformations, and use this to show that if the cost of splaying a sequence of items is an upper bound on the cost of splaying every subsequence thereof, then Splay is dynamically optimal. We call this the subsequence property. Building on this machinery, we show that if Splay is dynamically optimal, then with respect to optimal costs, its additive overhead is at most linear in the sum of initial tree size and number of requests. As a corollary, the subsequence property is also a necessary condition for dynamic optimality. The subsequence property also implies both the traversal [27] and deque [30] conjectures.The notions of simulation embeddings and bounding additive overheads should be of general interest in competitive analysis. For readers especially interested in dynamic optimality, we provide an outline of a proof that a lower bound on search costs by Wilber [32] has the subsequence property, and extensive suggestions for adapting this proof to Splay.
We introduce the zip tree, a form of randomized binary search tree. One can view a zip tree as a treap [8] in which priority ties are allowed and in which insertions and deletions are done by unmerging and merging paths (unzipping and zipping) rather than by doing rotations. Alternatively, one can view a zip tree as a binary-tree representation of a skip list [7]. Doing insertions and deletions by unzipping and zipping instead of by doing rotations avoids some pointer changes and can thereby improve efficiency. Representing a skip list as a binary tree avoids the need for nodes of different sizes and can speed up searches and updates. Zip trees are at least as simple as treaps and skip lists but offer improved efficiency. Their simplicity makes them especially amenable to concurrent operations. 1 Definition of Zip Trees A binary search tree is a binary tree in which each node contains an item, each item has a key, and the items are arranged in symmetric order : if x is a node, all items in the left subtree of x have keys less than that of x, and all items in the right subtree of x have keys greater than that of x. Such a tree supports binary search: to find an item in the tree with a given key, proceed as follows. If the tree is empty, stop: no item in the tree has the given key. Otherwise, compare the desired key with that of the item in the root. If they are equal, stop and return the item in the root. If the given key is less than that of the item in the root, search recursively in the left subtree of the root. Otherwise, search recursively in the right subtree of the root. The path of nodes visited during the search is the search path. If the search is unsuccessful, the search path starts at the root and ends at a missing node corresponding to an empty subtree. ∗Research at Princeton University partially supported by an innovation research grant from Princeton and a gift from Microsoft. †Department of Computer Science, Princeton University, and Intertrust Technologies; ret@cs.princeton.edu. ‡Program in Applied and Computational Mathematics, Princeton University, and Intertrust Technologies; cclevy@princeton.edu. 1Zip: “To move very fast.” 1 ar X iv :1 80 6. 06 72 6v 1 [ cs .D S] 1 8 Ju n 20 18 To keep our presentation simple, in this and the next section we do not distinguish between an item and the node containing it. (The data structure is endogenous [11].) We also assume that all nodes have distinct keys. It is straightforward to eliminate these assumptions. We call a node binary, unary, or a leaf, if it has two, one or zero children, respectively. We define the depth of a node recursively to be zero if it is the root, or one plus the depth of its parent if not. We define the height of a node recursively to be zero if it is a leaf, or one plus the maximum of the heights of its children if not. The left (resp. right) spine of a tree is the path from the root to the node of smallest (resp. largest) key. The left (resp. right) spine of x contains only the root and left (resp. right) children. We represent a binary search tree by storing in each node x its left child x.left , its right child x.right , and its key, x.key . If x has no left (resp. right) child, x.left = null (resp. x.right = null). A zip tree is a binary search tree in which each node has a numeric rank and the tree is (max)-heap-ordered with respect to ranks, with ties broken in favor of smaller keys: the parent of a node has rank greater than that of its left child and no less than that of its right child. We choose the rank of a node randomly when the node is inserted into the tree. We choose node ranks independently from a geometric distribution with mean 1: the rank of a node is non-negative integer k with probability 1/2. We denote by x.rank the rank of node x. We can store the rank of a node in the node or compute it as a pseudo-random function of the node (or of its key) each time it is needed. The pseudo-random function method, proposed by Aragon and Seidel [8], avoids the need to store ranks but requires a stronger independence assumption for the validity of our efficiency bounds, as we discuss in Section 3. To insert a new node x into a zip tree, we search for x in the tree until reaching the node y that x will replace; namely the node y such that y.rank ≤ x.rank , with strict inequality if y.key < x.key . From y, we follow the rest of the search path for x, unzipping it by splitting it into a path P containing all nodes with keys less than x.key and a path Q containing all nodes with keys greater than x.key . Along P from top to bottom, nodes are in increasing order by key and non-increasing order by rank; along Q from top to bottom, nodes are in decreasing order by both rank and key. Unzipping preserves the left subtrees of the nodes on P and the right subtrees of the nodes on Q. We make the top node of P the left child of x and the top node of Q the right child of x. Finally, if y had a parent z before the insertion, we make x the left or right child of z depending on whether its key is less than or greater than that of z, respectively (x replaces y as a child of z); if y was the root before the insertion, we make x the root. Deletion is the inverse of insertion. To delete a node x, do a search to find it. Let P and Q be the right spine of the left subtree of x and the left spine of the right subtree of x. Zip P and Q to form a single path R by merging them from top to bottom in non-decreasing rank order, breaking ties in favor of smaller keys. Zipping preserves the left subtrees of the nodes on P and the right subtrees of the nodes on Q. Finally, if x had a parent z before the insertion, make the top node of R (or null if R is empty) the left or right child of z, depending on
We present a method of extracting figures and images from the pages of scanned documents, especially from technical research articles. Our approach is novel in two key ways. First, we treat this as a computer vision problem, and train convolutional neural networks to recognize figures in scanned pages. Second, we generate our training data from 'born-digital' structured documents, allowing us to automatically produce labels for our training set using PDF figure extractors. This avoids the otherwise tedious task of hand-labelling thousands of document pages. Our convolutional neural networks achieve precision and recall of close to 85% in identifying figures from a test set consisting of modern journal papers and conference proceedings, and obtain precision and recall above 80% on an application data set comprised of historical technical documents scanned from the Bell Labs Records. Our results show that models trained on digital documents transfer very well to historical scans. Finally, it is easy to extend our models to identify other document elements such as tables and captions.
Chun-Nam John Yu合作论文数Dept. of Computer Science
Cornell University1