We present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos. We build a data engine, which improves model and data via user interaction, to collect the largest video segmentation dataset to date. Our model is a simple transformer architecture with streaming memory for real-time video processing. SAM 2 trained on our data provides strong performance across a wide range of tasks. In video segmentation, we observe better accuracy, using 3x fewer interactions than prior approaches. In image segmentation, our model is more accurate and 6x faster than the Segment Anything Model (SAM). We believe that our data, model, and insights will serve as a significant milestone for video segmentation and related perception tasks. We are releasing our main model, the dataset, an interactive demo and code.
We introduce the Segment Anything (SA) project: a new task, model, and dataset for image segmentation. Using our efficient model in a data collection loop, we built the largest segmentation dataset to date (by far), with over 1 billion masks on 11M licensed and privacy respecting images. The model is designed and trained to be promptable, so it can transfer zero-shot to new image distributions and tasks. We evaluate its capabilities on numerous tasks and find that its zero-shot performance is impressive – often competitive with or even superior to prior fully supervised results. We are releasing the Segment Anything Model (SAM) and corresponding dataset (SA-1B) of 1B masks and 11M images at segment-anything.com to foster research into foundation models for computer vision. We recommend reading the full paper at: arxiv.org/abs/2304.02643.
Invariance to a broad array of image corruptions, such as warping, noise, or color shifts, is an important aspect of building robust models in computer vision. Recently, several new data augmentations have been proposed that significantly improve performance on ImageNet-C, a benchmark of such corruptions. However, there is still a lack of basic understanding on the relationship between data augmentations and test-time corruptions. To this end, we develop a feature space for image transforms, and then use a new measure in this space between augmentations and corruptions called the Minimal Sample Distance to demonstrate a strong correlation between similarity and performance. We then investigate recent data augmentations and observe a significant degradation in corruption robustness when the test-time corruptions are sampled to be perceptually dissimilar from ImageNet-C in this feature space. Our results suggest that test error can be improved by training on perceptually similar augmentations, and data augmentations may not generalize well beyond the existing benchmark. We hope our results and tools will allow for more robust progress towards improving robustness to image corruptions. We provide code at https://github.com/facebookresearch/augmentation-corruption.
Vision transformer (ViT) models exhibit substandard optimizability. In particular, they are sensitive to the choice of optimizer (AdamW vs. SGD), optimizer hyperparameters, and training schedule length. In comparison, modern convolutional neural networks are easier to optimize. Why is this the case? In this work, we conjecture that the issue lies with the patchify stem of ViT models, which is implemented by a stride-p p×p convolution (p = 16 by default) applied to the input image. This large-kernel plus large-stride convolution runs counter to typical design choices of convolutional layers in neural networks. To test whether this atypical design choice causes an issue, we analyze the optimization behavior of ViT models with their original patchify stem versus a simple counterpart where we replace the ViT stem by a small number of stacked stride-two 3×3 convolutions. While the vast majority of computation in the two ViT designs is identical, we find that this small change in early visual processing results in markedly different training behavior in terms of the sensitivity to optimization settings as well as the final model accuracy. Using a convolutional stem in ViT dramatically increases optimization stability and also improves peak performance (by ∼1-2% top-1 accuracy on ImageNet-1k), while maintaining flops and runtime. The improvement can be observed across the wide spectrum of model complexities (from 1G to 36G flops) and dataset scales (from ImageNet-1k to ImageNet-21k). These findings lead us to recommend using a standard, lightweight convolutional stem for ViT models in this regime as a more robust architectural choice compared to the original ViT model design.
The BFSS matrix model provides an example of gauge-theory / gravity duality where the gauge theory is a model of ordinary quantum mechanics with no spatial subsystems. If there exists a general connection between areas and entropies in this model similar to the Ryu-Takayanagi formula, the entropies must be more general than the usual subsystem entanglement entropies. In this note, we first investigate the extremal surfaces in the geometries dual to the BFSS model at zero and finite temperature. We describe a method to associate regulated areas to these surfaces and calculate the areas explicitly for a family of surfaces preserving SO(8) symmetry, both at zero and finite temperature. We then discuss possible entropic quantities in the matrix model that could be dual to these regulated areas.
We consider CFT states defined by adding nonlocal multi-trace sources to the Euclidean path integral defining the vacuum state. For holographic theories, we argue that these states correspond to states in the gravitational theory with a good semiclassical description but with a more general structure of bulk entanglement than states defined from single-trace sources. We show that at leading order in large N , the entanglement entropies for any such state are precisely the same as those of another state defined by appropriate single-trace effective sources; thus, if the leading order entanglement entropies are geometrical for the single-trace states of a CFT, they are geometrical for all the multi-trace states as well. Next, we consider the perturbative calculation of 1/N corrections to the CFT entanglement entropies, demonstrating that these show qualitatively different features, including non-analyticity in the sources and/or divergences in the naive perturbative expansion. These features are consistent with the expectation that the 1/N corrections include contributions from bulk entanglement on the gravity side. Finally, we investigate the dynamical constraints on the bulk geometry and the quantum state of the bulk fields which must be satisfied so that the entropies can be reproduced via the quantum-corrected Ryu-Takayanagi formula.
The gravitational Gauss law requires any addition of energy to be accompanied by the addition of gravitational flux. The possible configurations of this flux for a given source may be called gravitational hair, and several recent works discuss gravitational observables ('gravitational Wilson lines') which create this hair in highly collimated 'combed' configurations. We construct and analyze time-symmetric classical solutions of 2 + 1 Einstein-Hilbert gravity such as might be created by smeared versions of such operators. We focus on the AdS(3) case, where this hair is characterized by the profile of the boundary stress tensor; the desired solutions are those where the boundary stress tensor at initial time t = 0 agrees precisely with its vacuum value outside an angular interval [- alpha, alpha]. At linear order in source strength the energy is independent of the combing parameter alpha, but nonlinearities cause the full energy to diverge as alpha -> 0. In general, solutions with combed gravitational flux also suffer from what we call displacement from their naive location. For weak sources and large alpha one may set the displacement to zero by further increasing the energy, though for strong sources and small alpha we find no preferred notion of a zero-displacement solution. In the latter case we conclude that naively expected gravitational Wilson lines do not exist. In the zero-displacement case, taking the AdS scale l to infinity gives finite-energy flux-directed solutions that may be called asymptotically flat.
We examine the defect gauge theory on two perpendicular D3-branes with a 1+1 dimensional intersection, consisting of U(1) fields on the D3-branes and charged hypermultiplets on the intersection. We argue that this gauge theory must have a magnetically charged soliton corresponding to the D-string stretched between the branes. We show that the hypermultiplets actually source magnetic as well as electric fields. The magnetic charges are confined if the hypermultiplet action is canonical, but considerations of periodicity of the hypermultiplet space in string theory imply a nontrivial Gibbons-Hawking metric, and we show that there is then the expected magnetic kink solution. The hypermultiplet metric has a singularity, which we argue must be resolved by embedding in the full string theory. Another interesting feature is that the classical field equations have logarithmic divergences at the intersection, which lead to a classical renormalization group flow in the action.
We develop the point of view that brane actions should be understood in the context of effective field theory, and that this is the correct way to treat classical as well as loop divergences. We illustrate this idea in a simple model. We then consider the implications for the dynamics of antibranes in flux backgrounds, focusing on the simplest case of a single antibrane. We argue that that effective field theory gives a valid description of the antibrane, and that there is no instability in this approximation. Further, conservation laws exclude possible nonperturbative decays, aside from the well-known NS5-brane instanton.
Recently, Almheiri, Dong, and Harlow have argued that the localization of bulk information in a boundary dual should be understood in terms of quantum error correction. We show that this structure appears naturally when the gauge invariance of the boundary theory is incorporated. This provides a new understanding of the nonuniqueness of the bulk fields (precursors). It suggests a close connection between gauge invariance and the emergence of spacetime.
We revisit the derivation of higher spin bulk theory using the renormalization group in the dual field theory. We argue that existing proposals have problems already at the level of linearized perturbations on AdS. This is due to the form of the cutoff, which must act on bilinears in the fundamental fields rather than on the fields themselves. For the light-cone collective field theory, we show that the RG produces the correct linearized perturbations. We relate this to the precursor formula of de Mello Koch, Jevicki, Jin and Rodrigues, and we also elaborate on that result. The covariant RG and bulk interactions remain problems for the future.
Strongly modified h gamma gamma and hgg couplings indicate new electroweak and color mediators, respectively, with a light mass and a significant coupling to the Higgs boson. We point out the Higgs boson could have a significant decay width into the mediators. This represents one new class of exotic Higgs decay possibilities: off-shell exotic Higgs decays. We then propose uncovering the hidden new physics through such exotic decays. A great advantage of this strategy is that we can directly probe the couplings between the Higgs boson and the mediators, which is hard to achieve by using other methods. Focusing on the electroweak mediators, we study a simplified model using as an example final states with tau leptons and neutrinos. Because one of the mediators is off shell and its decay products are extremely soft, it is challenging to make a discovery at the Large Hadron Collider. A Higgs factory such as the International Linear Collider, however, could serve as a discovery machine for such exotic Higgs decays even in an early stage.