Inference-time computation, analogous to human System 2 Thinking, has recently become popular for improving model performance. However, most existing approaches suffer from several limitations: they are modality-specific (e.g., working only in text), problem-specific (e.g., verifiable domains like math and coding), or require additional supervision/training on top of unsupervised pretraining (e.g., verifiers or verifiable rewards). In this paper, we ask the question “Is it possible to generalize these System 2 Thinking approaches, and develop models that learn to think solely from unsupervised learning?” We find the answer is yes, by learning to explicitly verify the compatibility between inputs and candidate-predictions, and then re-framing prediction problems as optimization with respect to this verifier. Specifically, we train Energy-Based Transformers (EBTs)---a new class of Energy-Based Models (EBMs)---to assign an energy value to every input and candidate-prediction, enabling predictions through energy minimization until convergence. To support this approach, we introduce several key techniques for stable and parallelizable training, which enable the emergence of strong System 2 Thinking capabilities and scalable EBMs. Across discrete and continuous modalities, we find EBTs outperform the Transformer++ approach, scaling up to 35% faster during pretraining, and improving inference-time performance by up to 29%. EBTs also surpass Diffusion Transformers on image denoising while requiring 99% fewer forward passes. Moreover, System 2 Thinking with EBTs yields larger performance gains on data that is farther out-of-distribution, and EBTs achieve better results than existing models on most downstream tasks despite achieving the same or worse pretraining performance, enabling EBTs to generalize better than existing approaches. Consequently, EBTs are a flexible and exciting new approach for scaling both the learning and thinking capabilities of models.
We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance of state-of-the-art proprietary models, e.g., DINOv2, CLIP, SigLIPv2, etc. Our approach is grounded in a transparent training pipeline inspired by Web-SSL and uses publicly available data: ImageNet-21K and a subset of ReLAION-2B. Beyond model release, we tackle critical limitations in SSL clustering methods. While modern models rely on assigning image features to large codebooks via clustering algorithms like Sinkhorn-Knopp, they fail to account for the inherent ambiguity in clustering semantics. To address this, we introduce a parameter-efficient, multi-head clustering projector based on nested Matryoshka representations. This design progressively refines features into increasingly fine-grained clusters without increasing the model size, enabling both performance and memory efficiency. Additionally, we propose a novel positional disentanglement strategy that explicitly removes positional biases from dense representations, thereby improving the encoding of semantic content. This leads to consistent gains on several downstream benchmarks, demonstrating the utility of cleaner feature spaces. Our contributions establish a new standard for transparent, high-performance vision models and open a path toward more reproducible and generalizable foundation models for the broader AI community.
We investigate whether synthetic question-answer (QA) data generated by large language models (LLMs) can serve as an effective proxy for human-labeled benchmarks when the latter is unavailable. We assess the reliability of synthetic benchmarks across two experiments: one varying retriever parameters while keeping the generator fixed, and another varying the generator with fixed retriever parameters. Across four datasets, of which two open-domain and two proprietary, we find that synthetic benchmarks reliably rank the RAGs varying in terms of retriever configuration, aligning well with human-labeled benchmark baselines. However, they do not consistently produce reliable RAG rankings when comparing generator architectures. The breakdown possibly arises from a combination of task mismatch between the synthetic and human benchmarks, and stylistic bias favoring certain generators.
Abstract Fentanyl is frequently used in the intensive care unit for procedural sedation, ventilatory synchrony, and analgesia. While common side effects of synthetic opioids are well recognized, a rare and underreported complication is fentanyl-induced chest wall rigidity, also known as "wooden chest syndrome."The largest published case series in critically ill adults identified 42 cases of suspected fentanyl-induced rigid chest syndrome but did not provide a denominator of all fentanyl-exposed ICU patients1, so a precise incidence rate is not currently available. We present a case of a patient with acute liver failure complicated by ventilator dependence who, following initiation of fentanyl, had the new onset of ventilator dyssynchrony with obvious chest wall and abdominal rigidity. A 59-year-old male with decompensated cirrhosis (decompensations included esophageal varices and hepatic encephalopathy on lactulose), and type 2 diabetes mellitus was transferred from an outside facility for liver transplant evaluation. He initially presented for acute liver failure and septic shock, with his course complicated by oliguric acute kidney injury requiring continuous renal replacement therapy. Initial ventilator settings onarrival to our hospital included pressure support of 5 cmH2O, PEEP of 5 cmH2O, and FiO2 of 35% while sedated with dexmedetomidine infusion. Fentanyl was initiated on the day of admission, starting at 50 mcg per hour and titrated to 100 mcg per hour within 15 minutes. Shortly thereafter, the patient developed ventilator dyssynchrony and intermittent apnea, with tidal volumes falling below 100 mL and elevated peak pressures. On exam, he was normotensive and unresponsive to noxious stimuli; he demonstrated vigorous chest wall and abdominal muscle activation, particularly during expiration. Despite ventilator and circuit adjustments, ventilation remained ineffective. ETCO2 rose from 27 to 40 mmHg. Point-of-care ultrasonography and chest radiography showed no acute pulmonary pathology, and bronchoscopy revealed diffusely collapsible airways but no other abnormalities. His medication list was examined for possible contributions to other common causes of rigidity in the ICU—including serotonin syndrome, neuroleptic malignant syndrome, malignant hyperthermia, and baclofen withdrawal— and none were identified. Fentanyl was discontinued and propofol was initiated, resulting in improved synchrony. Fentanyl was later reintroduced at lower doses without recurrence; naloxone was not administered as a reversal agent. This case highlights the importance of recognizing fentanyl-induced chest wall rigidity as a potential causeof ventilator dyssynchrony in critically ill patients. 1. Tammen AJ, Brescia D, Jonas D, Hodges JL, Keith P. Fentanyl-Induced Rigid Chest Syndrome in Critically IllPatients. J Intensive Care Med. 2023;38(2):196-201. doi:10.1177/08850666221115635 This abstract is funded by: None