Large frontier models such as GPT-5 and Gemini have demonstrated remarkable performance in a wide range of health application benchmarks. However, underneath the seemingly promising results lie salient growth areas, especially in cutting-edge frontiers such as multimodal reasoning. Here we systematically apply and integrate a series of adversarial stress tests to assess the robustness of flagship models and health benchmarks. Our study reveals prevalent brittleness in the presence of simple adversarial transformations: leading systems can guess the correct answer even with key inputs removed yet may get confused by the slightest prompt alterations while fabricating convincing but flawed reasoning traces. Using clinician-guided rubrics, we demonstrate that popular health benchmarks vary widely in what they truly measure. Our study reveals considerable gaps between benchmark performance and the robustness evidence needed to support claims about multimodal medical reasoning in health applications.
Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths by matching reward distributions rather than collapsing to dominant modes. Recent work shows promise on math and code, but scaling GFlowNet-style RL to modern post-training pipelines remains difficult: as model size, rollout horizon, reward noise, and distributed-systems complexity grow together, a learned prompt-conditional partition function becomes a source of gradient instability and engineering overhead rather than a useful normalizer. Through systematic analysis, we find that the learned partition function, previously treated as essential, can be replaced by an in-batch Monte Carlo estimate computed from the rollout group already required for training. We propose GFlowRL, a streamlined GFlowNet-style RL algorithm that removes the auxiliary partition network entirely while preserving the reward-distribution-matching objective, completed by two stabilizers: importance-sampling correction for rollout/trainer drift and asymmetric flow-gap clipping for outlier residuals. GFlowRL exceeds all counterparts on math, code, and adversarial red-teaming benchmarks, reaching a Codeforces rating of 2048 at the 14B scale (within 25 Elo of o3-mini) and attaining the highest average ASR@1 on AdvBench and HarmBench, outperforming the previous SOTA multi-turn attacker in a regime where FlowRL, a prior GFlowNet-style method, diverges. The same recipe transfers to all evaluated MoE configurations up to 235B parameters, where FlowRL again fails to converge. To our knowledge, GFlowRL is the first GFlowNet-style RL algorithm to scale stably across both dense and sparse architectures. Code will be at: https://github.com/microsoft/gflowrl
Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, especially in cutting-edge frontiers such as multimodal reasoning. In this paper, we introduce a series of adversarial stress tests to systematically assess the robustness of flagship models and medical benchmarks. Our study reveals prevalent brittleness in the presence of simple adversarial transformations: leading systems can guess the right answer even with key inputs removed, yet may get confused by the slightest prompt alterations, while fabricating convincing yet flawed reasoning traces. Using clinician-guided rubrics, we demonstrate that popular medical benchmarks vary widely in what they truly measure. Our study reveals significant competency gaps of frontier AI in attaining real-world readiness for health applications. If we want AI to earn trust in healthcare, we must demand more than leaderboard wins and must hold AI systems accountable to ensure robustness, sound reasoning, and alignment with real medical demands.
As early career researchers supported by modest grants from the National Science Foundation, we modeled the expected effects of semiconductor scaling trends on computer architectures and built compiler infrastructure to optimize for inevitable heterogeneity in computer architecture. This early work inspired the vision for the UT-TRIPS project, which was subsequently funded by the U.S. Department of Defense, the University of Texas at Austin, multiple computer companies, and a private foundation. This partnership enabled our team to develop novel computer architectures and compiler technologies that demonstrated the viability of highly parallel chip designs that were composed of distributed processing and memory systems components. This article traces the history and ultimate impact of the project.
AutoGen is an open-source framework that allows developers to build LLM applications via multiple agents that can converse with each other to accomplish tasks. AutoGen agents are customizable, conversable, and can operate in various modes that employ combinations of LLMs, human inputs, and tools. Using AutoGen, developers can also flexibly define agent interaction behaviors. Both natural language and computer code can be used to program flexible conversation patterns for different applications. AutoGen serves as a generic infrastructure to build diverse applications of various complexities and LLM capacities. Empirical studies demonstrate the effectiveness of the framework in many example applications, with domains ranging from mathematics, coding, question answering, operations research, online decision-making, entertainment, etc.
Digital systems are growing in importance and computing hardware is growing more heterogeneous. Hardware design, however, remains laborious and expensive, in part due to the limitations of conventional hardware description languages (HDLs) like VHDL and Verilog. A longstanding research goal has been programming hardware like software, with high-level languages that can generate efficient hardware designs. This paper describes Kanagawa, a language that takes a new approach to combine the programmer productivity benefits of traditional High-Level Synthesis (HLS) approaches with the expressibility and hardware efficiency of Register-Transfer Level (RTL) design. The language's concise syntax, matched with a hardware design-friendly execution model, permits a relatively simple toolchain to map high-level code into efficient hardware implementations.
Narrow bit-width data formats are key to reducing the computational and storage costs of modern deep learning applications. This paper evaluates Microscaling (MX) data formats that combine a per-block scaling factor with narrow floating-point and integer types for individual elements. MX formats balance the competing needs of hardware efficiency, model accuracy, and user friction. Empirical results on over two dozen benchmarks demonstrate practicality of MX data formats as a drop-in replacement for baseline FP32 for AI inference and training with low user friction. We also show the first instance of training generative language models at sub-8-bit weights, activations, and gradients with minimal accuracy loss and no modifications to the training recipe.
This paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-point and block floating-point. MX utilizes multiple levels of quantization scaling with ultra-fine scaling factors based on shared microexponents in the hardware. The effectiveness of MX is demonstrated on real-world models including large-scale generative pretraining and inferencing, and production-scale recommendation systems.
Una primera tarjeta de interfaz de red en linea, NIC (104), para indexar los flujos de red (106), comprendiendo la primera NIC en linea (104): un primer controlador de acceso a los medios, MAC (150); un segundo MAC (152); hardware de procesamiento (124) configurado para proporcionar la transmision de paso (154, 156) de los paquetes de los flujos de red (106) mediante la transmision de los paquetes del primer MAC (150) recibidos por el segundo MAC (152) y mediante la transmision de los paquetes del segundo MAC (152) recibidos por el primer MAC (150); un primer modulo (130) configurado para implementar un protocolo de transporte ligero, LTP; y un segundo modulo (132) configurado para comunicarse con una segunda NIC en linea arbitraria por medio del primer modulo (130) especificando una direccion de red correspondiente a la segunda NIC en linea para permitir que el primer modulo (130) establezca una conexion LTP (108) con puntos finales en la primera NIC en linea (104) y en la segunda NIC en linea, en donde la primera NIC en linea (104) se conecta a un primer anfitrion y a la red de datos, y la segunda NIC en linea se conecta a un segundo anfitrion y a la red de datos, en donde el hardware de procesamiento se configura para proporcionar conectividad de red entre el primer anfitrion y el segundo anfitrion, realizando una transmision de paso de los paquetes recibidos que se ha determinado que no son paquetes LTP entre NIC, y proporcionar conectividad LTP entre NIC entre la primera NIC en linea (104) y la segunda NIC en linea, y en donde los paquetes recibidos por la primera NIC en linea que se determina que son paquetes LTP entre IC se consumen por la primera NIC en linea (104) y no se reenvian al primer o segundo anfitrion mediante la primera NIC en linea, y en donde los paquetes LTP entre NIC se originan por la primera o segunda NIC en linea.
Un metodo para restaurar la aceleracion del servicio para un servicio, el metodo que comprende: determinar que la aceleracion del servicio para el servicio esta operando incorrectamente, la aceleracion del servicio proporcionada por un grupo de componentes de aceleracion de interoperacion (1301, 1302, 1303, 1304) y por papeles (1311, 1312, 1313, 1314) en cada componente de aceleracion en el grupo de componentes de aceleracion de interoperacion enlazados entre si para componer un grafico (1333), en donde los componentes de aceleracion estan en un plano de aceleracion de hardware; detectar que el rendimiento degradado en un componente de aceleracion (1303), incluido en el grupo de componentes de aceleracion de interoperacion, hizo que la aceleracion del servicio operase incorrectamente, el componente de aceleracion asignado para proporcionar un papel (1313) que esta enlazado a uno o mas de otros papeles (1312, 1314) en el grafico; seleccionar un componente de aceleracion de sustitucion (1306) de entre uno o mas de otros componentes de aceleracion para proporcionar el papel; y restaurar la aceleracion del servicio para el servicio asignando el componente de aceleracion de sustitucion para proporcionar el papel y enlazar el papel (1313) proporcionado por el componente de aceleracion de sustitucion con uno o mas de otros papeles (1312, 1314).
Apparatus and methods for training neural networks based on a performance metric, including adjusting numerical precision and topology as training progresses are disclosed. In some examples, block floating-point formats having relatively lower accuracy are used during early stages of training. Accuracy of the floating-point format can be increased as training progresses based on a determined performance metric. In some examples, values for the neural network are transformed to normal precision floating-point formats. The performance metric can be determined based on entropy of values for the neural network, accuracy of the neural network, or by other suitable techniques. Accelerator hardware can be used to implement certain implementations, including hardware having direct support for block floating-point formats.
Un metodo para reconfigurar parcialmente un componente de aceleracion de hardware programado con un rol (1303, 1304) y una interfaz (1306) de red, el rol enlazado a uno o mas de: un rol de flujo descendente en el componente (1301) de aceleracion vecino del flujo descendente y un rol de flujo ascendente en el componente (1301) de aceleracion vecino del flujo ascendente para componer un grafico, el metodo que comprende: detectar una razon (1321) para cambiar el rol durante la monitorizacion del componente de aceleracion para comportamientos incorrectos; detener (1322) el rol (1303, 1304) que incluye instrucciones al menos de: el rol del flujo descendente y el rol del flujo ascendente para parar la recepcion de datos desde el rol; reconfigurar parcialmente el componente (1301) de aceleracion de hardware mediante la escritura de una imagen (1312) para el rol (1303, 1304) desde una ubicacion de almacen de imagenes al componente (1301) de aceleracion; mantener la interfaz (1306) de red como operativa durante la reconfiguracion parcial del componente (1301) de aceleracion para permitir que un segundo rol programado en el componente de aceleracion intercambie comunicacion de red con uno o mas otros roles compuestos en otro grafico en otros componentes de aceleracion; y activar (1326) el rol (1303, 1304) en uno de los componentes (1301) de aceleracion del flujo ascendente o del flujo descendente despues de que la reconfiguracion parcial del componente (1301) de aceleracion se complete que incluye notificar a al menos uno de: el rol del flujo descendente y el rol del flujo ascendente de que el rol esta operativo.
Low-power potential of mixed-signal design makes it an alluring option to accelerate Deep Neural Networks (DNNs). However, mixed-signal circuitry suffers from limited range for information encoding, susceptibility to noise, and Analog to Digital (A/D) conversion overheads. This paper aims to address these challenges by offering and leveraging the insight that a vector dot-product (the basic operation in DNNs) can be bit-partitioned into groups of spatially parallel low-bitwidth operations, and interleaved across multiple elements of the vectors. As such, the building blocks of our accelerator become a group of wide, yet low-bitwidth multiply-accumulate units that operate in the analog domain and share a single A/D converter. The low-bitwidth operation tackles the encoding range limitation and facilitates noise mitigation. Moreover, we utilize the switched-capacitor design for our bit-level reformulation of DNN operations. The proposed switched-capacitor circuitry performs the group multiplications in the charge domain and accumulates the results of the group in its capacitors over multiple cycles. The capacitive accumulation combined with wide bit-partitioned operations alleviate the need for A/D conversion per operation. With such mathematical reformulation and its switched-capacitor implementation, we define a 3D-stacked microarchitecture, dubbed BIHIWE.
A system for block floating point computation in a neural network receives a block floating point number comprising a mantissa portion. A bit-width of the block floating point number is reduced by decomposing the block floating point number into a plurality of numbers each having a mantissa portion with a bit-width that is smaller than a bit-width of the mantissa portion of the block floating point number. One or more dot product operations are performed separately on each of the plurality of numbers to obtain individual results, which are summed to generate a final dot product value. The final dot product value is used to implement the neural network. The reduced bit width computations allow higher precision mathematical operations to be performed on lower-precision processors with improved accuracy.
Perplexity scores are computed for training data samples during ANN training. Perplexity scores can be computed as a divergence between data defining a class associated with a current training data sample and a probability vector generated by the ANN model. Perplexity scores can alternately be computed by learning a probability density function ("PDF") fitting activation maps generated by an ANN model during training. A perplexity score can then be computed for a current training data sample by computing a probability for the current training data sample based on the PDF. If the perplexity score for a training data sample is lower than a threshold, the training data sample is removed from the training data set so that it will not be utilized for training during subsequent epochs. Training of the ANN model continues following the removal of training data samples from the training data set.
Technology related to training a neural network accelerator using mixed precision data formats is disclosed. In one example of the disclosed technology, a neural network accelerator is configured to accelerate a given layer of a multi-layer neural network. An input tensor for the given layer can be converted from a normal-precision floating-point format to a quantized-precision floating-point format. A tensor operation can be performed using the converted input tensor. A result of the tensor operation can be converted from the block floating-point format to the normal-precision floating-point format. The converted result can be used to generate an output tensor of the layer of the neural network, where the output tensor is in normal-precision floating-point format.
In this paper, we explore the limits of Microsoft Floating Point (MSFP), a new class of datatypes developed for production cloud-scale inferencing on custom hardware. Through the co-evolution of hardware design and algorithms, MSFP achieves accuracy comparable to or better than industry standards Bfloat16 and INT8 at 3x and 4x lower cost, respectively. MSFP incurs negligible impact to accuracy (
Technology related to incremental training of machine learning tools is disclosed. In one example of the disclosed technology, a method can include receiving operational parameters of a machine learning tool based on a primary set of training data. The machine learning tool can be a deep neural network. Input data can be applied to the machine learning tool to generate an output of the machine learning tool. A measure of prediction quality can be generated for the output of the machine learning tool. In response to determining the measure of prediction quality is below a threshold, incremental training of the operational parameters can be initiated using the input data as training data for the machine learning tool. Operational parameters of the machine learning tool can be updated based on the incremental training. The updated operational parameters can be stored.
Changkyu Kim合作论文数Google18
Paul V. Gratz合作论文数Texas A&M University9