Multimodal large language models (MLLMs) often suffer from perceptual impairments under extended reasoning modes, particularly in visual question answering (VQA) tasks. We identify attention dispersion as the underlying cause: during multi-step reasoning, model's visual attention becomes scattered and drifts away from question-relevant regions, effectively “losing focus” on the visual input. To better understand this phenomenon, we analyze the attention maps of MLLMs and observe that reasoning prompts significantly reduce attention to regions critical for answering the question. We further find a strong correlation between model’s overall attention on image tokens and the spatial dispersiveness of model’s attention within the image. Leveraging this insight, we propose a training-free Visual Region-Guided Attention (VRGA) framework that selects visual heads based on an entropy–focus criterion and reweights their attention, effectively guiding the model to focus on question-relevant regions during reasoning. Extensive experiments on vision-language benchmarks demonstrate that our method effectively alleviates perceptual degradation, leading to improvements in visual grounding and reasoning accuracy, while offering interpretable insights into how MLLMs process visual information.
”Thinking with Images” has emerged as an effective paradigm for fine-grained visual reasoning: by explicitly zooming into relevant regions and reasoning over crops, models can access local evidence that is difficult to recover from a single global image. However, this benefit comes with redundant tool invocations and longer inference traces. Moreover, when such behaviors are learned mainly from outcome reward, the resulting intermediate crops or visual cues can be noisy or fail to faithfully capture task-relevant visual evidence. In this work, we ask whether the reasoning benefits of ”Thinking with Images” can be internalized through Thinking with Imagination: an internal process that decides where to look and imagines what visual cues closer inspection would reveal without actually invoking tools. We propose Imagine-OPD, an on-policy self-distillation framework in which a teacher plays the role of a ”Thinking with Images” reasoner during training: it receives privileged zoomed evidence views derived from annotated regions, and supervises the model's own imagination reasoning trajectories. Imagine-OPD does not require an external teacher or high-quality imagination demonstrations. Experiments on vision-centric benchmarks show that Imagine-OPD achieves the best average performance among compared models while significantly reducing inference overhead compared with ”Thinking with Images” methods.
GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging for the open-source community. Existing vision-language models rely on external tools for speech processing, while speech-language models still suffer from limited or totally without vision-understanding capabilities. To address this gap, we propose the EMOVA (EMotionally Omni-present Voice Assistant), to enable Large Language Models with end-to-end speech abilities while maintaining the leading vision-language performance. With a semantic-acoustic disentangled speech tokenizer, we surprisingly notice that omni-modal alignment can further enhance vision-language and speech abilities compared with the bi-modal aligned counterparts. Moreover, a lightweight style module is introduced for the flexible speech style controls including emotions and pitches. For the first time, EMOVA achieves state-of-the-art performance on both the vision-language and speech benchmarks, and meanwhile, supporting omni-modal spoken dialogue with vivid emotions.
Whole slide image (WSI) classification is one of the important fields of digital pathology, and is generally solved as a weakly supervised learning problem by adopting multiple instance learning (MIL). However, a common but crucial challenge faced by existing MIL models is their inability to provide convincing explanations that can win the trust of pathologists and be applied to clinical diagnosis. In addition, most attention-based MIL models use attention scores to represent the importance of each patch in the WSI rather than inferring patch probabilities directly, which does not accurately detect the critical patches. To address these two challenges, we propose a ProtoTree based MIL model for WSI classification, called ProtoTree-MIL, where ProtoTree is an interpretable model that combines the advantages of prototype-learning and decision tree. ProtoTree-MIL not only explains why some patches are important for the final prediction through prototype-learning, but also provides global and local explanation through decision tree. We also propose a method to infer patch probabilities and measure their importance under the framework of ProtoTree-MIL. By conducting various experiments on three public WSI datasets, Camelyon16, TCGA-NSCLC, and TCGA-RCC, we demonstrate that our proposed ProtoTree-MIL can achieve a competitive performance to the state-of-the-art MIL models but provide more persuasive explanations than them. Explicitly generating patch probabilities also makes ProtoTree-MIL more accurate to detect the key patches than other attention-based MIL models. Specially, by evaluating our model on a real clinical gastritis and gastric cancer dataset, we show the explanations provided by ProtoTree-MIL are significant and faithful.
End-to-end architectures in autonomous driving (AD) face a significant challenge in interpretability, impeding human-AI trust. Human-friendly natural language has been explored for tasks such as driving explanation and 3D captioning. However, previous works primarily focused on the paradigm of declarative interpretability, where the natural language interpretations are not grounded in the intermediate outputs of AD systems, making the interpretations only declarative. In contrast, aligned interpretability establishes a connection between language and the intermediate outputs of AD systems. Here we introduce Hint-AD, an integrated AD-language system that generates language aligned with the holistic perception-prediction-planning outputs of the AD model. By incorporating the intermediate outputs and a holistic token mixer sub-network for effective feature adaptation, Hint-AD achieves desirable accuracy, achieving state-of-the-art results in driving language tasks including driving explanation, 3D dense captioning, and command prediction. To facilitate further study on driving explanation task on nuScenes, we also introduce a human-labeled dataset, Nu-X. Codes, dataset, and models will be publicly available.
The need to explain the output of a deep neural network classifier is now widely recognized. While previous methods typically explain a single class in the output, we advocate explaining the whole output, which is a probability distribution over multiple classes. A whole-output explanation can help a human user gain an overall understanding of model behaviour instead of only one aspect of it. It can also provide a natural framework where one can examine the evidence used to discriminate between competing classes, and thereby obtain contrastive explanations. In this paper, we propose a contrastive whole-output explanation (CWOX) method for image classification, and evaluate it using quantitative metrics and through human subject studies. The source code of CWOX is available at https://github.com/vaynexie/CWOX.
Despite the popularity of Vision Transformers (ViTs) and eXplainable AI (XAI), only a few explanation methods have been designed specially for ViTs thus far. They mostly use attention weights of the [CLS] token on patch embeddings and often produce unsatisfactory saliency maps. This paper proposes a novel method for explaining ViTs called ViT-CX. It is based on patch embeddings, rather than attentions paid to them, and their causal impacts on the model output. Other characteristics of ViTs such as causal overdetermination are considered in the design of ViT-CX. The empirical results show that ViT-CX produces more meaningful saliency maps and does a better job revealing all important evidence for the predictions than previous methods. The explanation generated by ViT-CX also shows significantly better faithfulness to the model. The codes and appendix are available at https://github.com/vaynexie/CausalX-ViT.
Recently, neural networks based models have been widely used for recommender systems (RS). Unfortunately, the existing neural network based RS solutions are often treated as black-boxes, which gain little trust and confidence from users. Thus, there is an increasing demand of explainability. Several explainable recommendation methods have been introduced to RS. However, there is a trade-off between explainability and performance among these methods. In this paper, we propose a novel framework, the Tower Bridge Net (TB-Net), using the proposed bidirectional embedding propagation approach to achieve both superior recommendation and explainability performances. Extensive validation on three public datasets shows that the performance of TB-Net dominates the state-of-the-art models. We quantitatively evaluate the explainability by using numerical metrics and experimentally prove that TB-Net achieves a significant improvement on explainability compared with existing methods. More importantly, TB-Net has been deployed and offers explainable recommendation service for the largest bank in China, Industrial and Commercial Bank of China Limited (ICBC). Results on a billion-scale dataset (1.2 billion nodes and edges) from ICBC show that TB-Net can provide both accurate recommendations and semantic explanations, and is very effective and deployable in practice.
We are witnessing a fast development of Artificial Intelligence (AI), but it becomes dramatically challenging to explain AI models in the past decade. “Explanation” has a flexible philosophical concept of “satisfying the subjective curiosity for causal information”, driving a wide spectrum of methods being invented and/or adapted from many aspects and communities, including machine learning, visual analytics, human-computer interaction and so on. Nevertheless, from the view-point of data and knowledge engineering (DKE), a best explaining practice that is cost-effective in terms of extra intelligence acquisition should exploit the causal information and explaining scenarios which is hidden richly in the data itself. In the past several years, there are plenty of works contributing in this line but there is a lack of a clear taxonomy and systematic review of the current effort. To this end, we propose this survey, reviewing and taxonomizing existing efforts from the view-point of DKE, summarizing their contribution, technical essence and comparative characteristics. Specifically, we categorize methods into data-driven methods where explanation comes from the task-related data, and knowledge-aware methods where extraneous knowledge is incorporated. Furthermore, in the light of practice, we provide survey of state-of-art evaluation metrics and deployed explanation applications in industrial practice.
Some examples are easier for humans to classify than others. The same should be true for deep neural networks (DNNs). We use the term example perplexity to refer to the level of difficulty of classifying an example. In this paper, we propose a method to measure the perplexity of an example and investigate what factors contribute to high example perplexity. The related codes and resources are available at https://github.com/vaynexie/Example-Perplexity.
Deep learning has shown powerful performances in many fields, however its black-box nature hinders its further applications. In response, explainable artificial intelligence emerges, aiming to explain the predictions and behaviors of deep learning models. Among many explanation methods, counterfactual explanation has been identified as one of the best methods due to its resemblance to human cognitive process: to deliver an explanation by constructing a contrastive situation so that human may interpret the underlying mechanism by cognitively demonstrating the difference. In this tutorial, we will introduce the cognitive concept and characteristics of counterfactual explanation, its computational form, mainstream methods, and various adaptation in terms of different explanation settings. In addition, we will demonstrate several typical use cases of counterfactual explanations in popular research areas. Finally, in light of practice, we outline the potential applications of counterfactual explanations like data augmentation or conversation system. We hope this tutorial can help the participants get an overview sense of counterfactual explanations.
It has been long debated that eXplainable AI (XAI) is an important technology for model and data exploration, validation, and debugging. To deploy XAI into actual systems, an executable and comprehensive evaluation of the quality of generated explanation is highly in demand. In this paper, we briefly summarize the status quo of the quantitative metrics of different properties of XAI including evaluation on faithfulness, localization, sensitivity check, and stability. With an exhaustive experimental study based on them, we conclude that among all the typical methods we compare, no single explanation method dominates others in all metrics. Nonetheless, Gradient-weighted Class Activation Mapping (Grad-CAM) and Randomly Input Sampling for Explanation (RISE) perform fairly well in most of the metrics. We further present a novel utilization of the evaluation results to diagnose the classification bases for models. Hopefully, this valuable work could serve as a guide for future research.
In this paper we extend the Falicov-Kimball model (FKM) to the case where the "quasiparticles" entering the FKM are not "ordinary" fermions. As an example we first discuss how the FKM can be generalized to the case with spin-dependent hopping. Afterward, we discuss several cases where the quasiparticles entering the FKM are Majorana fermions [extended Majorana-Falicov-Kimball model (MFKM)]. Two examples of the extended MFKM are discussed in detail: (i) a p-wave BCS superconductor on a bipartite lattice and (ii) a BCS-Anderson model. We also discuss the most general forms of extended MFKM, including a brief discussion of the case where the Majorana fermions represent spins but not real fermion particles.
In this paper we extend the Falicov-Kimball model (FKM) to the case where the quasi-particles entering the FKM are not ordinary fermions. As an example we first discuss how the FKM can be generalized to the case with spin-dependent hopping. Afterward we discuss several cases where the quasi-particles entering the FKM are Majorana fermions (Majorana-Falicov-Kimball Model (MFKM). Two examples of MFKM are discussed in detail: (i) a $p$-wave BCS superconductor on a bipartite lattice and (ii) a BCS-Anderson model. We also discuss the most general forms of MFKM, including a brief discussion on the case where the Majorana fermions represent spins, but not real fermion particles.
We introduce in this Letter an exact solvable BCS-Hubbard model in arbitrary dimensions. The model describes a p-wave BCS superconductor with equal spin pairing moving on a bipartite (cubic, square, etc.) lattice with on-site Hubbard interaction U. We show that the model becomes exactly solvable for arbitrary U when the BCS pairing amplitude Δ equals the hopping amplitude t. The nature of the solution is described in detail in this Letter. The construction of the exact solution is parallel to the exactly solvable Kitaev honeycomb model for S=1/2 quantum spins and can be viewed as a generalization of Kitaev's construction to S=1/2 interacting lattice fermions. The BCS-Hubbard model discussed in this Letter is just an example of a large class of exactly solvable lattice fermion models that can be constructed similarly.
In this paper we study the properties of cold bosons in a two-dimensional optical lattice system where Bose-condensation occurs at a momentum point k with non-zero k-space Berry curvature. By combining results from both analytic and numerical approaches, we show that the boson system carries non-universal, temperature dependent equilibrium angular momentum and edge current at low temperatures.
A model for the accurate measurement of low frequency sound spectrum is put forward and a simplified experimental device is built. Based on the principle of mechanical resonance, the natural frequency of the system is controlled, which realizes the function of both measuring frequency spectrum and monitoring noises at low frequency. The device is able to measure the acoustic source at a single frequency below 100 Hz as well as to analyze the spectra of acoustic sources at mixed frequencies, with a relatively high resolution of 1 Hz beyond 40 Hz.