
Window-based attention mechanisms have been introduced to alleviate the excessive computational cost inherent in global attention mechanisms. In this paper, we propose a novel architecture named FwNet-ECA, which integrates the Fourier transform with learnable weight matrices to enhance spectral features of images. By performing filter enhancement after window-based attention, our method establishes a global receptive field, thereby overcoming the limited receptive field typically associated with windowed attention. Furthermore, we incorporate the existing Efficient Channel Attention module to improve interchannel information exchange. Unlike approaches that rely on physical window shifting, our method leverages frequency-domain enhancement to implicitly connect spatial regions. We evaluate our model on the iCartoonFace dataset and demonstrate competitive performance on fine-grained classification benchmarks. Experimental results show that, compared to shift-based window methods, our model achieves comparable accuracy with fewer parameters and lower computational overhead. Moreover, visualization analyses clearly indicate that the filter enhancement technique is particularly effective in the shallow layers of the network, where feature maps are relatively large. This work presents an effective solution to the limited receptive field problem in window-based attention mechanisms. The code is publicly available at https://github.com/qingxiaoli/FwNet-ECA .
Fine-grained visual classification (FGVC) is challenged by subtle interclass differences, which lead to both out-of-group and in-group errors, particularly when label hierarchies are only partially available. To leverage hierarchical labels for enhanced feature discriminability, this paper introduces two novel loss functions. First, a hierarchically discriminative loss () linearly combines fine-grained subcategory features with their ancestral superclass features. This fusion, enhanced by cross-channel max pooling and channel-wise dropout, strengthens feature discriminability to suppress out-of-group errors. Second, an in-group regularization loss () addresses visually similar categories within the same superclass. It introduces controlled interference by mixing target class features with randomly selected features from other in-group classes. This process increases learning difficulty without altering labels, thereby mitigating sample-specific overfitting and reducing in-group errors. Our proposed loss functions are easy to integrate into existing frameworks, require no extra annotations or complex network modifications, and support end-to-end training. Extensive experiments on five benchmark datasets, including BreakHis, CUB-200-2011, Butterfly-200, FGVC-Aircraft, and Stanford Cars, demonstrate that our method consistently improves top-1 accuracy, as well as weighted average precision and mean average precision. Furthermore, we also observe consistent improvements in -score. The performance gains are particularly significant in scenarios with low fine-grained annotation ratios.
In order to provide personalized recommendations, decision support methods need to take into account user preferences in the production of their outcomes. This is a particularly relevant issue in the case of argumentative approaches to decision support, which have often focused on the construction of user-agnostic arguments concerning the various options at stake. In particular, while gradual bipolar argumentation (GBA) has been successfully adopted as a formal basis for the realization of decision support in a variety of application domains, none of these previous works involved personalization aspects and, hence, the study of techniques to integrate user-provided preferences into GBA turns out to be an open research question. In this paper, we provide an initial contribution to this investigative direction from both a theoretical and a practical perspective. On the theoretical side, we introduce a property of local coherence to characterize the expected effects of user preferences on argument strength assessment in GBA and provide results concerning the relations between local coherence and global behaviors. On the practical side, we illustrate a preliminary experimentation in the context of a GBA-based review aggregation system extended with the handling of user preferences, which allows us to draw some considerations on the opportunities and challenges of putting the proposed approach into practice.
The challenge of reconstructing three-dimensional (3D) models from images lies in how to infer a complete 3D structure with detailed geometry from two-dimensional (2D) images. However, single-view reconstruction requires only one image to reason about the 3D shape, demonstrating significant application potential. Yet, most existing methods rely on fully supervised learning approaches, which demand large amounts of labeled data. To alleviate this issue, semi-supervised learning strategies have been proposed to reduce the dependence on annotated data, offering a more efficient solution for single-view 3D reconstruction. We propose a semi-supervised single-view 3D point cloud reconstruction framework that employs a teacher-student paradigm to leverage limited labeled data together with abundant unlabeled samples. To address the modality heterogeneity between 2D images and 3D point clouds, a Heterogeneous Feature Attention Mechanism is designed to align cross-modal features and embed image cues into 3D spatial structures, preserving both geometric and appearance information. Moreover, a Self-Attention Decoder captures global dependencies and salient regions, enabling fine-grained structural recovery. Our model demonstrates outstanding performance on both the ShapeNet dataset and the Pix3D dataset, achieving an L1-Chamfer distance (CD) value of 5.60 on ShapeNet and an L1-CD value of 6.29 on Pix3D. Furthermore, rigorous ablation studies provide additional confirmation of the remarkable effectiveness of our approach.
This study examines how artificial intelligence (AI) anthropomorphism influences user trust and decision logic in digital environments. Drawing on theories of social cognition, trust formation, and dual-process decision-making, a moderated mediation model is developed to explain how anthropomorphic design features (e.g., human-like voice, avatars, and expressions) affect users’ cognitive processing. A controlled experimental design, complemented by simulation-based validation, provides robust evidence for the proposed relationships. The findings reveal that anthropomorphic cues significantly enhance user trust, which in turn mediates their impact on decision logic. However, this effect is contingent on contextual factors: user experience strengthens the anthropomorphism–trust pathway, while task criticality attenuates the influence of trust on decision-making. These results underscore that anthropomorphism functions as a social heuristic, effective in routine or low-risk tasks but less influential in high-stakes contexts where analytic reasoning dominates. The study contributes to the fields of human–computer interaction and information systems by clarifying when and how anthropomorphism fosters trust and shapes cognition. Practically, it offers design guidelines for adaptive AI systems that dynamically adjust anthropomorphic cues to balance engagement, trust calibration, and decision quality.
Diffusion models have demonstrated impressive performance in text-to-image generation and image editing. However, in instruction-based image editing, they often encounter two challenges: (1) inaccurate localization of the editing targets and (2) unintended modifications in nontarget regions. These issues stem from the global processing of diffusion models due to attention mechanisms. To address these limitations, we conduct a systematic analysis of attention maps under editing instructions and design localization instructions to obtain the desired attention. We propose Instruction Attention Maps (IAM)-Edit, a localized image editing framework that explicitly decouples an editing pipeline into two stages: region localization followed by region-aware editing. Specifically, to localize the editing region, a mask is generated by clustering patches of self-attention maps and combining them with the focal points of cross-attention maps under the editing instruction. To preserve nonediting regions, we apply an attention modulation method that adjusts cross-attention weights at each denoising step based on the generated mask, enabling the denoising process to focus on the editing region. Experiments show that IAM-Edit outperforms state-of-the-art methods both qualitatively and quantitatively.
Representation learning is critical for multimodal methods; traditional consistency-based multimodal methods always constrain the disagreements among different modality embeddings or predictions as an extra regularization. However, these methods may appear to cause performance degeneration in open environments. This is mainly attributed to the interference of asymmetric information, that is, different modality information exists divergence, whereas consistency regularization prefers to simply minimize the divergence rather than optimal classifiers. Therefore, it is unsafe to directly use consistency regularization. To this end, we propose modality-specific subspace learning (MSSL). It learns the modality-specific subspace representations by treating modality divergence and consistency separately. In particular, MSSL is a semi-supervised framework that maps different modality feature embeddings into shared and independent subspaces. The shared subspace applies reliable consistency regularization by measuring intermodality structural similarities. The independent subspace uses a discriminative modality-separation network to emphasize modality complementary information. Finally, labeled instances from different modalities are classified with weighted predictions over concatenated embeddings. Consequently, MSSL improves both the single modal and ensemble classification results and acquires more robust mapping among different modalities. Empirical studies show the superior performance of MSSL on real-world datasets.
With the rapid development of deep learning technology, spiking neural networks (SNNs) and transformer models have made significant progress in computer vision. This article proposes an image recognition algorithm model that combines SNN and transformer networks and validates it using the Modified National Institute of Standards and Technology (MNIST), Fashion-MNIST, and CIFAR-10 datasets. During preprocessing, Poisson encoding transforms pixel intensities into spike sequences. In the hybrid network architecture constructed, the SNN is embedded into the transformer encoder, and its output features are transformed before being passed to the transformer decoder. The spike neural network is constructed with LIF (leaky integrate-and-fire) neurons to build the network layer, and feature learning and extraction are carried out using the spike-timing-dependent plasticity (STDP) learning rule. The decoder recognizes the features extracted by the encoder, and the cross-entropy loss function is used to compute the difference between the output results and the true labels. During the training phase, the encoder and decoder parameters are updated simultaneously through backpropagation. The model is evaluated using independent test sets and validation sets. The experimental results show that after 100 iterations, the hybrid model using only the LIF neuron network achieves validation accuracies of 0.9803, 0.9262, and 0.7217 on the MNIST, Fashion-MNIST, and CIFAR-10 datasets, respectively. Further research reveals that incorporating the STDP learning rule improves the model’s validation accuracy to 0.9977, 0.9315, and 0.7461, demonstrating that this approach can enhance recognition accuracy to some extent, offering new insights for further research in image recognition.
Conditional syntax splitting for inductive inference from conditional belief bases has been proposed as a generalization of syntax splitting, which also covers cases where the conditionals in the subbases may share some atoms. While p-entailment and system Z fail to satisfy conditional syntax splitting, two other inductive inference operators, lexicographic inference and system W, have been shown to satisfy this property for reasoning from strongly consistent belief bases. In this article, we introduce the concept of conditional semantic splitting of a belief base. For both syntax splitting and semantic splitting, we take not only strongly consistent belief bases into account, but also belief bases that are only weakly consistent, enforcing some worlds to be fully infeasible. We show that c-representations satisfy a core postulate relating conditional splittings on the syntax and the semantic level. Based on these findings, we investigate conditional syntax splitting for nonmonotonic inference with c-representations. Regarding single c-representations, we utilize the concept of selection strategies and show that a straightforward property of the selection strategy leads to inference operators satisfying conditional syntax splittings. Furthermore, we show that c-inference, taking all c-representations of a belief base into account fully complies with conditional syntax splitting, and we prove that credulous and weakly skeptical inference based on c-representations also satisfies conditional syntax splitting. In particular, we show that these syntax splitting properties hold for strongly consistent belief bases and also in the case of only weakly consistent belief bases.
Propositional logic (as other logics) can be seen as a way to represent sets of states of the world (by the sets of models of formulas). Representation of variations from a set of states of the world to another is motivated by studies on case-based reasoning (CBR): the comparison of two problems and the way a solution is transformed into another solution can be seen as variations from a set of states to another. This has led to a notation for representing these variations (a syntax) and this article studies how to associate to this notation a semantics based on model theory in which an interpretation is an ordered pair of interpretations in propositional logic. The article presents a classical study of this logic (syntax, semantics, properties, a formal system, etc.). It also explains how CBR can benefit from this logic, in particular through the representation of adaptation rules and to a case-based inference algorithm using such rules.
Single image super-resolution (SR) aims to reconstruct high-resolution (HR) images from their low-resolution (LR) counterparts. Although recent transformer-based SR methods have achieved impressive performance, their substantial computational complexity and memory requirements severely restrict practical deployment. To address these challenges, we propose the gated multi-scale interaction network (GMIN), a lightweight convolutional neural network architecture that effectively integrates transformer design principles. GMIN introduces the gated multi-scale interaction module, which comprises a spatially adaptive mixing layer (SML) and an enhanced spatial gated feed-forward network (EGSFN). The SML dynamically filters less informative features and aggregates multi-scale spatial information through innovative gating mechanisms, while EGSFN employs large-kernel convolutions with gating operations to capture rich spatial dependencies, significantly enhancing feature representation capabilities. Comprehensive experimental results demonstrate that GMIN achieves an exceptional balance between SR quality and computational efficiency, outperforming the transformer-based ESRT by 0.14 dB in peak signal-to-noise ratio while utilizing fewer parameters and requiring 77% fewer floating-point operations per second. These findings establish GMIN as a novel and practical solution for lightweight image SR, making a significant contribution to the development of efficient SR models suitable for resource-constrained environments.
This article describes the early years of artificial intelligence (AI) at the University of Edinburgh, roughly from the early 1960s to the mid-1980s. It covers the key founders and various administrative structures, research highlights and teaching developments from this period, notable women in early AI research, and discusses the national and international connections of the work carried out at Edinburgh. Motivated by a scarcity of historical documentation on this pivotal period, we employ a methodological blend of archival research and recent oral histories to draw a detailed picture. The study highlights the contributions of key figures such as Donald Michie and Christopher Longuet-Higgins, whose interdisciplinary work established Edinburgh as a European hub for AI research. Despite challenges like the "AI winter", sparked by the Lighthill Report, Edinburgh fostered advances in many areas of AI. The analysis also explores early ethical considerations in AI development, reflecting on technological neutrality and the anticipation of AI's societal impact. This historical account not only clarifies Edinburgh's critical role in AI's evolution but also offers enduring lessons on interdisciplinary collaboration and ethical responsibility, relevant for the continued advancement of AI technologies today.
In numerous real-world contexts, the prevalence of abnormal instances in comparison to normal instances is markedly low. This phenomenon, characterized by an imbalanced data distribution, results in an increased likelihood of misclassifying the minority class as the majority. The publicly accessible collection of coronavirus disease-2019 (COVID-19) chest X-ray images has become significantly imbalanced due to the ramifications of the pandemic. To address this challenge, the proposed methodology integrates Histogram of Oriented Gradients (HOG) with the Synthetic Minority Over-sampling Technique (SMOTE) and Kernel-based Extreme Learning Machine (KELM). This process is executed in three phases: initially, the images are subjected to preprocessing, followed by the extraction of features using the HOG algorithm. These extracted features facilitate the generation of synthetic minority class samples, thereby achieving a more balanced dataset. The SMOTE-augmented dataset is subsequently employed for training the KELM, which demonstrates superior performance relative to existing state-of-the-art models. A comprehensive experimental analysis was conducted on four datasets comprising chest X-ray images of COVID-19, pneumonia, tuberculosis, and Healthy lungs. The classification accuracy obtained are 96.72%, 96.38%, 97.15%, and 98.66% on Dataset-1, Dataset-2, Dataset-3, and Dataset-4, respectively.
Belief Change explores methods to handle potential changes in the beliefs of an agent. The Alchourr & oacute;n, G & auml;rdenfors, and Makinson model describes three change operations, expansion, revision, and contraction, which serve as the foundation for revision models. On the other hand, Ontology Repair searches for methods to change an ontology so that it does not have unwanted consequences. Recent works established connections between the two fields of Belief Change and Ontology Repair.In this paper, we analyze specific methods for building optimal repairs and compare them to a specific belief change operator, namely, partial meet pseudo-contraction. In addition, we present an implementation that makes pseudo-contraction closer to optimal repairs.
Recently, due to the excellent computational efficiency and interpretability, discriminative correlation filter (DCF)-based tracking methods have received extensive attention in the field of unmanned aerial vehicles (UAVs). However, existing methods are usually susceptible to interference from significant appearance changes of the target object or background occlusion, which leads to tracking failure. To effectively address these issues, we propose a distortion-aware correlation filter with target mask (DACFTM) for UAV, which introduces a target regularization term to enhance the target perception ability of the tracking model. Specifically, we construct a target mask matrix based on the highest peak of the response map of the previous frame, thereby leveraging prior reliable localization confidence, and multiply it with the current feature map to obtain a regularization term containing only target information, effectively distinguishing the target from the background. In addition, to deal with tracking failure caused by large appearance changes, we propose a distortion-aware mechanism. When the quality of the response map corresponding to the filter is higher than a set threshold, we consider the filter is reliable and adopt the filter fusion strategy; otherwise, the saved high-quality filter is selected for the tracking in the next frame. Finally, we comprehensively evaluate the performance of DACFTM on three mainstream UAV benchmark datasets, and experimental results demonstrate that the DACFTM achieves impressive tracking performance.
With the aim of providing a natural and expressive domain for defining change operators, we introduce the epistemic space of rational rankings. We show that this epistemic space is highly useful for understanding various aspects of belief dynamics, particularly those related to the improvement of new information. It enables us to define certain classes of improvement operators, a generalization of iterated revision operators, in a natural and intuitive way. A key feature of rational rankings is the possibility of defining improvement operators in such a way that the negation of the new information does not worsen. This is impossible within the frameworks of total preorders or ordinal conditional functions, two well-known epistemic spaces. Another notable aspect of this space is that the behavior of these operators can be characterized by a few simple equations and inequalities, whose meaning remains transparent. Additionally, there are no stationary states: The epistemic state resulting from applying an operator to a prior epistemic state and new information is always distinct from the prior state. Finally, we prove that this class of operators is indeed a subclass of improvement operators. Consequently, we show that these operators exhibit desirable behavior when iterated sufficiently, ultimately converging to Darwiche and Pearl revision operators.
Heterogeneous information networks significantly enhance the performance of recommendation systems by integrating rich structural and semantic information. However, most existing methods are designed for homogeneous networks and utilize techniques such as attention mechanisms and multi-layer architectures, which often introduce unnecessary parameter complexity. Furthermore, meta-path-based approaches typically rely on manually predefined meta-paths, which limit the models' generalization capabilities. This article conducts an in-depth investigation into meta-path identification and proposes an Automatic Meta-Path Identification Recommendation (AMPIRec) framework. The framework optimizes the semantic representation of meta-paths through matrix self-transformation properties using weighted aggregation, while simultaneously reducing model complexity. Additionally, AMPIRec employs a multi-head attention mechanism for the joint training and optimization of user and item features, with its primary advantages being numerical stability and high prediction accuracy. To enhance recommendation precision, we specifically design a tailored top-k smoothed loss function that aligns recommendation objectives with real-world requirements. Extensive experiments on multiple real-world heterogeneous graph datasets demonstrate that AMPIRec achieves outstanding performance in both accuracy and stability.
In this article, we investigate the issue of representing total preorders by ranking functions in belief revision scenarios. Both total preorders and Spohn's ranking functions are most popular semantic structures in nonmonotonic reasoning and belief revision. While each ranking function uniquely induces a total preorder, total preorders can usually be represented by infinitely many ranking functions which only differ in the position of their empty layers. In order to work towards representation invariance of total preorders by ranking functions, we first introduce the notion of revision equivalence which postulates that equivalence is preserved during (most general) revision operations. Moreover, we single out so-called linearly equivalent ranking functions as prototypes of ranking functions with regularly inserted empty layers. We show that, in general, revision equivalence is hard to achieve, which prompts us to take a more operator-focused perspective. We introduce the postulates Preservation of Equivalence (PoE) and Preservation of Scale (PoS) in order to axiomatize operators which can guarantee that the equivalence resp. linear equivalence including the scaling factor between ranking functions is preserved. Afterwards, we evaluate various iterated revision operators from the literature with respect to these postulates. While (PoE) and (PoS) do not generally hold for revision operators in the Darwiche-Pearl framework, we show that strategic c-revisions with suitably chosen strategies are scale-preserving with respect to linearly equivalent ranking functions. Furthermore, we present approaches to defining equivalence- and scale-preserving revision operators for ranking functions from arbitrary existing revision operators.
The creation of the Catalan Association for Artificial Intelligence (ACIA) was driven not only by practical and scientific objectives, but also by a powerful symbolic vision. It showcases the potential of Catalan society to produce innovative ideas and add value on a global scale. This paper illustrates the pivotal role of the ACIA in the development and consolidation of artificial intelligence (AI) research in Catalonia and across Europe. Founded in 1994, ACIA emerged from the need to create a cohesive AI research community in Catalan-speaking territories and to promote AI literacy. Over the past three decades, ACIA has made significant contributions to AI research through initiatives detailed in this paper, such as the International Conference on AI (CCIA), the magazine NODES, the Marc Esteva Vivanco Award for the Best PhD thesis in AI, and the donesIAcat working group. By presenting these initiatives, analyzing the evolution of AI research topics in Catalonia, and detailing ACIA's involvement in Europe-including its role within EurAI and its industrial and international impact-, this article highlights how local AI societies contribute to the advancement of AI in Europe while preserving their unique cultural and academic identities. Finally, the future evolution of AI in Europe is discussed.
This article presents a targeted narrative review establishing the historical and theoretical foundations for computational belief change implementation. Seeded by Doyle and London's foundational 1980 taxonomy, we trace the evolution of belief revision from computational origins through the theoretical transformation of the Alchourr & oacute;n, G & auml;rdenfors, and Makinson (AGM) framework to contemporary approaches. Our analysis demonstrates how pre-AGM computational pragmatism relates to AGM theoretical constructs, revealing both continuities and transformations across this evolution. We analyze how each taxonomical category evolved in the post-AGM era, identifying the theoretical foundations and historical precedents that inform contemporary implementation challenges. This foundation enables subsequent research into robust computational blueprints that synthesize historical insights with formal guarantees, providing the baseline for systematic implementation analysis and engineering-focused belief change research.