Text classification is a crucial and fundamental task in web content mining. Compared with the previous learning paradigm of pre-training and fine-tuning by cross entropy loss, the recently proposed supervised contrastive learning approach has received tremendous attention due to its powerful feature learning capability and robustness. Although several studies have incorporated this technique for text classification, some limitations remain. First, many text datasets are imbalanced, and the learning mechanism of supervised contrastive learning is sensitive to data imbalance, which may harm the model's performance. Moreover, these models leverage separate classification branches with cross entropy and supervised contrastive learning branches without explicit mutual guidance. To this end, we propose a novel model named SharpReCL for imbalanced text classification tasks. First, we obtain the prototype vector of each class in the balanced classification branch to act as a representation of each class. Then, by further explicitly leveraging the prototype vectors, we construct a proper and sufficient target sample set with the same size for each class to perform the supervised contrastive learning procedure. The empirical results show the effectiveness of our model, which even outperforms popular large language models across several datasets. Our code is available here.
Diffusion-based methods have shown strong potential for stochastic human motion prediction, but existing ap proaches typically formulate the task as a conditional diffusion process initialized from Gaussian noise. This introduces a substantial distribution mismatch between the Gaussian prior and the conditional future-motion manifold, increasing the difficulty of reverse inference and often requiring many denoising steps. Although standard Brownian bridge diffusion provides a more condition-aware alternative, it fixes the endpoint to a deterministic target, resulting in a zero-variance terminal endpoint that may restrict its ability to model the stochastic and multimodal nature of future human motion. To address these issues, we propose HumanBBD, a generalized Brownian bridge diffusion framework for stochastic human motion prediction. HumanBBD explic itly constructs a bridge process between the observed-motion domain and the future-motion domain, thereby reducing initialization mismatch and providing stronger condition-aware generation. Moreover, by relaxing the zero-variance terminal endpoint into a distributional endpoint, HumanBBD improves the flexibility and expres siveness of bridge diffusion for modeling multimodal futures. Built upon this formulation, we further develop a unified spatio-temporal generation backbone that jointly captures long-range temporal dynamics and structured spatial dependencies among joints. Extensive experiments on HumanEva-I, Human3.6M, and AMASS demonstrate that HumanBBD achieves strong prediction performance while requiring only 5-10 reverse steps, substantially improving sampling efficiency compared with conventional diffusion-based methods. Comprehensive empiri cal analyses further validate the effectiveness, efficiency, and robustness of the proposed framework. Overall, HumanBBD provides an effective and efficient alternative to standard conditional diffusion for stochastic human motion prediction.
Computational argumentation, a pivotal interdisciplinary field that simulates human reasoning, has grown rapidly. However, a systematic review that unifies its diverse tasks and methodologies under a coherent framework is lacking. This paper presents a comprehensive survey to address this gap. We first structure the field into three core subfields: (1) Argumentation Mining, (2) Argumentation Assessment, and (3) Argumentation Generation. Our primary contribution is a novel methodological framework that classifies existing techniques into three engineering paradigms: Feature Engineering, Model Engineering, and Adaptation Engineering. Applying this framework, we analyze and categorize 150 datasets and 148 methods from the last decade, providing a structured map of the landscape. Based on our analysis, we identify and discuss key future research directions, including the push towards multilingual and multimodal argumentation, the transformative potential of Large Language Models (LLMs), and the increasing demand for robust real-world applications. This survey not only offers a structured reference for researchers but also provides a clear roadmap for future innovation in the field. To facilitate continuous updates and resource sharing, we have established a GitHub repository at: https://github.com/shilida/Computational-Argumentation.
The Semantic Gap Problem (SGP) in Computer Vision (CV) arises from the misalignment between visual and lexical semantics leading to flawed CV dataset design and CV benchmarks. This paper proposes that classification principles of S.R. Ranganathan can offer a principled starting point to address SGP and design high-quality CV datasets. We elucidate how these principles, suitably adapted, underpin the vTelos CV annotation methodology. The paper also briefly presents experimental evidence showing improvements in CV annotation and accuracy, thereby, validating vTelos.
In real-world deployment, LLMs are often adapted continually across tasks to keep LLMs up-to-date in production, where new fine-tuning should preserve previously learned skills. However, indiscriminately mixing tasks can dilute task specialization, while sequential fine-tuning (full-parameter or low rank adaptation) often causes catastrophic forgetting due to destructive overwriting. Replay-based continual tuning and maintaining separate task-specific adapters can mitigate forgetting, but introduce additional compute, storage, and management overhead. Recognizing the redundancy of LLM parameters for any single task, we reframe continual task adaptation as task-specific parameter discovery via adaptation-aware probing: a short warm-start probe exposes a task's adaptation trace, enabling us to identify and isolate the small subset of parameters essential for each task to mitigate catastrophic forgetting. Building on this view, we introduce TRACE, a novel approach for discovering Task-specific paRameters via Adaptation-aware probing for Continual finE-tuning. We perform a short warm-start fine-tune to derive task-specific core parameters by comparing the warm-started and pre-trained models. Core parameters are identified via two strategies: importance scoring (L_2 norm and Fisher Information) and specificity analysis (cosine similarity of parameter updates). In continual fine-tuning settings, only the active task's core parameters are updated while others remain frozen, preserving prior knowledge. We conduct extensive experiments across multiple standard benchmarks to demonstrate the superior performance of our proposed method. Additionally, we validate the generalization of our method through a cross-model and scale transferability study, demonstrating a "small-to-large" paradigm that guides the fine-tuning of large-scale models under resource constraints.
The paper explores how "culture" can be operationalised in Natural Language Processing (NLP) and what this reveals about the possibilities and limits of considering a plurality of cultural backgrounds in technological design. It proposes that cultural alignment cannot be achieved only by adding more examples of "other cultures", rather it requires plural epistemologies: allowing multiple, locally grounded ways of knowing. To analyze how this plurality of knowing can be addressed in NLP, the paper uses a socio-technical model of language technology (LT) design, the five layers of technological activity model, for collecting and systematizing approaches to culture in NLP. The analysis shows that while NLP research has made progress toward more culturally sensitive systems, many approaches remain partial, addressing "culture" primarily at the level of output or representation while leaving deeper questions of power, governance, and social context unresolved. The paper concludes that operationalising culture requires much more than technical adaptation; it suggests a reflexive and plural socio-technical approach that navigates potentials and limits of computational formalisation for accounting multiple linguistic and socio-cultural backgrounds.
Modern Large Language Models (LLMs) have shown impressive performances in user-facing tasks such as question answering, as well as consistent improvements in reasoning capabilities. Still, the way these models encode knowledge seems inherently flawed: by design, LLMs encode world-knowledge within their parameters. This way of representing knowledge is inherently opaque, difficult to debug and update, and prone to hallucinations. On the other hand, Knowledge Graphs can provide human-readable and easily editable world knowledge representations, and their application in knowledge-intensive tasks has consistently proven beneficial to downstream performance. Nonetheless, current integration techniques require extensive retraining or finetuning. To overcome this issue, we introduce KoRe, a methodology to encode 1-hop sub-graphs into compact discrete knowledge tokens and inject them into a LLM backbone. We test the proposed approach on three established benchmarks, and report competitive performances coupled with a significant reduction (up to 10x) in token usage. Our results show that compact discrete KG representations can efficiently and effectively be used to ground modern LLMs.
The purpose of this paper is to introduce a Material Passport Ontology (MPO), which is designed to enable the interoperability of material passport data across stakeholders, including manufacturers, suppliers, collectors, and recyclers. The MPO provides a novel unifying ontological infrastructure for representing the properties of products and recycling processes. These properties are essential for estimating the potential recycled yield and enabling the stakeholders to make informed decisions about the feasibility of recycling. The ontology was co-created with industrial stakeholders, following a pragmatic ontology development method. The MPO is organised into facets describing the physical properties, composition and circularity, biological and other properties of products, components and materials. A set of constraints was developed to address the non-conformant syntactic or value range inconsistencies identified by the industrial stakeholders. These inconsistencies occurred in properties such as international product identification numbers and mass-related properties represented in a material passport knowledge graph. The information provided by the knowledge graph enables the assessment of the circularity of products, components and materials. The MPO was validated using reasoning tools and in collaboration with domain experts from industrial partners representing two use cases: the manufacture of components for motor vehicles and blades for wind turbines. These use cases demonstrate its effectiveness and applicability in identifying recyclable materials, maximising resource reuse, and enhancing sustainability practices, thereby facilitating the transition to a circular economy.
In-context learning (ICL) possesses the remarkable ability to ignite the reasoning capabilities of large language models (LLMs) with mere handfuls of samples; yet its efficacy hinges heavily on the quality of demonstrations. Consequently, several approaches have been devised to bolster the performance of ICL through sophisticated demonstration retrieval techniques. However, in out-of-distribution (OOD) scenarios, even the most advanced retrieval strategies encounter formidable hurdles, as it is arduous to extract pertinent test-related knowledge from disparate demonstrations. To address these challenges, this paper introduces a novel context-aware retrieval framework that effectively mitigates the adverse effects of OOD discrepancies on various tasks. This framework leverages LLMs to generate domain-pertinent data and retrieves demonstrations from the newly generated corpus. Given the inherent challenge of acquiring labeled data in real-world applications, we further propose an innovative retrieval method that meticulously balances sample confidence, similarity, and diversity, which ensures the judicious utilization of unlabeled samples generated by LLM. Extensive practical experiments conducted on multiple LLMs and within the realm of natural language processing have unequivocally validated the OOD robustness of our proposed framework. Our code is available at https://github.com/songruiecho/Ralood .
Articulated object manipulation is a unique challenge for service robots. Existing methods employ end-to-end policy learning, visionmotion planning, and large-language/visual-language model (LLM/VLM), but often overlook the diversity of articulated objects and the complexity of interactions between end-effector and handle, leading to limited generalization and destructive collisions. To address this, we propose GSAM, a generalizable and safe robotic framework for articulated object manipulation. Specifically, a vision-based perceiver generates the kinematic parameters. Considering that pre-trained markers in perceiver yield raw estimations that may deviate from commonsense, we present a f ine-tuned VLM-based refiner, using chain-of-thought (COT) commonsense reasoning to refine perception. To prevent destructive collisions, we design an interaction constraint function generator, integrating articulated object, interaction pose, and obstacle avoidance knowledge into a base. LLM then functionalize these constraints and apply them to trajectory and posture planning. A kinematic-aware manipulation planner verifies reachability for trajectory and posture. Experiments on 50 hinge tasks across 5 object categories and 50 randomly initialized end-effectorhandle configurations show that GSAM reduces standard deviation by 3.1
Graph few-shot learning, which aims to classify nodes from novel classes with only a few labeled examples, is a widely studied problem in graph learning. However, existing methods often face two key limitations. First, the predominant graph few-shot learning paradigm relies on supervised tasks, failing to leverage the vast number of unlabeled nodes in the graph. Second, many approaches require complex task adaptation or fine-tuning during inference, limiting their efficiency and applicability. Inspired by the powerful in-context learning capabilities of large language models, we propose a novel model named VISION for adVancIng graph few-Shot learning via In-cOntext LearNing to address these challenges. Our model reframes graph few-shot learning as a fine-tuning-free sequence reasoning problem. At its core is a context-aware network that initializes nodes with role embeddings and employs a dual-context fusion module to synergistically integrate local topological structures and global task-level dependencies. This allows our model to dynamically generate class-aware representations for the query set conditioned on the support set context in a single forward pass. To effectively train our model, we introduce an unsupervised task generator that creates structure-adaptive features and constructs diverse pseudo-tasks from abundant unlabeled data. Our method unifies unsupervised meta-learning with graph in-context learning, achieving efficient inference. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our model. Our public code can be found
Social Media and the Internet have catalyzed an unprecedented potential for exposure to human diversity in terms of demographics, language, culture, knowledge, opinions, talents and the like. Access to people's diversity gives us the possibility of exploiting skills and competences that we do not have, that we may not even know they exist, some the so-called unknown unknowns, but which, however, could be exactly what we need when looking for help in the solution to the our current problem. However, this potential has not come with new, much needed, instruments and skills to harness the complications which rise when trying to exploit the diversity of people. This paper presents our vision of the "Internet of Us (IoU)," a new type of online diversity-aware social interactions capable of promoting richer and deeper social relations. We discuss the multiple facets of diversity in social settings as well as the multidisciplinary work that is required to reap the benefits of the IoU, toward a IoU-enabled diversity-aware hybrid human-AI society.
Although studies have demonstrated that Large Language Models (LLMs) can perform well on Out-of-Distribution (OOD) tasks, their advantage tends to diminish as the distribution shift becomes more severe. Consequently, researchers aim to retrieve distributionally similar and informative demonstrations from the available source domain to boost the inference capabilities of LLMs. However, in practical scenarios where the target domain is inaccessible, evaluating the unknown distribution is challenging, which indirectly impacts the quality of the selected demonstrations. To address this problem, we propose DOPA , a demonstration search framework that incorporates an OOD proxy to approximate the inaccessible target domain and guide the retrieval process. Building on proxy-based evaluation, DOPA further introduces a Mahalanobis distance-based global diversity constraint to ensure sufficient diversity among the retrieved demonstrations. Experimental results on multiple LLMs and tasks demonstrate that DOPA effectively enhances robustness in OOD settings.
Graph anomaly detection (GAD) aims to distinguish anomalies from the majority of normal nodes in graph-structured data. Due to its extensive real-world applications, GAD has garnered increasing attention from both academia and industry. Recently, Graph Neural Networks (GNNs) have been integrated into GAD frameworks, yielding promising results by effectively characterizing structural information. However, existing GNN based methods suffer from several critical limitations: (I) the difficulty of learning discriminative representations for anomalies in the feature space; (II) the structural sparsity caused by the lack of essential connections between anomalies; and (III) the camouflage effect resulting from redundant edges between anomalies and normal nodes. To address these challenges, we propose a novel framework, Generate and Filter graph learning for Graph Anomaly Detection (GFGAD). Specifically, GFGAD first generates a diverse set of synthetic anomalies with enriched feature and structural information to balance the data distribution. Subsequently, these generated anomalies are strategically connected to original ones to compensate for missing structural patterns, while a filtering mechanism is employed to eliminate redundant connections and mitigate camouflage. Extensive experiments on several benchmark datasets demonstrate that GFGAD significantly outperforms state-of-the-art baselines.
Graph few-shot learning, which focuses on effectively learning from only a small number of labeled nodes to quickly adapt to new tasks, has garnered significant research attention. Despite recent advances in graph few-shot learning that have demonstrated promising performance, existing methods still suffer from several key limitations. First, during the meta-training phase, these methods typically perform node representation learning in Euclidean space, which often fails to capture the inherently hierarchical structure existing in real-world graph data. Second, during the meta-testing phase, they usually fit an empirical target distribution derived from only a few support samples, even when this distribution significantly deviates from the true underlying distribution. To address these issues, we propose IMPRESS, a novel framework that IMproves graPh few-shot learning with hypeRbolic spacE and denoiSing diffuSion. Specifically, our model learns node representations in a hyperbolic space and enriches the support distribution through denoising diffusion mechanisms. Theoretically, IMPRESS achieves a tighter generalization bound. Empirically, IMPRESS consistently outperforms competitive baselines across multiple benchmark datasets.
The rapid advancement of large language models (LLMs) creates new research opportunities in stance classification. However, existing studies often lack a systematic evaluation and empirical analysis of the performance of mainstream large models. In this paper, we systematically evaluate the performance of 5 SOTA large language models, including LLaMA, DeepSeek, Qwen, GPT, and Gemini, on stance classification using 13 benchmark datasets. We explore the effectiveness of two strategies - random selection and semantic similarity selection - within the framework of in-context learning. By comparing these approaches through cross-domain and in-domain experiments, we reveal how they impact model performance and provide insights for future optimization. Overall, this study clarifies the influence of different models and sampling strategies on stance classification performance and suggests directions for further research. Our code is available at: https://github.com/shilida/In-context4Stance.
Any digital personal assistant, whether used to support task performance, answer questions, or manage work and daily life, including fitness schedules, requires high-quality annotations to function properly. However, user annotations, whether actively produced or inferred from context (e.g., data from smartphone sensors), are often subject to errors and noise. Previous research on Skeptical Learning (SKEL) addressed the issue of noisy labels by comparing offline active annotations with passive data, allowing for an evaluation of annotation accuracy. However, this evaluation did not include confirmation from end-users, the best judges of their own context. In this study, we evaluate SKEL's performance in real-world conditions with actual users who can refine the input labels based on their current perspectives and needs. The study involves university students using the iLog mobile application on their devices over a period of four weeks. The results highlight the challenges of finding the right balance between user effort and data quality, as well as the potential benefits of using SKEL, which include reduced annotation effort and improved quality of collected data.
Over the past decade, deep learning has made significant progress and has become a prominent area of artificial intelligence. During the training process of deep neural networks, a large amount of diverse historical data is accumulated, including gradients, model weights, features, logits, probability distributions, and losses. These historical data can be used to help train the current model. Recent studies have tried to use the last epoch/batch's probability distribution as soft labels to supervise the training of the current model. However, the last epoch/batch's probability distribution may not be optimal for the current training, resulting in limited performance improvements. In this paper, we introduce a simple and universal method called Learn From the Best (LFB). We introduce the concept of "the best", which is reflected in two levels: the best historical model weights and dynamic supervision. To achieve sample-level logits supervision, we introduce a dynamic loss function called Historical Logits Contrastive (HLC) loss. Our method has been extensively evaluated on various benchmark datasets, demonstrating its universality and effectiveness. Compared to various self-distillation and regularization methods, the LFB method has achieved state-of-the-art performance. Furthermore, additional experimental analysis has shown that this method achieves rapid convergence and exhibits remarkable anti-overfitting and anti-noise capabilities.
Recent advances in data-centric artificial intelligence highlight inherent limitations in object recognition datasets. One of the primary issues stems from the semantic gap problem, which results in complex many-to-many mappings between visual data and linguistic descriptions. This bias adversely affects performance in computer vision tasks. This paper proposes an image annotation methodology that integrates knowledge representation, natural language processing, and computer vision techniques, aiming to reduce annotator subjectivity by applying visual property constraints. We introduce an interactive crowdsourcing framework that dynamically asks questions based on a predefined object category hierarchy and annotator feedback, guiding image annotation by visual properties. Experiments demonstrate the effectiveness of this methodology, and annotator feedback is discussed to optimize the crowdsourcing setup.
Pavel Shvaiko合作论文数Trentino Digitale61