
Counterfactual explanations (CFEs) are essential for interpreting machine learning model predictions by illustrating how minimal input changes can alter outcomes in counterfactual scenarios. Effective CFEs should satisfy two key properties: feasibility and diversity. Feasibility ensures that the generated CFEs are realistic and applicable across various contexts, while diversity provides multiple perspectives on model predictions. However, existing methods face significant limitations, such as difficulty in generating multiple CFEs and reliance on a convex loss landscape, which hinders adaptability to complex scenarios and diverse applications. To overcome these challenges, we propose a novel approach for generating diverse CFEs applicable to both convex and non-convex machine learning models. Our method leverages a Farthest Point Sampling-enhanced Genetic Algorithm to improve diversity, while a sample-based Manhattan distance regularization enhances feasibility. Experimental results demonstrate that our approach outperforms state-of-the-art methods in generating CFEs that are both feasible and diverse.
Machine Learning as a Service (MLaaS) systems face a fundamental tension between regulatory demands for explainability and security vulnerabilities from model extraction attacks. While global regulations increasingly mandate transparent AI systems, emerging evidence suggests increasing model explainability can increase the vulnerability to model extraction attacks. This paper advocates for investigating the underexplored relationship between model extractability and explainability. Like how bulletproof glass provides visibility without vulnerability, we envision theoretical frameworks enabling MLaaS systems to achieve human-level explainability without enabling extraction attacks. Impeding this vision, we identify three critical technical gaps: the absence of general frameworks to predict extraction risk, standardized extractability metrics, and predictive security approaches. Institutional fragmentation across model owners, security teams, explainability researchers, and regulators further blocks coordinated solutions. The urgency of this agenda stems from rapid MLaaS adoption, expanding regulatory mandates, and increasing AI-related security incidents. Without an explainability-extractability framework, organizations must choose not to comply with regulatory transparency requirements or face exploitation through model theft and training data extraction. Success in this area can extend to complex systems beyond AI, from financial markets to government operations, wherever transparency is desirable but potentially dangerous.
The proliferation of open-sourced Large Language Models (LLMs) and diverse downstream tasks necessitates efficient model selection, given the impracticality of fine-tuning all candidates due to computational constraints. Despite the recent advances in LLM selection, a fundamental research question largely remains nascent: how can we model the dynamic behaviors of LLMs during fine-tuning, thereby enhancing our understanding of their generalization performance across diverse downstream tasks? In this work, we propose a novel theoretical framework that provides a proper lens to assess the generalization capabilities of LLMs, thereby enabling accurate and efficient LLM selection for downstream applications. In particular, we first derive a PAC-Bayesian Generalization Bound that unveils fine-tuning dynamics of LLMs and then introduce LensLLM, a Neural Tangent Kernel (NTK)-based Rectified Scaling Model that enables accurate performance predictions across diverse tasks while maintaining computational efficiency. Extensive empirical results on 3 large-scale benchmarks demonstrate that our model achieves up to 91.1% accuracy and reduces up to 88.5% computational cost in LLM selection, outperforming 5 state-of-the-art methods. We open-source our proposed LensLLM model and corresponding results at LensLLM.io.
Intrusion Detection Systems (IDS) are critical in defending modern networks against increasingly sophisticated cyber threats. This paper introduces a comprehensive framework for automatic IDS rule generation using machine learning (ML). We trained decision tree-based classifiers, such as XGBoost, to classify and detect malicious network traffic. The developed models are then analyzed to extract key attack signatures, which are then used to generate Lua-enabled Suricata IDS rules. These custom rules enhance Suricata's capability to detect a broad spectrum of attacks in live network environments. Experimental results demonstrate the effectiveness of the proposed framework in both classifying network traffic and generating deployable IDS rules. This work provides a reproducible, automated pipeline for transforming raw network data into actionable IDS rules, contributing to developing more adaptive and robust network security systems.
Current solar flare predictions often lack precise quantification of their reliability, resulting in frequent false alarms, particularly when dealing with datasets skewed towards extreme events. To improve the trustworthiness of space weather forecasting, it is crucial to establish confidence intervals for model predictions. Conformal prediction, a machine learning framework, presents a promising avenue for this purpose by constructing prediction intervals that ensure valid coverage in finite samples without making assumptions about the underlying data distribution. In this study, we explore the application of conformal prediction to regression tasks in space weather forecasting. Specifically, we implement full-disk solar flare prediction using images created from magnetic field maps and adapt four pre-trained deep learning models to incorporate three distinct methods for constructing confidence intervals: conformal prediction, quantile regression, and conformalized quantile regression. Our experiments demonstrate that conformalized quantile regression achieves higher coverage rates and more favorable average interval lengths compared to alternative methods, underscoring its effectiveness in enhancing the reliability of solar weather forecasting models.
Influenza is a seasonal respiratory virus with the potential to evolve into pandemic strains. With the unprecedented COVID-19 pandemic taking the lives of millions, enhancing viral surveillance is critical for future pandemic preparedness. In this study, the authors have applied the Differential Population Growth Rate (DPGR), a sliding-window and data-driven pairwise comparison method to measure relative transmission fitness among circulating influenza A/H1N1, A/H3N2, and B strains across several time periods and geographic locations. DPGR considers variants as internal controls to reduce sampling bias. DPGR is suitable for time intervals where the logarithmic ratio between two variants follows an approximately linear trend. This study marks the first cross-species application of the DPGR framework beyond SARS-CoV-2. DPGR estimates fitness by analyzing weekly log-transformed growth ratios between two variants. The authors also have constructed distance and difference matrices to reveal epidemiological relationships and fitness dynamics. The results identify specific periods and regions where particular influenza strains exhibit stronger fitness advantages, offering insights into regional viral dominance. Finally, the authors have presented a transmission-fitness model for influenza that captures shifting evolutionary trends.
Biomedical research now generates more than a million articles annually, overwhelming researchers and hindering discovery. This surge has sparked interest in biomedical hypothesis generation (HG), which aims to uncover implicit patterns among biomedical concepts. Most existing methods focus on pairwise link prediction, overlooking the complex, multi-concept relationships underlying many breakthroughs. We introduce HyHG, a temporal Hypergraph contrastive learning framework for biomedical Hypothesis Generation, which redefines hypotheses as hyperedges-sets of co-mentioned concepts in an article. By representing articles as hyperedges and organizing them into a temporal hypergraph, HyHG captures the evolution of scientific ideas over time. A transformer-based architecture learns from historical hyperedge sequences to predict future hyperedges-sets of concepts likely to co-occur in future literature. To distinguish genuine hypotheses from misleading ones, HyHG employs a time-anchored contrastive loss and hard negative sampling based on minimal edits to real hyperedges. We demonstrate stateof-the-art performance on three biomedical datasets. Our code and data are available at: https://github.com/amirhassan25/Temporal-Hypergraph-Contrastive-Learning
Discovering Granger causality from time series data is fundamental to understanding dynamic systems, yet most existing methods struggle with unknown intervention targets or causal structures in real-world scenarios. In this paper, we propose DiffuGC, a novel diffusion-based framework that unifies observational and interventional causal discovery through a generative denoising process. By introducing diffusive interventions, which apply progressive interventions without any prior knowledge, DiffuGC amplifies causal signals while preserving structural information. Furthermore, we introduce a denoising NoiFormer with adaptive attention to both short- and long-term causal dependencies, which disentangles trend and seasonal components to enable accurate reconstruction of causal structures from interventional data. To the best of our knowledge, we are the first to integrate diffusion models with interventional Granger causal discovery. Extensive experiments on synthetic, quasi-real, and real-world benchmarks demonstrate that DiffuGC consistently outperforms state-of-the-art baselines in both observational and interventional data. Moreover, we introduce an intriguing notion, Causality Acceleration, characterized by the early emergence of informative causal patterns within the diffusion path, which may open up promising directions for future research on efficient and adaptive causal discovery.
High-fidelity simulation and reconstruction of physical fields are essential in both scientific research and engineering, yet classical solvers can be prohibitively expensive at resolutions needed to capture fine-scale structures. We propose FNODiffSR, a data-driven super-resolution framework that couples a residual-guided diffusion model with an Adaptive Weighted Fourier Neural Operator (AWFNO). AWFNO models longrange spectral dependencies while selectively emphasizing highfrequency components, and the diffusion module employs a conditional probability-flow ODE instead of stochastic sampling to deterministically bridge low- and high-fidelity representations. Final reconstructions are obtained by integrating this ODE with an adaptive time-stepping solver. Experiments on quasigeostrophic turbulence across varied upsampling and sparsesampling regimes show that FNODiffSR consistently surpasses interpolation and learning-based baselines in reconstruction fidelity, structural similarity, and physical consistency (as assessed by a dimensionless equation-residual), while offering predictable runtime and scalability. These qualities make FNODiffSR a strong candidate for high-quality scientific data recovery and downstream analysis.
Real-Time Bidding (RTB) is a cornerstone of digital advertising, yet existing methods often struggle with slow convergence and poor adaptability in its dynamic, non-stationary environment. These challenges frequently arise from inefficient exploration and learning bidding strategies from scratch. To address this, we introduce eXploratory Calibrated Bidding (XCalBid), a novel three-step framework designed to accelerate learning and enhance performance. Our core innovation is a hybrid offlineonline approach that leverages a robust, pre-trained model to provide an effective initialization for online fine-tuning. XCalBid begins by collecting a diverse offline dataset using an exploratory bidding policy. This rich data then facilitates a two-phase pretraining regimen, which uses Soft Actor-Critic with Calibrated Q-Learning principles to learn a well-calibrated and robust value function. Finally, the pre-trained model is efficiently fine-tuned online, leveraging both historical and new experiences. Extensive experiments on the iPinYou dataset show that XCalBid achieves state-of-the-art performance, outperforming baselines in click maximization. It also demonstrates faster convergence, marking a leap in training efficiency for RTB systems. Our code is available at https://github.com/zhaoningyuan/XCalBid
Good rationale quality from large language models (LLMs) is essential for reliability and interpretability. However, the rationales produced by existing LLMs still have shortcomings, such as the lack of informativeness and faithfulness, which affect their practical applications. Reproduction experiments for enhancing rationale generation present significant challenges due to several factors. To study rationale quality in an agnostic manner, we develop a novel framework, Metric-guided Rationale Enhancement Framework (MREF), that re-weighs training instances based on multiple aspects of rationale quality. Specifically, MREF fine-tunes an LLM at hand on two benchmark multiple-choice question (MCQ) datasets, ECQA and MedMCQA, to generate answers and rationales. In the fine-tuning process, it exploits metrics from ROSCOE to evaluate the produced rationales across five dimensions: faithfulness, informativeness, coherence, repetition, and grammar, and uses these metric scores to guide re-weighting of training instances, hence encouraging the LLM to emphasize rationales of higher quality. Comprehensive experimental results demonstrate that this metric-guided re-weighting strategy significantly improves rationale quality across all evaluated ROSCOE metrics over the baselines without re-weighting, leading to more reliable and understandable outputs. MREF can be seamlessly integrated with existing LLMs for various NLP tasks beyond MCQs. Our code and datasets will be made available upon acceptance.
Graphs can model complex relational data, which makes them invaluable in numerous machine-learning applications. While the graph's structure can be efficiently represented and stored, the associated feature memory is substantial. The feature memory can be terabytes or petabytes for web-scale graphs, especially at companies like Pinterest and Google. The extra feature memory can require external storage devices with slower I/O times for data loading, making training even slower. To address the storage issue with feature memory, we aim to reduce the number of training nodes needed through sampling, reducing the storage requirement. However, the problem is that simply training a model with fewer nodes will result in worse performance. To solve that problem, we propose BLB-HGNN, a training algorithm based on the Bag of Little Bootstraps. BLB-HGNN independently trains several replicas of the architecture on different subsamples of the data. For each training epoch, our blb-sampler creates a bootstrap resampling of the data for the replica to train on. The trained replicas are merged using parameter averaging and then fine-tuned for inference. We conduct experiments with the OGB_MAG and MAG240M datasets to demonstrate the effectiveness of BLB-HGNN over simple training. We also conduct experiments on the impact of different sampling methods and model merging techniques. With almost no additional runtime cost, BLB-HGNN consistently provides a performance boost of up to 5% compared to standard training with the same training budget. Applying a non-uniform sampling method, such as Personalized PageRank or Spread Sampling, further improves performance. Furthermore, BLB-HGNN can achieve performance close to full dataset training with less than 50% of the training data on specific models. To our knowledge, this is the first work addressing the storage problem and uses Bag of Little Bootstraps for HGNN training.
Federated Learning (FL) offers a promising approach for collaborative model training in healthcare while preserving data privacy. However, existing FL methods often fall short in addressing two critical challenges: client-level fairness and compounded uncertainty from data heterogeneity and privacy-preserving mechanisms. We propose fair-LDP, a fairness-aware Local Differential Privacy framework that promotes fairness and privacy via uncertainty-guided aggregation in federated healthcare AI. fair-LDP leverages evidential neural networks (ENNs) to quantify predictive uncertainty and introduces a novel strategy that uses uncertainty-driven local differential privacy to guide fairness-aware updates while preserving data privacy. This ensures equitable performance across clients with varying data quality while mitigating the influence of unreliable or outlier updates. fair-LDP incorporates an adaptive mechanism that adjusts each client's privacy budget based on model performance, balancing fairness, privacy, and accuracy. We evaluate fair-LDP on real-world healthcare datasets under both IID and non-IID settings. Our experimental results show that it consistently outperforms state-of-the-art fairness-aware and privacy-preserving FL baselines, with no added computational overhead, while maintaining privacy guarantees comparable to homomorphic encryption and secure multiparty computation. By integrating uncertainty modeling, fairness-aware aggregation, and adaptive local differential privacy, fair-LDP provides a practical and principled solution for responsible, equitable, and privacy-preserving federated learning in healthcare.
Learning from time series data, particularly univariate time series data (UTSD), is important in various fields. Learning accurately from UTSD is challenging, since the data may be very granular and noisy and thus may contain time points which do not necessarily contribute to, and may even prevent, the ability to attain high generalization capabilities. Downsampling, in which a pared down version of the original UTSD that contains its most valuable parts is preserved, is commonly performed to cope with highly granular or noisy UTSD. However, there is a lack of effective downsampling methods, along with the appropriate evaluation metrics for them. We propose four new UTSD downsampling methods and evaluate their performance on nine commonly used and publicly available datasets. To enable proper evaluation of the downsampling methods, we have proposed three new evaluation metrics. Our evaluation demonstrates that our downsampling methods outperform state-of-the-art (SOT A) methods in two aspects: (1) our methods better preserve the original UTSD; and (2) more importantly, downsampled UTSD provided by our methods allow machine learning (ML) models to attain higher generalization capabilities compared to UTSD that SOTA downsampled. Particularly, our methods obtained better results in almost all the examined datasets (nearly 90%) in comparison to SOTA.
Federated Learning (FL) enables collaborative model training across distributed clients without requiring the exchange of raw data. However, existing One-Shot FL (OSFL) methods, designed for communication efficiency by reducing federated rounds to one, suffer substantial performance degradation when faced with highly non-IID data across clients, primarily due to critical distribution shifts: label shift, feature shift, and concept shift. In this paper, we introduce POSFed, a new personalized three-stage approach, to systematically address these fundamental limitations: (1) Each client locally generates robust and label-agnostic synthetic datasets via self-supervising learning, ensuring essential knowledge is captured despite local distribution shifts; (2) The server aggregates all synthetic datasets to train a global feature extractor, capturing generalizable and transferable representations across heterogeneous client data; and (3) Each client efficiently adapts the feature extractor by learning a personalized classification head on its own data, enabling effective local customization and mitigating both feature and concept shifts. Extensive experiments across multiple benchmarks demonstrate that POSFed significantly outperforms state-of-the-art methods, achieving performance comparable to multi-round personalized approaches while using only one communication round. By ensuring both superior personalization and practical communication efficiency, POSFed establishes a feasible paradigm for FL under more realistic and heterogeneous conditions. Code is available at https://github.com/1643204431/POSFed.
Anomaly detection involves finding unusual data instances of interest. This task has often been compared to looking for a needle in a haystack, especially in high dimensions. One strategy for improving anomaly detection is to leverage expert feedback in a human-in-the-loop process to label anomalies and nominal points. For high-dimensional data, the majority of existing anomaly detection approaches that incorporate expert feedback do not scale and often result in algorithms that are too slow to operate in an interactive setting with a human expert. We introduce a new anomaly detection algorithm intended for high dimensional data that can efficiently incorporate expert feedback on each round of querying. Our approach uses ideas from metric learning to perform an efficient incremental update when the user labels a data instance. We demonstrate through extensive experiments that our work provides the best tradeoff between performance and running time among existing approaches.
Recovering Markov boundary-the minimal set of variables that maximizes predictive performance for a response variable-is crucial in many applications. While recent advances improve upon traditional constraint-based techniques by scoring local causal structures, they still rely on nonparametric estimators and heuristic searches, lacking theoretical guarantees for reliability. This paper investigates a framework for efficient Markov boundary discovery by integrating conditional entropy from information theory as a scoring criterion. We design a novel masked autoregressive network to capture complex dependencies. A parallelizable greedy search strategy in polynomial time is proposed, supported by analytical evidence. We also discuss how initializing a graph with learned Markov boundaries accelerates the convergence of causal discovery. Comprehensive evaluations on real-world and synthetic datasets demonstrate the scalability and superior performance of our method in both Markov boundary discovery and causal discovery tasks.
Sharpness-aware minimization (SAM) is widely rec-ognized for its ability to improve the generalization of deep neural networks by transforming the optimization problem into a minimax problem, aiming to minimize the maximum loss caused by adversarial parameter perturbations within a neighborhood. However, existing work almost exclusively focuses on the original minimization optimization, with very little attention paid to the minimax optimization. In this paper, we introduce a novel algorithm, VaSSO-SGDAM, by leveraging Variance-Suppressed Sharpness-aware Optimization (VaSSO) for deep AUC maximization. We provide a theoretical convergence analysis of this algorithm, marking it as the first work to achieve such significant theoretical outcomes for this kind of problem. Lastly, we implement our method for optimizing the AUC maximization problem, and the experimental findings validate the efficacy of our approach.
In the era of rapid development of social media, social recommendation systems as hybrid recommendation systems have been widely applied. Existing methods capture interest similarity between users to filter out interestirrelevant relations in social networks that inevitably decrease recommendation accuracy, however, limited research has a focus on the mutual influence of semantic information between the social network and the user-item interaction network for further improving social recommendation. To address these issues, we introduce a social recommendation model with robust graph denoising-augmentation fusion and multi-semantic Modeling(Burger). Specifically, we firstly propose to construct a social tensor in order to smooth the training process of the model. Then, a graph convolutional network and a tensor convolutional network are employed to capture user's item preference and social preference, respectively. Considering the different semantic information in the user-item interaction network and the social network, a bi-semantic coordination loss is proposed to model the mutual influence of semantic information. To alleviate the interference of interest-irrelevant relations on multi-semantic modeling, we further use Bayesian posterior probability to mine potential social relations to replace social noise. Finally, the sliding window mechanism is utilized to update the social tensor as the input for the next iteration. Extensive experiments on three real datasets show Burger has a superior performance compared with the state-of-the-art models.
Despite driving record performance, the increasing reliance of deep learning on ever-larger datasets has led to prohibitively high storage and management costs that threaten continued progress. While coreset selection offers a promising solution to this challenge, existing methods often rely on expensive iterative optimization procedures or fail to select samples that allow strong generalization across tasks. In this work, we introduce ECHO, a coreset construction and augmentation strategy that leverages the relational properties inherent to a dataset to find its most representative samples. Unlike prior methods, our approach constructs a structured graph that encodes intrinsic dataset patterns, based on which influential samples are identified and augmented to maximize generalization performance. Extensive experiments across five benchmark datasets and against eighteen different coreset selection baselines show that ECHO achieves up to 60% accuracy gains under extreme compression, while being orders of magnitude faster than state-of-the-art alternatives. These results establish a new benchmark for data-efficient learning, particularly under tight coreset budgets, and showcase the benefits of structured coreset selection for effective generalization.