
Spiking Neural Networks (SNNs) offer energy-efficient processing suitable for edge applications, but conventional sensor data must first be converted into spike trains for neuromorphic processing. Environmental sound–including urban soundscapes–poses challenges due to variable frequencies, background noise, and overlapping acoustic events, while most spike-based audio encoding research has focused on speech. This paper analyzes three spike encoding methods–Threshold Adaptive Encoding (TAE), Step Forward (SF), and Moving Window (MW)–across three datasets: ESC-10, UrbanSound8K, and TAU Urban Acoustic Scenes. Our multi-band analysis shows that TAE consistently outperforms SF and MW in reconstruction quality, both per frequency band and per class across datasets. Moreover, TAE yields the lowest spike firing rates, indicating superior energy efficiency. For downstream environmental sound classification with a standard SNN, TAE also achieves the best performance among the compared encoders. Overall, this work provides foundational insights and a comparative benchmark to guide the selection of spike encoders for neuromorphic environmental sound processing.
Cybersecurity taxonomies comprise complex relational dependencies and detailed textual descriptions, forming heterogeneous hierarchical knowledge graphs. Modeling such systems requires jointly leveraging semantic content and graph topology, particularly when new entities are continuously introduced. This work addresses inductive link prediction in dynamic cybersecurity taxonomies, focusing on associating previously unseen software weaknesses with established attack patterns under limited or absent relational context. In such scenarios, new weaknesses appear as isolated nodes, constraining structural information at inference. Although recent approaches rely on language models for textual encoding, they often underutilize the graph’s relational structure. To address this, the proposed method adopts a two-stage strategy: candidate links are generated through textual similarity to existing weaknesses, then an adapted SEAL+ framework learns topology-aware representations enriched with PCA-reduced semantic embeddings to discriminate links. Experimental evaluation across two inductive settings–absent and limited relational context—demonstrates that jointly modeling structural dependencies and textual semantics consistently outperforms text-only baselines, with ROC-AUC improvements of 3.2–4.5
Tabular classification is a central supervised learning problem in high-stakes domains. While tree ensembles are strong baselines and deep tabular models sometimes help, standard ensembles typically use global fusion (e.g., uniform averaging), implicitly assuming constant expert reliability across the input space. We propose a lightweight, architecture agnostic dynamic gating mechanism that performs instance-wise mixture weighting using raw covariates, expert probability vectors, and two confidence features (entropy and margin). Across 14 benchmark datasets, dynamic gating improves mean-ensemble accuracy from 0.9047 to 0.9149 in a heterogeneous pool ( +0.0102 ) and from 0.9082 to 0.9126 in a tree-only pool ( +0.0044 ), with 11/14 dataset wins in both settings. Aggregate paired Wilcoxon tests confirm significance ( p=2.44× 10^-4 for heterogeneous, p=1.07× 10^-2 for tree-only), showing that compact instance-wise weighting improves robustness without architecture-specific fusion.
Contextual audience construction has long relied on rule-based keyword and taxonomy matching, often applied as touch-based frequency rules. As third-party tracking weakens, contextual audience construction gains prominence, while embedding-based semantic cohorting emerges as a plausible alternative to represent intent under privacy constraints. In most workflows, advertisers provide targeting terms, and platforms determine how inputs are translated into deliverable cohorts. This motivates a practical product question: do keyword/touch-based cohorts reflect stated intent, and can embedding-based cohorting improve suitability without unacceptable scale loss? Using the Microsoft News Dataset (MIND), we compare classical touch-based cohorting versus semantic cohorting through controlled offline experiments. Under strict leakage prevention and chronological holdout, we audit resulting audiences using structural diagnostics (constraint violation rate, topic drift, coverage, concentration) and a blind held-out click proxy. Touch-based cohorts maximize reach, but exhibit a higher negative-topic association. Centroid-based semantic cohorting reduces negative-topic association by 27 K=10,000 while reducing inventory coverage. Brief-based intent embeddings can increase topical drift, whereas centroid cohorts improve alignment at larger audience sizes. Human-authored briefs under-perform empirical seed-based definitions. The contribution is primarily conceptual and evaluative: it frames contextual targeting as an audience-construction problem and introduces a structural audit framework for comparing the cohorts produced by different intent-encoding choices.
The rapid progress of large language models (LLMs) has unlocked significant capabilities across diverse domains, yet has also raised critical concerns regarding the trustworthiness and safety of these systems. Although many studies address isolated risks such as hallucinations, bias, or vulnerability to adversarial prompts, comprehensive frameworks for systematically evaluating LLM trustworthiness remain scarce. This research proposes a novel evaluation framework to benchmark LLMs along key dimensions: honesty, bias mitigation, calibration, consistency, and resistance to deception. Drawing inspiration from recent safety benchmarks, we design specialized evaluation protocols for each dimension and compute aggregate trustworthiness scores using weighted combinations of individual metrics. Experiments on prominent LLMs demonstrate significant variability in performance across dimensions, highlighting strengths and weaknesses of current models. Our results underscore the need for multidimensional safety evaluations and provide practical tools for developers, policymakers, and researchers looking to build and deploy more trustworthy AI systems.
Real-time interactive systems require decision-making mechanisms that simultaneously ensure high behavioral quality and low response latency. The choice of artificial intelligence architecture critically affects system responsiveness, scalability, and the realism of autonomous agent behavior. The paper presents an experimental comparative analysis of three decision-making approaches that have been utilized in interactive environments: namely, finite state machines, goal-oriented planning, and utility-based models. Each approach was implemented within a unified real-time simulation framework to ensure methodological consistency. The evaluation protocol assessed both behavioral effectiveness and computational performance, with particular emphasis on decision latency under increasing agent populations. Experimental results indicate that utility-based models produce the most adaptive and behaviorally optimal outcomes, albeit with a higher computational cost. Finite state machines demonstrate superior time efficiency and scalability, though at the expense of behavioral flexibility. Goal-oriented planning exhibits intermediate characteristics in terms of both adaptability and computational overhead. The findings provide a structured analysis of the trade-offs between behavioral optimality and computational efficiency, contributing to a deeper understanding of the scalability of decision-making architectures in real-time intelligent systems.
Accurately predicting credit risk is essential in the financial industry to ensure stability and effective risk management. Many existing credit risk models still rely on structured data streams such as financial ratios, credit scores, and repayment histories, which are effective for historical analysis but often inadequate for identifying early warning signals of potential default. Even though unstructured data sources, which include financial news, market analyst reports, and earnings call transcripts, provide forward-looking and contextual insights, their integration with structured data remains limited due to heterogeneity, noise, redundancy, and uncertainty in information quality. This paper addresses these challenges by proposing a novel FinFusion framework for early credit risk insight. The proposed framework models structured financial attributes using a Multilayer Perceptron (MLP), while unstructured textual data are converted into dense semantic representations. To ensure reliable interaction between these modalities, a Quality-Gated Cross-Attention (QGCA) mechanism is introduced, which dynamically evaluates the relevance and reliability of unstructured signals and selectively assigns higher attention weights to risk-informative content. This gating strategy reduces the influence of noisy and irrelevant textual information and enhances alignment between structured indicators and contextual signals. Experimental results demonstrate that the proposed approach consistently outperforms conventional credit risk models and standalone deep learning methods across multiple evaluation metrics, particularly in early-stage default identification.
We propose a hybrid framework for knowledge extraction from large text corpora that combines aspect-conditioned LLM summarization with clustering-based topic modeling. The approach selects the most semantically stable prompting strategy via entropy minimization, generates aspect-conditioned summaries, and applies QA-based filtering before clustering. Applied to aviation incident reports, decomposed chain-of-thought prompting substantially reduces semantic instability compared to zero-shot generation while preserving macro-level thematic structure. The resulting topics are more causally oriented, improving interpretability while maintaining structural fidelity to the underlying corpus.
Neural network models are used across different fields as efficient tools to approximate nonlinear functions and dynamics, also in relation to forecasting. Accurate forecasts and a sound management of prediction uncertainty is of great importance at various decision-making levels. One of these contexts is life insurance, with respect to the actuarial valuations that rely on probabilistic assumptions about the future behaviour of financial and demographic phenomena. We focus on risk of longevity underestimation and show how the assessment of this risk is sensitive to model choices. In the context of the Lee–Carter model, being a benchmark mortality model, we use both autoregressive neural network model with exogenous variables and linear time series model to obtain both point and forecast distributions of death rates. We perform a numerical application to esimate the risk of longevity underestimation implied by both ARIMA and neural network models and monetize it in actuarial terms, based on Swedish mortality data.
While standard Retrieval-Augmented Generation (RAG) is typically deployed alongside massive, cloud-based Large Language Models, this study introduces SHEA-SLM: a Sustainable High-Efficiency Architecture tailored for Small Language Models (SLMs) running locally under limited computing resources. Unlike conventional pipelines, SHEA-SLM is explicitly designed to address the noise sensitivity and limited parametric knowledge of SLMs in offline, privacy-preserving environments. The system leverages a domain-specific Wikipedia corpus, dense retrieval, strict reranking, and grounded response generation to support accurate factual question answering. We evaluate three instruction-tuned Qwen2.5 models (1.5B, 3B, and 7B) on a manually curated, domain-specific benchmark of 200 questions covering various retrieval scenarios. The results show that SHEA-SLM significantly improves answer quality for the smaller models, enabling a 1.5B model to outperform a 7B base model in certain settings. However, these gains are accompanied by higher latency, highlighting a clear trade-off between answer quality and computational efficiency. The benefits are less pronounced for the largest model, suggesting that increased parametric capacity reduces reliance on external retrieval. Overall, the findings demonstrate that tailored retrieve-and-rerank methods can make small, locally deployed models highly accurate and effective for sustainable, domain-specific, and resource-constrained AI applications.
India’s judiciary faces a backlog exceeding 44 million pending cases, with a significant fraction attributable to petitions filed in inappropriate judicial forums due to the absence of automated routing guidance at the point of electronic filing. We present LeX-Route, an automated petition routing system that classifies legal petitions into six judicial court categories using a hybrid feature representation combining semantic embeddings from LegalBERT, structured legal code multi-hot vectors (IPC sections and statutory Acts), and eight argument-mined scalar features. Evaluated on 2,467 Supreme Court petitions from the Indian Legal Documents Corpus (ILDC) under stratified 5-fold cross-validation with SMOTE-based oversampling, LightGBM achieves 92.50
Reinforcement learning (RL) agents often operate as black boxes, making it difficult to understand their decision-making in dynamic environments. This study proposes a novel framework for explainable RL based on structural causal models (SCMs). Here, the approach learns an SCM of the environment dynamics and reward process in a mobile network simulator (mobile-env), and uses this causal model to generate counterfactual explanations and perform interventions to understand agent behavior. The approach demonstrates that the learned SCM can closely approximate the environment’s transition dynamics while remaining interpretable. By leveraging do-calculus and counterfactual reasoning, our framework explains the long-term effects of actions through causal chains and highlights key influential factors. Experiments on a wireless network control task show that our method provides meaningful explanations for agent decisions (e.g., why a given action yields a higher reward), with minimal loss in policy performance. The study also presents comparative evaluations against baseline explanation approaches and discusses how our SCM-based explanations improve transparency and trust in RL policies.
This study investigates the predictive and allocative performance of multiple machine learning (ML) models in the context of quarterly equity return forecasting and portfolio optimization. Using exclusively publicly accessible data from the AlphaVantage API, the research aims to provide a transparent and reproducible comparison of ML-based predictive frameworks under conditions that approximate the informational environment of a real-world investor. The analysis incorporates twenty fundamental financial ratios capturing profitability, valuation, leverage, liquidity, and efficiency dimensions, which serve as explanatory variables for predicting three-month ahead stock returns. To ensure methodological rigor, the study employs a rolling-window design with 12–20 quarter training periods and 4-quarter out-of-sample testing windows, combined with realistic two-month reporting lags to mitigate look-ahead bias. The evaluated models include Linear Regression, Ridge Regression with cross-validated regularization, Random Forest, XGBoost, and a Stacked Ensemble integrating these approaches. Portfolios are constructed by selecting the top 5–20
QA systems often fail on ambiguous questions because they assume a single correct answer. We propose CenterDistill, a weakly supervised framework that learns semantic center distributions from clustered question embeddings to guide inference-time behaviour: answer, clarify, or present alternatives. Unlike prior work, it requires no manual interpretation labels and derives supervision directly from data. The model jointly predicts center distributions and answer spans, using the predicted distribution to select behaviour at inference. In English–Spanish cross-lingual QA, CenterDistill achieves 91.2 https://github.com/hacky1997/Centerdistill .
Vehicle–pedestrian collisions pose a significant risk to Vulnerable Road Users (VRUs), with head injuries being among the most severe outcomes. Accurate prediction of the pedestrian head impact location and collision time is essential for the effective deployment of safety mechanisms such as pedestrian airbags and active hood systems. Traditional crash analysis relies on high-fidelity Finite Element Analysis (FEA) simulations, which are computationally expensive and time-consuming. In this work, we investigate deep learning models as surrogate predictors for key collision parameters. Using a dataset generated from virtual crash simulations covering multiple vehicle profiles, impact velocities, and pedestrian configurations, we train models to predict the three-dimensional head impact coordinates and the collision time. We compare the performance of a Multilayer Perceptron (MLP) and a Kolmogorov–Arnold Network (KAN). Experimental results show that KAN significantly outperforms MLP, achieving lower test losses for both collision time (21.0 vs. 106.7) and head coordinate prediction (1149.9 vs. 6982.2). These results demonstrate the potential of KAN-based surrogate models to accelerate crash analysis and support the development of improved pedestrian safety systems.
Federated Learning enables privacy-preserving model training across distributed clients, but heterogeneous client distributions for Non-IID data often lead to unreliable predictions without uncertainty guarantees. Conformal Prediction provides distribution-free coverage guarantees under exchangeability, an assumption violated in heterogeneous federated environments. We propose FedCP, a federated conformal prediction framework that provides reliable uncertainty quantification across client heterogeneity. Each client performs local calibration using Inductive Conformal Prediction with likelihood-ratio-weighted nonconformity scores to correct distribution shift between local and global data. The server aggregates client statistics through a weighted quantile rule to obtain a global prediction threshold. We show theoretically that FedCP achieves coverage of at least 1-α , with degradation bounded by the Earth Mover’s Distance between client and population distributions. Experiments across multiple datasets and varying Non-IID levels demonstrate that FedCP maintains near-nominal coverage ( ≥ 90% ) while producing compact prediction sets. The proposed approach builds on established conformal prediction principles while addressing the non-trivial challenges introduced by heterogeneous federated data distributions.
Motivated by the growing demand for computational storytelling systems in domains such as digital entertainment and accessibility for visually impaired readers, ComicX converts static sequential art into dynamic, immersive media. Beginning with heterogeneous PDF-based comic inputs, the framework operationalizes panel detection and segmentation to isolate visual primitives, followed by panel sequencing to preserve discourse continuity. Through a dual-stage detection process, character instances are localized, after which character re-identification assigns consistent identity embeddings across panels, enabling narrative coherence in voice assignment and temporal tracking. Speech bubble regions and contextual explanation regions are simultaneously localized, enabling robust OCR-based text extraction for dialogue retrieval. A character–speech mapping module uses contextual cues to map utterances to speaker identities, while the text-to-speech module supports prosody transfer. In addition, onomatopoeic tokens are identified to provide realistic sound-effects. With the combination of computer vision, natural language processing, and generative speech synthesis constituting the architecture, ComicX enables multimodal alignment and audiovisual reconstruction of sequential art.
Data drift, that is temporal shifts in the underlying data distribution, constitutes a major threat to model reliability and long-term performance of machine learning systems operating in dynamic environments. We propose ADAZOR a drift detection algorithm focusing on the distribution of outliers relative to a learned reference, detected via a one-class SVM, which maps the original covariate stream into a sequence of binary indicators. This mapping yields a Bernoulli model, allowing the application of Z-test–based procedures for drift detection. This offers a principled, interpretable, and model-agnostic solution, facilitates data-driven maintenance decisions and provides explicit statistical guarantees, which many existing approaches lack. We compare ADAZOR against a common threshold-based technique from the literature, where drift is signaled once a monitored statistic surpasses a fixed bound. Experiments on diverse synthetic and real-world datasets with both abrupt and gradual drift indicate that our approach markedly cuts unnecessary retraining events while preserving accuracy comparable to the threshold-based approach.
Modern code intelligence tools struggle to achieve a unified understanding of code semantics and structure, limiting their applicability to advanced refactoring, optimization, and evolution tasks. This paper proposes a novel hybrid architecture that combines sequential transformers for capturing semantic relationships in code with Graph Neural Networks (GNNs) for modeling its structural properties. A hybrid architecture is proposed in this paper where Transformer and GNN layers alternate, enhancing both context-aware representation and structural reasoning. The alternating arrangement between Transformer and GNN layers enables the model to iteratively refine both semantic and structural representations. Each pass enhances the other’s context awareness—Transformers benefit from graph-level dependencies, while GNNs gain richer token-level semantics—resulting in more precise and comprehensive code understanding. The proposed model is evaluated on the 150k Python Dataset for the task of source code vulnerability detection. Our results demonstrate a 92.3
The exponential growth of digital media has led to significant information overload, making it necessary the creation of robust recommendation systems (RS) that personalize content search and assist users in decision-making. Although traditional collaborative filtering (CF) models provide robust, computationally efficient baselines, and Deep Learning (DL) architectures perform very well at capturing complex, non-linear interactions, the challenge of the optimization of predictive accuracy still remains open. To study and bridge the gap between linear baselines and sophisticated neural representations, this study conducts a comprehensive evaluation of traditional CF algorithms (SVD, NMF, k-NN, Co-Clustering) alongside a custom TensorFlow-based Neural Collaborative Filtering model (DLRecom), through a unified Ensemble Learning framework. Based on the MovieLens 1M benchmark dataset, our purpose is to reduce model variance and mitigate overfitting, with predictive performance being evaluated rigorously using Mean Absolute Error (MAE) and Root Mean Square Error (RMSE). Experimental results reveal that, through bagging, all algorithmic paradigm’s predictive accuracy is enhanced, with the bagged SVD ensemble pointed out as the optimal configuration, achieving a test MAE of 0.6841 and an RMSE of 0.8666, thus outperforming the best single-model baseline. Furthermore, this research provides a reproducible experimental protocol, highlighting how ensemble techniques deliver robust performance gains over isolated recommendation algorithmic architectures.