
The quantification of uncertainty in flood forecasting, especially in small catchments, presents significant challenges due to the inherent uncertainty in precipitation now- and forecasts. Several authors have investigated uncertainty quantification for time series forecasting in general and for flood forecasting specifically. To our knowledge, none have focused on small catchments, tidal influences, large hourly forecast horizons, or the incorporation of classic forecasting tools such as differencing. We implement and evaluate approaches using Conformal Prediction, Monte Carlo Dropout, Ensembles and direct distribution forecasting to quantify the uncertainty of LSTM networks trained to predict the change in water level for the next 48 h in three catchments in Northern Germany. We analyze the performance of the different approaches regarding the width and accuracy of the prediction intervals. Our study shows that the incorporation of differencing strongly influences which uncertainty quantification methods are suitable, with direct distribution forecasting ignoring correlations between the forecasting steps and Conformal Prediction being the most suitable for our specific datasets.
Cardiovascular diseases (CVDs) remain a leading cause of global mortality, underscoring the need for reliable and interpretable predictive models to support early diagnosis and clinical decision-making. This study proposes a hybrid ensemble framework that integrates SMOTEENN resampling with a metric-weighted Voting Classifier to address class imbalance and improve prediction robustness. Using the publicly available Kaggle CVD dataset ( ≈ 70,000 records), the dataset was split into 80
Understanding and classifying vehicle driving behavior is critical for the development of intelligent transportation systems (ITS) and road safety analytics. In this paper, we propose a novel frequency-domain descriptor called the Spectral Curvature Signature (SCS), which characterizes the temporal dynamics of vehicle trajectory curvature through spectral and fractal features. The proposed framework processes raw (x, y) trajectories using a Savitzky–Golay filter, computes curvature over time, and applies a Fourier Transform to extract a set of compact, interpretable descriptors, including spectral energy, centroid frequency, Hurst exponent, and Band Energy Ratios (BER). We evaluate the effectiveness of SCS on both synthetic trajectories generated from a kinematic model and real-world data from the US Highway 101 Dataset (NGSIM). After applying random undersampling to address label imbalance, the classification model achieves an accuracy of 99.8
Many existing Deep Learning (DL)-based approaches for cervical cancer diagnosis lack comprehensive architectural comparisons and fail to effectively integrate Explainable AI (XAI) methods, such as Grad-CAM. This study proposes a cervical cancer classification framework that utilizes several state-of-the-art DL models and employs Grad-CAM to generate heatmaps highlighting the specific image regions that influenced the model’s predictions. These visual explanations improve transparency and support clinical interpretation, thereby enhancing trust in AI-assisted diagnostic systems. The experimental results demonstrate that the proposed framework not only achieves high classification accuracy but also provides valuable visual insights, contributing to the development of more interpretable and reliable AI tools for medical image diagnosis.
This paper presents a novel traffic signal control methodology that combines Lyapunov Drift-Plus-Penalty (DPP) optimization with a shock-aware phase re-service mechanism and LSTM-based traffic forecasting to achieve both queue stability and delay minimization at urban intersections. The proposed controller dynamically reallocates green time in response to evolving queue lengths, particularly prioritizing oversaturated approaches to prevent congestion build-up and mitigate traffic shockwaves. LSTM networks are employed to capture real-time temporal patterns in traffic flow, enabling more responsive and data-driven decision-making. The framework is grounded in queueing theory and incorporates Kingman’s delay approximation to model system dynamics, while theoretical analysis ensures strong stability and provable near-optimal delay under admissible traffic demand. Extensive simulations were conducted across two scenarios. In the first, the proposed method was compared with traditional fixed-time control, achieving a 17.21 · slots), with 2,064 targeted shock re-allocations contributing to improved traffic stability. These results confirm the method’s effectiveness in both proactive congestion management and reactive traffic surge mitigation. This work contributes a theoretically grounded and practically viable adaptive signal control strategy, enhanced by real-time sequence modeling, with strong potential for scalable deployment in smart urban traffic systems.
Recommendation systems provide personalized content suggestions to users based on various approaches. Recommendation systems based on Graph Convolutional Networks (GCNs) are the newest and have yielded promising results. However, most existing GCN-based methods focus only on content or context features, lacking integration of both. This study proposes a new recommendation method using GCN models to learn and combine the content and context features of users and items (called User-Item-Context Graph Convolutional Network), aiming to enhance recommendation quality. The proposed method includes the following main layers: (i) Data representation layer to create initial feature vectors for each entity; (ii) A GCN layer that learns high-level representations for each entity from the initial feature vectors; and (iii) Interaction prediction layer to predict interactions between entities and provide recommendations to users. The proposed method was experimented on MovieLens 100k and MovieLens 1M datasets using common metrics such as Precision, Recall, F1-score, AUC-ROC, MAE and RMSE.
This paper proposed two deep learning approaches for saliency retargeting while preserving aesthetic quality in images, addressing the limitations of conventional methods that focus solely on maximizing object saliency. The first approach, “OperatorNet”, estimates the intensity of multiple image editing operators to enhance the main object, utilizing EfficientNet-lite3 and multitask learning. The second approach, “GeneratorNet”, directly transforms images at the pixel level using the U-Net architecture. Both quantitative evaluations and subjective human assessments were conducted on images generated by the proposed methods. Experiments using the MS-COCO dataset confirmed that the proposed approaches achieved superior saliency retargeting while considering aesthetic quality compared to existing methods.
The increasing demand for personalized, emotionally expressive synthetic voices in Text-to-Speech (TTS) systems presents the challenge of generating voices that sound genuinely friendly, especially with limited data and computational resources. This study proposes a voice cloning method for data-scarce scenarios, requiring only 10–30 minutes of speech and optimized for low-complexity environments like Google Colab. By leveraging So-VITS-SVC, a robust voice conversion model, and the acoustic precision of Parselmouth (Praat), we extract and refine vocal attributes such as intonation, pitch, and timbre for more accurate and emotionally resonant voice generation. The method was deployed in an automated answering system at HCMUT, where it generates voices that are friendly and relatable for students. Preliminary evaluations show that the system outperforms baseline models with lower F0 RMSE (19.4 Hz), higher MOS-N (4.4), MOS-I (4.7), and reduced WER (3.2
Autonomous systems have become an important fixture in modern day life. Certification of ethical decision-making in autonomous system is necessary to minimize harm. In this work, a formalism of ethical decision making under uncertainty is described. The ethical principles are incorporated using Deontic logic rules of obligations, forbidden actions, and permissible actions. The uncertainty in the environment is modeled using interval discrete-time Markov chain. A tractable probabilistic model checking on interval discrete time Markov chain model is constructed for evaluation ethical decision making in a system. The results from experiments are performed using probabilistic model checking on a tractable model of interval discrete time Markov chain and statistical inference is conducted on a prototype of a moving aircraft is presented.
Retrieval-Augmented Generation (RAG) systems increasingly leverage structured knowledge graphs to enhance factual accuracy and interpretability. However, scaling such systems introduces challenges in routing queries efficiently across partitioned knowledge sources, particularly under memory and bandwidth constraints. We propose ScatterRAG, a lightweight and scalable routing framework for distributed RAG over partitioned knowledge graphs. ScatterRAG employs Bloom filters for compact indexing and fast negative lookups, combined with fuzzy matching to address lexical variation and noisy queries. This approach enables efficient query filtering and routing without centralized control or exhaustive broadcast. Experimental evaluations on the Natural Questions benchmark show that ScatterRAG achieves acceptable memory usage while maintaining scalability in distributed environments. Although its current prototype yields slower inference and lower retrieval accuracy compared to centralized baselines, ScatterRAG demonstrates a practical balance between scalability and resource efficiency, providing a promising foundation for future research on decentralized and efficient RAG architectures.
The tool condition plays a critical role in CNC machining processes and requires rigorous monitoring to ensure machining quality and minimize production costs. This paper presents a triaxial fusion network for tool wear prediction using Gramian Angular Summation Field (GASF) encoded force signals. The method transforms one-dimensional cutting forces (Fx, Fy, Fz) into two-dimensional images via GASF encoding, enabling the use of convolutional neural networks while preserving temporal correlations. In this way, specialized features from each force component can be extracted and fused for efficient prediction of three wear states: break-in, steady-state, and severe wear. Experiments on the PHM 2010 dataset with EfficientNet-B0 backbone for triaxial components demonstrate superior performance, achieving prediction accuracies of 90.48
Accurate estimation of forest carbon stocks from remote-sensing imagery is critical for climate monitoring and ecosystem management. We propose HAF R-CNN (Height-perceptual Attention Fusion R-CNN), a Faster R-CNN extension that fuses RGB (red-green-blue) and canopy height model (CHM) features via multi-level cross-attention to improve tree crown detection. Unlike previously proposed RGB-CHM fusion approaches that rely mainly on simple concatenation or late fusion, our cross-attention design enables richer bidirectional structural-spectral interactions. We further develop a tree-level carbon estimation pipeline combining crown structural descriptors and vegetation indices with a Random Forest regressor. On the NEON dataset, HAF R-CNN outperforms baselines in detection, and the pipeline achieves reliable carbon estimation at the OSBS site ( R^2 = 0.67, MAE = 59.6 kg), though performance decreases in denser forests (MLBS). These results highlight both the promise and current limitations of multimodal detection for scalable carbon monitoring.
In the context of the increasing need for accurate image analysis at the pixel level, the problem of semantic image segmentation has been posed to separate different meaningful regions without the need for labeling data. An unsupervised learning method has been applied, in which the boundary information is enriched with an energy distance map and then highlights are detected through the Points of Interest (PoI) technique to improve boundary discrimination. The model is carried out through K-means clustering in normalized characteristic space, combining energy characteristics and POI coordinates to increase accuracy in complex regions. The experimental results have shown that the semantic regions are more clearly separated, the number of POIs is satisfactory, and the overall accuracy is improved compared to the method of using only basic characteristic clustering.
Vision Transformers (ViTs) perform well in computer vision but need high computational power. This paper explores Grouped Query Attention (GQA) to reduce parameters in TinyViT models with competitive accuracy. We compare GQA settings with query-to-key ratios of 2:1 to 10:1 on CIFAR-10 and CIFAR-100 benchmarks. Findings show that GQA decreases attention parameters by a maximum of 42.8
Continuous Integration (CI) servers are widely-used and essential for automated software testing and deployment nowadays. To serve customer’s services smoothly and reliably, CI servers are expected to be fault-tolerant, leading to a critical need of their reliability and thus availability. On the other hand, they are complicated with many metrics that need to be monitored in real time from many different perspectives. Addressing this issue, several existing works have taken into account early fault detection in CI servers. However, it remains unsolved due to their proposed incomplete solutions and the challenges from the complex, dynamic nature of server metrics, prompting the development of our work. In this paper, we propose a comprehensive solution to early fault detection in CI servers. Our solution is based on machine learning and statistical methods for metric value prediction and then fault detection. For the first part, we define a new hybrid model by integrating AutoRegressive Integrated Moving Average (ARIMA)’s linear modeling with Random Forest’s capability to capture non-linear interactions. The novelty of our model is reflected by the stacking mechanism which effectively enhances its model components and further makes the model yield better prediction results than the others. For the latter, a combined statistical method is defined to accurately identify future faults associated with the server metrics. As a result, our solution is more effective while conserving computational resources for CI servers as evaluated on the real-world datasets including the metric values from CI server nodes across multiple global sites.
The emergence of Transformer-based Pre-trained Language Models (PLMs) has had a significant impact across a variety of Natural Language Processing (NLP) domains. Pre-training language models on curated legal corpora can assist researchers in developing models to improve performance on downstream legal NLP tasks. This paper reports experiments with pre-training and fine-tuning language models on British and Irish Legal Information Institute (BAILII) data, addressing specifically Appeal Court judgments (these being, in effect, meta-judgments on previous legal judgments). We pre-train BERT-based language models on this corpus and evaluate their effectiveness on BAILII-based domain-specific tasks such as named entity recognition (NER), multi-label classification (MLC), and question answering (QA). The performance of this is then compared to baseline RoBERTa (Robustly Optimized BERT Pretraining Approach) and DistilRoBERTa models with two publicly available PLMs specifically designed for legal text (i.e., LegalBERT and CaseLawBERT), all pre-trained on the BAILII dataset. Pre-training on BAILII improves the performance of the PLMs on downstream tasks, and domain-specific pre-training enables a relatively smaller model such as BERT to achieve performance at par with a larger model such as RoBERTa. The pre-trained PLMs are now publicly available for downstream tasks on BAILII.
Motor imagery electroencephalography (MI-EEG) decoding is challenged by noisy, non-stationary signals, high variability, and limited labeled data. We propose FA-GPNet, a unified framework that integrates Filter Bank Common Spatial Pattern (FBCSP) for multi-band spectral–spatial feature extraction, a deep autoencoder for nonlinear compression and noise reduction, and a Gaussian Process Classifier (GPC) for probabilistic, uncertainty-aware predictions. Unlike conventional FBCSP pipelines that rely on manual feature selection and deterministic classifiers, FA-GPNet replaces heuristic ranking with data-driven latent representation learning and leverages GP’s Bayesian framework for calibrated outputs. Under within-subject evaluation, FA-GPNet achieves 78.19
Grape leaf diseases pose a significant threat to vineyard productivity and quality, underscoring the importance of early and accurate diagnosis. This paper introduces AFC-Net, an Attention-Guided Fusion Convolutional Network designed to enhance classification performance through the integration of multiple backbone architectures. Our model combines ResNet50, EfficientNetB0, and MobileNetV2 to extract various and complementary features from input images. These features are fused using an attention-based fusion mechanism that dynamically learns the contribution weights ( α _1 , α _2 , α _3 ) for each backbone, allowing the model to emphasize the most informative representations. We evaluated our approach in the Grape disease dataset, which includes four classes: Black Rot, ESCA, Healthy, and Leaf Blight. AFC-Net achieves a classification accuracy of 99.83
Blind image inpainting aims to restore images degraded by unknown corruption, where the locations and shapes of missing regions are unspecified at inference time. Existing methods typically separate mask estimation and image restoration into sequential stages, which can lead to error propagation and poor integration of structural priors, especially when dealing with complex or diverse artifacts. In this paper, we present Dual-Prior Diffusion (DPDiff), a novel framework that addresses blind inpainting by jointly predicting corruption masks and reconstructing edge maps. This simultaneous prediction of dual priors complementarily rectifies one another, significantly reducing error accumulation and enabling robust, structure-aware restoration. Leveraging these learned priors, DPDiff guides a diffusion model to generate restored images that both preserve geometric fidelity and blend seamlessly with uncorrupted content. Extensive experiments demonstrate that DPDiff sets a new state-of-the-art on multiple blind inpainting benchmarks using only four diffusion timesteps, and generalizes strongly to related image restoration tasks such as image deraining and watermark removal.
We propose a Classificatory Topos, a mathematical framework to model the dynamic evolution of knowledge within a finite system of interacting learning machines. Following guidelines of category theory, the construction establishes a Grothendieck topos, Sh(C_learn, J) , as a mathematical universe for this problem domain. By defining a base site on a category of epistemic states with causal morphisms, and equipping it with a Grothendieck topology that formalizes a logic of justification, the framework provides a rich, non-linear model of system evolution. The use of sheaves ensures causal consistency, while the internal logic of the topos, governed by a subobject classifier, provides the machinery to trace, verify, and explain the refinement of classifications.