
This paper proposes a novel class of fractional-order octonion-valued fuzzy bidirectional associative memory neural networks (FOOVFBAMNNs). To address the analytical challenges arising from the inherent nonassociativity and noncommutativity of octonion algebra, we employ the Cayley–Dickson construction to decompose the original octonion-valued system into four coupled fractional-order complex-valued subsystems. Two inequalities are established in the complex-valued fuzzy logic domain, which are specifically tailored to handle complex-valued activation functions, connection weights, and state variables, providing tighter bounds for synchronization criterion derivation. A simple yet effective linear feedback controller and a properly constructed Lyapunov–Krasovskii functional (LKF) are designed for the transformed complex-valued error systems. By integrating fractional-order Lyapunov stability theory with the proposed inequalities, we derive sufficient conditions for global Mittag-Leffler synchronization of FOOVFBAMNNs under both general activation functions and linear threshold activation functions. Numerical simulations are conducted to verify the correctness and effectiveness of the obtained theoretical results.
Adaptive optimization algorithms are fundamental to modern deep learning; however, the global organization of neural-network training regimes induced by optimizer hyperparameters remains insufficiently understood. In particular, the influence of the Adam moment coefficients on the stability and qualitative behavior of the learning process has not been systematically investigated through parameter-space regime mapping. In this work, we introduce an observable-based empirical regime-mapping framework for analysing neural-network training under the Adam optimizer in the two-dimensional hyperparameter space defined by the exponential decay coefficients of the first and second gradient moments, (β1, β2). Neural-network optimization is treated as an iterative parameter-update process evolving in a high-dimensional parameter space, while its behavior is characterized through low-dimensional observable fields derived from neuron-wise training-error dynamics. Rather than relying on a single observable, the proposed framework combines seven complementary empirical descriptors that characterize regime organization, temporal stability, alignment, anisotropy, and training evolution. Experiments are performed on datasets of increasing complexity, including printed-digit patterns, Fashion-MNIST, CIFAR-10, and CIFAR-100, using multilayer neural-network architectures of varying width and depth. The resulting regime maps reveal fragmented hyperparameter landscapes containing stable, oscillatory, slow-learning, and irregular observable regimes. Increasing dataset complexity and network capacity is generally associated with smaller coherent stable regions and increased sensitivity to the Adam moment coefficients. The geometric complexity of the regime boundaries is quantified using box-counting analysis. The estimated dimensions approach D ≈ 1.9 for several investigated configurations, indicating highly irregular and nearly space-filling boundaries at the available numerical resolution. These values are interpreted as empirical measures of boundary complexity rather than as evidence of exact mathematical fractality. Additional robustness experiments performed using 300, 500, and 1000 optimizer steps demonstrate that the large-scale organization of all seven observable fields remains largely preserved, whereas the principal changes are concentrated near transition boundaries. This persistence indicates that the detected regime structures are reproducible and are not solely artifacts of short optimization histories. The proposed methodology provides an empirical computational framework for visualizing and comparing optimization regimes in the Adam hyperparameter space. It facilitates the identification of comparatively stable hyperparameter regions and offers a complementary observable-based perspective on the complex behavior of adaptive neural-network optimization.
While unsupervised skill discovery based on mutual information (MI) has shown promise for learning reusable robot behaviors, existing methods often discover static skills with limited state coverage. This can limit their usefulness in robot–object interaction settings where meaningful object-state transitions are difficult to induce through the robot’s motion. In this work, we propose DDOI (Decomposed Skill Discovery for Object Interaction) for a robot agent interacting with an object. Under the MI-based skill discovery framework, the latent skill is decomposed into an object skill and a robot skill. The object skill is learned via an object skill discriminator constrained by either a Euclidean or controllability-aware distance, encouraging far-reaching and hard-to-achieve transitions in object states. The robot skill is trained through a robot skill discriminator that conditions on both robot and object states as well as the object skill, enabling the robot to acquire behaviors that help realize desired object-state transitions. We tested DDOI in Ant, Ant-Box, and Humanoid MuJoCo environments as planar single-object interaction benchmarks, using object-state coverage as the main evaluation metric and downstream goal-reaching success rates as an additional metric. Across these environments, DDOI variants demonstrated more consistent object-state coverage and stronger downstream performance than the compared baselines.
This study explores the optimality criteria for a class of fractional optimization problems characterized by interval-valued objective functions under constraints. These problems involve curvilinear fractional integral cost functionals that are path-independent and fundamentally connected to fractional calculus, particularly through Riemann–Liouville fractional integrals. The application of fractional calculus is particularly significant in modeling complex systems exhibiting memory effects, as the fractional derivatives and integrals naturally incorporate the system’s historical behavior. An important outcome of this research is the development and demonstration of a core optimality condition—rooted in fractional calculus—ensuring that a solution locally LR-optimal for the associated variational control problem governed by PDE and PDI constraints is also globally LR-optimal. To underscore the practical relevance of these findings, an illustrative example is provided, demonstrating how fractional calculus and memory dynamics within complex systems can be effectively utilized to model and optimize real-world dynamic processes, specifically in the control of artificial neural networks.
The RBF neural network (RBFNN) is proposed to address the challenges of modeling nonlinear dynamic systems. However, the modeling effectiveness is often affected by the insufficient adjustment of the structure and parameters. To address this problem, a collaborative optimization algorithm based on activity density (AD-RBFNN) is proposed in this paper for training the RBFNN. First, an activity density (AD) is introduced to characterize the contribution of neurons, which represents the average ratio of the output to the width of the neuron across the entire sample space. Second, the parameters and structure of the RBFNN are tuned based on the activity density to achieve an effective tuning result collaboratively. Third, a compensation mechanism is designed to eliminate the errors caused by the adjustment of the structure. Finally, the analysis of error-boundedness theory is provided to guarantee the stability of this method, ensuring the successful deployment of the AD-RBFNN. To validate the effectiveness of the designed AD-RBFNN, it is applied to nonlinear function approximation and industrial wastewater treatment process modeling, and compared against other mainstream algorithms. Results demonstrate that the AD-RBFNN exhibits significant advantages in both model accuracy and generalization capability, offering an effective intelligent modeling for complex industrial processes.
As a comprehensive art form integrating body movements, music rhythm, and emotional expression, dance has unique value in cultivating emotion-related competencies. However, traditional training methods have obvious deficiencies in objective assessment and personalized guidance. This study proposes a dance emotion-related performance training system based on multimodal perception and knowledge graph. Through integrating visual motion capture, physiological signal monitoring, and audio feature extraction, the system achieves multi-dimensional recognition of dancers’ emotional expression. A hierarchical attention mechanism is employed to adaptively fuse three modalities of information, achieving an emotion recognition accuracy of 89.6
Deep learning has achieved remarkable success across various fields; however, most architectural studies still emphasize depth expansion more than systematic width exploration. Inspired by the divergent-convergent organization observed in biological neurons, we investigate whether a width-first multi-branch design can improve representation quality without relying on very deep backbones. To this end, we propose the Neuron Bundle Network (NB-Net), a multi-branch architecture built around repeated Neuron Bundle Layers and a two-stage 1 × 1 convolution fusion module for progressive feature integration. In the standard configuration, all parallel branches share the same input tensor and use grouped convolutions with different kernel sizes. The proposed two-stage fusion module first performs moderate channel compression and then completes feature integration, which improves stability compared with a single 1 × 1 merge. We further analyze branch design, residual connections, and width settings. Experiments on CIFAR-10 and ImageNet show that NB-Net provides competitive accuracy with controlled parameter growth, supporting the value of width-oriented design and staged multi-branch fusion.
Images captured in low-light environments often suffer from various complex degradations. The main task of low-light image enhancement (LLIE) is to improve the brightness and restore the details of the image, thereby producing images that are more in line with human perception. Retinex-based methods decompose images into illumination and reflectance components, and enhance images by adjusting illumination and restoring reflectance details. However, the presence of noise in images usually leads to inaccurate decomposition, producing results that deviate from realistic appearances. In order to prevent the influence of noise on the subsequent decomposition tasks, we first denoise the low-light images, perform Retinex decomposition on the denoised images, and try to establish paired low-light images to maintain the consistency of the reflectance components. Specifically, we propose a Clean Decomposition Network (CDNet), an unsupervised learning method that first performs denoising on low-light image pairs to prevent loss of details and color deviation, and then performs Retinex decomposition, and guides network optimization through a carefully designed denoising network and loss function. Extensive experiments on eight public datasets, including LOL and SICE, demonstrate that our method achieves superior enhancement performance, is highly competitive with state-of-the-art unsupervised methods, and achieves comparable levels to supervised methods on certain metrics and datasets.
Cross-embodiment generalization in embodied reinforcement learning is challenging. Changes in morphology, such as limb geometry, torque limits, mass distribution, or friction, induce shifts in the transition kernel that often cause learned policies to overfit to embodiment-specific shortcuts, leading to degraded performance under out-of-distribution conditions. We propose Causal Morphology-Invariant Reinforcement Learning (CMIRL), an interventional structural world-modeling framework for morphology-aware model-based reinforcement learning. CMIRL implements a causal morphology-invariance principle: morphology modulates latent transitions via a dedicated pathway, while a separate mechanism captures invariant interaction dynamics. The model factorizes a morphology embedding m_e=g_ψ (ϕ _e) and mechanism representation d_t=ℳ_θ (z_t,a_t) to predict the next latent state z_t+1 through ℱ_θ . CMIRL is trained jointly with latent dynamics prediction, an IRM-style cross-morphology invariance penalty, and Morphology Counterfactual Consistency (MCC) using controlled simulator interventions do(ϕ ) . Across ManiSkill2, DMControl, and MuJoCo-Embodied benchmarks under a unified ϕ _1/ϕ _2/ϕ _3 protocol, CMIRL achieves 98.4± 0.8% normalized performance on the baseline ϕ _1 and 96.2± 1.3% / 96.8± 1.1% on shifted morphologies, improving over the strongest baseline by +1.1 – +1.3 points. Residual error reduction, morphology shift gap, AUC, and structural dependency indices confirm that mechanism separation provides robust and sample-efficient generalization. Design sensitivity ablations indicate stability across encoder architectures, embedding dimensions, and regularization strengths.
In financial time series classification, concept drift adaptation is crucial to maintain model performance, as data distributions evolve over time. To handle concept drift, there are mainly two methods: detection-based, which uses historical data, and non-detection-based, which relies on more immediate, smaller volume real-time data. The former preserves historical patterns but comparatively exhibits hysteresis, while the latter is opposition. Thus, it remains a challenge to achieve continuous model updating while preserving historical patterns. To address this issue, we propose a novel non-detection-based method. This method retains valuable historical patterns and achieves continuous model updating by observing distribution variations, which simultaneously controls different hyperparameters on a data. Specifically, we develop the Adaptive Financial Time Series Classification Model (AFinSeqClass), which integrates the Complementary Margin Support Vector Machine (C-Margin SVM) and the Inverse Derivation Algorithm (InvDA) to handle sample and feature selection. For sample selection, C-Margin SVM is a dual-distribution approach that skillfully utilizes soft and hard margin theories to generate two distinct data distributions from one dataset. These two distributions generate a dynamic signal-to-noise ratio margin—complementary margin, which is the set difference between the soft margin and hard margin. We then select low-noise and information-rich samples from the complementary margin. For feature selection, InvDA reverses the forward derivation of financial features using cooperative co-evolution strategies to break down complex problems into smaller sub-problems, corresponding to homologous feature groups to gradually refine the feature set, minimizing the redundancy among features. Through the cooperation of these two algorithms, across a series of stocks, the AFinSeqClass model achieves a classification accuracy exceeding 60
Standing long jump performance optimization has long been a focal point in sports biomechanics research. This study introduces an adaptive framework integrating I-JEPA (Image-based Joint Embedding Predictive Architecture) world models with Particle Swarm Optimization (PSO) for performance prediction and feature space exploration. Through self-supervised masked prediction mechanisms, the proposed framework learns motion dynamics representations employing an architecture comprising multi-scale temporal stem, Transformer encoder, and attention statistical pooling modules. Further, PSO is introduced as a hypothesis generation tool for heuristic exploration in the trained model’s feature space. Experimental validation on a collected sports biomechanics dataset (11 athletes, 35 jump trials in total; 10 athletes with 34 trials used for training and cross-validation, 1 athlete reserved for external validation) reveals that, under an athlete-level grouped cross-validation protocol, the model achieves Mean Absolute Error (MAE) of 0.157 ± 0.053 m, outperforming Transformer without JEPA pretraining (0.199 m), TCN (0.207 m), and Graph Skeleton (0.221 m) baselines. Repeated resampling experiments further validate model robustness. Systematic ablation experiments verify module effectiveness: removing attention pooling increases MAE by 19.9% , and removing reconstruction loss increases MAE by 21.6% . PSO explores a sample’s predicted score from 1.542 m to 1.583 m in the feature space; this result represents a model-space hypothesis rather than a guaranteed real-world performance improvement. This work, as a proof-of-concept for future personalized training systems, combines world models with heuristic exploration, providing an exploratory analysis framework for athlete training.
Deploying Large Language Models on edge platforms with mobile-oriented resource constraints faces challenges of limited resources and hallucination issues. While Retrieval-Augmented Generation (RAG) mitigates hallucinations through external knowledge, existing RAG systems on such platforms suffer from poor retrieval quality and lack standardized protocols. We propose a RAG framework for edge platforms with mobile-oriented resource constraints based on the Model Context Protocol (MCP), enabling plug-and-play access to heterogeneous knowledge bases. Our framework introduces a weighted voting fusion ranking mechanism integrating BERT-Recall, F1, Relaxed Exact Match (REM), and Query Relevance scores to enhance retrieval accuracy, combined with model quantization and few-shot learning for efficient on-device operation. Experiments on SQuAD, HotpotQA, and TriviaQA demonstrate that our framework achieves higher accuracy and lower error rates than state-of-the-art methods while maintaining low latency.
The growing usage of federated learning in healthcare opens new prospects for privacy- preserving analytics over heterogeneous data. This work introduces a new Federated Autoencoder-boosted Fuzzy C-Means (FCM) clustering architecture for health risk stratification for non-independent and identically distributed (non-IID) client settings. A synthetic 500-sample dataset generated in line with the UCI Obesity dataset was split evenly across five federated clients based on BMI quantiles to simulate true- world heterogeneity. The autoencoder was used to compress lifestyle and anthropometric high-dimensional features into a two-dimensional latent space, in which FCM was used to get soft, interpretable risk clusters. The proposed method consistently achieves higher clustering scores than the baseline methods across repeated runs. That is, it had the best Silhouette score of 0.600 compared to 0.327 with PCA + KMeans, having more compact and clearly separated clusters. Likewise, it achieved the minimum Davies–Bouldin Index of 0.549, indicating better compactness and minimum overlap compared to PCA + KMeans (0.909) and PCA + FCM (0.954). The Calinski–Harabasz Index also established strength with a value of 1182.0, significantly greater than PCA + KMeans (323.3) and using raw features as the basis for clustering (31.8). Client-wise analysis revealed uniform clustering quality, with Silhouette values between 0.573 and 0.659 across all five clients, demonstrating stable clustering performance under the simulated heterogeneous client distributions considered in this study. These results affirm that the combination of nonlinear latent feature extraction and fuzzy clustering results in high-quality, interpretable risk stratification while preserving data locality within the federated learning framework.
Partial differential equations (PDEs) serve as crucial tools for modeling complex real-world phenomena. The physics-informed neural network (PINN) is an advanced method within artificial neural network (ANN)-based approaches for solving a wide range of PDE problems. The restarting strategy of PINN (r-PINN) is a novel modification that has demonstrated a significant reduction in computational time compared to its predecessor. However, the complexity of the r-PINN architecture poses challenges, particularly in accurately solving PDE problems with complex geometries, resulting in a large number of trainable parameters that affect computational performance and memory requirements. To address these challenges, this study integrates truncated singular value decomposition (TSVD) into both the PINN and r-PINN frameworks to uncover the underlying structure of the trainable parameter matrices, enabling their factorization into low-rank components. This process yields a compressed trainable parameter matrix, improving training efficiency while preserving accuracy. The proposed approach aims to obtain numerical solutions for various benchmark PDE problems, including up to three-dimensional cases and a PDE system. This approach is applicable not only to the novel r-PINN but also to the basic PINN, allowing for an analysis of the improvements gained through TSVD. Comparative analyses with non-TSVD approaches will be presented, and experimental results will demonstrate the efficacy of TSVD in enhancing computational performance across all benchmark problems, warranting further investigation.
Fine-grained image classification remains challenging due to high intra-class variation and strong inter-class similarity, which make it difficult to obtain well-structured output distributions. While most existing approaches address this problem through spatial modeling or feature-level representation learning, output-level processing remains relatively underexplored despite its scalability across backbone architectures. To address this limitation, this paper proposes a self-calibrated mutual learning framework that improves performance in fine-grained contexts by reshaping output distribution geometry through the interaction between cross-model consistency and self-calibration. Cross-model consistency is achieved via deep mutual learning, while self-calibration is induced using online label smoothing. Unlike a simple combination of two loss terms, the proposed framework jointly enforces sample-wise consistency and class-wise target regularization within a unified probability space. Experiments on three benchmark datasets across multiple backbone architectures demonstrate that the proposed method consistently improves classification accuracy over existing output-level mutual learning methods and produces a complementary effect beyond their individual contributions.
Modern manufacturing increasingly relies on robotics to achieve high throughput and quality, especially as production lines become more flexible and parts more customized. Robotic inspection is a critical enabler for quality assurance as it supports repeatable measurements while reducing human workload and variability. Recent vision-language-action (VLA) models have advanced robotic manipulation by integrating visual perception and language understanding for autonomous control. However, the application to robotic inspection, which requires accurate movement without altering the environment, remains underexplored. This work investigates the feasibility of adapting manipulation-pretrained VLA models to an inspection-oriented feature-following task and presents the following contributions: Tailored to the requirements of inspection problems, two approaches for assessing VLA performance are introduced: A trajectory-based evaluation metric to quantify performance in rollouts as well as an action-level metric, useful during the fine-tuning process. In addition, an open-source, manipulation-pretrained VLA model is fine-tuned for a feature-following task. This task represents a simplified 2D inspection setting, designed to capture core aspects of inspection problems encountered in domains such as manufacturing and infrastructure. The model successfully executes these complex feature-following trajectories with competitive performance relative to a human operator in a real-world robotic setup, demonstrating effective transfer from manipulation to this class of inspection tasks. While the study is limited in scale, the results provide initial evidence that VLA models can be extended beyond manipulation to support feature perception and motion generation in automated, robotic inspection. This suggests their potential to support more consistent and automated inspection processes, motivating further investigation into robustness and generalization.
Continual Learning (CL) has proven an effective method of avoiding Catastrophic Forgetting (CF) when sequentially training neural networks. This improves network efficiency while facilitating the Knowledge Transfer (KT) between tasks. CL serves as an ideal setting for studying KT and network behavior. In particular, pruning methods for CL train subnetworks to handle the sequential tasks which allows us to take a structured approach to investigating KT. This subnetwork approach allows for the promotion of Forward KT (FKT) and complete prevention of negative Backward KT (BKT) and CF. Understanding which weights to share between tasks is crucial for FKT as sharing all weights can worsen accuracy. This paper demonstrates how the information flow (IF) within the network can reflect task similarity and usefulness, informing optimal sharing decisions to improve FKT between diverse tasks. We leverage IF to determine which weights to share to promote FKT. By doing this, we match the optimal sharing decisions for each implemented dataset, and compare the resulting accuracies against several benchmark methods which emphasize promoting FKT and BKT. Experiments are run on three CL datasets designed to emphasize variation in task complexity and similarity, using both ResNet-18 and VGG-16.
Wearable sensor-based human activity recognition (HAR) methods utilize multimodal data from sensors such as accelerometers, gyroscopes, and magnetometers to infer human activities. Recent advancements employ deep learning techniques, leveraging deep neural networks to automate feature extraction for activity classification. However, existing deep hybrid networks face challenges in distinguishing between similar activities and handling complex scenarios, particularly when integrating data from diverse sensor types. To address these limitations, this study introduces a novel hybrid deep learning architecture, termed MultiLevel Convolutional Transformer (MLConvTrans), aimed at enhancing the efficiency and accuracy of sensor-based HAR systems. MLConvTrans integrates two core components: the Multi-Level Convolutional Network (MLConvNet) and the n-stacked Transformer Encoder (TransEncoder). MLConvNet captures local features from individual sensors and performs multimodal fusion to obtain comprehensive global feature representations, effectively extracting both low-level and high-level information. These features are then processed by the TransEncoder block, which integrates global features and models long-term dependencies for activity classification. The proposed architecture is extensively validated on multiple public datasets, including UCI-HAR, MotionSense, HAPT, KU-HAR, SHL2018, and PAMAP2. Experimental results demonstrate that MLConvTrans significantly outperforms state-of-the-art methods in wearable sensor-based HAR while maintaining computational efficiency.
Multi-view clustering aims to integrate complementary information from multiple views to achieve better performance than single-view clustering. However, in practical scenarios, the quality of each view is often inconsistent, with some views containing substantial noise or redundant information, which may adversely affect the overall clustering performance. Moreover, existing contrastive learning techniques typically employ overly simplistic strategies for negative sample selection, making them prone to local optima during training and compromising model effectiveness. To address these challenges, this paper proposes a novel information fusion-based deep multi-view contrastive clustering algorithm, termed ECMVC. The proposed method explicitly models both consistency and complementarity among views and leverages a feature fusion network to enhance the stability and accuracy of clustering in noisy and redundant environments. In addition, we further propose a curriculum-guided contrastive learning approach, where an entropy-driven dynamic scheduler adaptively selects informative negative samples and progressively increases the training difficulty. This curriculum-guided mechanism enables faster convergence and more stable optimization. Experiments on multiple benchmark datasets demonstrate the effectiveness of the proposed method.