Multi-resolution analysis has been widely used to alleviate the limitations of single-resolution approaches in polyphonic sound event detection. However, existing multi-resolution methods usually rely on category-agnostic fusion and simple interpolation-based alignment, which cannot effectively exploit category-dependent resolution preference and may introduce temporal inconsistency across resolutions. To address these issues, this paper proposes a category-guided multi-resolution fusion approach(CGMRF). Specifically, a Category-Guided Fusion (CGF) method is designed to learn category-dependent fusion weights over different resolutions, enabling more targeted and interpretable use of complementary time-frequency information. In addition, a Frame-Overlap Weighted Alignment (FOWA) method is introduced to align multi-resolution outputs by explicitly modeling the overlap relationship between adjacent STFT frames, thereby better preserving temporal continuity than linear interpolation. Overlap averaging is further employed as a lightweight post-processing step to suppress prediction jitter and improve temporal smoothness. Experiments on the TUT Sound Events 2017 and DESED datasets show that the proposed approach generally outperforms averaging fusion and several attention-based fusion baselines. On TUT Sound Events 2017, CGF improves the F1 score by 1.6
Spatial cues preserving speech enhancement is crucial for achieving both intelligibility and a "being-there" impression in speaker-dominant Ambisonics audio communication. Data-driven speech enhancement for Ambisonics input-output systems faces two key challenges. First, designing the target signal for reverberation shaping merely from a temporal perspective, as was done in single-channel scenarios, tends to degrade spatial perception. Second, the suitability of various filter matrix formulations for different target signals has not been systematically studied. To address the design of the target signal, we formulate it as an Ambisonics room impulse response shaping problem, and we propose a spatial shaping based on maximum directivity, as well as a variant that losslessly passes the omnidirectional component. To estimate these target signals, we establish a neural filtering framework, encompassing both the spherical harmonic domain and the plane wave domain, with three filter matrix parameterizations: mask, beamform-and-project, and unconstrained matrix. The experiments show that the proposed spatio-temporal reverberation shaping yields a more natural spatial auditory impression of the target signal and further enhances the spatial release from masking, where the performance of neural filtering primarily depends on the suitability of the filter matrix's rank for signal spatial covariance matrices rather than the spatial domain transformation.
In complex electromagnetic environments, the scale effect of Unmanned Aerial Vehicle (UAV) swarm presents significant potential for enhancing cooperative effectiveness. However, the accuracy of Time Difference of Arrival (TDOA)-based localization for non-cooperative emitters using UAV swarm is significantly affected by the coupling of multi-source errors, which mainly include UAV position error (UPE), clock synchronization error (CSE), and TDOA measurement error (TME). To address the challenges of evaluating cooperative effectiveness under multi-source errors coupling and balancing localization accuracy with computational efficiency, a cooperative utility of information (CUoI) optimization approach is proposed.First, a TDOA observation uncertainty model is constructed by integrating multi-source errors. Then, the information gain of target position estimation is derived to build the CUoI evaluation model. Next, the characteristic of Dueling Deep Q-Network (Dueling DQN) that decouples state value from action advantage is leveraged, enabling precise evaluation of the potential benefits of different hyperparameter adjustment strategies. This characteristic facilitates adaptive tuning of key hyperparameters in Particle Swarm Optimization (PSO). Finally, a dynamic PSO framework based on Dueling DQN is proposed to effectively balance localization accuracy and computational efficiency. Numerical experiments demonstrate that the proposed algorithm achieves reductions in average localization RMSE of 19.1%, 6.0%, and 1.4%, respectively, compared to Semidefinite Relaxation-TDOA (SDR-TDOA), Grey Wolf Optimizer (GWO), and Multi-swarm Discrete Quantum-inspired Particle Swarm Optimization with Adaptive Simulated Annealing (MDQPSO-ASA).
Primary-ambient extraction (PAE) is a technique to enhance the user listening experience in spatial audio reproduction. This is achieved by extracting the primary and ambient components from the sound scene. The PAE approach of ambient phase estimation with a sparsity constraint (APES) leverages the magnitude consistency of ambient components and the sparsity of the primary components to refine the PAE performance. This approach demonstrates an improved extraction accuracy when the ambient component is relatively strong. However, APES suffers from severe extraction errors when the primary amplitudes are equal in two channels of a stereo signal, which is a common sound scene in stereo signals. In this paper, the limitations of APES are analyzed, and a novel ambient phase estimation method is proposed under the joint constraints of sparsity and independence, called APESI. This method uses the independence between the primary component and the ambient component to correct the ambient phase estimation condition. Both objective and subjective experimental results demonstrate that the proposed APESI outperforms the APES and other traditional approaches in terms of extraction accuracy and ambient spatial accuracy, especially when the primary amplitudes are equal.
Sound source localization (SSL) is an essential task in many applications involving speech separation and enhancement. As such, speaker localization with microphone arrays has received significant research attention. While traditional signal processing (SP)-based SSL methods provide analytic solutions under specific signal and noise assumptions, recent deep learning (DL)-based methods have significantly outperformed them. However, their success depends on extensive training data and substantial computational resources. Moreover, they often rely on large-scale annotated spatial data and may struggle to adapt to reverberation and noise. To mitigate these challenges, we propose a novel embedding and beamforming network (EaBNet)-based SSL approach, termed SSL-EaBNet, which computes the spatial covariance matrix (SCM) of the observed signal after masking, compares the error with the target SCM, and performs error backpropagation to update the model weights. In particular, EaBNet has high computational efficiency and real-time inference capabilities, and its offline training process meets the requirement for real-time application. Empirical results in the public localization and tracking (LOCATA) dataset demonstrate the superior performance of our proposal, achieving a localization accuracy (ACC) of 91.49%, surpassing other competitive methods.
Human–computer interaction (HCI) is a cutting-edge and useful research frontier. This study introduces a novel HCI framework called Kin-LeapK. The framework aims to address three critical limitations of conventional single-sensor systems. These limitations are performance degradation under adverse conditions, insufficient spatial feature extraction, and low robustness in multi-sensor interaction signal processing. As for the interaction information recognition methods, our key contributions based on improved AdaBoost methods are: (1) hand gesture-based interaction: a dimension-by-dimension reverse processing (DDRP) strategy refines the Cuckoo Search Algorithm (CSA) for Support Vector Machine (SVM) hyperparameter optimization, establishing the CSA-SVM-AdaBoost framework to address multi-scale gesture variability. (2) Action-based interaction: incorporating a Gaussian mutation factor into Lévy flight-distributed SVM kernels produces the Lévy-SVM-AdaBoost classifier. It enhances robustness to occlusions and posture ambiguities. (3) Speech-based interaction: a modified Sparrow Search Algorithm (SSA) optimizes Bi-directional Long Short-Term Memory (BiLSTM) temporal dependencies, forming the SSA-BiLSTM-AdaBoost ensemble to improve noisy speech signal discrimination significantly. Experimental validation demonstrates superior performance: the framework achieves a 98.80
In scenes with noise and overlapping speakers, directionally extracting audio tracks corresponding to individual speakers is crucial for immersive and interactive spatial audio systems. Although neural networks have been successful in this task, existing steering approaches for adjusting the direction of neural speech extraction mainly target spatial audio directly collected by microphone arrays, while directional speech extraction with Ambisonics spatial audio is less well studied. Therefore, to encode the target directional information as input for the neural network, this paper proposes two Ambisonics directional features based on the spatial feature difference and beamforming principle: the relative harmonic difference and the directional signal enhancement ratio. Using the special property of Ambisonics' rotation transform, a rotary steering pre-processing is also proposed to align the target speaker's direction with a fixed reference by inversely rotating the sound field, thereby simplifying multi-directional extraction to fixed-directional extraction. Finally, we integrate these proposed approaches with the existing temporal-spectral-spatial filtering neural networks to establish a generalized framework for steerable speech extraction and conduct experiments on a simulated Ambisonics dataset containing multiple speakers and noise sources. The experiments show that the proposed approaches outperform existing conditional steering and can be applied to various existing neural network architectures.
In full-duplex underwater wireless optical communication(UWOC)systems,self-interference caused by photon backscattering significantly impacts system performance.To address this issue,this study proposes a data-assisted self-interference cancellation algorithm applicable to multiple-input multiple-output(MIMO)-UWOC systems.The algorithm improves the accuracy of self-interference channel estimation by splicing the shared pilot of the transmission and self-interference channels with partial data in the half-duplex mode,and using it as the self-interference channel pilot.The Henyey-Greenstein scattering phase function is employed to simulate actual underwater self-interference channels,and system simulations are conducted using the Monte Carlo method.The results demonstrate that the proposed algorithm achieves a maximum self-interference cancellation depth of approximately 20.4 dB,thereby outperforming traditional algorithms in terms of cancellation effectiveness while enhancing system transmission efficiency.Furthermore,when the received useful signal power degrades to a level comparable to the self-interference signal power owing to transmission distance variations,the proposed algorithm maintains a system error rate on the order of 10-4 under a signal-to-noise ratio as low as approximately 9 dB after self-interference cancellation.
By fully exploring the edge computing "supply-demand" relationship between the mobile-edge computing (MEC) servers and the differentiated application requests, the computing pricing (i.e., "supply") and allocating (i.e., "demand") can be coordinated well for the practical network consisting of heterogeneous users and MEC operator. In this article, the fair-aware computing pricing, beneficial offloading (i.e., obtaining positive utility) and local computing adjustment are jointly discussed under a pricing-enabled MEC. By considering heterogeneous application requests, fair service demand and limited computing provisioning, a multiobjective composite utility optimization is developed to maximize the user utility and the MEC operator profit simultaneously. Therein, the fair service condition is proposed, under which each user can experience a similar chance to obtain beneficial offloading. In order to solve the goal problem with undetermined objective function and conditions, a fair service enabled pricing and allocating algorithm (FS_PAA) with extremely low complexity is proposed by exploiting classification discussion method and convex optimization. Our FS_PAA reveals the explicit relationship between the optimal offloading decision and computing pricing, and the explicit relationship between the optimal computing pricing and the maximum computing provisioning, which helps to provide an effective reference for practical edge computing deployment. Simulation results show that our FS_PAA can 1) ensure fair offloading services for practical differentiated requests; 2) provide green offloading service for more users; 3) greatly improve the utilization of edge computing resource.
To more effectively prevent unauthorized eavesdropping, this paper proposes a distributed spectrum division transmission (DSDT) scheme for secure wireless communications. The original signal is divided into multiple independent subband signals, each carrying partial confidential information. These subband signals are mapped to spatially distributed transmitter arrays and then directed to the target receive area. Spectrum division disrupts information integrity, while focusing spatial beams shrinks the effective receive area. Both make it difficult for eavesdroppers outside the target area to correctly receive and reassemble all subband signals for complete information recovery. Considering a practical scenario where a powerful eavesdropper intercepts a single subband signal, other missing subband signals are modeled as equivalent interference to derive an accurate characterization of signal-to-interference-plus-noise ratio (SINR). Simulations demonstrate that the proposed DSDT scheme can effectively defend against high-threat eavesdroppers.
Gait is considered a valuable biometric feature, and it is essential for uncovering the latent information embedded within gait patterns. Gait recognition methods are expected to serve as significant components in numerous applications. However, existing gait recognition methods exhibit limitations in complex scenarios. To address these, we construct a dual-Kinect V2 system that focuses more on gait skeleton joint data and related acoustic signals. This setup lays a solid foundation for subsequent methods and updating strategies. The core framework consists of enhanced ensemble learning methods and Dempster-Shafer Evidence Theory (D-SET). Our recognition methods serve as the foundation, and the decision support mechanism is used to evaluate the compatibility of various modules within our system. On this basis, our main contributions are as follows: (1) an improved gait skeleton joint AdaBoost recognition method based on Circle Chaotic Mapping and Gramian Angular Field (GAF) representations; (2) a data-adaptive gait-related acoustic signal AdaBoost recognition method based on GAF and a Parallel Convolutional Neural Network (PCNN); and (3) an amalgamation of the Triangulation Topology Aggregation Optimizer (TTAO) and D-SET, providing a robust and innovative decision support mechanism. These collaborations improve the overall recognition accuracy and demonstrate their considerable application values.
Objective Compact underwater communication devices face stringent constraints in surface area and structural design, which necessitate densely arrayed elements in optical multiple-input-multiple-output (MIMO) systems. As a result, strong spatial correlation, or even homogeneity, often arises among sub-channels. This imposes rigorous requirements on spatial interference suppression techniques. Space-frequency block coding (SFBC) is an effective technique for mitigating spatial correlation interference in underwater LED-based MIMO systems. Although conventional two-dimensional SFBC schemes, such as Alamouti coding, preserve linear detection simplicity and full-rate advantages, their extension to 4x4 and larger configurations in optical MIMO systems faces fundamental limitations. High-dimensional orthogonal space-frequency block coding (HD-OSFBC) can be achieved via three primary approaches: rate compromise, codeword dimension reduction, and phase rotation-based precoding, each subject to distinct technical constraints. Rate compromise inherently degrades coding efficiency. In the case of codeword dimension reduction, spatial modulation-based schemes suffer from compromised antenna constellation identification under strong sub-channel spatial correlation, while optical antenna-based anti-aliasing strategies exhibit reduced tolerance to optical path misalignment. Similarly, the phase rotation approach leads to rank deficiency in the Gram matrix of the equivalent channel matrix, due to strong spatial correlation among sub-channels. In this paper, we present an HD-OSFBC scheme derived from Tarokh-Bharadia-Hochwald (TBH) codes for underwater LED-based optical MIMO systems. By strategically integrating orthogonalized phase rotation with optimized power distribution and common channel frequency response (CFR)-based equalization, the proposed scheme not only effectively addresses the inherent spatial correlation between sub-channels in optical MIMO systems but also delivers robust symbol error rate (SER) performance under both time-varying turbulent and static non-flat fading channel conditions. Methods In this paper, we present a comprehensive theoretical analysis of the quasi-orthogonality properties of TBH codes and fundamental limitations of conventional phase-rotation-based precoding approaches. Building on these insights, we propose a novel 2q-dimensional (q >= 2) full-rate OSFBC architecture that employs optimized phase rotation precoding. The effectiveness of the proposed dimension-expansion approach for the phase rotation matrix is rigorously verified through mathematical induction. In underwater LED-based compact-array MIMO, strong sub-channel spatial correlation causes conventional phase rotation to generate rank-deficient Gram matrices of the equivalent channel matrix, thus fundamentally limiting SER performance. To address this limitation, we propose an optimized power allocation precoding scheme that preserves full-rank Gram matrix properties through exhaustive-search optimization of predetermined power distribution coefficients. Moreover, in both turbid underwater conditions and high-speed transmission regimes, the channel demonstrates non-flat fading, which substantially degrades the orthogonality of OSFBC. Leveraging the inherent sub-channel spatial homogeneity of LED-based compact-array MIMO systems, we apply unified equalization based on the averaged CFR across sub-channels as a common reference. This compensation strategy effectively transforms non-flat fading channels into approximate flat fading conditions, thereby enabling the OSFBC scheme to overcome its intrinsic transmission rate limitations. With reduced channel coherence time (CCT) requirements and effective channel equalization, the proposed HD-OSFBC exhibits enhanced robustness in complex environments over both conventional high-dimensional orthogonal space-time block coding (HD-OSTBC) and space-time trellis coding (STTC). Results and Discussions The proposed HD-OSFBC scheme is evaluated under both time-varying turbulent channels and static non-flat fading channels. Specifically, the time-varying turbulent channel is modeled by combining the Lambertian radiation pattern with turbulence models of varying intensities, namely the Gamma-Gamma and log-normal distributions, under clean water conditions. Meanwhile, the channel impulse response (CIR) of the static non-flat fading channel is derived through Monte Carlo photon-tracing simulations in turbid harbor water environments. The results demonstrate that the sub-channels of the underwater LED-based compact-array MIMO system exhibit remarkable homogeneity across varying communication spacings (Fig. 2). Under uniform power allocation, the high-dimensional phase rotation scheme produces rank deficiency in the Gram matrix of the equivalent channel matrix, which in turn degrades SER performance (Fig. 4). In contrast, non-uniform power allocation increases the minimum diagonal element of the Gram matrix of the equivalent channel matrix (Fig. 3), thus preserving full rank and yielding superior SER performance. Even when employing non-uniform power allocation, the SER performance associated with the optimal and suboptimal power allocation coefficient vectors remains distinct (Fig. 5). Under time-varying turbulent channel conditions with CCTs ranging from 10 ms to 0.01 ms, the proposed HD-OSFBC scheme exhibits superior performance compared to the HD-OSTBC scheme, which requires longer CCT, and the STTC scheme with limited traceback length (Fig. 7). Moreover, the HD-OSFBC scheme demonstrates robust performance without SER bottlenecks, even under challenging conditions such as low symbol rates, high codeword dimensions, and strong turbulence intensities (Figs. 6-8). Under static non-flat fading channels, the HD-OSFBC scheme effectively overcomes the SER performance bottleneck through channel equalization by common CFR (Fig. 11). However, under identical total transmit power, the SER performance of HD-OSFBC remains slightly inferior to that of HD-OSTBC across various antenna configurations and symbol rates, due to the convexity of the complementary error function (Fig. 12). Overall, HD-OSFBC demonstrates superior adaptability to both communication scenarios. Conclusions Through optimized phase rotation, strategic power pre-allocation, and common CFR-based equalization, the rank deficiency problem of the Gram matrix of the equivalent channel matrix, caused by high sub-channel spatial correlation in underwater LED-based compact-array MIMO, can be effectively addressed within the quasi-orthogonal SFBC framework. This integrated approach enables low-complexity implementation of both HD-OSTBC and HD-OSFBC schemes. In time-varying turbulent channels, HD-OSTBC, being highly sensitive to CCT, is prone to SER bottlenecks or even severe performance degradation. This tendency is further aggravated by larger array scales, longer OFDM symbol durations, and higher turbulence intensities. In contrast, HD-OSFBC exhibits significantly lower sensitivity to CCT, demonstrating consistently favorable and stable SER performance across a broad range of CCT (0.01-10 ms). In low-quality water environments, the underwater optical channel exhibits deteriorating spectral uniformity with higher symbol rates and extended transmission distances. When the symbol rate reaches the GBaud level, the SER performance of HD-OSFBC becomes notably constrained. By leveraging the common CFR derived from the spatial homogeneity of sub-channels in LED-based compact-array MIMO, channel equalization allows the HD-OSFBC scheme to achieve SER performance close to that of HD-OSTBC in static non-flat fading channels, with minimal observable bottlenecks. In underwater LED-based compact-array MIMO systems, the proposed HD-OSFBC scheme demonstrates superior reliability, stability, and computational efficiency. As a complex-coding scheme, HD-OSFBC exhibits notable adaptability not only for optical MIMO configurations but also for various MIMO systems characterized by high sub-channel spatial correlation.
Objective Federated Learning(FL)represents a distributed learning framework with significant potential,allowing users to collaboratively train a shared model while retaining data on their devices.However,the substantial differences in computing,storage,and communication capacities across FL devices within complex networks result in notable disparities in model training and transmission latency.As communication rounds increase,a growing number of heterogeneous devices become stragglers due to constraints such as limited energy and computing power,changes in user intentions,and dynamic channel fluctuations,adversely affecting system convergence performance.This study addresses these challenges by jointly incorporating assistance mechanisms and reducing device overhead to mitigate the impact of stragglers on model accuracy and training latency. Methods This paper designs an FL architecture integrating joint edge-assisted training and adaptive sparsity and proposes an adaptively sparse FL optimization algorithm based on edge-assisted training.First,an edge server is introduced to provide auxiliary training for devices with limited computing power or energy.This reduces the training delay of the FL system,enables stragglers to continue participating in the training process,and helps maintain model accuracy.Specifically,an optimization model for auxiliary training,communication,and computing resource allocation is constructed.Several deep reinforcement learning methods are then applied to obtain the optimized auxiliary training decision.Second,based on the auxiliary training decision,unstructured pruning is adaptively performed on the global model during each communication round to further reduce device delay and energy consumption. Results and Discussions The proposed framework and algorithm are evaluated through extensive simulations.The results demonstrate the effectiveness and efficiency of the proposed method in terms of model accuracy and training delay.The proposed algorithm achieves an accuracy rate approximately 5%higher than that of the FL algorithm on both the MNIST and CIFAR-10 datasets.This improvement results from low-computing-power and low-energy devices failing to transmit their local models to the central server during multiple communication rounds,reducing the global model's accuracy(Table 3).The proposed algorithm achieves an accuracy rate 18%higher than that of the FL algorithm on the MNIST-10 dataset when the data on each device follow a non-IID distribution.Statistical heterogeneity exacerbates model degradation caused by stragglers,whereas the proposed algorithm significantly improves model accuracy under such conditions(Table 4).The reward curves of different algorithms are presented(Fig.7).The reward of FL remains constant,while the reward of EAFL_RANDOM fluctuates randomly.ASEAFL_DDPG shows a more stable reward curve once training episodes exceed 120 due to the strong learning and decision-making capabilities of DDPG and DQN.In contrast,EAFL_DQN converges more slowly and maintains a lower reward than the proposed algorithm,mainly due to more precise decision-making in the continuous action space and an exploration mechanism that expands action selection(Fig.7).When the computing power of the edge server increases,the training delay of the FL algorithm remains constant since it does not involve auxiliary training.The training delay of EAFL_RANDOM fluctuates randomly,while the delays of ASEAFL_DDPG and EAFL_DQN decrease.However,ASEAFL_DDPG consistently achieves a lower system training delay than EAFL_DQN under the same MEC computing power conditions(Fig.9).When the communication bandwidth between the edge server and devices increases,the training delay of the FL algorithm remains unchanged as it does not involve auxiliary training.The training delay of EAFL_RANDOM fluctuates randomly,while the delays of ASEAFL_DDPG and EAFL_DQN decrease.ASEAFL_DDPG consistently achieves lower system training delay than EAFL_DQN under the same bandwidth conditions(Fig.10). Conclusions The proposed sparse-adaptive FL architecture based on an edge-assisted server mitigates the straggler problem caused by system heterogeneity from two perspectives.By reducing the number of stragglers,the proposed algorithm achieves higher model accuracy compared with the traditional FL algorithm,effectively decreases system training delay,and improves model training efficiency.This framework holds practical value,particularly for FL deployments where aggregation devices are selected based on statistical characteristics,such as model contribution rates.Straggler issues are common in such FL scenarios,and the proposed architecture effectively reduces their occurrence.Simultaneously,devices with high model contribution rates can continue participating in multiple rounds of federated training,lowering the central server's frequent device selection overhead.Additionally,in resource-constrained FL environments,edge servers can perform more diverse and flexible tasks,such as partial auxiliary training and partitioned model training.
Underwater sonar image instance segmentation is a crucial component for tasks such as marine resource exploration, underwater communication and navigation, and marine environmental monitoring. Existing sonar image segmentation models primarily rely on traditional machine learning methods, which face challenges in adapting to the complex characteristics of underwater sonar images. Underwater sonar image targets often exhibit complex scale variations and geometric deformations, which can pose challenges for segmentation models to accurately capture the details of all targets, especially for small objects. Additionally, these targets may possess low contrast and blurry edges, leading to foreground and background mixing during the training process, resulting in the loss of target features. Moreover, there is a class imbalance issue between the target and background categories, which can lead to the generation of blurry or inaccurate segmentation results by the segmentation models. To address the aforementioned issues, this paper proposes DBN-YOLO: a dynamic sparse attention-based deformable underwater sonar image instance segmentation model, built upon the existing YOLOv5-seg instance segmentation model. Firstly, the Deformable Convolution Network (DCN) is introduced. By incorporating learnable offsets, DCN can adaptively adjust the sampling positions of convolution kernels to better accommodate deformations and geometric variations of the targets. This helps improve the segmentation model's ability to accurately capture object boundaries, thereby enhancing the segmentation effectiveness. Next, the BiFormer attention module is introduced. It employs a dual-layer routing attention module to compute key regions at a coarse granularity and perform attention interaction at a fine granularity. This enhances the foreground weights and suppresses the background weights, enabling the model to focus more on the targets and improve segmentation accuracy and effectiveness. Lastly, the Normalized Wasserstein Distance (NWD) loss is introduced to alleviate the sensitivity of the original Intersection over Union (IOU) metric to positional deviations of small targets and mitigate the impact of imbalanced positive and negative samples on target segmentation. The experimental results illustrate that the proposed DBN-YOLO mode improves the recall rate by 3.9% and the mAP@0.5 by 3.1% on the SCTD-seg dataset.
The multicast routing problem in software-defined networking (SDN) is an NP-hard problem. The existing solution methods based on deep strength learning suffer from the problems of branch redundancy, an excessively large action space and slow convergence of the intelligent models. In this paper, an intelligent multicast routing algorithm based on deep hierarchical reinforcement learning is proposed to circumvent the aforementioned problems. First, the optimal multicast tree problem is decomposed into two subproblems: fork node selection and the construction of an optimal path from a fork node to a destination node. Second, a multichannel matrix is designed as the state space for the internal and external controllers of hierarchical reinforcement learning based on the global network-aware information characteristics of SDN. Then, different action spaces are designed for the upper and lower subproblems, four action selection policies are designed for constructing multicast paths, and different reward policies are designed at different levels. Finally, a series of experiments and their results show that the designed algorithm not only searches the multicast tree efficiently but also converges faster and without redundant branches, with better performance in terms of bandwidth, delay and packet loss rate than the current mainstream solution algorithms. The codes for DHRL-FNMR are open and available at https://github.com/GuetYe/DHRL-FNMR.
Message transmission and message synchronization for multicontroller interdomain routing in software-defined networking (SDN) have long adaptation times and slow convergence speeds, coupled with the shortcomings of traditional interdomain routing methods, such as cumbersome configuration and inflexible acquisition of network state information. These drawbacks make it difficult to obtain a global state information of the network, and the optimal routing decision cannot be made in real time, affecting network performance. This paper proposes a cross-domain intelligent SDN routing method based on a proposed multiagent deep reinforcement learning. First, the network is divided into multiple subdomains managed by multiple local controllers, and the state information of each subdomain is flexibly obtained by the designed SDN multithreaded network measurement mechanism. Then, a cooperative communication module is designed to realize message transmission and message synchronization between root and local controllers, and socket technology is used to ensure the reliability and stability of message transmission between multiple controllers to realize the real-time acquisition of global network state information. Finally, after the optimal intradomain and interdomain routing paths are adaptively generated by the agents in the root and local controllers, a prediction mechanism for the network traffic state is designed to improve the awareness of the cross-domain intelligent routing method and enable the generation of the optimal routing paths in the global network in real time. Experimental results show that the proposed cross-domain intelligent routing method can significantly improve the network throughput, reduce the network delay and the packet loss rate compared to the Dijkstra and OSPF routing methods.
The noncoherent multiple-input multiple-output (MIMO) system is inevitably subject to cross-product interference among different spatial channels, arising the problem on solution of nonlinear model. In this letter, we propose a common difference processing framework based on dual-frequency complementary (DFC) waveform, aiming to generate an opposite cross-product term to eliminate interference. Thus, it can turn the nonlinear MIMO model into a linear one to facilitate the use of linear detectors. In our framework, two linear models are derived for MIMO with energy detection (ED) and differential detection (DD). Finally, through the complexity analysis of detectors and performance simulation results, the proposed framework is demonstrated to be a proper candidate for low-complexity implementation of noncoherent spatial multiplexing.
In the research of spatial audio, binaural reproduction of Ambisonics can be categorized into nonparametric methods (e.g., classical Ambisonic decoding and All-Round Ambisonic Panning and Decoding (ALLRAD)) and parametric methods (e.g., COMPASS). While vector-based amplitude panning (VBAP) provides accurate source orientation information for both, the specific perceptual properties of the spatial audio it affects are unclear. We designed a subjective listening experiment based on ITU-R recommendations for subjective assessment of perceptual attributes. The results show that VBAP mainly affects "Sound colour" and "Acoustic balance", while its effects on "Homogeneity of spatial sound" and "Sound attack" are not obvious. Among the main perceptual attributes, VBAP has a slightly higher impact on "Spatial impression" than "Timbre".
As the communication distance changes, the received signal strength of an underwater optical communication system will change, and the range of its variation may not only exceed the dynamic range of the photoelectric detection device but also cause the reliability of communication to change due to the change in the received signal-to-noise ratio. In order to maintain better communication over a longer distance, this paper proposes a rate-adaptive method for underwater optical communication with joint control in the photoelectric domain. In the optical domain, the incident light’s power is adaptively adjusted by controlling the transmittance of the liquid crystal light valve to reduce saturation distortion. In the electrical domain, the constellation distribution is optimized according to the desired probability mass function, and the modulation order is adjusted in real time by estimating the received signal-to-noise ratio of the link. The simulation results show that under the forward error correction (FEC) threshold, the proposed method increases the dynamic range of the photomultiplier tube (PMT) by about 10 dB and expands the dynamic range of the system’s communication distance.