To overcome the limitations of single-modality imaging in maritime ship detection, this paper proposes a multiscale active–passive fusion framework. We develop a Laplacian-pyramid-based model to synergistically exploit complementary scattering and radiation information from heterogeneous active and passive data. The framework incorporates a soft ROI mask and a sea clutter suppression mechanism to enhance target saliency and suppress background interference. A multidimensional evaluation framework with Pareto frontier analysis is employed to determine the optimal operating point. Experimental results demonstrate that the proposed method not only outperforms single-modality imaging and representative methods including DWT, LP, CNN, and Transformer in target enhancement, boundary fidelity, and background suppression, but also shows clear advantages when compared with PCA-FLF from the literature on SAR and microwave data fusion. Furthermore, we apply the proposed fusion method to a downstream binary ship classification detection task based on the YOLO26 network, further validating its effectiveness in real detection scenarios. Ablation studies also confirm the contribution of each module within the proposed framework. The proposed active–passive fusion framework and analytical methodology provide a robust theoretical foundation for maritime ship detection under complex environments and adverse sea conditions.
A quasi-steady pressure decrease flow-radiation framework was established to investigate the evolution of rocket exhaust plumes under chamber depressurization conditions. In the simulations, the chamber pressure was reduced from 70 atm to 10 atm at several prescribed depressurization rates. Infrared (IR) radiative transfer within the plume was then evaluated using the discrete ordinate method (DOM) combined with a statistical narrow-band (SNB) model. The proposed approach was validated by comparison with experimental plume spectra and reference line-by-line (LBL) solutions. The flow dynamics and infrared thermal radiation characteristics of a typical solid rocket motor (SRM) plume under prescribed chamber-pressure decay were analyzed. The results show an approximately linear dependence between plume infrared radiation and the temporal evolution of chamber pressure. This relationship indicates that the chamber depressurization rate can be used to estimate the infrared thermal radiation of the plume during the engine shutdown process.
Recent advances in speech-driven facial animation have attracted significant interest across computer graphics, human-computer interaction systems, and immersive virtual reality applications. However, existing methods remain constrained by dependencies on specific reference videos or proprietary face mesh structures, limiting their applicability across diverse production pipelines and reducing compatibility with industry-standard animation workflows. To overcome these fundamental limitations in generalization and deployment flexibility, we propose Speech2Blend-an end-to-end hybrid convolutional-recurrent network that directly learns nonlinear speech-to-blendshape parameter mappings. This novel approach enables markerless speech-driven facial animation generation without restrictive inputs like video references or specialized facial rigs. Trained on the largest available digital human dataset (BEAT) and rigorously evaluated using three benchmark datasets with photorealistic visualization tools, Speech2Blend achieves state-of-the-art performance. It delivers superior audio-visual synchronization through learned temporal dynamics and reduces lip vertex error by 30% compared to existing baseline methods. These advances significantly lower production costs for virtual human speech animation while enabling cross-platform compatibility with common game engines and animation software.
Generating realistic and controllable aerial images is important for building and evaluating remote sensing recognition systems, especially when real samples of rare aircraft types or dense airport layouts are limited. However, airplane synthesis remains challenging for generic generative models. Aircraft have rigid and symmetric structures, and airport scenes often contain many closely spaced instances; as a result, existing models tend to produce distorted wings and fuselages or merge adjacent airplanes into ambiguous shapes. To address these issues, we propose AirplaneGen, a skeleton-guided latent diffusion framework for multi-airplane remote sensing image generation. AirplaneGen represents each airplane with an editable eight-keypoint skeleton and uses skeleton-derived soft masks to separate instance-level refinement from background-context modeling during denoising. To support this task, we construct MARS20, a benchmark with 2778 high-resolution aerial scenes and 16,673 airplane instances annotated with skeletons, categories, and contextual descriptions. Experiments on MARS20 show that AirplaneGen improves image fidelity, geometric consistency, and instance separation over representative controllable generation methods.
Predicting the complex spectral responses of metasurfaces is critical for their design, whereas conventional simulations are computationally expensive, and existing neural models often lack generalization and fail to maintain phase-amplitude consistency. Thus, a unified complex residual neural network (Uni-CRN) for forward modeling of diverse all-dielectric metasurfaces, including cylindrical and H-shaped structures is proposed in this paper. Uni-CRN integrates complex-valued operators with residual modules in a three-stage architecture-comprising input projection, stacked complex residual blocks, and output prediction-enabling direct learning in the complex domain while preserving gradient stability in deep networks. This unified framework allows the same model to handle multiple metasurface types with minimal modification. Experiments demonstrate that Uni-CRN achieves a composite mean squared error of 3.2 x 10-4, with amplitude and phase prediction fidelities of 95.50 % and 99.37 % on the cylindrical dataset, outperforming previous methods. The results highlight Uni-CRN as an efficient and general approach for metasurface spectral modeling, providing a robust foundation for inverse design and cross-structure transfer learning.
The prevalence of algorithm-driven personalization often traps users in information cocoons, creating an urgent need for tools that can break these constraints by synthesizing diverse perspectives. Inspired by the cognitive principles of distributed intelligence and social collaboration, this article presents a novel framework for multi-stance opinion summarization. Specifically, we first propose a novel large language model (LLM) agent collaboration framework for summarizing public opinions from multiple stances in the form of multi-turn conversations, where the LLM as controller and different agents as sub-tasks executors. Furthermore, fusing the online text content and visual multi-modal information based on image description to prompt LLM generate a comprehensive and stance-aware summary about public opinions. Finally, the efficacy of the overall framework is demonstrated through quantitative experiments on stance classification and extensive qualitative case studies on trending topics, confirming its strong capability in understanding and summarizing complex online public opinions. The results showcase a cognitively-inspired approach to information condensation, providing technical support for balanced decision-making for both individuals and organizations.
Most existing music recommendation systems struggle to perceive users’ implicit emotional states and fail to adapt dynamically to evolving preferences in emotionally rich, context-sensitive scenarios. To address this limitation, we propose an emotion-aware conversational music recommender built on a multiagent system. The system incorporates specialized agents for emotion recognition, semantic intent analysis, and contextual understanding. It distinguishes between explicit emotions, which are directly expressed by the user (e.g., “I feel anxious”), and implicit emotions inferred from contextual cues such as time, environment, or behaviors the user may not be fully aware of. A dual-memory mechanism models long-term musical preferences using a linear decay function and captures short-term, emotion-driven preferences using exponential decay. To enrich music content understanding, multisource information fusion combines streaming platform suggestions with rich metadata from external repositories. The system employs a large language model (LLM) to conduct multiturn dialogues and generate personalized, explainable recommendations. Experimental results show that the proposed approach significantly outperforms existing platforms (e.g., Spotify and Last.fm) in recommendation accuracy, ranking performance, and Hit Ratio@K. These findings underscore the effectiveness of integrating multiagent collaboration, emotion modeling, memory-augmented user profiling, and multisource data fusion for adaptive, user-centric music recommendation.
In this letter, a scattering center model for targets illuminated by vortex electromagnetic (EM) waves is presented. By employing the vector angular spectrum decomposition approach, the incident vortex beam is decomposed into a series of plane waves. Based on the geometric relationship between the target and the incident field, the whole scattering field is represented as a superposition of the scattering contributions from individual plane waves originating from all scattering centers of the target. The proposed approach is applied to three standard targets and a composite model to compute their bistatic scattering responses. The proposed scattering center model for vortex electromagnetic waves is validated using high-frequency approximation methods and the method of moments (MoM). The numerical results exhibit good agreement with the simulation data, demonstrating the effectiveness and accuracy of the proposed method. This method offers a promising tool for rapid scattering characteristics prediction in vortex EM wave radar detection and imaging applications.
The high volatility, seasonality and complex trading environment of agricultural futures markets present significant challenges to the dynamic adaptability of algorithmic trading systems. Traditional methods rely on fixed-scale image analysis and lack adaptability, making them unsuitable for varying volatility conditions. Therefore, this study proposes an AI agent for agricultural futures trading decisions. First, an agricultural futures adaptive volatility rate serves as a computational tool to dynamically identify high-volatility regions and generate finer-scale candlestick charts. Second, the agent uses Vision-Language Large Model with robust image comprehension to analyze the multiscale candlestick charts and extract key features such as trend direction and technical patterns. Subsequently, a Large Language Model with advanced natural language understanding and logical reasoning serves as the “decision-making brain”, evaluating market trends and making buy/sell decisions. Finally, the agent refines its decision logic through a multimodal feedback mechanism that combines numerical and textual information from ongoing interactions with the environment, thereby enhancing system adaptability and robustness. Experimental results indicate that this framework significantly improves the accuracy, stability, and risk control of trading strategies, offering valuable insights for human decision-making. In addition, it demonstrates potential applicability in other financial markets such as stock indices, energy, and major commodities, providing an innovative solution for intelligent trading in complex market conditions.
Social fintech envelopes social networks within financial concepts and tokenization, and its complexity requires innovative technological solutions and theories to meet user demands and achieve sustainability. The first thing is the mapping principle between physical real society and virtual digital world. Therefore, this article uses online game Nova Empire as a case study, aiming at using multiplex networks to understand the physical-digital structural consistency for social fintech sustainable development. Specifically, we first proposed an eight-layer multiplex network to model the complex gaming behaviors for players. Furthermore, we analyze the structural properties and social balance of the social networks. Particularly, the layers in multiplex networks composed of positive behaviors has higher reciprocity than the layers composed of negative behaviors, the out-degree distributions of nodes in the layers composed of negative behaviors basically conform to the power-law distribution, and the small-world phenomenon is also common in virtual game. The experimental results prove the structural consistency between physical and digital. Finally, new solutions for social fintech and mobile Internet industry are proposed based on the mapping principle, which will provide technical supports for the realization of sustainable development and social responsibility of social fintech.
A modified physical optics (PO) scattering algorithm for analyzing the high-frequency electromagnetic (EM) scattering of vortex wave fields is proposed in this Letter. While traditional PO methods struggle to accurately model edge diffraction effects in orbital angular momentum (OAM)-carrying beams, especially for complex targets, we address this limitation by integrating the equivalent edge current (EEC) theory into the conventional PO framework. By decomposing the Bessel vortex wave into a spectrum of plane wave components, we independently compute the PO and EEC scattering contributions for each spectral element. The coherent superposition of these contributions yields the total scattered field, significantly improving modeling accuracy. Numerical results demonstrate a 20%-30% enhancement in radar cross section prediction accuracy over the standard PO method. The proposed hybrid PO and EEC algorithm offers a computationally efficient solution for vortex EM scattering analysis, with direct applications in OAM-based radar systems, computational imaging, and next-generation communications.
By integrating unmanned aerial vehicle (UAV) remote sensing with advanced deep object detection techniques, it can achieve large-scale and high-throughput detection and counting of maize tassels. However, challenges arise from high sunlight, which can obscure features in reflective areas, and low sunlight, which hinders feature identification. Existing methods struggle to balance real-time performance and accuracy. In response to these challenges, we propose DLMNet, a lightweight network based on the YOLOv8 framework. DLMNet features: (1) an efficient channel and spatial attention mechanism (ECSA) that suppresses high sunlight reflection noise and enhances details under low sunlight conditions, and (2) a dynamic feature fusion module (DFFM) that improves tassel recognition through dynamic fusion of shallow and deep features. In addition, we built a maize tassel detection and counting dataset (MTDC-VS) with various sunlight conditions (low, normal, and high sunlight), containing 22,997 real maize tassel targets. Experimental results show that on the MTDC-VS dataset, DLMNet achieves a detection accuracy AP50 of 88.4%, which is 1.6% higher than the baseline YOLOv8 model, with a 31.3% reduction in the number of parameters. The counting metric R2 for DLMNet is 93.66%, which is 0.9% higher than YOLOv8. On the publicly available maize tassel detection and counting dataset (MTDC), DLMNet achieves an AP50 of 83.3%, which is 0.7% higher than YOLOv8, further demonstrating DLMNet’s excellent generalization ability. This study enhances the model’s adaptability to sunlight, enabling high performance under suboptimal conditions and offering insights for real-time intelligent agriculture monitoring with UAV technology.
The analysis of marine environmental parameters plays an important role in areas such as sea surface simulation modeling, analysis of sea clutter characteristics, and environmental monitoring. However, ocean observation remote sensing satellites typically deliver large volumes of data with limited spatial resolution, which often does not meet the precision requirements of practical applications. To overcome challenges in constructing high-resolution marine environmental parameters, this study conducts a systematic comparison of various interpolation techniques and deep learning models, aiming to develop a highly effective and efficient model optimized for enhancing the resolution of marine applications. Specifically, we incorporated adaptive global attention (AGA) mechanisms and a spatial gating unit (SGU) into the model. The AGA mechanism dynamically adjusts the weights of different regions in feature maps, enabling the model to focus more on critical spatial features and channel features. The SGU optimizes the utilization of spatial information by controlling the information transmission pathways. The experimental results indicate that for four types of marine environmental parameters from ERA5, our model achieves an overall PSNR of 44.0705, an SSIM of 0.9947, and an MAE of 0.2606 when the resolution is increased by a upscale factor of 2, as well as an overall PSNR of 35.5215, an SSIM of 0.9732, and an MAE of 0.8330 when the resolution is increased by an upscale factor of 4. These experiments demonstrate the model’s effectiveness in enhancing the spatial resolution of satellite-derived marine environmental parameters and its ability to be applied to any marine region, providing data support for many subsequent oceanic studies.
The measurement and analysis of the interaction between Bessel vortex electromagnetic (EM) and several standard targets are presented in this paper. With the aid of the angular spectrum expansion (ASE) method and physics optics (PO) theorem, scattering results on the plates (metal and dielectric) and a sphere could be derived. Furthermore, plane near-field scanning and near-far field conversion methods were implemented to compare the theoretical radar cross section (RCS). In the experiment, the quasi Bessel vortex wave was generated by a holographic metasurface antenna, and the whole measurement was performed in an anechoic chamber. The results of both the theory and measurement show that the scattered fields of the plate and sphere still had characteristics of the vortex EM wave, and the scientificity and accuracy of the measured RCS were verified. Our work involved a vortex scattering experiment in the microwave frequency band, which provides strong support for the application of vortex waves in radar detection and target recognition.
The widespread use of Internet has accelerated the explosive growth of data, which in turn leads to information overload and information confusion. This makes it difficult for us to communicate effectively in social groups, thereby intensifying the demands for emotional companionship. Therefore, we propose a novel social group chatting framework based on Large Language Model (LLM) powered multiple autonomous agents collaboration in this article. Specifically, BERTopic is used to extract topics from history chatting content for each social group everyday, and then multiple topics tracking is realised through multi-level association by adaptive time sliding-window mechanism and optimal matching. Furthermore, we use topic tracking architecture and prompts to design and implement an AI Chatbot system with different characters that can conduct natural language conversations with users in online social group. LLM, as the controller and coordinator of the whole AI Chatbot for sub-tasks, allows different AI Agents to autonomously decide whether to participate in current topic, how to generate response, and whether to propose a new topic. Each AI Agent has their own multi-store memory system based on the Atkinson-Shiffrin model. Finally, we construct a verification environment based on online game that is consistent with real society. Subjective and objective evaluation methods were deployed to perform qualitative and quantitative analyses to demonstrate the performance of our AI Chatbot system.
The smart healthcare system not only focuses on physical health but also on emotional health. Music therapy, as a non-pharmacological treatment method, has been widely used in clinical treatment, but music selection and generation still require manual intervention. AI music generation technology can assist people in relieving stress and providing more personalized and efficient music therapy support. However, existing AI music generation highly relies on the note generated at the current time to produce the note at the next time. This will lead to disharmonious results. The first reason is the small errors being ignored at the current generated note. This error will accumulate and spread continuously, and finally make the music become random. To solve this problem, we propose a music selection module to filter the errors of generated note. The multi-think mechanism is proposed to filter the result multiple times, so that the generated note is as accurate as possible, eliminating the impact of the results on the next generation process. The second reason is that the results of multiple generation of each music clip are not the same or even do not follow the same music rules. Therefore, in the inference phase, a voting mechanism is proposed in this paper to select the note that follow the music rules that most experimental results follow as the final result. The subjective and objective evaluations demonstrate the superiority of our proposed model in generation of more smooth music that conforms to music rules. This model provides strong support for clinical music therapy, and provides new ideas for the research and practice of emotional health therapy based on the Internet of Things.
Remote Sensing Image Generation (RSIG) offers a viable solution to the high data collection costs by facilitating the generation of large datasets. However, it is often hindered by the complexity of tasks and the diverse requirements of generated images. This paper presents the RS-Agent system, a novel approach that harnesses the capabilities of Large Language Models (LLMs) within an innovative agent solution, effectively addressing these issues. The system comprises a Task Agent, a Prompt Agent, and a LoRA Agent, each performing crucial roles in task decomposition, prompt generation, and scene-specific fine-tuning, respectively. Anchored by Diffusion models and employing natural language dialogues for interaction, the system aligns closely with user intent and produces high-quality results. Experimental results demonstrate the efficacy of the RS-Agent system in managing diverse RSIG tasks, adapting to various input forms, and generating high-quality results.
Vortex electromagnetic (EM) wave, which carries orbital angular momentum (OAM), possesses a spiral wavefront phase distinct from traditional plane waves. As a signal source, it can convey more information about the target, presenting novel possibilities for radar detection and target recognition. With the development of electromagnetic vortex radar systems and electromagnetic vortex imaging techniques, research on the backward scattering characteristics of vortex EM waves has become a focal point. This study focuses on the investigation of the backward scattering characteristics of symmetric targets under vortex EM wave. Utilizing the angular spectrum expansion method, the Bessel form of vortex EM waves of symmetrically incident targets are decomposed into a series of plane sub-waves propagating in different directions, maintaining symmetrical relationships in their incident directions. Subsequently, the target scattering characteristics of vortex EM waves are analyzed through the theory of plane waves. Combining the physical optics approximation, the scattering of individual plane sub-waves is calculated and superimposed to obtain the total backward scattering field of the target. The analysis indicates that, unlike plane waves, due to the unique phase distribution, the backward scattering results for axisymmetric targets illuminated by the same vortex EM wave are unequal. However, symmetrically irradiating an object with the vortex EM waves possessing of opposite topological charges yields identical backward scattering results. Finally, the backscattering characteristics of vortex EM wave on symmetric targets were simulated and validated using electromagnetic simulation software CST, confirming the accuracy of the conclusions. This research provides great significance for target detection and recognition in vortex EM wave radar applications.
Accurate modeling of sea clutter amplitude distribution plays a crucial role in enhancing the performance of marine radar. Due to variations in radar system parameters and oceanic environmental factors, sea clutter amplitude distribution exhibits multiple distribution types. Focusing solely on a single type of amplitude prediction lacks the necessary flexibility in practical applications. Therefore, based on the measured X-band radar sea clutter data from Yantai, China in 2022, this paper proposes a multi-task one-dimensional convolutional neural network (MT1DCNN) and designs a dedicated input feature set for the joint prediction of the type and parameters of sea clutter amplitude distribution. The results indicate that the MT1DCNN model achieves an F1 score of 97.4% for classifying sea clutter amplitude distribution types under HH polarization and a root-mean-square error (RMSE) of 0.746 for amplitude distribution parameter prediction. Under VV polarization, the F1 score is 96.74% and the RMSE is 1.071. By learning the associations between sea clutter amplitude distribution types and parameters, the model’s predictions become more accurate and reliable, providing significant technical support for maritime target detection.