
As deep neural networks (DNNs) are becoming vital to numerous safety-critical applications, ensuring their fault reliability is crucial. This paper presents a hybrid framework that combines genetic algorithms with fault injection and analytical methods to identify vulnerable neurons and layers in DNNs. Validated on LeNet-5, AlexNet, and VGG-11, our approach significantly enhances model accuracy in harsh environments by protecting critical components, achieving reliability improvements (accuracy drop of the model compared to the golden network after fault injection) of $69.86 \%$ for LeNet, and 99.65% for AlexNet at BER=1E-4, and 91.37% improvement for VGG at $\mathrm{BER}=1 \mathrm{E}-6$. Furthermore, our framework reduces computational costs by requiring fewer inferences than traditional analytical methods, highlighting its potential to improve DNN accelerators’ reliability and contribute to their safety.
In 3D reconstruction, creating accurate models from a single 2 D image remains a complex challenge in computer vision and virtual environments. Acknowledging this complexity, our study presents a research approach that aims to advance the field beyond conventional reconstruction methods. We introduce a novel method designed to reconstruct 3D models from single damaged 2D images while tackling the added challenge of restoring missing details. Our method employs an expandedscale stable diffusion model for inpainting the input 2D image, restoring missing information and enhancing depth estimation for sparse-view 3D reconstruction. By integrating inpainting within a scalable diffusion framework, we achieve improved fidelity to the original structure, even with limited sparse mesh data. Additionally, our framework incorporates multi-view stereo processing to optimize available viewpoint information, enhancing reconstruction accuracy while maintaining computational efficiency.Experimental results indicate a significant advancement in transforming single 2 D image inpainting into effective 3 D object reconstruction. By bridging inpainting and 3D reconstruction, our approach recovers missing or occluded details in the 2D image and enhances the accuracy of 3D modeling from sparse viewpoints. The inpainting technique is important here, as it restores essential details within the 2D input, directly improving the quality and fidelity of the 3D reconstruction output. Our integrated method of inpainting and sparse-view reconstruction outperforms existing approaches, achieving higher accuracy in 3D object representation from limited data.
Chemical experimentation remains fundamental to material discovery and compound synthesis, yet conventional methods face critical limitations. Manual laboratory workflows require repetitive operations that prolong research timelines, while robotic automation platforms inadvertently shift the burden to chemists through complex programming requirements. Existing systems also struggle to adapt to novel experimental designs, constraining scientific innovation. To bridge these gaps, we present a large language model (LLM)-driven robotic framework that redefines automated experimentation. Our system interprets natural language instructions through advanced LLM processing, dynamically integrates real-time environmental data, and autonomously generates executable code utilizing a modular action library. A self-correcting validation mechanism ensures operational precision, systematically resolving execution errors through iterative feedback. Empirical validation across two benchmark experiments including the iodine clock reaction kinetics and the salt purification demonstrates transformative advantages: 1) intuitive interaction; 2) extended workflow handling efficiency gain in multi-step protocols; 3) adaptive code generation. This paradigm shift may enable chemists to focus on hypothesis-driven research rather than procedural implementation, establishing a scalable infrastructure for next-generation autonomous laboratories.
With the development of artificial intelligence technology, large language models have shown tremendous potential in professional document generation. However, in document generation tasks with high professionalism and strict content requirements, such as project proposals, ensuring the accuracy and professionalism of generated content remains challenging. This paper proposes a project proposal generation method based on modular knowledge-enhanced large language models. First, this method divides the complex proposal generation task into interrelated submodules such as project background, technical approach, and implementation plan through modular decomposition, and designs a complete intermodule relationship processing mechanism. Second, it constructs a knowledge injection framework based on professional literature, providing professional domain knowledge support for the model through multi-level literature retrieval, knowledge extraction, and organization. Finally, it implements an automatic proposal generation system based on microservice architecture, integrating core functional modules such as literature processing, content generation, and quality control. Experimental results show that this method can significantly improve the professionalism and accuracy of proposal generation, providing new ideas and methods for AI-assisted professional document writing.
Coronary artery disease (CAD) is a leading cause of death worldwide. X-ray coronary angiograms are a diagnostic imaging technique, considered as the benchmark for identifying vascular anomalies. The two-dimensional nature of coronary angiograms limits spatial understanding that can complicate the diagnosis and cause misinformation. This paper presents a multi-stage pipeline leveraging deep learning techniques for reconstructing three-dimensional coronary vasculature from twodimensional X-ray angiograms. The pipeline involves various stages such as pre-processing, keyframe selection, image registration, segmentation and three-dimensional reconstruction, using image processing and deep learning techniques. Additionally, a synthetically generated coronary tree dataset is used to train the neural network for reconstruction and ensure accurate centerline predictions. The method is successful in generating precise reconstructions, overcoming the spatial limitations of twodimensional imaging.
This study explores the role of artificial intelligence (AI) and computing techniques in enhancing customer satisfaction in mobile banking through best practices, cybersecurity, and data privacy. Critical elements influencing customer satisfaction—system reliability, technological innovation, privacy policies, security measures, and customer engagement—were examined using data collected from bank customers and employees via a validated questionnaire. Ordinal logistic regression a machinelearning approach was used to develop the model regarding the relationships between the independent variables and their influence on customer satisfaction. The results highlighted AI-driven technological innovations that strongly focus on strategies centered on customer needs that enhance customer satisfaction. Moreover, strong cybersecurity measures and data privacy policies are necessary to foster customer trust and loyalty. The study indicated no significant differences in the perceptions of cybersecurity and data privacy as perceive by bank employees and bank clients. This indicated the need to have a comprehensive mobile banking strategy The practical application of machine learning in mobile banking was illustrated in this research, showing how ordinal logistic regression which is a technique of machine-learning should be integrated into AI systems specifically in the analysis of complicated data about customer satisfaction. By using AI powered analytics, banks can enhance their capability to address dynamic customer needs. This study contributes to the body of knowledge in areas such as customer experience, enhances customer delivery, strengthens security and emphasizes the potential to drive data-informed decision-making.
Optical Coherence Tomography Angiography (OCTA) is a non-invasive imaging technique widely used in clinical practice for retinal blood vessel imaging. Indicators of retinal blood vessel such as density, diameter, and curvature are closely associated with various retinal diseases, including Diabetic Retinopathy (DR) and Age-related Macular Degeneration (AMD). Obtaining these indicators relies on accurate blood vessel segmentation. Therefore, blood vessel segmentation serves as the foundation for automated diagnosis of retinal diseases. However, blood vessel segmentation from OCTA images is often affected by some noises, such as Gaussian noise and shadow noise. This study proposes a novel deep learning model called the comparative constraint model. “Comparative” means comparing the outputs of multiple models, while “constraint” means restricting the model produce the same output for images with different noises. The key idea is to train the model to produce consistent segmentation results for images with and without added noise, thereby enhancing the noise resistance of model. The effectiveness of the proposed method has been validated on a publicly available dataset, OCTA-500-6M, achieving an average Dice of $\mathbf{8 8. 5 8 \%}$ and an average $\mathbf{I o U}$ of $\mathbf{7 9. 6 0 \%}$. These results are significantly better than those obtained using a standard U-Net.
This paper investigates the potential of Probabilistic Models combined with Generative Artificial Intelligence (GenAI) in mass media and entertainment, with a specific focus on its application in radio broadcasting. We present a methodology that applies Markov models to develop an Agentic framework capable of generating engaging radio show segments. The proposed system aims to produce engaging audio commentary, recommend and play relevant music, and integrate demographically contextual daily news and weather updates, ensuring both content coherence and the organic structure of the generated programme.
This paper explores the use of a Deep Reinforcement Learning (DRL) model for dynamic portfolio management in the financial market. With the help of deep neural networks and the Twin-Delayed Deep Deterministic Policy Gradient (TD3) algorithm, the framework is able to process high dimensional market data and dynamic environment. The TD3 algorithm incorporates transaction costs and risk aversion constraints in order to simulate the environment of real-world investments. It uses features such as the Moving Average ConvergenceDivergence (MACD) and Relative Strength Index (RSI) to construct its decision-making state space. The performance of the model was assessed using the historical data of six NASDAQ stocks, starting from 2021 to 2023. The results obtained were then compared with two other methods, namely Proximal Policy Optimization (PPO) and Deep Deterministic Policy Gradient (DDPG). Some of the performance measures of the portfolio include cumulative return, annual volatility, Sharpe ratio and the maximum drawdown. The TD3 algorithm produced better results in terms of cumulative and risk return, where it got a $51.28 \%$ cumulative return as compared to a cumulative return of ${2 5. 9 1 \%}$ by DDPG and $17.56 \%$ by PPO. However, the TD3 managed portfolio was accompanied by high annual volatility and drawdown suggesting a risk-return paradox. From the results, it is shown that the TD3 policy is able to produce high returns while maintaining a certain level of risk and outperform static strategies like buy and hold. However, it also included some drawbacks, where the model was not able to forecast the short-term movements of the market and was based on lagging indicators.
Federated Learning (FL) is rapidly becoming a popular cooperative and distributed approach, utilized by edge devices to develop machine learning models. In this research, we present a high-efficiency FL network designed for analyzing healthcare data, leveraging VPN technology and implementing a cross-silo methodology across a wireless backhaul network. Our detailed evaluation revealed that the FedProx algorithm combined with mmWave technology. It significantly enhances accuracy and decreases convergence time from 55 to 38 seconds, highlighting the benefits of high-bandwidth communication links. Furthermore, we established a comprehensive three-tier security strategy. This strategy starts with the integration of a private network into the telecom framework. Powered by licensed frequency channels and reinforced by VPN-based protections which provides the extensive security for the FL network and its critical data.
When outdoors, nowadays it is common to rely on Global Navigation Satellite Systems (GNSS) based navigation to find a location. However, for indoor environments no common solution exists, as GNSS positioning is not available indoors. While many substitute technologies rely on infrastructure installed in buildings, e.g., beacons, in this paper we use the magnetic field characteristics of buildings as a solution that is available everywhere. In the implementation, a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) model are used for classification of the characteristics and regression, respectively. The approaches are evaluated against each other on a public dataset showing that the magnetic field can be a robust ubiquitous solution for indoor localization. The regression with LSTM shows the highest precision, while the error of a classification approach is constrained by the building boundaries and enables the usage of class confidence values for further processing.
The studies of Unmanned aerial vehicles (UAVs) have received much attention due to their wide applications in the courier delivery, target detection, precision agriculture, and disaster rescue and others. Path planning is one of the core technologies in the UAV field, aiming to provide feasible, safe, and optimal flight paths for UAVs to meet the flying requirements in complex environments. Effective path planning can improve the adaptability of UAVs in various complex environments. Traditional path planning methods, although capable of finding reasonable paths, often suffer from slow convergence speed and the tendency to fall into local optima. To address these shortcomings, this paper proposes an improved Genetic Algorithm (MLGA) combining Meta-learning and local search mechanisms for UAV target detection task path planning. The algorithm can dynamically adjust parameters based on the performance at different stages, allowing it to self-optimize according to task changes, while enhancing its local search capability. This leads to improved solution quality and faster convergence. Simulation results show that MLGA outperforms several other algorithms in solving path planning problems.
With the advent of the Industry 4.0 era, the demand for precise time synchronization and high-speed data communication has been steadily increasing across various industries. In this context, Time-Sensitive Networking (TSN), as a significant enhancement of traditional Ethernet protocols, is gradually becoming a core enabling technology in multiple fields due to its technological advantages, including accurate time synchronization, high bandwidth, low latency, and high determinism. This paper first reviews the major protocol standards of Time-Sensitive Networking (TSN) and its current state of development, providing a detailed summary of the key characteristics of TSN technology. Next, it focuses on analyzing the latest research advancements in the field of TSN testing, both in academia and industry, covering the development of theoretical frameworks, the design of verification schemes, and the construction of testing tools and experimental platforms. Furthermore, this paper summarizes and reviews the technical features and application practices of TSN testing products developed by several companies, highlighting the efforts and achievements of the industry in advancing the maturity of TSN testing technologies. Finally, the paper concludes by summarizing the key points and proposing potential directions for future research.
AI represents a crucial information development that is vital to human innovation and progress. Agent AI further integrates AI into human life. As networks transform into AI, humans may develop an addiction-like dependence on AI. This study primarily investigates the impact of human AI addiction on decision quality and job satisfaction. The sample consists of ${1 4 0}$ firm surveys, with hypotheses verified using SPSS 23.0. Results indicate that Agent AI addiction affects decision quality and job satisfaction. In particular, decision quality in AAI serves as a key mediating factor. This enhances organizational theory regarding AI usage behavior while highlighting the importance of decision quality in AI design.
Alignment in Artificial Intelligence systems (i.e., behavior that is agreed upon by its designers) is a challenging aspect of such systems, requiring contextual understanding of actions and how they fit within a world-model. In this paper, we show that Large Language Models fail to correctly infer behavioral context for certain actions as these are decidable only by embodied agents; i.e., context depends on physical properties that are not completely captured by language alone. We investigate whether prompting techniques can mitigate this limitation by augmenting a Large Language Model with specific alignment prompts, and evaluate behavior on the Machiavelli benchmark. Results suggest this approach is promising, especially when combined with properly curated training data for alignment on actions that can be contextualized purely on language. We highlight further research directions on the alignment problem as we move toward fully-embodied Artificial Intelligence.
Efficient decision-making in Cyber Physical Systems is hindered by their inherent non-deterministic nature. Statistical methods are required to obtain well performant solutions, especially in the context of multi-agent systems that interfere with one another. We experimentally evaluate Multi-Armed Bandits, a class of statistical methods that balances the trade-off between exploration and exploitation, on Cyber Physical Games, a specific instance of emergent properties of multi-agent Cyber Physical Systems. We show Multi-Armed Bandits can efficiently identify adequate strategies for decision-making, outperforming baseline approaches.
Intelligent driving without relying on offline HDmaps has become one of the hot spots in the industry, and many scholars recently focus on building environmental maps online. In this paper, we effectively combine the color, texture information of the image and the position, intensity information in the LiDAR point cloud to realize the multimodal fusion detection of road elements, and propose the FusionMapper algorithm. In the proposed method, we design a dual multi-scale temporal feature fusion (DMTFF) module, which combine feature information of past frames and different scales to improve the continuity and robustness of the network for linear target detection such as lane lines and road curbs. Further, we use the Sparse Map Query based Transformer decoding layer, combined with novel Lane key-point query, Curb key-point query and crosswalk key-point query to achieve inference of three types of target key points, while effectively eliminating the time-consuming postprocessing of existing algorithms. The quantitative results show that the proposed algorithm has strong robustness and excellent accuracy. Specifically, our method reached $70.16 \% \mathrm{mAP}$ on nuScenes map dataset, which outperforms some industry-leading algorithms.
There are multiple advantages that multi-robot mapping has over single robot mapping: faster exploration speed, higher redundancy, and a possibility to use multiple robots in the same mapping framework. However, merging when the maps are not homogeneous is a challenging problem. In this paper a map fusion method for heterogeneous occupancy grid maps is proposed that can fuse different resolution and quality maps by considering that the maps may have varying region quality and that the resulting maps must be separate for each robotic system.
Multimodal social relation extraction requires effective feature fusion to recognize relationships across various targets. However, existing work often struggles to model the finer correlations among various modalities and ignores the higherorder complex relationships between multimodal information. To overcome current limits, a multimodal relationship extraction method based on hypergraph attention neural networks is proposed. The method can encode the higher-order data correlations in the hypergraph structure, obtain the higher-order complex relationships between multimodal information, and reduce the noise generated in the process of hypergraph encoding and enhance the correlation with text semantics through the cross-diffusion attention mechanism. Ultimately, it is combined with the original unimodal text and visual input to enhance the inference ability. Experimental results on three character social relationship datasets, Dream of the Red Chamber (DRC-TF), Water Margin (OM-TF), and Four Classics (FC-TF), in which TF indicates the inclusion of both textual and facial image data, clearly show the advantages of our and proposed methodology.
The process of converting scientific research papers into patent documents is crucial for the commercialization of research outcomes. However, this process typically requires specialized knowledge and a significant amount of time. This study presents an AI-powered system designed to automatically generate patent documents from research papers. The system utilizes large language models to generate patent documents. By designing specific prompts and calling APIs, the system can convert research papers into patent documents that meet the requirements. Experimental results show that the patent documents generated by the system have significantly improved in terms of quality and coherence, potentially reducing the time and effort required for manual document preparation.