Real-world optimization challenges frequently involve computationally expensive evaluations, necessitating efficient optimization strategies. To address the demands of medium-scale expensive optimization problems, this research proposed a novel Surrogate Assisted Evolutionary Algorithm (SAEA): the Weighted Committee-Based Surrogate-Assisted Differential Evolution Framework (WCBDEF). This framework combines principles from active learning and ensemble learning, iteratively interrogating the most ambiguous and high-fidelity solutions to ensure judicious allocation of evaluation resources. WCBDEF employs a dual sampling criterion, with offline optimization dedicated to exploration and online optimization focused on exploitation. Comparison experiments with 4 classical and 4 state-of-the-art SAEAs were conducted on 18 benchmark functions to verify the effectiveness of WCBDEF on medium-sized expensive optimization problems respectively. Moreover, its application in optimizing operational parameters for two Enhanced Geothermal Systems (EGS) models has resulted in a significant reduction in the Levelized Cost of Electricity (LCOE), surpassing existing algorithmic solutions. The results show that WCBDEF is a competitive alternative for medium-scale expensive optimization problems.
In recent years, sparse large-scale multiobjective optimization problems (LSMOPs) have found widespread application in real-world scenarios and have become a focus of evolutionary computing research. Due to the high dimensionality of decision variables in LSMOPs, evolutionary algorithms (EAs) often struggle to efficiently find optimal solutions. In an effort to settle this difficulty, we raise an enhanced sparse multiobjective evolutionary algorithm (ESMOEA) that uses the strongly convex sparse (SCSparse) operator to optimize the decision variables, which can further enhance the sparsity of solutions. Additionally, to consider the sparsity property of solutions during variable grouping, the parameter in the sparse operator that represents whether the solution becomes sparse is ingeniously incorporated into the proposed sparse grouping technique. To evaluate the performance of the proposed ESMOEA, a set of experiments is carried out on both benchmark and real-world problems. The experimental results indicate that the proposed ESMOEA achieves superior performance compared to existing large-scale multiobjective evolutionary algorithms (MOEAs).
Structured illumination microscopy (SIM) has emerged as a pivotal super-resolution technique in biological imaging. This review aims to introduce the fundamental principles of SIM, primarily focuses on the latest developments in super-resolution SIM imaging, such as the light illumination and modulation devices, and the image reconstruction algorithms. Additionally, the application of deep learning (DL) technology in SIM imaging is explored, which is employed to enhance image quality, accelerate imaging and reconstruction speed or replace the current image reconstruction method. Furthermore, the key evaluation metrics are proposed and discussed for assessment of deep-learning neural networks, especially for their employment in SIM. Finally, the future integration of artificial intelligence (AI) with SIM system and the perspective of smart microscope are also discussed.
Spectral imagery object detection involves identifying and localizing specific targets in multispectral (MSI) or hyperspectral (HSI) imagery. Leveraging more spectral bands than visible light, spectral images provide superior recognition capabilities in complex environments such as low visibility or adverse weather, benefiting applications in remote sensing, medical imaging, agriculture, and security. Visible light imaging delivers high-resolution spatial details, including fine textures and object shapes, while spectral imaging captures thermal radiation and material properties, exhibiting strong robustness under low-light conditions. Fusing these two modalities can significantly improve detection accuracy. However, existing methods often suffer from spectral redundancy, weak inter-modal interactions, and limited multi-scale local correlations, leading to performance degradation. To address these issues, we propose FA-YOLO, which adopts a Feature Aggregation and Attention Network (FAANet) as its backbone. FAANet integrates a Feature Interaction Transformer (FIT) module that explicitly reduces spectral redundancy by reconstructing and refining modality-specific features, thereby enhancing cross-modal information exchange under high dimensionality and modality discrepancies. In addition, we introduce the Multi-Scale Hybrid Attention Module (MSHAM) to robustly handle objects of varying scales. MSHAM combines spatial and channel attention across multiple receptive fields, strengthening local correlations and multi-scale feature representation, effectively capturing cross-modal interactions and establishing long-range dependencies. This integration improves both feature extraction and detection accuracy. Experimental results show that in low-visibility scenarios such as nighttime and foggy conditions, FA-YOLO achieves a 7.7
Nowadays fashion image retrieval is one of the key tasks in the e-commerce sphere. In recent years, several different approaches have emerged to help solve this problem. In this paper, the algorithm for fashion image retrieval in the e-commerce sphere using YOLOv8 and attention model is presented. The innovation of the proposed algorithm is its modular architecture, which allows for more efficient extraction of image features and their further analysis using an attention model. One more advantage of the algorithm is the possibility to use its modules independently so that not the only image retrieval task could be solved. The algorithm is validated in the real e-commerce application and compared with existing e-commerce retrieval algorithms proving to successfully solve the task. Experiments showed that the developed algorithm achieved state-of-the-art results, reaching an accuracy of 94
In this paper, we propose a pipeline algorithm for analysis of human gait in video for orthopedic diagnostic tasks. It is based on 3D human pose estimation that is used to detect human skeletons in 3D space. Images have depth blur in the third dimension, so obtaining accurate 3D pose estimates is more difficult. In this experiment, pre-training masks the data and uses Transformer based on the addition of parallel temporal and spatial feature modules to complete the 2D pose estimation of 3D pose. For such representations, there is a possibility to calculate angles at the articulation of the joint lines of the skeleton, changing such angles during human motion, and define the pathological direction of motion. This analysis calculates the changing of angles between key points of the constructed skeleton during the walking cycle. Angles are defined as dynamical characteristics. Such characteristics are very important for the diagnostics of orthopedic diseases, the definition of the treatment pipeline, and the determination of the type of corrective footwear.
Stochastic partial differential equations (SPDEs) are commonly encountered in the realms of engineering and computational science. Solving SPDEs can be regarded as quantifying the impact of stochastic inputs on system responses or quantities of interest, which constitutes performing uncertainty quantification (UQ) for SPDEs. Recently, the application of neural networks to solve SPDEs has attracted considerable attention due to their potential to outperform traditional numerical solvers in computational efficiency. However, the challenge of enhancing the accuracy of neural network approaches for UQ in SPDEs remains largely unresolved. In this study, we develop neural networks capable of flexibly addressing Neumann boundary conditions while simultaneously relaxing the smoothness requirements. By avoiding the need for higher-order derivatives in the loss function, our approach demonstrates clear advantages. Numerical experiments have confirmed that our method substantially surpasses several established neural network approaches to improve the accuracy of UQ.
This paper proposes an approach for tracking the behavior of people in a group on video by using convolutional neural networks. At the beginning, definitions of group movement of people are given, and features for accompaniment are defined that can be used to analyze people’s behavior. Next, an algorithm is proposed for calculating the distance between people in video, which includes three stages: detection and tracking of objects, coordinate transformation, calculation of the distance between people and detection of distance violations. The results of experimental studies and comparison with known algorithms are presented, which confirms the effectiveness of the algorithm.
Fuzzy systems have been employed in many fields and achieved remarkable results. However, due to the inherent characteristics, there is still a challenge for them to deal with high-dimensional data which includes a large number of features and becomes very common in the current era of big data. Most of the literature employ the product, minimum, softmin and its adaptive versions to compute the firing strengths in fuzzy system modeling, where only the first two are standard T-norms conforming to the theoretical basis of fuzzy logic reasoning. But easy to cause numeric underflow and nondifferentiability are their disadvantages, respectively. To comply with fuzzy theory basis and alleviate the aforementioned predicament of dimensionality, this article investigates how to design high-dimensional Takagi-Sugeno-Kang (TSK) fuzzy system based on Dombi T-norm, which has not been explored for fuzzy system modeling to the best of our knowledge. We first build Dombi T-norm based TSK (DombiTSK) model, then an adaptive strategy is designed for the index parameter of Dombi T-norm to enhance the performance of DombiTSK, which results in so-called adaptive DombiTSK (ADMTSK) fuzzy system. To further improve the design and performance of ADMTSK, we give a novel membership function with positive lower bound to match the use of adaptive Dombi T-norm, which is constructed based on Gaussian membership function (GMF) and named as composite GMF (CGMF). The experiments are conducted on high-dimensional classification datasets with feature dimensions varying from 1024 to 120432 to evaluate the proposed methodologies. Experimental results verify the advantages of adaptive Dombi T-norm and CGMF. The comparison results and statistical tests among ADMTSK and other state-of-the-art fuzzy algorithms demonstrate that our proposed model outperforms its rivals.
Urban green space (UGS) vegetation plays an important role in mitigating the urban heat island effect by improving the environment and quality of life. Hence, there is a dire necessity for urban planning and management to precisely obtain the spatial distribution and structural information employing high-resolution data. Nevertheless, the limitations of remote sensing (RS) data and the complexity of urban landscapes pose significant challenges, so this study aims to introduce a method to classify UGS vegetation more precisely by integrating high spatial resolution multi-spectral and oblique photography images captured by unmanned aerial vehicle (UAV). A novel canopy height model (CHM) method is proposed to generate UGS vegetation information for urban areas while addressing the errors associated with traditional approaches in estimating non-ground vegetation heights, achieving a total Mean Absolute Error (MAE) of 0.17 m and an overall accuracy of 95.03 %. The proposed UGS mapping method combines spectral features, canopy height information, vegetation indices (VIs), and texture features to evaluate the impact of various characteristics on classification accuracy. The obtained experimental results show that by incorporating canopy height information classification accuracy is significantly improved and achieve overall accuracy of 93.82 % and Kappa coefficient of 0.91. Moreover, the proposed method not only precisely reflects the structure and distribution of UGS vegetation by showing specific advantages in complex environments but also offers a new arena for UGS vegetation classification based on the integration of multiple features.
In road infrastructure monitoring, the demand for automated pavement damage detection is growing, but traditional methods relying on manual inspection or expensive equipment struggle with large-scale, real-time detection. Deep learning-based object detection offers an efficient solution, yet challenges remain in computational constraints, environmental variations, and diverse damage types. An ideal model must balance accuracy and efficiency for deployment in embedded devices, drones, and edge computing. Compared to two-stage models like Faster R-CNN, YOLO series models, particularly YOLOv9, optimize performance with PGI, GELAN structures, and reversible functions, making them suitable for constrained environments. While YOLOv9 has higher computational overhead than YOLOv8, its superior detection accuracy enhances its potential in resource-limited settings. To improve adaptability, we integrate transfer learning and semisupervised learning, reducing training complexity and enhancing generalization. Our improved YOLOv9-based method achieves efficient, high-precision detection with low computational costs. Experiments on the China Motorcycle and Japan datasets show that the YOLOv9s-TLS model improves mAP50 by 0.5 % and F 1-score by 1.4 % , validating the effectiveness of transfer learning in cross-environment detection.
The paper proposes the noninvasive image egg growing monitoring method based on an illumination and transfer learning. During the egg growing, the size of egg air cell is increased. The segmentation is performed to extract cells and segmentation parameters are adjusted and trained on an air cell datasets by transfer learning to separate air cells with high light transmittance from the background. The improved DeepLabV3+ network model for image egg monitoring is proposed. The network embeds coordinate attention in the lightweight network MobilenetV2. The decoder feature fusion method is improved to a semantic embedding branch structure. The middle-level features that have been newly introduced are merged with the high-level features and low-level features. The results show that the mean intersection over union of the model reaches 89.06% and that the mean pixel accuracy rate reaches 94.66%. The method can effectively segment the air cell part of the eggs. The feasibility of the method was verified by measuring the air cells of egg growing process from the 7th to the 19th day.
The paper proposes a new approach for crowd movement type estimation in video by combining convolutional neural network and integral optical flow. At first, main notions of crowd detection and tracking are given. Secondly, crowd movement features and parameters are defined. Three rules are proposed to identify direct crowd motion. Signs are presented for identifying chaotic crowd movement. Region movement indicators are introduced to analyze the movement of a group of people or a crowd. Thirdly, an algorithm of crowd movement types estimation using convolutional neural network and integral optical flow is proposed. We calculate crowd movement trajectories and show how they can be used to analyze behavior and divide crowds into groups of people. Experimental results show that with the help of convolutional neural network and integral optical flow crowd movement parameters can be calculated more accurately and quickly. The algorithm demonstrates stronger robustness to noise and the ability to get more accurate boundaries of moving objects.
Recently, theory-guided neural networks have attracted significant attention in solving partial differential equations due to their minimal data requirements and alignment with physical laws. However, selecting the penalty coefficient by incorporating physical laws as penalizing terms in the loss function undoubtedly affects the model's performance. In this paper, we utilize physical laws to guide U-Net neural networks and employ bilevel programming to tune the loss function's hyperparameters, thereby improving the overall model performance. The upper-level variables within the framework are optimized using a water flow optimizer based on the golden spiral eddying function, which enhances population diversity and broadens the search range in the turbulent flow operator. Experimental results confirm the effectiveness of this approach in handling stochastic partial differential equations.
The Buckley-Leverett (BL) equation of plane radial flow describes the flow state of two phases of oil and water in an ideal planar porous medium. It is a nonlinear hyperbolic conservation partial differential equation (PDE) with a jump value, making it difficult to solve numerically. This paper presents a physics-informed neural network (PINN) with entropy constraints to solve the BL equation of plane radial flow. Specifically, the Oleinik entropy condition is used to limit the change of the relationship between water content and saturation in accordance with the physical law, which does not change the structure of the target PDE to be fitted by the model. Benefit from the interpretability of PINN and its fit in solving PDE problems, the model can be trained in an unsupervised way, thus eliminating the effort of obtaining sample labels. The experimental results show that the root mean square error fluctuates between the interval [0.03436,0.16054], indicating a good fitting effect. Especially the jump value, which refers to the position of the leading edge of water saturation, can be clearly output by the model.
Particle swarm optimization algorithms play an important role in various optimization problems by simulating the collaboration and information sharing among individuals in natural systems and using the local optimal solutions of individuals in the population to guide the global search. In this paper, based on the particle swarm algorithm with pyramid topology, the adaptive pyramid particle swarm algorithm (APPSO) is proposed by innovating the update of particles through the use of parameter adaptive competition and cooperation strate-gies. The algorithm adds adaptive strategies to improve the accuracy performance of the algorithm and introduces the concept of iterative stagnation to maintain the diversity of the population. In this paper, the proposed algorithm (APPSO) is compared with six algorithms on the CEC2022 test set, and extensive experimental results show that APPSO has the ability to accurately converge to the global optimum.
The article describes the results of Belarusian scientists in the field of image recognition and analysis obtained over the last 30 years. The main groups of scientists working at the National Academy of Sciences of Belarus and universities are shown. Conferences held by Belarusian scientists are shown, with special attention paid to the most popular conference Pattern Recognition and Information Processing . Information is provided on the development of the knowledge-intensive IT sector and companies working in the field of computer vision. Data on the training of scientific personnel in this field are provided.
The main scientific results of the 16th International Conference on Pattern Recognition and Information Processing (PRIP-2023), Minsk, Republic of Belarus, October 2023, are reviewed and analyzed. The history of this series of conferences is outlined, and its significant role in the development of the theory and practice of image analysis, pattern recognition, and artificial intelligence is indicated. A list of articles in the special issue is provided, prepared from reports selected by the PRIP-2023 Program Committee.
Broad learning system (BLS) has emerged as an effective alternative to deep learning. Due to the linear output characteristics of BLS, the generalized inverse method based on ridge regression is commonly used to solve for the hidden layer weights. However, when dealing with large matri-ces, the generalized inverse algorithm often underperforms. Moreover, it considers all features simultaneously during computation, resulting in redundancy and inefficiency. To address these issues, we first extend the RCD algorithm for multi-output problems, enabling parallel computation of multiple systems. Furthermore, the introduction of a novel iterative update strategy allows for the incorporation of randomness while capturing important low-probability feature information, thereby achieving superior results. Thus, we propose an enhanced parallel randomized coordinate de-scent (EPRCD) algorithm for weight calculation. Numerical experiments demonstrate that the newly proposed algorithm performs well on hlah-dimcnsional datasets.
Physics-informed neural networks (PINNs) incorporate physical constraints into their loss functions, allowing them to efficiently solve Partial Differential Equations (PDEs). In this work, we introduce an innovative network structure based on PINNs, specifically designed to address fluid-structure interaction (FSI) problems in porous media, referred to as M-PINN. Our approach is inspired by fully coupled traditional numerical methods, leveraging a multi-task learning (MTL) framework with soft parameter sharing to enhance the interactions across different physical fields. This alignment leads to an improved representation of multi-physics coupling phenomena. Experiments on the constant pressure production problem demonstrate that our method provides better performance than both traditional and sequential PINN approaches.