
This study is an extension of our previously published manuscript [2]. Building upon the foundations laid in our earlier work, this research further refines and extends the transformer-based approach to enhance its diagnostic capabilities. The focus remains on the integration of multi-modal information, combining both textual clinical narratives and imaging data from multi slice CT scans, to provide a more comprehensive and accurate diagnosis of brain strokes. In addition, we introduce a new ensemble learning approach based on logic gates. This innovative method combines information from clinical narratives and CT scans to enhance the diagnostic capabilities of our framework, marking a significant evolution in our approach to advancing precision in brain stroke detection.
We discuss a generalization of geometric algebras known as ternary Clifford algebras. In these objects, we have a fixed ternary form instead of a quadratic form as in ordinary geometric algebras. Basis-free definitions of the determinant, trace, and characteristic polynomial in ternary Clifford algebra are introduced. Explicit formulas are presented for all coefficients of the characteristic polynomial and inverse in ternary Clifford algebra. The operation of Hermitian transpose (Hermitian conjugation) in ternary Clifford algebra is introduced without using the corresponding matrix representation. We present a natural realization of the unitary Lie group SU(3) , which is important for physical applications, using only operations in ternary Clifford algebra. An explicit basis of the corresponding Lie algebra 𝔰𝔲(3) is presented. We present an explicit connection with the well-known Gell-Mann basis of 𝔰𝔲(3) . The results can be used in physics, computer science, and engineering.
The goal of this research is to identify surgical movements using RGB videos. Prior methods for recognizing surgical gestures mainly concentrate on robot-assisted surgery, but less on surgical gesture training. In this paper, we introduce a novel surgical gesture dataset that comprises the human surgical gestures of experienced surgeons captured from multiple viewpoints. Moreover, we propose a spatial-temporal transformer-based surgical gesture recognition model called TransSG to identify the user’s surgical gestures better. Our framework comprises three key elements: (1) the importance-weight layer that learns the significance of segmented patches before passing them to the spatial encoder, (2) the aggregated temporal feature extractor that facilitates computationally efficient processing, and (3) a novel loss function that integrates a range of effective loss functions to enhance recognition quality. Our experiments demonstrate that TransSG attains excellent Top-1 accuracy (83.6
To capture 3D geometric details, point cloud understanding primarily employs convolution, dynamic map, or transformer methods to build complex local geometry extractors. When dealing with sparse and noisy point clouds in real-world environments, the structural design of these methods tends to be complex and perform poorly. In this paper, we propose a lightweight point cloud understanding framework based on Laplacian feature convolution, which has high accuracy on normal-sized point cloud datasets and performs well when dealing with sparse and noisy point cloud inputs. The adaptive Laplacian convolution module combined with the vector attention mechanism can effectively capture the local details of point cloud geometry. The framework employs a multi-scale feature fusion mechanism between modules to reduce the number of sampling points step by step, reducing redundant computations and achieving a lightweight effect. In addition, we propose a parameter-free module for interpolation to enhance the training part. Experimental results on object-level datasets show that the lightweight method proposed in our paper has similar accuracy to the current best method in terms of accuracy (Acc) and mean intersection over union (mIoU), and has better robustness in handling sparse and noisy point clouds.
Enhancing immersive environments in an efficient way is essential across various industries. However, the task of selecting and placing suitable objects to decorate 3D scenes often requires a substantial investment of time and expertise. We present a novel semi-automatic approach to accelerate this process by fusing state-of-the-art generative text-to-image models with object detection. We use a rendering of a plain 3D scene as input for Stable Diffusion and generate an image depicting a decorated version. With object detection, we identify the location, size and class of decorative elements in the generated image. Based on this information, we populate the original 3D scene with matching assets from a database. Our results demonstrate the capabilities of our proposed approach in enhancing 3D scenes through the incorporation of decorative models across different scenarios. We show that our method seamlessly integrates into existing 3D content creation tools highlighting its applicability within established workflows.
Segmentation of regions of interest (ROIs) from medical images is considered as a fundamental requirement for many imaging-based clinical decision support systems. Deep learning based automatic segmentation methods have emerged as the state-of-the-art. However, their performance is heavily reliant on using a large number of annotated training dataset to encompass all the possible variations of ROIs, and when the coverage of these variations is inadequate, these methods have difficulties in segmenting the ROIs. Semi-automatic segmentation methods, which fuse user-inputs with high-level semantic image features derived from convolutional neural networks (CNNs) or vision transformers (ViTs) offer an alternative to overcome the limitations of automatic segmentation methods. Unfortunately, CNNs are restricted by the limited receptive fields that cannot capture long-range dependencies, while ViT usually rely on the self-attention mechanism to capture the global context and therefore have high computational complexity. In this study, we propose a Semi-Mamba method for semi-automatic medical image segmentation. The novelty we introduce is to leverage the capability of Mamba to efficiently capture long-range dependency and combine this together with user-clicks to segment challenging ROIs, e.g., ROIs with fuzzy boundaries and inhomogeneous textures. Our experiments with three well benchmarked medical image datasets across different modalities showed that our method consistently outperformed existing automatic and semi-automatic segmentation, which demonstrates strong generalizability.
With the rapid increase of complexity and volume of 3D building models, industries such as digital games and computer-aided design face considerable challenges. One feasible solution is to generate low-polygon models for complex building meshes in a relatively short time to apply in a level of detail (LOD) algorithm. We propose an efficient building model simplification algorithm with opening detection, aiming to realize an extremely low polygon count while preserving their visual features. Our algorithm has three steps. We detect and construct the opening structures presented in the building model to incorporate in the subsequent carving process; improve the approach to generate a visual hull that represents the fundamental boundary of the building model and carve out details with high efficiency; and apply planarization and simplification algorithms to obtain the simplified model. Our algorithm can generate highly simplified building models while preserving important opening details to serve as the roughest level in the LOD hierarchy. Experimental results demonstrate that our algorithm efficiently handles building models with complex opening structures and outperforms other recent simplification algorithms.
With the ongoing evolution of intelligent sensors and artificial intelligence, wearable devices have achieved remarkable breakthroughs. Smart clothing, encompassing a diverse array of types, has found widespread applications across various life domains, particularly in the healthcare sector. However, despite these advancements, current smart clothing still faces limitations such as limited interaction capabilities, significant application constraints, and inadequate monitoring functionalities. To overcome these challenges and further enhance the user experience, this paper proposes a new smart clothing system. The system has the capability to monitor multiple physiological parameters of the human body. Furthermore, it incorporates a new algorithm called ARDN specifically designed for initial arrhythmia detection. To validate the effectiveness of the arrhythmia detection algorithm, an experimental analysis was conducted using the MIT-BIH arrhythmia database. Furthermore, to facilitate real-time monitoring and effective human-computer interaction, the system incorporates digital twin (DT) technology. This integration of intelligent sensors and terminal equipment through DT technology enables the system to conduct a comprehensive analysis of the users' physical condition, thereby providing an initial health diagnosis. The system presents novel opportunities for advancing the development of personalized healthcare, facilitating seamless integration into everyday life, and enhancing users' well-being.
Network structures in images, resembling graphs with nodes and edges, are easily recognized by humans but challenging to extract computationally. Traditional methods for road extraction involve thinning semantic segmentation of road network images, followed by curve segmentation and linear approximation to form a graph. These methods struggle with image interference and lack precision in identifying intersections and turns. Supervised machine learning shows promise but requires extensive sample creation and labeling. This study introduces a conventional computer vision method using variable-sized windows for extracting topological structures from binary network images, demonstrating strong resistance to interference. We validated our method through a comparative analysis with a Harris corner detection approach, using binary road structure images from Hangzhou, China. Our results highlight the superior anti-interference capabilities of the variable window recognition method, especially in node recognition, compared to the Harris corner detection method.
A robotic arm equipped with a vision system can significantly improve production efficiency. However, due to the complexity of the robotic arm’s working environment, the useful information about the objects recognized by the visual system is significantly reduced. To solve this problem, we propose a real-time semantic segmentation network to achieve high-precision and efficient robotic arm grasp. First, we propose an enhanced feature extraction module for the encoding stage, which can improve feature extraction capabilities without increasing or even reducing the amount of model calculations and parameters. Then, we propose an enhanced feature reconstruction module for the decoding stage, which can better preserve essential features in the feature map restoration stage. To validate the efficacy of our proposed approach, experiments are conducted on the rigid dataset Jacquard and the flexible dataset Electric Wires. Experimental results show that our method can quickly and accurately identify target objects, achieving the best trade-off between accuracy and inference speed. Our method achieves an mIOU of 94.66
Category-level object pose estimation plays a crucial role in a wide range of practical applications by accurately predicting the poses and sizes of unseen objects within a specific category. However, accurately estimating object poses remains a significant challenge due to substantial shape variations within the same category. To address this issue, this paper introduces a novel learning network for object pose estimation that is guided by a shape descriptor. By capturing the geometric information of an object's shape, the shape descriptor provides valuable input for subsequent feature learning, effectively handling shape variations. Moreover, our framework incorporates a confidence-based pose estimator, which assigns confidence scores to each pose prediction. This integration allows for the acquisition of more accurate poses with higher confidence by penalizing poses with low confidence. Experimental results on the CAMERA25 and REAL275 datasets demonstrate the superiority of our approach over state-of-the-art methods.
The rapid development of artificial intelligence has brought many innovative achievements to fields such as drug design and drug discovery [1]. Combining traditional graphics methods with deep learning methods can significantly improve the accuracy of drug screening. Hypoxia refers to the process in which tissues or cells in the body undergo abnormal changes in morphology, physiological functions, and metabolism due to insufficient oxygen supply or oxygen utilization obstacles [2]. In the previous study, Enlargement of the cell nucleus is a recognizable morphological feature of hypoxia cells [3]. Based on a deep learning multi-cell image classification model, this study constructed a high-throughput compound screening system for discriminating the anti-hypoxia activity of thousands of compounds. By simultaneously performing prediction scoring using AC16 and H9C2 models, the anti-hypoxia activity of thousands of compounds was predicted, and some compound molecules with anti-hypoxia effects were successfully screened.
In this work we report recently obtained results on the full geometric product of two oriented points in conformal geometric algebra Cl(4, 1) that models Euclidean space. We analyze its Cl(3, 0) multivector coefficients and their relationships. Moreover, we derive explicit formulas for extracting from the geometric product of two oriented points both the center position of the point pair and the radius vector of the point pair (half of the point to point distance vector). Finally, we introduce a simple direct method to obtain the two standard conformal points from the geometric product of two oriented points.
Realistic image super-resolution (RISR) has been a challenging research topic in image restoration, aiming to address complex degradation. However, existing methods often struggle to handle various unknown degradation factors presented in low-quality images, limiting their effectiveness to simplify degradation models. This gap between current RISR methods and real-world scenarios hinders their ability to generate realistic details. In this paper, we propose a novel realistic image super-resolution method based on stable diffusion to enhance degradation perception and detail generation. The framework comprises a Degradation-aware Module (DAM), a Detail-enhanced Module (DEM), and a General Restoration Module (GRM). DAM adjusts the weights of deep features in different channels using the Residual Composite Attention Network (RCAN) to remove perceived degradation, producing a smooth image containing only essential information. DEM employs a feature enhancement structure from coarse to fine to transform low-dimensional feature data obtained by downsampling the image data into corresponding high-resolution image data. GRM utilizes a pre-trained stable diffusion model based on the inverse diffusion process, and we add Advanced Nets in the diffusion stage to provide additional semantic information for noise restoration. Experimental results demonstrate the superiority of our method over existing state-of-the-art methods on both synthetic and real-world datasets.
The automatic generation of implants holds significant importance in cranial repair surgery. Recent studies have attempted to use diffusion models to complete defective skull point clouds and obtain implants through a voxelization network. However, its multi-step sampling in point cloud diffusion is slow ( ∼ 1,000 s), often requiring tens of inference steps to get satisfactory results. In this paper, we explore a recent method called Rectified Flow, which straightens the trajectories of probability flows, and enables one-step generation while maintaining high quality. Moreover, we finetune the voxelization network to enhance its adaptability to the output of the point cloud completion network, thereby reducing the iteration of training required. Leveraging our new pipeline, each implant requires ∼ 56 s to generate, whereas point cloud completion alone consumes just ∼ 0.75 s. To evaluate the effectiveness of our method, we conducted experiments on SkullBreak and SkullFix datasets. The results, measured using the Dice score (DSC), the 10mm boundary DSC (bDSC), and 95 percentile Hausdorff Distance (HD95) metrics showed a clear performance advantage compared to other proposed methods.
While rice pests and diseases significantly impact crop yields, existing deep learning methods for their detection face challenges with accuracy and deployment complexity. Addressing these issues, this study proposes the YOLOv8-HSFPN, an advanced detection framework. Firstly, it features an innovative High-level Select Feature Pyramid Network (HSFPN) neck network that effectively integrates high-level and low-level feature sets for enhanced feature fusion. Secondly, the addition of a deformable self-attention module further refines the model’s adaptability to the varying shapes and locations of targets, dynamically adjusting to the salient features. The proposed model has undergone comparative and ablation studies alongside YOLOv8, YOLOv9, and YOLOv5, confirming its improved accuracy and streamlined deployment. This integration results in a robust detection model that not only marks a significant leap in accuracy, evidenced by a 3
Despite the high accuracy achieved by current pedestrian detection technologies, their performance is significantly hindered in foggy weather conditions due to image color distortion and an increased false alarm. This often leads to false positives, false negative, and other related issues. To address these limitations, we propose an enhanced FDW-YOLOv8 (Feature-Enhancing-Attention Dark Channel Prior Wise-IOU YOLOv8) model based on the YOLOv8 algorithm. This model incorporates FEAttention(Feature-Enhancing-Attention) mechanism into the original YOLOv8 network to improve the discriminability and robustness of extracted features, thus reducing the miss rate(MR). Furthermore, we employ the WiseIOU(WIoU) loss function to expedite model convergence and optimize the accuracy of anchor frames, effectively handling challenges posed by occlusions and scale variations in foggy environments. Prior to model prediction, we utilize the Dark Channel Prior(DCP) defogging algorithm to preprocess input images, effectively extracting pedestrian information while suppressing irrelevant details, thus minimizing the false positive rate (FPR). To evaluate the performance of our model, we create a specialized Foggy-Pedestrian dataset, focusing on pedestrian detection in foggy conditions. Our experiments demonstrate that the proposed FDW-YOLOv8 model significantly outperforms the benchmark model in terms of reducing FPR and MR, exhibiting high generalization and practical applicability for real-world applications in foggy environments.
Interacting with object instances is crucial in specific embodied tasks such as rescue and industrial services, not just navigating to the target. Previous research on Object Goal Navigation (ObjNav) has only addressed navigating to the object's vicinity. However, the interface between approaching the object and finely localizing and grasping it has yet to be adequately investigated. This paper presents the Finding and Grasping (FaG) framework, which is based on the mobile manipulator, to address the last-mile ObjNav problem. The framework incorporates an object localization module and a sub-target extraction module into a traditional autonomous exploration and navigation pipeline. The object localization module uses a grasp estimation method based on point cloud completion to determine the global goal. In another module, a sub-target extraction algorithm is used to break down global goals into navigation and grasping sub-goals. The proposed framework effectively solves the ObjNav last-mile problem and is superior to the baseline in all settings, as demonstrated by sufficient experiments. Furthermore, real robot grasping experiments were conducted to verify the method's feasibility. A demo video is available at: https://youtu.be/AGlIO-DBM1g
Large-scale computational fluid dynamics methods rely on efficient and accurate mesh generation algorithms, especially when handling complex geometries, as mesh generation in such cases can be time-consuming. This paper presents a new approach to mesh generation, integrating the multi-layer lattice Boltzmann method with moving boundaries. This technique minimizes the impact of machine errors on computational geometry, thereby improving the precision of geometric object recognition within the mesh. Furthermore, it significantly cuts down the time needed for mesh generation. Theoretical analyses and numerical experiments reveal that the time required for mesh generation in this study has been reduced by an order of magnitude, facilitating the rapid creation of multi-layer grid nodes numbering in the tens of millions within seconds. By combining moving grids with moving boundaries, this method verifies airflow around airfoils in various grid setups. Moreover, we have extended its application to three-dimensional space, effectively simulating submarine movements. In scenarios with moving boundaries, where mesh generation is necessary at each time step, this method particularly shines. It requires minimal additional computation time, which is a significant advantage. This makes it a practical solution for mesh generation in large-scale simulations. Such simulations often use the multi-layer lattice Boltzmann method and involve moving boundary flow.