In modern medicine, X-ray images are widely used as an important diagnostic tool for detecting and evaluating a variety of diseases such as bone fractures, lung diseases, and tumors. However, traditional X-ray image recognition relies on the experience and expertise of doctors, which involves the risk of misdiagnosis and missed diagnosis, especially in areas with limited medical resources or inexperienced doctors. In order to improve the accuracy and efficiency of diagnosis, the application of deep learning techniques becomes increasingly important. In this study, a capsule network-based fracture recognition model for X-ray images is proposed, and the advantages of capsule networks in capturing complex features of images are verified by training and testing the model on X-ray images from different anatomical regions. The experimental results show that the CapsNet model achieves an accuracy of 84.2% in fracture recognition, which is better than the traditional neural network model, demonstrating its potential and application value in medical imaging diagnosis.
With the development of compression ignition (CI) engines, a detailed investigation of the in-cylinder combustion process is needed under strengthened regulation on the emission. The development of a tomographic algorithm makes it possible for further investigation of the 3D structure of the in-cylinder flame. In this study, the Full-field Cross-Interfaces Computed Tomography algorithm (FCICT) is utilized to reconstruct the 3D chemiluminescence distribution of the in-cylinder flame in an optical engine. The reconstructed flame provides 3D information on exhibiting flame equivalence radius variation, as well as clarifying the effect of injection pressure on flame development. The study reveals the 3D structure of the in-cylinder flame and makes it possible for further investigation of the 3D information of the in-cylinder combustion process in the future.
With the growing demand for sustainable animal products, intelligent pig farming has become key to improving the efficiency of livestock agriculture. This paper employs the advanced Real-Time Detection Transformer (RTDETR) algorithm, based on Transformer technology, aiming to enhance the accuracy of pig behavior recognition. To adapt to the variable challenges of real-world environments, we have applied various data augmentation techniques to the Edinburgh Pig Behavior Video Dataset, including random cropping, horizontal flipping, RandomBrightnessContrast, and perspective transformation. By training and comparing the RTDETR with the YOLOv5 model, we observed the exceptional performance of RTDETR in complex environments, achieving a mean Average Precision (mAP) of 0.82. This research not only provides a robust algorithmic foundation for intelligent pig farming but also emphasizes the significant advantages of the Transformer-based RTDETR algorithm in improving the accuracy of pig behavior recognition in challenging environments.
Neural Style Transfer (NST) has exerted algorithms to generate animation images in computer vision for decades. The Convolution Neural Network (CNN) applied to image content and styles in the NST has improved the extraction of functionalities and the calculation of the convergence speed to recognize and generate high quality structure images, but unpredictable loss elements are inadequate to iterate human learning ability of unique artists’ paintings or styles. This paper offers a chaotic VGG10 NST model based on CNN, ReLU and Lee-Oscillator. The proposed ReLU-Oscillator dynamically relies on activation functions in a chaotic state that dynamically improves high-frequency iterative training of high-quality image and high-speed time optimization. In addition, the best parameters recognizing and determining a personalized painting style from the search for ReLU-Oscillator by corresponding to a set of parameters from proposed Optimal Oscillator Parameter Search Algorithm. Experimental results showed that the stylized image generated by the Chaotic VGG 10 model with high-frequency oscillation succeeded in reducing the training time in magnitude models with the smallest Params and FLOPs in model performance and image quality with the lowest content loss for preserving semantic information and moderate style loss for style similarity balanced with comparison to 8 state-of-the-art models in visual perception evaluation. Chaotic NST has a unique identification for each artist supplemented with a set of oscillator parameters to evaluate the loss performance based on the relative error between the famous painting and its imitation, indicating that the authentication of paintings can be detected through specific ReLU-Oscillator parameters for each stylization, estimated from the loss value performance.
Three-dimensional (3D) tomographic reconstruction in confined-space requires a mapping relationship which considers the refraction distortion caused by optical walls. In this work, a tomography method, namely full-field cross-interface computed tomography (FCICT), is proposed to solve confine-space problems. The FCICT method utilizes Snell’s law and reverse ray-tracing to analytically correct imaging distortion and establishes the mapping relationship from 3D measurement domain to 2D images. Numerical phantom study is first employed to validate the FCICT method. Afterwards, the FCICT is applied on the experimental reconstruction of an illuminated two-phase jet flow which is initially generated inside an optical cylinder and then gradually moves outside. The comparison between accurately reconstructed liquid jet by FCICT and coarse result by traditional open space tomography algorithm provides a practical validation of FCICT. Based on the 3D liquid jet reconstructions at different time sequences, the distributions of surface velocity and 3D curvatures are calculated, and their correspondences are systematically analyzed. It is found that the velocity of a surface point is positively correlated with the mean curvature at the same point, which indicates the concavity/convexity of liquid jet surface is possibly in accordance with the surface velocity. Moreover, the surface velocity presents monotonical increasing trend with larger Gaussian curvature for elliptic surface points only, due to the dominated Brownian motion as the liquid jet develops.
Sign language is a process by which people with speech and hearing disabilities to communicate with the world. It uses body movements to simulate syllables and form corresponding words to convey information. With the introduction of the CLIP model, which makes multi-modal tasks possible, there is a new solution for sign language recognition and translation. This paper proposes a novel framework for sign language translation, CL-MarianMT, which can effectively combine a sign language recognition model with a SOTA model for translation. In the part of sign language recognition, the video feature vectors are extracted by the Video Encoder of CLIP4clip architecture, and then the feature vectors are input into the Transformer model, and the recognized sentences are sent into the fine-tuned translation model MarianMT to realize the translation from English to Chinese. The experiment shows that the fine-tuned MarianMT improves the Chinese translation ability of sign language. The study in this paper can serve the purpose of enabling deaf-mute people to communicate in different languages and facilitate hearing-impaired people in language communication. At the same time, it can also allow hearing-impaired people to communicate with people with different language systems, which have a certain social value.
Isolated sign language recognition has been an important part of breaking down communication bottlenecks for deaf-mute and others. While facing this problem, the purpose of this paper is to classify American isolated sign language video by modeling pose, hands and face keypoints representation. Specifically, this paper introduces a novel framework whose main components are the altered Dense Predictive Coding (DPC) pre-trained model and the Encoder pre-trained model. The DPC model is trained using self-supervised learning to obtain representation of pose and hands keypoints. The Encoder model is trained using supervised learning to obtain representation of face keypoints. Combining the altered DPC model with image inductive biases and the Encoder model with a self-attention mechanism, the final combined model achieves 0.81 on the test set of the ISAL dataset, outperforming the current open-source solution by a significant margin.
Verifying individual Chinese handwritten signatures is an essential biometric technology that is widely used in banking, finance, and legal business. The forging of signatures for the purpose of cheating is a serious detriment to the interests of these industries. This paper proposes a Siamese network verification signature based on image domain transfer. The Siamese network uses a convolutional neural network as a sub-network to build the structure of Siamese network by combining the genuine signature network and the unauthenticated signature network. Each sub-network transfers ImageNet weights for training in the Chinese recognition task so that the new weighted image domain is suitable for the Chinese signature image domain. This Siamese network is trained with new weights to determine the authenticity of the signature. Currently, there is no publicly available dataset of Chinese handwritten signatures. This paper develops the Chinese signatures dataset from 295 induvial persons, 885 different persons participated, including approximately 9,000 pictures. The Siamese network achieves 92.75% accuracy on the test set, and the verification time for a single Chinese handwritten signature is 0.32 seconds. Finally, the Siamese network model achieves more than 90% accuracy on the public handwritten signature datasets CEDAR, BHSig-B and BHSig-H in three different languages, and the experiments demonstrate the good generalization of the proposed method.
This work reports an improved tomography method to solve three-dimensional (3D) reconstructions in confined space with enhanced calculation efficiency and accuracy compared to other similar approaches. Confined-space tomography methods are designed to correct the image distortion on recorded target images caused by light refraction through optical walls, such as optical engine cylinders. However, past confined space tomography methods have shortcomings in reconstruction accuracy and time efficiency, since they usually involve time-consuming iterations or numerical interpolation during calculating the mapping relationship from 3D measurement domain to 2D imaging planes. Therefore, based on the improvement and innovation of our existing confined space tomography methods, the present method developed in this work directly calculates the mapping relationship by performing reverse ray-tracings originated from imaging planes, then decides the intersection volumes with the discretized measurement domain. Numerical and experimental demonstrations of present method are, respectively, performed based on multiple simulated phantoms and a two-branch laminar flame contained inside an optical cylinder. Compared to past confined space tomography algorithms, the present method consumes ~ 40% of the computational time under the voxel size of 0.5 mm, along with slightly enhanced accuracy. Moreover, the present method becomes more efficient under smaller voxel sizes. The robustness of present method and its endurance on measurement errors are then systematically analyzed and demonstrated.
Abstract This work reports an optimized tomography method, termed Direct-Mapping Cross-Interfaces Computed Tomography (DMCICT), with enhanced calculation efficiency and accuracy for three-dimensional (3D) reconstruction in confined space. Confined-space tomography methods are designed to correct the image distortion on recorded target images caused by light refraction through optical walls, such as optical engine cylinders. However, past confined-space tomography methods have shortcomings in reconstruction accuracy and time efficiency, since they usually involve time-consuming iterations or numerical interpolation during calculating the mapping relationship from 3D measurement domain to 2D imaging planes. There, DMCICT is developed in this work to directly calculating the mapping relationship by performing reverse ray-tracings originated from imaging planes, then decide the intersection volumes with discretized measurement domain. Numerical and experimental validations of DMCICT are respectively performed based on multiple simulated phantoms and a two-branch laminar flame contained inside an optical cylinder. Compared to past confined-space reconstructions, DMCICT can reduce more than 50% of the computational time in majority of tested cases, while the reconstruction accuracy is also significantly enhanced. Moreover, DMCICT demonstrates the robustness under different spatial resolution conditions and presents solid endurance on measurement errors.
Practical applications of computed tomography (CT) in optical engines require an advanced algorithm that can correct the light refraction via optical windows and reconstruct the 3D signal field partially blocked by structural obstacles. In this work, an advanced CT algorithm is designed for optical engines to simultaneously eliminate the imaging distortion by refraction and diminish the reconstruction errors using partial signal blocking. By combining the pinhole model and Snell's law, the ray tracings from discretized 3D voxels in the measurement domain to 2D pixels in the imaging planes are accurately calculated, thus restoring the distortion in recorded projections. Besides, by deciding the locations and numbers of voxels that actually participate in iterative CT calculation, the iterative update process of voxel intensity becomes independent of the blocked rays, reducing the reconstruction errors. The algorithm is then numerically validated by reconstructing a simulated signal phantom inside an optical cylinder with a lightproof obstacle between the phantom and a recording camera, which imitates the refraction and blocking conditions in practical optical engines. Moreover, experimental demonstration is performed by reconstructing practical premixed flames inside optical engines. Both the simulation and the experiment present significantly enhanced flame chemiluminescence reconstruction by applying the optimized CT algorithm compared to the original algorithm utilized in open space applications.
In 2021, the World Health Organization estimates that there are approximately 70 million deaf mutes in the world. At present, the method that facilitates the communication between normal people and deaf mutes is still not widely available. In the era of rapid development in the field of artificial intelligence, sign language recognition technology based on deep learning and mining human visual and cognitive laws has become an effective tool. In this paper, a Transformer based end-to-end continuous sign language sentence recognition model (TrCLR) is established. The CLIP4Clip video retrieval method is used for feature extraction, and the overall model framework uses an end-to-end Transformer structure. The sign language data set (CSL data set) is used as the data of this experiment. Nine sign language recognition models are used for experimental comparison on this data set. The experimental results show that the accuracy of TrCLR reaches 96.3%, which is 13.9% improvement over the best results of other models. Our model promotes the communication between normal people and deaf-mute people, and contributes to the establishment of a barrier free society.
With around 1.5 billion people worldwide suffering from hearing impairment, it is particularly important to communicate between non-disabled people and people with hearing or speech impairment and to build a barrier-free society. Multi-modal learning provides an excellent artificial intelligence channel for this purpose. In this article, we create an End-to-end Chinese Lip-Reading Recognition System based on multi-modal fusion to implement Chinese lip translation in order to facilitate communication between individuals with hearing impairment. Our system adopts the End-to-end Audio-visual feature fusion Lip-reading Recognition Architecture (EALRA), with feature extraction based on a MobileNet0.25 tuned CNN skeleton and the encoder back-end using the Conformer self-attentive convolution encoder for modelling. The largest Chinese Mandarin Lip-Reading (CMLR) was selected as the dataset for the empirical study, and the performance metric for Chinese lip recognition was the character error rate (CER). The results of our experiments show that the CER metric of EALRA in the lip-recognition model is 8.0, which is on average 23.74% lower than the CER metrics of other lip-recognition models, indicating that EALRA performs better in fusing image features and audio features.
This work reports the development and validation of a new imaging sensor arrangement optimization method for volumetric tomography, namely effective voxel corrections maximization (EVCM). Different from past optimization approaches that only studied the influence of sensor orientations on the reconstruction accuracy, the EVCM method considers the impact of both sensor orientations and measured target distribution. Combining all recorded target projections, the EVCM first determines the number of effective voxels, indicating the voxels that participate in tomographic calculation. The method then calculates the number of corrections that modifies the value of effective voxels within a single reconstruction iteration step. The ratio between numbers of corrections and effective voxels (R-E) is established as a new criterion for the sensor arrangement optimization to achieve improved reconstruction accuracy. Both numerical simulations on signal phantoms and controlled experiments on lab-scale flames are employed to validate the EVCM. Comparisons between EVCM and other sensor arrangement optimization methods are also performed based on the 3D reconstructions both numerically and experimentally. Results show that EVCM turns out to be a more accurate way to decide the optimal arrangements specified for different 3D targets.
This work reports the modification and optimization of a computed tomography (CT) algorithm to become capable of resolving an optical field with internal optical blockage (IOB) present. The IOB-practically, the opaque mechanical parts installed inside the measurement domain-prevents a portion of emitted light from transmitting to optical sensors. Such blockage disrupts the line-of-sight intensity integration on recorded projections and eventually leads to incorrect reconstructions. In the modified algorithm developed in this work, the positions of the obstacle are measured a priori, and then the discretized optical fields (i.e., voxels) are classified as those that participate in the CT process (named effective voxels) and those that are expelled, based on the relative positions of the imaging sensors, IOB, and light signal distribution. Finally, the effective voxels can be iteratively reconstructed by combining their projections on sensors that provide direct observation. Moreover, the impact of IOB on reconstruction accuracy is discussed under different sensor arrangements to provide hands-on guidance on sensor orientation selection in practical CT problems. The modified algorithm and sensor arrangement strategy are both numerically and experimentally validated by simulated phantoms and a two-branch premixed laminar flame in this work.
This study reports a new, to the best of our knowledge, view registration method that can achieve high-quality tomographic reconstruction in spite of a large view registration (VR) error. The correlation-based view registration (CBVR) method is a directional orientation modification method based on the cross-correlation between measured projections and ray-tracings generated from the reconstruction, which can reduce the gross VR error to moderate levels by iterations. In the CBVR method, a traditional multi-camera VR process is first performed, based on the sensitivity of the projections to the VR error, and are evaluated and quantified for all cameras. Afterward, the orientation of each camera is iteratively updated based on the cross-correlation of the measured projections and the ray-tracings generated from the reconstruction calculated through all other cameras. The CBVR is consecutively validated by numerical and experimental studies. Through a numerical study on a controlled phantom introduced with 2% Gaussian noise, the CBVR method is proved to be able to reduce the large VR error (up to 4.8°) to 0.2° as well as to reduce the reconstruction error to ∼6.7% in 12 rounds of iterations, which is very close to that obtained without any VR error (6% caused by Gaussian noise only). The CBVR method is then demonstrated and validated by reconstructing a two-branch laminar flame. By implementing the method, the initial projection orientations are optimized from traditional multi-camera VR results within a range of ±3∘, leading to effectively improved tomographic reconstruction of flame chemiluminescence distribution.
Tomographic approaches in confined space require advanced imaging algorithms that can properly consider the refractive distortion as the imaging rays pass through the optical wall. Our previous work established an algorithm (cross-interfaces computed tomography, CICT) and practically solved tomographic problems in confined space. However, critical restriction was found in CICT, which is that images simulated at small azimuth angles are contaminated with noticeable signal loss and become unusable. Based on this recognition, this work has developed an improved tomography approach, namely, full-field cross-interfaces computed tomography (FCICT), to extend the available view angles to all perspectives. The key to this approach involves the 3D domain discretization using voxel parallelepipeds instead of traditional voxel layers to establish the ray-tracing relationship between imaging planes and the measurement domain. The imaging process of FCICT is first validated by quantitatively comparing the grid imaging locations in measured and simulated projections of a calibration plate. By evenly distributing the view angles in the whole azimuth angle range, the FCICT reconstruction is then numerically validated by reconstructing a simulated double-cone flame phantom. The reconstruction presents a high correlation coefficient of ${\sim}{98}\%$ with the original phantom. Finally, the FCICT is employed to reconstruct an ethylene-air premixed flame. Comparisons show that re-projections generated by the FCICT reconstruction are in accordance with measured flame images, with the mean correlation coefficients of more than 95%.
This work reports the development and validation of a new tomography approach, termed cross-interfaces computed tomography (CICT), to address confined-space tomography problems. Many practical tomography problems require imaging through optical walls, which may encounter light refractions that seriously influence the imaging process and deteriorate the three-dimensional (3D) reconstruction. Past efforts have primarily focused on developing open-space tomography algorithms, but these algorithms are not extendable to confined-space problems unless the imaging process from the 3D target and its line-of-sight two-dimensional (2D) images (defined as “projections”) is properly adjusted. The CICT approach is therefore proposed in this work to establish an algorithm describing the mapping relationship between the optical signal field of the target and its projections. The CICT imaging algorithm is first validated by quantitatively comparing measured and simulated projections of a calibration plate through an optical cylinder. Then the CICT reconstruction is numerically and experimentally validated using a simulated flame phantom and a laminar cone flame, respectively. Compared to reconstructions formed by traditional open-space tomography, the CICT approach is demonstrated to be capable of resolving confined-space problems with significantly improved accuracy.