IntroductionGastrointestinal diseases (GIDs) are a major global health burden, but traditional endoscopic screening faces limitations due to its invasiveness. Tongue image analysis offers a non‑invasive alternative but is subjective, relying on clinical experience.MethodsTo address this, we propose a novel computational framework that synergistically integrates clinical prior knowledge with data-driven learning for non-invasive screening via tongue image analysis. Our framework introduces two key innovations: a TransNeXt hybrid backbone to extract comprehensive local and global tongue features, and a Channel Attention Gated Fusion module that performs asymmetric feature recalibration and prior-guided adaptive integration.ResultsEvaluated on a multicenter gastrointestinal disease tongue image dataset, our framework demonstrates robust performance, with an SVM classifier achieving an accuracy of 0.874, a Macro-F1 score of 0.879, and an AUC of 0.898.DiscussionAblation studies confirm the critical synergy between multimodal feature fusion and data augmentation. This work establishes an effective and objective computational solution for tongue diagnosis, with strong potential for clinical translation as an auxiliary screening tool.
Objective assessment of schizophrenia using eye-tracking data remains challenging, particularly in heterogeneous stimulus settings, where cross-task variability may undermine the stability of subject-level representations. Existing methods often rely on low-level spatiotemporal features or coarse fusion schemes, limiting their ability to capture complementary information across tasks. To address this issue, we propose MIFNet, a multi-information fusion framework for cross-paradigm eye-tracking analysis. MIFNet unifies heterogeneous eye-tracking inputs, including directly recorded gaze coordinates and video-derived gaze trajectories, and jointly models visual exploration behavior from semantic, spatial, and temporal perspectives. It further incorporates a center-based fusion strategy, in which SFP extracts an initial compact central representation, BlockLite models cross-stimulus dependencies, and ACR iteratively refines the central representation to improve stability. Experiments on two datasets demonstrate the effectiveness of the proposed framework.
Objective. Accurate depression classification using fNIRS signals is critical for objective auxiliary diagnosis, yet many existing methods separately model static characteristics or temporal dynamics, limiting their ability to capture coordinated brain activity.Approach. This study proposes a static-dynamic fusion network for fNIRS-based depression classification that jointly models stable hemodynamic and network properties together with time-varying functional connectivity. A static feature encoder captures global and channel-wise traits, while a dynamic feature encoder employs sequential modeling with attention mechanisms to identify temporally salient neural variations. A channel-aware fusion module is further introduced to adaptively integrate static and dynamic features based on their diagnostic relevance.Main results. Validated on 352 subjects undergoing verbal fluency tasks, static and dynamic feature-fusion network achieved 91.2% accuracy on 90 s signal segments and 81.5% on 160 s segments.Significance. The proposed framework demonstrates improved classification performance over existing methods by explicitly modeling complementary static and dynamic neural features within an interpretable architecture for depression analysis.
With the rapid aging of the global population, constipation has become a major gastrointestinal concern among elderly individuals. Bowel sounds provide a non-invasive acoustic signal for assessing gastrointestinal function, but their automatic analysis remains challenging due to sparsity and non-stationarity. This study proposes a two-stage bowel sound analysis framework based on continuous abdominal recordings. First, a Convolutional Neural Network-MAMBA (CNN-MAMBA) model was used for salient bowel sound detection. Second, a patient-level constipation classification model was developed using multi-view spectral representations and a Convolutional Neural Network-Conformer-Multiple Instance Learning (CNN-Conformer-MIL) architecture. On a held-out test set, the detection model achieved an accuracy of 0.87, an F1-score of 0.78, and a ROC-AUC of 0.93. For patient-level classification under binary Bristol Stool Form Scale (BSFS) grouping, five-fold cross-validation yielded a mean accuracy of 0.665 and an F1-score of 0.755. All BSFS labels were annotated by clinical physicians and temporally aligned with bowel sound recording. Given the modest improvement and cross-validation variability, the patient-level results should be interpreted as preliminary feasibility evidence. These findings suggest that bowel sound analysis may serve as an auxiliary screening or longitudinal monitoring tool rather than a stand-alone diagnostic system.
This study proposes an automated tongue analysis system that combines deep learning with traditional Chinese medicine to enhance the accuracy and objectivity of tongue diagnosis. The system includes a hardware device to provide a stable acquisition environment, an improved semi-supervised learning segmentation algorithm based on U2net, a high-performance colour correction module for standardising the segmented images, and a tongue image analysis algorithm that fuses different features according to the characteristics of each feature of the TCM tongue image. Experimental results demonstrate the system’s performance and robustness in feature extraction and classification. The proposed methods ensure consistency and reliability in tongue analysis, addressing key challenges in traditional practices and providing a foundation for future correlation studies with endoscopic findings.
Midline shift is a critical biomarker of mass effect in traumatic brain injury, directly influencing neurosurgical decisions and patient prognosis. Current MLS quantification relies heavily on manual measurements, which are timeconsuming, subjective, and prone to inter-observer variability. This study presents a fully automated, training-free geometric pipeline for robust MLS detection from CT images. The framework avoids manual annotation or machine learning by using principal component analysis to define the brain's symmetry axis and operates in three stages: selecting the optimal slice, identifying ventricles to establish the reference axis, and detecting the midline shift magnitude and direction. Tested on 20 TBI patient datasets comprising 420 CT slices, the method showed strong correlation with expert annotations ($\mathbf{r}=0.869$, $\mathbf{p}<$ 0.001) and clinically acceptable mean absolute error (2.25 mm). It achieved 80% accuracy in classifying clinical severity ($<5 ~\text{mm}$, $5-10 ~\text{mm},>10 ~\text{mm}$). These results demonstrate the feasibility of accurate, reproducible, and training-free MLS quantification for TBI assessment.
Background Neonatal intestinal diseases often have an insidious onset and can lead to poor outcomes if not identified early. Early assessment of abnormal bowel function is critical for timely intervention and improving prognosis, underscoring the clinical importance of reducing mortality related to these conditions through rapid diagnosis and treatment. Bowel sounds (BSs), produced by intestinal contractions, are a key physiological indicator reflecting intestinal function. However, manual clinical assessment of BSs has limitations in terms of consistency and interpretative accuracy, which restricts its clinical application. This study aims to develop an machine learning-based diagnostic model for neonatal intestinal diseases using BS analysis and to compare its diagnostic accuracy with that of manual clinical assessment.Methods and analysis This diagnostic study employs a cross-sectional design. The case group includes neonates diagnosed with intestinal diseases (using clinical diagnosis as the gold standard), such as neonatal necrotising enterocolitis (NEC), food protein-induced allergic proctocolitis, and other intestinal conditions (eg, intestinal obstruction, midgut volvulus, congenital megacolon). The control group will be established using frequency matching, stratified by gestational age and postnatal age. Based on the distribution of each stratum in the case group, neonates without intestinal diseases who were hospitalised during the same period will be randomly selected in proportion from the corresponding strata. BSs will be collected using a 3M stethoscope (Littmann 3200). The study will occur in two phases. In the first phase (July 2024 to July 2025), participants from West China Second University Hospital will be randomly divided into a training cohort (for model development with 10-fold cross-validation) and an internal validation cohort in a 7:3 ratio. The second phase (July 2025 to July 2026) will involve external validation, with patients from Sichuan Provincial Children’s Hospital and Shenzhen Children’s Hospital. Clinical diagnosis will serve as the gold standard, and diagnostic outcomes between the machine learning-based model and manual clinical assessment by physicians of varying experience levels will be compared.Ethics and dissemination Ethical approval has been obtained from the Medical Ethics Committee of West China Second University Hospital (Registration No.: 2023SCHH0021), Sichuan Provincial Children’s Hospital (Registration No.: 2021YFC2701704) and the Shenzhen Children’s Hospital (Registration No.: 2021015). Written informed consent will be collected from all participants prior to BS collection. Study findings will be disseminated through conferences and publications in peer-reviewed journals.Trial registration number ChiCTR2400086713.
Objective.This paper presents a novel dual-branch framework for estimating blood pressure (BP) using photoplethysmography (PPG) signals. The method combines deep learning with clinical prior knowledge and models different time periods (morning, afternoon, and evening) to achieve precise, cuffless BP estimation.Approach.Preprocessed single-channel PPG signals are input into two feature extraction branches. The first branch converts PPG dimensions to 2D and uses pre-trained Mobile Vision Transformer-v2 (MobileViTv2) and Visual Geometry Group19 (Vgg19) backbones to extract deep PPG features based on the different mechanisms of systolic blood pressure (SBP) and diastolic blood pressure (DBP) formation. The second branch calculates multi-dimensional feature parameters based on the relationship between PPG waveforms and factors affecting BP. We fuse the features from both branches and consider diurnal BP variations, using AutoML strategy to construct specific SBP and DBP estimation models for the different periods. The algorithm was developed on the human resting state PPG and BP dataset (HRSD) and validated on the MIMIC-IV dataset for generalization performance.Main results.The mean absolute error (MAE) for BP estimation is 6.42 mmHg SBP and 4.96 mmHg DBP in the morning, 4.84 mmHg (SBP) and 3.73 mmHg (DBP) in the afternoon, and 2.65 mmHg (SBP) and 2.56 mmHg (DBP) in the evening. Performance on the MIMIC-IV database was 4.34 mmHg (SBP) and 3.11 mmHg (DBP). The method meets the standards of the Association for the Advancement of Medical Instrumentation and achieves Grade A of the British Hypertension Society (BHS) standards.Significance. This indicates that it is an accurate and reliable non-invasive BP monitoring technology, applicable for continuous health monitoring and cardiovascular disease prevention.
Modeling and visualization of glioma growth could assist in cancer diagnosis, tumor progression prediction, and clinical treatment outcome improvement. However, most studies either failed to make patient-specific predictions or could only display information about tumor size and shape, lacking the capability to characterize the impact of tumor growth on surrounding tissues. In this study, a method (HybrSyn) combining tumor growth model and deformable image registration technique for synthesizing MRIs at arbitrary time point after the detection time has been proposed. Through the tumor growth model, tumor growth process for consecutive time point has been predicted according to the characteristics of tumor cell diffusion and proliferation within the brain. The glioma deformable image registration model was employed to obtain the deformation fields between the tumors at detection time and simulations at subsequent time points. These fields were then mapped to the patient's initial MRI scans to generate the synthetic MRIs corresponding to that time points. To validate the HybrSyn, various experiments were conducted on the BraTS19 and the internal dataset collected from Zigong First People's Hospital. The quantitative results demonstrated a structural similarity of 80.93% between the synthesized MRIs and the patients' MRI scans. The qualitative results indicated that the HybrSyn could effectively capture changes during tumor progression and provide a global view. From the clinical point of view, synthesized longitudinal brain MRIs could potentially aid in presenting the impact of glioma growth on surrounding functional areas, and identifying target regions for personalized treatment planning.
Cardiovascular disease is a major health threat closely associated with blood pressure levels. While continuous monitoring is essential, traditional cuff-based devices are inconvenient for long-term use. Current methods often fail to balance deep learning capabilities with interpretability, limiting further accuracy improvements. To address this problem, we propose a novel two-branch deep learning framework combining Residual Networks (ResNet) and Bidirectional Long Short-Term Memory (BiLSTM) for photoplethysmography (PPG)-based cuffless blood pressure estimation. The ResNet branch processes 60 features selected by Support Vector Machine-Recursive Feature Elimination (SVM-RFE) from manually extracted features, including our newly proposed trend features, while the BiLSTM branch processes complete PPG waveforms. Testing on 220 waveform segments from 218 patients in the MIMIC-IV dataset, our method achieves mean absolute errors of 3.47 mmHg and 2.81 mmHg, with standard deviations of 5.06 mmHg and 4.11 mmHg for systolic and diastolic blood pressure. This performance meets the Association for the Advancement of Medical Instrumentation (AAMI) standards and achieves an A rating according to British Hypertension Society (BHS) standards.
Attention deficit hyperactivity disorder (ADHD) is a prevalent neurodevelopmental disorder among children and adolescents. Behavioral detection and analysis play a crucial role in ADHD diagnosis and assessment by objectively quantifying hyperactivity and impulsivity symptoms. Existing video-based action recognition algorithms focus on object or interpersonal interactions, they may overlook ADHD-specific behaviors. Current keypoints-based algorithms, although effective in attenuating environmental interference, struggle to accurately model the sudden and irregular movements characteristic of ADHD children. This work proposes a novel keypoints-based system, the Multi-cue Feature Fusion Network (MF-Net), for recognizing actions and behaviors of children with ADHD during the Test of Variables of Attention (TOVA). The system aims to assess ADHD symptoms as described in the DSM-V by extracting features from human body and facial keypoints. For human body keypoints, we introduce the Multi-scale Features and Frame-Attention Adaptive Graph Convolutional Network (MSF-AGCN) to extract irregular and impulsive motion features. For facial keypoints, we transform data into images and employ MobileVitv2 for transfer learning to capture facial and head movement features. Ultimately, a feature fusion module is designed to fuse the features from both branches, yielding the final action category prediction. The system, evaluated on 3801 video samples of ADHD children, achieves 90.6% top-1 accuracy and 97.6% top-2 accuracy across six action categories. Additional validation experiments on public datasets NW-UCLA, NTU-2D, and AFEW-VA verify the network’s performance.
Lung cancer has the highest mortality rate among cancers. The commonly used clinical method for diagnosing lung cancer is the CT-guided percutaneous transthoracic lung biopsy (CT-PTLB), but this method requires a high level of clinical experience from doctors. In this work, an automatic path planning method for CT-PTLB is proposed to provide doctors with auxiliary advice on puncture paths. The proposed method comprises three steps: preprocessing, initial path selection, and path evaluation. During preprocessing, the chest organs required for subsequent path planning are segmented. During the initial path selection, a target point selection method for selecting biopsy samples according to biopsy sampling requirements is proposed, which includes a down-sampling algorithm suitable for different nodule shapes. Entry points are selected according to the selected target points and clinical constraints. During the path evaluation, the clinical needs of lung biopsy surgery are first quantified as path evaluation indicators and then divided according to their evaluation perspective into risk and execution indicators. Then, considering the impact of the correlation between indicators, a path scoring system based on the double spherical constraint Pareto and the importance-correlation degree of the indicators is proposed to evaluate the comprehensive performance of the planned paths. The proposed method is retrospectively tested on 6 CT images and prospectively tested on 25 CT images. The experimental results indicate that the method proposed in this work can be used to plan feasible puncture paths for different cases and can serve as an auxiliary tool for lung biopsy surgery.
Objective . The percutaneous puncture lung mass biopsy procedure, which relies on preoperative CT (Computed Tomography) images, is considered the gold standard for determining the benign or malignant nature of lung masses. However, the traditional lung puncture procedure has several issues, including long operation times, a high probability of complications, and high exposure to CT radiation for the patient, as it relies heavily on the surgeon’s clinical experience. Approach. To address these problems, a multi-constrained objective optimization model based on clinical criteria for the percutaneous puncture lung mass biopsy procedure has been proposed. Additionally, based on fuzzy optimization, a multidimensional spatial Pareto front algorithm has been developed for optimal path selection. The algorithm finds optimal paths, which are displayed on 3D images, and provides reference points for clinicians’ surgical path planning. Main results. To evaluate the algorithm’s performance, 25 data sets collected from the Second People’s Hospital of Zigong were used for prospective and retrospective experiments. The results demonstrate that 92% of the optimal paths generated by the algorithm meet the clinicians’ surgical needs. Significance. The algorithm proposed in this paper is innovative in the selection of mass target point, the integration of constraints based on clinical standards, and the utilization of multi-objective optimization algorithm. Comparison experiments have validated the better performance of the proposed algorithm. From a clinical standpoint, the algorithm proposed in this paper has a higher clinical feasibility of the proposed pathway than related studies, which reduces the dependency of the physician’s expertise and clinical experience on pathway planning during the percutaneous puncture lung mass biopsy procedure.
Necrotizing enterocolitis (NEC) is a severe gastrointestinal emergency in neonates, marked by its complex etiology, ambiguous clinical manifestations, and significant morbidity and mortality, profoundly affecting long-term pediatric health outcomes. The prevailing diagnostic approaches for NEC, including traditional manual auscultation of bowel sounds, suffer from limited sensitivity and specificity, leading to potential misdiagnoses and delayed treatment. In this paper, we introduce a groundbreaking NEC diagnostic framework employing machine learning algorithms that utilize multi-feature fusion of bowel sounds, significantly improving the diagnostic accuracy. Bowel sounds from NEC patients and healthy newborns are meticulously captured using a specialized acquisition system, designed to overcome the inherent challenges associated with the low amplitude, substantial background noise, and high variability of neonatal bowel sounds. To enhance the diagnostic framework, we extract mel-frequency cepstral coefficient (MFCC), short-time energy (STE), and zero-crossing rate (ZCR) to capture comprehensive frequency and time domain features, ensuring a robust representation of bowel sound characteristics. These features are then integrated using a multi-feature fusion technique to form a singular feature vector, providing a rich, integrated dataset for the machine learning algorithm. Employing the support vector machine (SVM), the algorithm achieved an accuracy (ACC) of 88.00%, sensitivity (SEN) of 100.00%, and an area under the receiver operating characteristic (ROC) curve (AUC) of 97.62%, achieving high accuracy in diagnosing NEC. This innovative approach not only improves the accuracy and objectivity of NEC diagnosis but also shows promise in revolutionizing neonatal care through facilitating early and precise diagnosis. It significantly enhances clinical outcomes for affected neonates.
Deformable Image registration is a fundamental yet vital task for preoperative planning, intraoperative information fusion, disease diagnosis and follow-ups. It solves the non-rigid deformation field to align an image pair. Latest approaches such as VoxelMorph and TransMorph compute features from a simple concatenation of moving and fixed images. However, this often leads to weak alignment. Moreover, the convolutional neural network (CNN) or the hybrid CNN-Transformer based backbones are constrained to have limited sizes of receptive field and cannot capture long range relations while full Transformer based approaches are computational expensive. In this paper, we propose a novel multi-axis cross grating network (MACG-Net) for deformable medical image registration, which combats these limitations. MACG-Net uses a dual stream multi-axis feature fusion module to capture both long-range and local context relationships from the moving and fixed images. Cross gate blocks are integrated with the dual stream backbone to consider both independent feature extractions in the moving-fixed image pair and the relationship between features from the image pair. We benchmark our method on several different datasets including 3D atlas-based brain MRI, inter-patient brain MRI and 2D cardiac MRI. The results demonstrate that the proposed method has achieved state-of-the-art performance. The source code has been released at https://github.com/Valeyards/MACG.
Tooth instance segmentation, which separates teeth from surrounding tissues and structures, provides the basis for preoperative planning and research. Existing methods have demonstrated breakthrough results on partial dental segmentation tasks. However, tooth instance segmentation based on cone-beam computed tomography in patients with alveolar clefts is still challenging due to tooth shape variation, high interdental similarity and adjacent tooth occlusion issues. In this study, we propose a new indicator called the tooth descriptor to assist in localizing tooth instances and guide segmentation tasks. The proposed two-stage network first predicts the tooth descriptors and the centroids used to generate the tooth cubes, and then the segmentation network extracts the complete tooth from the tooth cube guided by the morphology of the tooth descriptor. The experimental results demonstrate that the proposed algorithm exhibits excellent robustness in the task of segmenting CBCT (Cone-Beam Computed Tomography) tooth instances of patients with alveolar clefts, with a Dice accuracy of 94.4%, which outperforms state-of-the-art dental segmentation methods.
Schizophrenia is a common mental disease worldwide. To assist doctors in diagnosing schizophrenia through objective biological indicators, an automatic detection method for schizophrenia is proposed in this work. The proposed automatic detection method, which is based on the abnormal eye movements of schizophrenic patients in reading tasks, includes three parts: valid reading state trajectory extraction, eye movement trajectory extraction, and model construction based on abnormal eye movements in schizophrenic patients. The valid reading state trajectory extraction algorithm is proposed to divide the reading process into a valid reading state and an invalid reading state based on the visual perception of the subjects in the reading process. For the eye movement trajectory extraction, an eye center positioning algorithm based on the double ring of the limbic boundary is proposed for the valid reading state. Based on the extracted eye movement trajectory, models from three aspects of abnormalities of schizophrenic patients during reading are built: the merged eye-head movement, relative eye movement, and comprehensive reading performance. The features extracted from the three models are combined with a pattern recognition classifier to realize the automatic diagnosis of schizophrenia (the support vector machine, random forest, and adaptive boosting are used in this work). The video dataset used in this experiment was recorded using 40 subjects (20 schizophrenic patients and 20 healthy controls) from the Psychiatry Department of the Mental Health Center. The experimental results show that the detection accuracy of the automatic detection method for schizophrenia can reach 96.25%, which indicates that the proposed method can be used as a computer-aided diagnosis method for schizophrenia.
Velopharyngeal insufficiency (VPI) is a type of pharyngeal function dysfunction that causes speech impairment and swallowing disorder. Speech therapists play a key role on the diagnosis and treatment of speech disorders. However, there is a worldwide shortage of experienced speech therapists. Artificial intelligence-based computer-aided diagnosing technology could be a solution for this. This paper proposes an automatic system for VPI detection at the subject level. It is a non-invasive and convenient approach for VPI diagnosis. Based on the principle of impaired articulation of VPI patients, nasal- and oral-channel acoustic signals are collected as raw data. The system integrates the symptom discriminant results at the phoneme level. For consonants, relative prominent frequency description and relative frequency distribution features are proposed to discriminate nasal air emission caused by VPI. For hypernasality-sensitive vowels, a cross-attention residual Siamese network (CARS-Net) is proposed to perform automatic VPI/non-VPI classification at the phoneme level. CARS-Net embeds a cross-attention module between the two branches to improve the VPI/non-VPI classification model for vowels. We validate the proposed system on a self-built dataset, and the accuracy reaches 98.52%. This provides possibilities for implementing automatic VPI diagnosis.
Schizophrenia is a severe mental disease that affects patients' thoughts, feelings, and behaviors. Speech signal has proven to be a biomarker in the early diagnosis of schizophrenia. Previous studies on schizophrenic speech detection are mainly based on manual feature extraction engineering, which requires domain knowledge for researchers and has difficulties extracting effective features. This work proposes an end-to-end architecture, called Axial-attention-based Network using Wideband and Narrowband Spectrograms (WNSA-Net), to detect schizophrenia. Specifically, we adopt both wideband and narrowband spectrograms as inputs to represent speech signals using fine time and frequency structures. Then dilated convolution blocks are employed to capture detailed and long-range information in spectrograms. Axial-attention blocks are introduced to augment the information in feature maps along the time and frequency axes. In addition, we employ a gate mechanism to fuse the output feature maps from all channels. Experimental results on the Schizophrenia dataset and its subdatasets show that schizophrenic patients have difficulties in expressing emotions. To validate the performance of our WNSA-Net, experiments are conducted on Schizophrenia dataset and open-access TORGO database, achieving 97.37% and 98.16% accuracy in detecting schizophrenia and dysarthria, respectively. The results show promise for the proposed method in the diagnosis of disordered speech.