
This retrospective case-control study investigated clinical correlates and gut microbiota signatures associated with reduced bone mineral density (BMD) through the gut-bone axis, a concept describing microbiota-related regulation of bone remodeling through immune, metabolic, intestinal barrier, and endocrine pathways. Using retrospective criterion-based sampling, eligible records from adults undergoing BMD testing and health screening between January 2024 and January 2025 were selected according to predefined criteria, yielding 60 participants with osteoporosis, 60 with osteopenia, and 58 controls. Clinical indicators, fasting blood markers, stool 16S ribosomal RNA (16S rRNA) sequencing profiles, differential taxa, predicted Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways, and correlations with BMD were compared among groups. Group comparisons used appropriate parametric, nonparametric, or chi-square ([Formula: see text]) tests; graphical and statistical presentations included principal coordinate analysis (PCoA), linear discriminant analysis effect size (LEfSe), Spearman correlation, and receiver operating characteristic (ROC) curves. Baseline characteristics were comparable (all [Formula: see text]), whereas BMD, inflammatory indices, bone-turnover markers, 25-hydroxyvitamin D, and selected cytokines differed significantly (all [Formula: see text]). Alpha diversity was similar among groups, but reduced bone status was associated with microbial dysbiosis, including a lower Firmicutes/Bacteroidetes ratio, decreased Megamonas, Agathobacter, and Faecalibacterium/Proteus ratio, and increased Escherichia and Aeromonas. LEfSe and predicted functional analyses identified disease-associated taxa and metabolic pathway shifts. Serum uric acid (SUA), erythrocyte sedimentation rate (ESR), and alkaline phosphatase (ALP) correlated negatively with BMD (all [Formula: see text]), whereas transforming growth factor-[Formula: see text] (TGF-[Formula: see text]) correlated positively ([Formula: see text]). Osteoporosis was associated with gut microbial dysbiosis and altered inflammatory and metabolic profiles, supporting the gut-bone axis and providing targets for future validation.
Evidence on effective gait rehabilitation strategies for patients with Parkinson's disease (PD) who experience acute functional deterioration due to medical complication remains limited. We report a case of torque-assisting hip wearable robot training in a patient with PD during acute post-infectious gait deterioration. An 80-year-old woman with PD was admitted with acute gait deterioration following urosepsis. Upon admission, she was unable to stand or ambulate. In the levodopa ON state, her Berg Balance Scale (BBS) score was 31 of 56, her gait speed on the 10-m walk test was 0.61 m/s, and the 6-minute walk test distance was 178 m with moderate assistance. She underwent conventional rehabilitation and overground gait training using a torque-assisting hip-wearable robot for 11 sessions. Spatiotemporal gait parameters and hip joint kinematics improved during robot-assisted gait training regardless of medication status. Plantar pressure analysis demonstrated improved heel strike on the symptom-dominant left side. At 1-month follow-up, the patient ambulated independently without an assistive device, and her BBS score, gait speed, and 6-minute walk distance improved to 52, 0.85 m/s, and 455.4 m, respectively. Early gait rehabilitation with a torque-assisting hip wearable robot may support functional recovery adaptive gait changes in patients with PD experiencing acute functional deterioration.
The accurate and objective recognition of mental health conditions, such as anxiety and depression, remains a significant challenge in biomedical engineering and clinical practice, where traditional assessment methods often suffer from subjectivity, inefficiency, and a lack of quantitative biomarkers. In this study, a hierarchical multimodal convolutional neural network (CNN) is proposed. By integrating heterogeneous biomedical data sources from different public datasets, including texts, facial images, and physiological signals, a multimodal learning framework is constructed for mental health status recognition. Special convolution and pooling structures are designed for different modalities, fine-grained spatial and temporal features are extracted from corresponding data sources, and the complementary information between different modal combinations is learned through a dynamic fusion mechanism. Experimental evaluations are conducted on the publicly available Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) dataset and the Affect, Mood, and Personality Database (AMIGOS). On the DAIC-WOZ dataset, the proposed model achieved an accuracy of 0.818, a recall of 0.801, an F1 score of 0.809, an area under the curve (AUC) of 0.868, and a precision-recall (PR)-AUC of 0.842. Compared with the best-performing baseline (Multimodal Transformer), it increases by 6.5%, 7.2%, 6.9%, 6.1%, and 8.5%, respectively. The model's effectiveness in multimodal feature extraction and fusion is verified on the AMIGOS dataset. The proposed model is evaluated in the emotion recognition task. The accuracy under short and long video conditions reaches 0.791 and 0.801, respectively, and the corresponding AUC values are 0.837 and 0.849. Ablation experiments confirm that multimodal feature fusion significantly enhances recognition performance compared to single-modality approaches. In contrast, robustness tests demonstrate stable performance across varying hyperparameter configurations. These results indicate that the proposed hierarchical dynamic multimodal CNN framework can effectively integrate text, visual, and physiological information based on the modal features of different datasets; thus, it can provide a repeatable, objective, and clinically feasible technical approach for intelligent mental health assessment and personalized intervention.
Traditional somatosensory music therapy predominantly employs fixed-parameter protocols, which lack the capacity for dynamic adjustment based on the real-time physiological and emotional states of patients. This limitation results in significant individual variability in therapeutic outcomes, making personalized and precise interventions difficult to achieve. This study aims to develop a Channel Attention Multi-modal Long Short-Term Memory (CA-Multi-LSTM) model, as well as an Intelligent Vibroacoustic Music Therapy Feedback System (IVMT-FS). The system achieves dynamic adaptive optimization of therapeutic parameters in response to the real-time physiological and emotional status of patients. To capture both fine-grained temporal fluctuations and long-term emotional changes, a two-layer stacked LSTM architecture is implemented with parallel branches established for electroencephalogram (EEG), heart rate variability (HRV), and galvanic skin response (GSR) modalities. A Modality Channel Attention (MCA) mechanism is introduced to enable adaptive dynamic weighted fusion based on the contributions of each modality. Operating on the CA-Multi-LSTM model, the IVMT-FS executes a real-time closed-loop cycle every 30 seconds. This process sequentially encompasses physiological signal collection, emotion recognition, therapeutic parameter generation, and composite vibration waveform output. Model training and validation are performed using three publicly available datasets: the Database for Emotion Analysis using Physiological Signals (DEAP), the Database for Emotion Recognition through EEG and electrocardiographic (ECG) signals from Wireless Low-cost Off-the-Shelf Devices (DREAMER), and A Dataset for Affect, Personality and Mood Research on Individuals and Groups (AMIGOS). Furthermore, a 12-week randomized controlled trial (n=90) evaluates the intervention efficacy of the system in patients with mild-to-moderate Alzheimer's disease. The results demonstrate that CA-Multi-LSTM achieves accuracies of 95.3%, 92.1%, and 94.2%, and F1-scores of 94.1%, 91.3%, and 93.9% on the DEAP, DREAMER, and AMIGOS datasets, respectively, significantly outperforming the baseline models (p < 0.001). Ablation experiments indicate that the MCA mechanism contributes the most to model performance, while the EEG modality branch plays a critical role in supporting recognition accuracy. Compared with the fixed-parameter mode, the intelligent feedback mode of the IVMT-FS reduces the time to reach the target state by 43.8%, increases the target state maintenance rate by 49.9%, and decreases the number of triggered physiological abnormalities by 61.1%. Following the 12-week clinical intervention, the experimental group exhibits a 118.5% increase in EEG a-band power (p < 0.001), a 72.3% increase in the standard deviation of HRV (p < 0.001), and a 67.0% reduction in GSR levels (p < 0.001) compared with the control group. Through adaptive multimodal fusion via MCA, this study significantly enhances the accuracy of affect recognition from physiological signals and realizes real-time adaptive closed-loop regulation of therapeutic parameters. The proposed system provides an efficient, personalized, and safe non-pharmacological intervention strategy for individuals with Alzheimer's disease, demonstrating substantial potential for clinical application.
Radiotherapy is a widely used method in the treatment of head-and-neck cancers; however, irradiation of surrounding oral structures often results in clinical complications such as dental caries, oral mucositis, osteoradionecrosis, dysphagia, dysgeusia, xerostomia and trismus. To address such issues, a patient specific hybrid double-layer orthodontic mouthguard has been proposed to simultaneously enhance the radiation shielding and biomechanical comfort in treatments. The customized device integrates radiation-shielding dental covers, a compliant soft-liner layer, a tongue depressor, oral positioning stents and cheek displacement stents to safeguard critical oral structures. Finite element analysis has been done to quantify stress transmissions to the mandibular dentition under static oral-stent loading. An initial single-layer configuration without the soft cove has been analysed to demonstrate the stresses without using any intermediate layer. Subsequently, a double-layer configuration incorporating a 0.5-mm cushioning layer was investigated using six biocompatible materials: Polylactic Acid (PLA), Polyethylene Terephthalate (PET), Polyetheretherketone (PEEK), Thermoplastic Polyurethane (TPU), Polytetrafluoroethylene (PTFE) and Polymethyl Methacrylate (PMMA). Among the materials evaluated, TPU offered the most effective stress mitigation, lowering the peak equivalent stress by nearly 39% compared to the single-layer configuration. Subsequent thickness optimization further highlighted the importance of the soft inner lining: a 2-mm TPU layer delivered the best performance, reducing the maximum stress to 14.98 MPa, a drop of about 64% relative to the original design. It demonstrated that a double-layer mouthguard not only reduces stress concentrations more efficiently across the dentition but also maintains reliable radioprotective functionality. As a result, this proposed double-layer mouthguard design provides strong advancement for enhancing patient comfort and protecting oral tissue and provides precision treatment during head-and-neck radiotherapy, supporting its potential for future clinical adoptions.
Background: Early-stage tumor tissues exhibit microstructural and optical property abnormalities, face challenges in achieving label-free, high-resolution, and accurate identification. Multiphoton microscopy (MPM) and polarization imaging can reflect tissue morphology, metabolism, and microscopic polarization characteristics, offering new approaches for early diagnosis of tumor abnormalities. Objectives: This study aimed to investigate the application value of multiphoton microscopy combined with polarization imaging in the diagnosis of early-stage tumor tissue abnormalities. Methods: A single-center retrospective study was conducted. A total of 180 archived breast tissue specimens were collected, including 58 cases of early-stage breast cancer, 64 cases of breast precancerous lesions, and 58 cases of normal breast tissue. MPM and polarization imaging were performed to extract morphological, metabolic, and polarization optical features. Using pathological diagnosis as the gold standard, one-way ANOVA, Spearman correlation analysis, and multivariate logistic regression analysis were employed to evaluate the diagnostic performance of dual-modality imaging for early-stage breast cancer. Results: Compared with precancerous lesions and normal breast tissues, early-stage breast cancer tissues showed significantly increased nucleus-to-cytoplasm ratio, cell area, and total depolarization, as well as significantly decreased collagen density, alignment regularity, and polarization anisotropy (P<0.05). Spearman analysis indicated significant correlations between major optical parameters and the degree of malignancy. Multivariate logistic regression identified nucleus-to-cytoplasm ratio, collagen alignment regularity, and total depolarization as independent diagnostic indicators for early-stage breast cancer (P<0.05). The diagnostic accuracy, sensitivity, specificity, and AUC of the combined MPM and polarization dual-modality imaging were significantly superior to those of single-modality imaging (P<0.05). Conclusion: Multiphoton microscopy combined with polarization imaging can effectively identify early-stage tumor tissue abnormalities, providing a reliable approach for label-free and accurate optical diagnosis of early tumors.
Traditional lower limb rehabilitation robots suffer from low motion intention recognition accuracy and poor human-robot compliance, which may cause secondary injury to athletes with lower limb fractures. To solve this problem, this study proposes a hybrid control system combining sparrow search algorithm optimized back propagation neural network (SSA-BP) and adaptive impedance control (AIC). The SSA-BP model is adopted to extract features from surface electromyography (sEMG) signals and recognize human motion intention, while the AIC strategy dynamically adjusts robot stiffness and damping to realize compliant human-robot interaction. Validated on a public lower limb multimodal dataset, the proposed system reduces mean squared error by 61.95% compared with conventional BP methods. Meanwhile, the R 2 is 0.952, which is superior to all methods. Under muscle fatigue and electrode displacement interference, the joint angle tracking root mean square error is only 1.52°, and the steady-state recovery time is 22.4 milliseconds (ms). The robot maintains low interaction torque and smooth mechanical output during the whole rehabilitation process. Under the three conditions of high-frequency noise, random phase shift, and sudden load interference, the SSA-BP-AIC model also maintains high performance. This proposed control framework exhibits high accuracy, anti-disturbance ability and real-time performance. It can ensure training safety and comfort for fractured athletes, and provides a reliable mechanical control solution for sports rehabilitation engineering.
To meet the growing demand for sports motion recognition in training assistance, rehabilitation assessment, and intelligent monitoring, this study developed an auxiliary action-recognition evaluation framework incorporating a Big Generative Adversarial Network (BigGAN)-based data augmentation mechanism. An experimental subset was first constructed from representative sports actions in the Nanyang Technological University RGB+D120 (NTU RGB+D120) dataset, including running, jumping, throwing, and gymnastics movements. BigGAN was subsequently employed to generate augmented training samples, while truncation strategies and class-embedding mechanisms were introduced to increase pose variability and visual diversity. The generated samples were used exclusively during model training, whereas only real samples were retained in the validation and test sets to ensure an unbiased performance evaluation. The quality of the generated data was assessed using Fréchet Inception Distance (FID), Inception Score (IS), Learned Perceptual Image Patch Similarity (LPIPS), nearest-neighbor retrieval, and truncation-threshold sensitivity analysis. Samples exhibiting semantic inconsistencies, structural distortions, or near-duplicate characteristics were removed through quality control procedures. Experimental results showed that the proposed framework achieved an accuracy of 89.8%, a recall of 89.2%, and an F1-score of 89.5% on the test set. The generated samples yielded an average FID of 38.80, an average LPIPS value of 0.331, and a nearest-neighbor duplication rate of 3.5%. In addition, the BigGAN-enhanced model demonstrated more consistent classification performance across different sports categories. Running achieved the highest recognition accuracy, whereas throwing remained the most challenging category. The augmentation strategy improved the model's robustness to intra-class variability in several action categories. Nevertheless, the proposed framework provides only auxiliary support for action recognition and cannot replace professional assessments conducted by coaches, clinical practitioners, or biomechanical analysis systems. These findings offer a reproducible foundation for data augmentation, action classification, and intelligent feedback in sports motion monitoring applications.
With the continuous evolution of digital interaction design toward intelligence and affective computing, the ability to accurately and stably perceive users’ emotional states during interaction has become critical for enhancing user experience and system adaptability. Electroencephalogram (EEG) signals directly reflect neural activity in the human brain and therefore hold significant potential for emotion recognition tasks. However, EEG signals are characterized by strong noise interference, substantial inter-subject variability, and continuously evolving emotional states over time, which pose challenges to spatial feature representation, long-term dependency modeling, and prediction stability. To address these issues, this paper proposes a dynamic EEG-based emotion recognition model for digital interaction design, termed interactive dynamic emotion perception network (IDEP-Net). Motivated by the requirements of dynamic affect perception in interactive scenarios, the proposed model integrates spatial heterogeneity modeling, local spatiotemporal embedding, and low-complexity sparse global dependency learning into a unified architecture. Specifically, a channel-adaptive modulation (CAM) is introduced to explicitly capture the heterogeneous contributions of different brain regions to emotion recognition. A convolutional spatiotemporal embedding module is then employed to extract local rhythmic patterns from EEG signals. Furthermore, an importance-guided sparse attention mechanism is incorporated to model long-range dependencies across temporal segments while effectively reducing computational complexity. To enhance prediction stability in interactive scenarios, a temporal continuity constraint is added to the optimization objective, encouraging smoother and more consistent emotional predictions over time. Extensive experiments conducted on the SEED and SEED-IV datasets under cross-subject evaluation protocols demonstrate that the proposed IDEP-Net achieves superior recognition performance compared to multiple baseline methods, while significantly improving temporal stability and generalization capability. These results validate the effectiveness and robustness of the proposed model for EEG-based emotion recognition and provide a feasible and efficient technical solution for affect-aware digital interaction design.
This study investigates heat transfer characteristics in the biphase flow of an Ellis fluid model within a sinusoidal wavy horizontal symmetric channel. A two-phase mathematical model incorporating the stress tensor contribution of the Ellis fluid is developed to analyze peristaltic motion under the influence of a uniform heat source. By employing the assumptions of long wavelength and low Reynolds number, the governing equations are simplified into a tractable form, and exact solutions for velocity and temperature distributions are obtained using Mathematica 14.2. The results reveal that the first and second Ellis fluid parameters significantly reduce the velocity profiles in the core region of the channel, while the temperature distribution exhibits an inverse relationship with these parameters. The suspension of particles and the heat source parameter enhance the heat transfer rate by up to 29% and 38%, respectively, for two successive parameter increments. Furthermore, the fluid phase velocity is found to be lower than the particle phase velocity. The findings of this study provide valuable insights into improving blood flow in biological vessels and offer practical implications for enhancing drug delivery efficiency and optimizing the design and performance of biomedical devices involving two-phase non-Newtonian fluid transport.
Ejection Fraction (EF) is a key clinical parameter for assessing cardiac function and guiding the management of heart failure. Traditional EF estimation methods depend on manual segmentation of the left ventricle (LV) from echocardiographic images, which are time-consuming, labor-intensive, and subject to inter-observer variability. To address these limitations, we propose a novel and efficient two-stage deep learning framework for fully automated EF estimation from echocardiographic video data. In the first stage, we employ the nnU-Net architecture to perform high-accuracy segmentation of the LV in both end-diastolic (ED) and end-systolic (ES) frames. In the second stage, we introduce lightweight 2D and 3D ResNet-based regression models that directly predict EF from the segmented LV regions, bypassing the need for explicit volume measurements. This approach allows for rapid, end-to-end inference while requiring only EF labels for training, significantly reducing dependence on noisy and variable clinical annotations. Extensive experiments demonstrate that our models achieve state-of-the-art accuracy across multiple evaluation metrics, with the 3D ResNet model reaching an R2 score of 0.7589. Moreover, our framework delivers substantial improvements in computational efficiency: it achieves real-time inference with mean processing times as low as 0.0151 s, making it over 60 times faster than the widely used EchoNet-Dynamic model. These results highlight the potential of our method as a robust, accurate, and clinically viable solution for real-time EF estimation, enabling faster and more consistent cardiac assessments in routine practice.
Electromyography (EMG) is the established standard for evaluating neuromuscular activation, yet its clinical utility is constrained by skin impedance changes, electrical noise, and motion artefacts. More recently, mechanomyography (MMG), which measures the mechanical behavior and oscillations produced during muscle contraction, has emerged as a promising complementary or alternative modality with advantages in signal stability and resistance to interference, particularly in rehabilitation settings. This systematic review synthesizes findings from twenty-five studies comparing MMG and EMG under matched experimental conditions or which provided indirect evidence on their relative performance. The analysis revealed consistent agreement between modalities in detecting muscle activation and fatigue, with MMG often demonstrating greater robustness in an electrically noisy environment and during neuromuscular electrical stimulation. MMG also provided distinctive mechanical insights during concentric contractions and relaxation phases, where EMG performance was more variable. However, substantial heterogeneity in MMG sensor types, frequency-domain processing, and reporting standards represents a central barrier to its broader clinical adoption. Overall, the evidence positions MMG as a valuable adjunct, and in selected contexts, a practical alternative to EMG for muscle function assessment. Standardized acquisition protocols and consensus-based guidelines are needed to establish methodological consistency and maximize the translation potential of MMG in rehabilitation and clinical practice.
This study aimed to preliminarily explore the current application status and potential value of Digital Twin Technology (DTT) that integrates Artificial Intelligence (AI) and medical imaging recognition in orthopedic surgical simulation education for bone tumors, and to analyze its effectiveness in surgical planning, biomechanical analysis, postoperative management, and training, through a systematic review conducted in accordance with PRISMA guidelines. Through a review of relevant literature, the progress in the application of AI-enhanced digital twins in bone tumor orthopedic surgery in recent years is analyzed. A structured literature search and review method was adopted, searching databases such as PubMed, Web of Science, and Scopus for literature published between 2015 and 2024. The inclusion criteria were studies involving orthopedic medical students or surgeons, with content related to the integration of DTT with AI and image recognition technologies applied to surgical planning, biomechanical analysis, postoperative management, and training. The results indicate that DTT, by combining AI-driven image analysis and biomechanical modeling with high-precision three-dimensional modeling and scenario simulation, significantly improves the accuracy of surgical planning, the precision of biomechanical analysis, and the scientific rigor of postoperative management. In training, the integration of DTT with virtual reality technology and AI-based simulation engines provides students and novice surgeons with a high-fidelity virtual surgical environment, significantly enhancing their operational skills and confidence. Although most current studies are small-scale experiments, the potential of AI-enhanced DTT in orthopedic medicine has been validated. In the future, with further integration of artificial intelligence, image recognition, and biomechanical simulation technologies, “personalized scenario generation” can be achieved, and a standardized scenario library for bone tumor surgical education can be established. This brief review aimed to map the preliminary landscape of the technology’s applications and explore its future potential. Its findings were expected to help address the scarcity of bone tumor cases and the uneven distribution of training resources, thereby providing critical technological support for the standardized education and remote skill training in bone tumor surgery.
A computational study is described of the structural performance of a bumblebee-inspired wing for a Micro Aerial Vehicle (MAV) during the initial downstroke phase of take-off using Finite Element Analysis (FEA). A detailed three-dimensional model of the forewing and hindwing was developed to replicate the morphological and structural features of the insect wing. Static FEA was performed within the ANSYS Workbench to assess stress distribution, strain, and deformation under aerodynamic loading conditions experienced during the first wing stroke, based on a wingspan of 18mm and the species-specific flapping frequency of 130Hz. The results indicate that the natural venation pattern of the wing provides effective load distribution, reducing stress concentrations along the wingspan. Peak Von Mises stresses were localized along the leading edge, following the primary branch of the vein structure, while maximum deformation occurred at the wingtip, consistent with experimentally observed honeybee wing deformation. The static pressure case simulations reveal that maximum deformation (total displacement) is observed to consistently decrease with increasing Young's Modulus. At 2GPa, the wing exhibits the highest displacement of 1189 mu m, and this is strongly reduced to 793 mu m at 3GPa, 595 mu m at 4GPa, and 476 mu m at 5GPa. The Von Mises elastic strain varied significantly with stiffness. The strain decreases when Young's modulus increases for chitin, with values of 3.68 & times;10-3 at 2GPa, 2.45 & times;10-3 at 3GPa, 1.84 & times;10-3 at 4GPa, and 1.47 & times;10-3 at 5GPa. Notably, the Von Mises stress results remained nearly constant across all Young's Modulus values, hovering around 4.93MPa. For the acceleration case, at 2GPa, the maximum total deformation was recorded as 1094 mu m. When the modulus was increased to 3GPa, the displacement reduced to 723 mu m, followed by 547 mu m at 4GPa and 438 mu m at 5GPa. As with the static simulation, in all cases, the displacement was concentrated toward the wing tip, with the largest deformation occurring along the outer regions of the bumblebee wing structure, away from the fixed support. Furthermore, the Von Mises strain recorded when the Young's Modulus was set to 4GPa exceeds that computed in the other three cases (5, 3, and 2GPa). As with the static pressure simulations, the Von Mises stress sustains a consistent value across all the Young's Modulus values, maintaining a peak value of approximately 4.27MPa. Peak Von Mises stress is observed at the leading edge, where the leading vein changes direction to follow the wing contours, which are most susceptible to loading and are subjected to the highest forces resulting from flight. The findings also demonstrate the potential of biomimetic venation layouts to enhance the strength-to-weight ratio in small-scale UAV wings, thereby supporting efficient lift production with minimal weight penalties.
Accepted 22 August 2025; Article withdrawn 14 May 2026 During the proof stages of preparing this accepted article for publication the publisher became aware that this article had recently been published in another journal. Therefore, this article has been withdrawn.
Pulmonary ground-glass nodules (GGNs) are important imaging biomarkers for early detection and diagnosis of lung tumors. Accurate identification and classification of GGN density remain challenging because of the subtle radiological features and variability in the clinical presentation of GGNs. This study aimed to develop a new recognition model based on the GGNs–YOLO (You Only Look Once) v11 neural network to improve detection accuracy and robustness in medical imaging. The model enhanced YOLOv11s by integrating a Mixed Aggregation Network into its backbone to replace the traditional C3k2 module. This integration substantially improved the model’s ability and accuracy to detect GGNs. Additionally, we proposed using the efficient up-convolution block to replace traditional upsampling and introducing the Inner-Complete Intersection over Union loss function to boost bounding box regression precision. These modifications improved the model’s performance in detecting small lung nodules. Finally, we compared the GGNs–YOLOv11 model with other existing models to validate our proposed enhancements. The results showed that the improved GGNs–YOLOv11 model outperformed other models in GGN detection tasks. Specifically, its mAP50 on the dataset reached 93.6%, with a recall of 88.6% and accuracy of 89.9%, which were 4.7%, 5.1%, and 3% higher than those of the baseline model, respectively. This approach not only improved diagnostic reliability but also provided a practical tool to assist radiologists in clinical decision-making. This study highlights the potential of deep learning-based detection frameworks to advance intelligent medical imaging and support early lung cancer screening.
Acute kidney injury (AKI) is a severe complication in hypertensive patients admitted to the intensive care unit (ICU), characterized by subtle early symptoms and the limited efficacy of traditional early warning methods. This study aims to develop and validate a deep learning model for predicting AKI risk within the next 6[Formula: see text]h in hypertensive ICU patients, identify the optimal algorithm, and determine key predictive features. A total of 2856 hypertensive ICU admissions were retrospectively extracted from the MIMIC-III database (2008–2019). After applying inclusion criteria (age 18–89 years, confirmed hypertension) and exclusion criteria (ICU stay [Formula: see text] [Formula: see text]h, [Formula: see text] missing data), 1947 valid cases were included, exhibiting significant class imbalance with an AKI incidence rate of 28.4%. The dataset was randomly split into training (70%) and test (30%) sets using stratified sampling to preserve class distribution. Static features (e.g., age, gender, ethnicity, and admission care type) and 14[Formula: see text]h dynamic features (e.g., heart rate, blood pressure, serum creatinine, and blood urea nitrogen; 14 clinical indicators in total) were extracted. The XGBoost algorithm was then employed to select 21 key features with a cumulative importance [Formula: see text] from the initial feature set. Seven binary classification prediction models were constructed and compared: S5, GRU, LSTM, RWKV, Transformer, Logistic Regression, and Random Forest. For the deep learning models, key training hyperparameters were uniformly set: A batch size of 128 and a learning rate of 1e−4. Model performance was comprehensively evaluated using metrics including accuracy, precision, recall, F1-score, and AUC, with statistical significance assessed via paired t-tests ([Formula: see text]). The reliability of the results was further validated through XGBoost feature importance analysis and paired t-tests ([Formula: see text]). After data preprocessing and feature selection, the S5 model demonstrated the best performance among the seven models. The S5 model achieved the best overall performance (accuracy 0.7206, precision 0.7343, recall 0.7353, F1-score 0.7246, and AUC 0.7565), significantly outperforming the other models ([Formula: see text]). Notably, although the logistic regression model achieved an accuracy of 0.8328, its precision was only 0.0985, indicating severe overfitting. The Transformer model had a recall of merely 0.1912, limiting its clinical utility. Additionally, XGBoost feature importance analysis revealed that neurosurgical services, age, ethnicity, serum creatinine, blood urea nitrogen, and blood pressure change rate are core predictors for AKI. The lightweight S5 model offers optimal balance and clinical adaptability for real-time 6 h AKI prediction in hypertensive ICU patients. Its high recall and robustness can effectively reduce the under diagnosis of AKI, while the identified core predictive features provide a clear basis for screening high-risk populations and developing personalized intervention strategies. Collectively, these attributes position the S5 model as a promising candidate for integration into ICU bedside early warning systems to facilitate timely intervention.
Digital Visual Interaction (DVI) systems, including virtual reality, augmented reality, and adaptive multimedia interfaces, demand real-time, lightweight emotion recognition to enable responsive and personalized user experiences. Multimodal approaches combining Electroencephalography (EEG) and eye movement signals offer complementary neural and behavioral cues, yet existing Transformer-based methods restrict cross-modal interaction to final fusion layers, while heavier CNN-based alternatives incur prohibitive computational costs for interactive deployment. We propose the Cross-Modal Feature Calibration Multimodal Adaptive Emotion Transformer (CMFC-MAET), a lightweight framework tailored for DVI scenarios. CMFC-MAET extends MAET with two innovations: (1) Cross-Modal Feature Calibration (CMFC) modules inserted at intermediate Transformer layers, which generate channel-wise scaling and shifting factors from EEG features to calibrate eye movement representations; and (2) A Lightweight Gated Attention Fusion (LGAF) mechanism that incorporates feature statistics for adaptive multimodal integration. The model maintains an embedding dimension of 32 with only three Transformer layers and introduces only about 4.5[Formula: see text]K additional parameters over MAET, enabling real-time inference on consumer-grade hardware commonly used in DVI applications. Extensive experiments on SEED, SEED-IV, and SEED-VII datasets demonstrate consistent improvements. On SEED-VII (7-class, subject-dependent), CMFC-MAET achieves 73.86% accuracy (a 2.58% gain over MAET). Cross-subject evaluation yields 43.12% accuracy (a 2.22% gain). Ablation studies confirm that CMFC contributes a 1.66% gain and LGAF contributes a 0.87% gain independently, while their combination produces the full 2.58% improvement with only about 4.5[Formula: see text]K additional parameters. CMFC-MAET achieves a strong accuracy-efficiency trade-off among compared multimodal emotion recognition methods, demonstrating that effective cross-modal calibration can be realized within a lightweight Transformer architecture suitable for real-time digital visual interaction systems.