Despite the widespread success of Graph Neural Networks (GNNs), understanding the reasons behind their specific predictions remains challenging. Existing explainability methods face a trade-off that gradient-based approaches are computationally efficient but often ignore structural interactions, while game-theoretic techniques capture interactions at the cost of high computational overhead and potential deviation from the model's true reasoning path. To address this gap, we propose FSX (Message Flow Sensitivity Enhanced Structural Explainer), a novel hybrid framework that synergistically combines the internal message flows of the model with a cooperative game approach applied to the external graph data. FSX first identifies critical message flows via a novel flow-sensitivity analysis: during a single forward pass, it simulates localized node perturbations and measures the resulting changes in message flow intensities. These sensitivity-ranked flows are then projected onto the input graph to define compact, semantically meaningful subgraphs. Within each subgraph, a flow-aware cooperative game is conducted, where node contributions are evaluated fairly through a Shapley-like value that incorporates both node-feature importance and their roles in sustaining or destabilizing the identified critical flows. Extensive evaluation across multiple datasets and GNN architectures demonstrates that FSX achieves superior explanation fidelity with significantly reduced runtime, while providing unprecedented insights into the structural logic underlying model predictions–specifically, how important sub-structures exert influence by governing the stability of key internal computational pathways.
BACKGROUND:Parkinson's disease (PD) diagnosis typically occurs after motor symptom onset, highlighting the need for biomarkers predicting phenoconversion from the prodromal parkinson's disease (pPD). Free-water (FW) imaging, reflecting extracellular microstructural changes, may capture early extrastriatal pathology. OBJECTIVE:To evaluate regional free-water alterations from whole-brain mapping as biomarkers predicting imminent pPD-to-PD phenoconversion. METHODS:Whole-brain FW mapping derived from diffusion tensor imaging was obtained for 51 healthy controls (HCs), 83 pPD, and 202 de novo PD (dnPD) subjects from the Parkinson's Progression Markers Initiative. The pPD cohort underwent ≥4-year longitudinal follow-up. Whole-brain FW values were compared across groups using false discovery rate correction. Predictors of phenoconversion to PD were identified via LASSO regression from baseline FW metrics and clinical covariates. A multivariable Cox proportional hazards model was subsequently constructed using the selected predictors, with model performance evaluated through Kaplan-Meier survival analysis and receiver operating characteristic (ROC) curves. RESULTS:Cross-sectionally, elevation in FW values in the left insula and left middle cingulate gyrus were observed in pPD compared to HCs, alongside lower values than in dnPD. Longitudinally, LASSO-Cox analysis identified six predictors: FW in left insula, left mid-cingulate cortex, left hippocampus, right orbitofrontal cortex, RBDSQ, and STAI scores. Left insula FW (HR = 1.403,P = 0.004) and left mid-cingulate FW (HR = 1.325,P = 0.029) independently predicted conversion. The combined model integrating FW and clinical metrics demonstrated superior AUCs (0.755). CONCLUSION:Free-water changes particularly in limbic regions predict phenoconversion from prodromal to clinical Parkinson's disease.
Background:Breast ultrasound (BUS) is widely used for breast cancer (BC) screening and diagnosis, yet accurate breast lesion segmentation remains challenging. Although You Only Look Once (YOLO) and its variants have shown strong performance in object segmentation, their effectiveness on BUS lesion segmentation has not been systematically explored. This study aims to benchmark twelve YOLO variants from four families (YOLOv5, YOLOv8, YOLOv9, and YOLO11) for BUS lesion segmentation under same-database and cross-database settings. Methods:Twelve YOLO variants spanning nano to extra-large scales were fine-tuned and evaluated on two public BUS datasets, the Breast Ultrasound Images (BUSI) dataset (n=647) and the breast ultrasound lesion segmentation dataset from the University of Castilla-La Mancha (BUS-UCLM) dataset (n=264). Each dataset was split into training (80%), validation (10%), and testing (10%) subsets with stratified random partitioning, and experiments were repeated across eight random seeds. Six evaluation metrics, including Dice coefficient, intersection over union (IoU), precision, recall, F1 score (F1S), and mean average precision at IoU threshold 0.5 (mAP@0.5), were used. In addition, U-Net and DeepLabV3+ were compared under the same protocol. Results:Under same-database evaluation, all variants achieved strong performance, with mean Dice scores ≥0.81 on BUSI and ≥0.87 on UCLM. On BUSI, yolov5s achieved the highest mean Dice (0.93±0.012) and IoU (0.88±0.014). On UCLM, yolov8m attained the highest mean Dice (0.98±0.006) and IoU (0.96±0.007). In cross-database evaluation, however, performance degraded substantially, with Dice scores decreasing by approximately 0.20 or more. For BUSI→UCLM, yolov5s achieved the highest mean Dice (0.71±0.032); and for UCLM→BUSI, yolov9c achieved the highest mean Dice (0.60±0.038). Among all variants, yolo11n demonstrated competitive performance across both same- and cross-database evaluations (BUSI Dice 0.91±0.014, and BUSI→UCLM Dice 0.69±0.034; UCLM Dice 0.96±0.009, and UCLM→BUSI 0.59±0.039) while maintaining low computational cost (training time <290 s; inference latency ~64 ms). Notably, yolo11n substantially outperformed U-Net and DeepLabV3+ under both same- and cross-database settings. Conclusions:Fine-tuned YOLO variants achieve strong same-database performance for BUS lesion segmentation, with yolo11n offering the most favorable balance among segmentation accuracy, cross-database competitiveness, and computational efficiency. However, the substantial performance degradation in cross-database settings highlights the critical need for improving domain generalization performance in future work.
We study feature-level and node-level explanations for graph neural networks (GNNs) through the lens of Aumann-Shapley attribution. Path-integral methods such as Integrated Gradients provide an axiomatic formulation of attribution, but their practical use in deep GNNs typically relies on finite-sample numerical approximations to the path integral, requiring a trade-off between quadrature error and computational cost. This paper proposes APEX, a model-attribution co-design framework that makes the attribution integral exactly computable under a polynomial GNN architecture. The key component is PolyGIN, a GIN-style graph network whose message-passing, normalization, and transformation operations preserve a bounded multivariate polynomial form for scalar model scores, such as pre-softmax logits. We show that, for a PolyGIN with L polynomial transformation blocks, the derivative along the attribution path has degree at most 2^L-1. Therefore, Gauss–Legendre quadrature can evaluate the Aumann–Shapley path integral exactly, up to floating-point precision, with 2^L-1 deterministic evaluation points. The resulting attributions can be computed at the feature level and then aggregated into node-level scores while preserving completeness. Experiments on synthetic and real-world graph benchmarks show that PolyGIN maintains competitive predictive performance, while the complete APEX framework achieves higher attribution fidelity than the compared baselines and substantially reduces the number of evaluations required for path integration.
Artificial intelligence (AI) is transforming modern agriculture from experience-driven practices to data-driven production paradigms. To provide an in-depth analysis of AI technologies in intelligent agriculture, we retrieved literature from Web of Science, IEEE Xplore, Google Scholar and Scopus, covering publications from 2015 to 2025, and 85 articles remained after screening 1867 relevant publications. These articles are grouped into three stages from perception, to decision making, to execution (PDE) in a closed-loop framework. At the perception level, we highlight progress in intelligent sensing systems, such as unmanned aerial vehicle (UAV) and multi-modal monitoring platforms, for crop disease and pest detection, growth monitoring and abiotic stress assessment. At the decision making level, integration of heterogeneous data sources, including meteorological records, soil measurements, remote sensing (RS) imagery and market information, supports advanced analytics, such as yield prediction, pest and disease warning, irrigation and fertilization planning, and crop management optimization. At the execution level, agricultural robots equipped with simultaneous localization and mapping (SLAM) and deep reinforcement learning (RL) facilitate precision spraying, autonomous harvesting, and unmanned field operations. Overall, AI technologies demonstrate substantial potential in the PDE pipeline of agricultural production. However, several challenges remain, including heterogeneous data fusion, limited generalization across diverse environments, complex system integration, and high hardware and deployment costs. Future directions are discussed from the perspectives of lightweight model design, cross-platform standardization, enhanced human–machine collaboration, and a deeper integration of emerging AI paradigms to support scalable, robust, and autonomous agricultural intelligence systems.
Ultra-high-definition (UHD) blind image quality assessment (BIQA) is challenging because native-resolution inference is computationally expensive, whereas common strategies of resizing or patching may suppress scale-sensitive distortions and weaken the relationship between local artifacts and global scene context. To address this challenge, a Global-Local graph representation learning framework for UHD image Quality prediction (GLUQ) is proposed, which models structural dependencies among patches rather than treating them as independent views. Specifically, it samples aspect ratio-aligned patches from each UHD image, encodes these patches as graph nodes and constructs a hybrid k-nearest-neighbor graph via weighted spatial proximity and feature similarity. Residual graph convolution is used to propagate contextual information across regions, and gated attention pooling is used to aggregate patch-level evidence into image-level quality prediction. Besides, an exponential moving average normalized multi-objective loss function is adopted to stabilize the joint optimization of regression, correlation, and ranking objectives. Experiments on the UHD-IQA benchmark database show that GLUQ achieves the lowest RMSE among the compared methods with competitive PLCC and SRCC, indicating strong absolute-score calibration for UHD image quality prediction. The results suggest modeling global-local graph relations enhances quality prediction for UHD images with extremely high-resolution visual content.
Blind image quality assessment (BIQA) for ultrahighdefinition (UHD) images remains challenging because native-resolution inference is computationally expensive, whereas aggressive resizing or isolated cropping may suppress scale-sensitive distortions and weaken the relationship between local artifacts and global scene context. This paper aims to improve UHD-BIQA by explicitly modeling the structural dependencies among sampled image regions rather than treating them as independent views, and a graph representation learning framework UHD-GCN-BIQA is proposed. The framework samples aspect-ratio-aligned patches from each UHD image, encodes them as graph nodes, and constructs a hybrid k-nearest-neighbor graph using spatial proximity and feature similarity. Residual graph convolution is used to propagate contextual information across regions, and gated attention pooling aggregates patchlevel evidence into an imagelevel quality prediction. An exponential moving average normalized multiobjective loss function is adopted to stabilize the joint optimization of regression, correlation, and ranking objectives. Experiments on the UHD-IQA benchmark show that UHD-GCN-BIQA achieves PLCC = 0.7784, SRCC = 0.8019, and RMSE = 0.0519, obtaining competitive correlation performance and the lowest RMSE among the compared methods. These results indicate that graph-based region relation modeling is effective for UHD image quality assessment, particularly for improving absolute quality score estimation under high-resolution visual content.
Wildlife detection is challenging due to various sizes, distance variability, and appearance similarity. To enhance the wildlife detection performance, the YOLOv11 network is updated with an ADown module in the trunk and neck network for efficient sampling of feature maps, a CBAM attention module at the feature fusion endpoint to focus on salient regions, and an ASFFHead module in the detection head for adaptive multi-scale fusion. Consequently, ACA-net, a modified YOLOv11 network, is formed. Extensive experiments on the NTLNP database show that the ACA-net achieves increase on the mAP@0.5(2.6% ↑) and Recall $(4.4 {\%} \uparrow)$ metrics compared to the baseline architecture. Network upgrading with efficient modules could further improve the baseline performance and benefit wildlife detection.
Kolmogorov-Arnold Network (KAN) has attracted growing interest for its strong function approximation capability. In our previous work, KAN and its variants were explored in score regression for blind image quality assessment (BIQA). However, these models encounter challenges when processing high-dimensional features, leading to limited performance gains and increased computational cost. To address these issues, we propose TaylorKAN that leverages the Taylor expansions as learnable activation functions to enhance local approximation capability. To improve the computational efficiency, network depth reduction and feature dimensionality compression are integrated into the TaylorKAN-based score regression pipeline. On five databases (BID, CLIVE, KonIQ, SPAQ, and FLIVE) with authentic distortions, extensive experiments demonstrate that TaylorKAN consistently outperforms the other KAN-related models, indicating that the local approximation via Taylor expansions is more effective than global approximation using orthogonal functions. Its generalization capacity is validated through inter-database experiments. The findings highlight the potential of TaylorKAN as an efficient and robust model for high-dimensional score regression.
Deep learning-based respiratory sound classification (RSC) has emerged as a promising non-invasive approach to assist clinical diagnosis. However, existing methods often face challenges, such as sub-optimal feature representation and limited model expressiveness. To address these issues, we propose an Attention-based Dual-stream Feature Fusion Network (ADFF-Net). Built upon the pre-trained Audio Spectrogram Transformer, ADFF-Net takes Mel-filter bank and Mel-spectrogram features as dual-stream inputs, while an attention-based fusion module with a skip connection is introduced to preserve both the raw energy and the relevant tonal variations within the multi-scale time–frequency representation. Extensive experiments on the ICBHI2017 database with the official train–test split show that, despite critical failure in sensitivity of 42.91%, ADFF-Net achieves state-of-the-art performance in terms of aggregated metrics in the four-class RSC task, with an overall accuracy of 64.95%, specificity of 81.39%, and harmonic score of 62.14%. The results confirm the effectiveness of the proposed attention-based dual-stream acoustic feature fusion module for the RSC task, while also highlighting substantial room for improving the detection of abnormal respiratory events. Furthermore, we outline several promising research directions, including addressing class imbalance, enriching signal diversity, advancing network design, and enhancing model interpretability.
Score prediction is crucial in evaluating realistic image sharpness based on collected informative features. Recently, Kolmogorov-Arnold networks (KANs) have been developed and witnessed remarkable success in data fitting. This study introduces the Taylor series-based KAN (TaylorKAN). Then, different KANs are explored in four realistic image databases (BID2011, CID2013, CLIVE, and KonIQ-10k) to predict the scores by using 15 mid-level features and 2048 high-level features. Compared to support vector regression, results show that KANs are generally competitive or superior, and TaylorKAN is the best one when mid-level features are used. This is the first study to investigate KANs on image quality assessment that sheds some light on how to select and further improve KANs in related tasks.
Respiratory diseases present significant global health challenges. Recent advances in respiratory sound analysis (RSA) have shown great potential for automated disease diagnosis and patient management. The International Conference on Biomedical and Health Informatics 2017 (ICBHI2017) database stands as one of the most authoritative open-access RSA datasets. This review systematically examines 135 technical publications utilizing the database, and a comprehensive and timely summary of RSA methodologies is offered for researchers and practitioners in this field. Specifically, this review covers signal processing techniques including data resampling, augmentation, normalization, and filtering; feature extraction approaches spanning time-domain, frequency-domain, joint time–frequency analysis, and deep feature representation from pre-trained models; and classification methods for adventitious sound (AS) categorization and pathological state (PS) recognition. Current achievements for AS and PS classification are summarized across studies using official and custom data splits. Despite promising technique advancements, several challenges remain unresolved. These include a severe class imbalance in the dataset, limited exploration of advanced data augmentation techniques and foundation models, a lack of model interpretability, and insufficient generalization studies across clinical settings. Future directions involve multi-modal data fusion, the development of standardized processing workflows, interpretable artificial intelligence, and integration with broader clinical data sources to enhance diagnostic performance and clinical applicability.
In the Ultra-High-Definition (UHD) domain, blind image quality assessment remains challenging due to the high dimensionality of UHD images, which exceeds the input capacity of deep learning networks. Motivated by the visual discrepancies observed between high- and low-quality images after down-sampling and Super-Resolution (SR) reconstruction, we propose a SUper-Resolved Pseudo References In Dual-branch Embedding (SURPRIDE) framework tailored for UHD image quality prediction. SURPRIDE employs one branch to capture intrinsic quality features from the original patch input and the other to encode comparative perceptual cues from the SR-reconstructed pseudo-reference. The fusion of the complementary representation, guided by a novel hybrid loss function, enhances the network’s ability to model both absolute and relational quality cues. Key components of the framework are optimized through extensive ablation studies. Experimental results demonstrate that the SURPRIDE framework achieves competitive performance on two UHD benchmarks (AIM 2024 Challenge, PLCC = 0.7755, SRCC = 0.8133, on the testing set; HRIQ, PLCC = 0.882, SRCC = 0.873). Meanwhile, its effectiveness is verified on high- and standard-definition image datasets across diverse resolutions. Future work may explore positional encoding, advanced representation learning, and adaptive multi-branch fusion to align model predictions with human perceptual judgment in real-world scenarios.
Freezing of gait (FOG) greatly impacts the daily life of patients with Parkinson’s disease (PD). However, predictors of FOG in early PD are limited. Moreover, recent neuroimaging evidence of cerebral morphological alterations in PD is heterogeneous. We aimed to develop a model that could predict the occurrence of FOG using machine learning, collaborating with clinical, laboratory, and cerebral structural imaging information of early drug-naïve PD and investigate alterations in cerebral morphology in early PD. Data from 73 healthy controls (HCs) and 158 early drug-naïve PD patients at baseline were obtained from the Parkinson’s Progression Markers Initiative cohort. The CIVET pipeline was used to generate structural morphological features with T1-weighted imaging (T1WI). Five machine learning algorithms were calculated to assess the predictive performance of future FOG in early PD during a 5-year follow-up period. We found that models trained with structural morphological features showed fair to good performance (accuracy range, 0.67–0.73). Performance improved when clinical and laboratory data was added (accuracy range, 0.71–0.78). For machine learning algorithms, elastic net-support vector machine models (accuracy range, 0.69–0.78) performed the best. The main features used to predict FOG based on elastic net-support vector machine models were the structural morphological features that were mainly distributed in the left cerebrum. Moreover, the bilateral olfactory cortex (OLF) showed a significantly higher surface area in PD patients than in HCs. Overall, we found that T1WI morphometric markers helped predict future FOG occurrence in patients with early drug-naïve PD at the individual level. The OLF exhibits predominantly cortical expansion in early PD.
Breast cancer is a global threat to women’s health. Three-dimensional (3D) automated breast ultrasound (ABUS) offers reproducible high-resolution imaging for breast cancer diagnosis. However, 3D-input deep networks are challenged by high time costs, a lack of sufficient training samples, and the complexity of hyper-parameter optimization. For efficient ABUS tumor classification, this study explores 2D-input networks, and soft voting (SV) is proposed as a post-processing step to enhance diagnosis effectiveness. Specifically, based on the preliminary predictions made by a 2D-input network, SV employs voxel-based weighting, and hard voting (HV) utilizes slice-based weighting. Experimental results on 100 ABUS cases show a substantial improvement in classification performance. The diagnosis metric values are increased from ResNet34 (accuracy, 0.865; sensitivity, 0.942; specificity, 0.757; area under the curve (AUC), 0.936) to ResNet34 + HV (accuracy, 0.907; sensitivity, 0.990; specificity, 0.864; AUC, 0.907) and to ResNet34 + SV (accuracy, 0.986; sensitivity, 0.990; specificity, 0.963; AUC, 0.986). Notably, ResNet34 + SV achieves the state-of-the-art result on the database. The proposed SV strategy enhances ABUS tumor classification with minimal computational overhead, while its integration with 2D-input networks to improve prediction performance of other 3D object recognition tasks requires further investigation.
Medical imaging description and disease diagnosis are vitally important yet time-consuming. Automated diagnosis report generation (DRG) from medical imaging description can reduce clinicians’ workload and improve their routine efficiency. To address this natural language generation task, fine-tuning a pre-trained large language model (LLM) is cost-effective and indispensable, and its success has been witnessed in many downstream applications. However, semantic inconsistency of sentence embeddings has been massively observed from undesirable repetitions or unnaturalness in text generation. To address the underlying issue of anisotropic distribution of token representation, in this study, a contrastive learning penalized cross-entropy (CLpCE) objective function is implemented to enhance the semantic consistency and accuracy of token representation by guiding the fine-tuning procedure towards a specific task. Furthermore, to improve the diversity of token generation in text summarization and to prevent sampling from unreliable tail of token distributions, a diversity contrastive search (DCS) decoding method is designed for restricting the report generation derived from a probable candidate set with maintained semantic coherence. Furthermore, a novel metric named the maximum of token repetition ratio (maxTRR) is proposed to estimate the token diversity and to help determine the candidate output. Based on the LLM of a generative pre-trained Transformer 2 (GPT-2) of Chinese version, the proposed CLpCE with DCS (CLpCEwDCS) decoding framework is validated on 30,000 desensitized text samples from the “Medical Imaging Diagnosis Report Generation” track of 2023 Global Artificial Intelligence Technology Innovation Competition. Using four kinds of metrics evaluated from n-gram word matching, semantic relevance, and content similarity as well as the maxTRR metric extensive experiments reveal that the proposed framework effectively maintains semantic coherence and accuracy (BLEU-1, 0.4937; BLEU-2, 0.4107; BLEU-3, 0.3461; BLEU-4, 0.2933; METEOR, 0.2612; ROUGE, 0.5182; CIDER, 1.4339) and improves text generation diversity and naturalness (maxTRR, 0.12). The phenomenon of dull or repetitive text generation is common when fine-tuning pre-trained LLMs for natural language processing applications. This study might shed some light on relieving this issue by developing comprehensive strategies to enhance semantic coherence, accuracy and diversity of sentence embeddings.
Background: Early detection of Parkinson's disease (PD) patients at high risk for mild cognitive impairment (MCI) can help with timely intervention. White matter structural connectivity is considered an early and sensitive indicator of neurodegenerative disease. Objectives: To investigate whether baseline white matter structural connectivity features from diffusion tensor imaging (DTI) of de novo PD patients can help predict PD-MCI conversion at an individual level using machine learning methods.Methods: We included 90 de novo PD patients who underwent DTI and 3D T1-weighted imaging. Elastic net based feature consensus ranking (ENFCR) was used with 1000 random training sets to select clinical and structural connectivity features. Linear discrimination analysis (LDA), support vector machine (SVM), K-nearest neighbor (KNN) and naive Bayes (NB) classifiers were trained based on features selected more than 500 times. The area under the ROC curve (AUC), accuracy (ACC), sensitivity (SEN) and specificity (SPE) were used to evaluate model performance.Results: A total of 57 PD patients were classified as PD-MCI nonconverters, and 33 PD patients were classified as PD-MCI converters. The models trained with clinical data showed moderate performance (AUC range: 0.62-0.68; ACC range: 0.63-0.77; SEN range: 0.45-0.66; SPE range: 0.64-0.84). Models trained with structural connectivity (AUC range, 0.81-0.84; ACC range, 0.75-0.86; SEN range, 0.77-0.91; SPE range, 0.71-0.88) performed similar to models that were trained with both clinical and structural connectivity data (AUC range, 0.81-0.85; ACC range, 0.74-0.85; SEN range, 0.79-0.91; SPE range, 0.70-0.89).Conclusions: Baseline white matter structural connectivity from DTI is helpful in predicting future MCI conversion in de novo PD patients.
Objective.Low-coupling seamless integration of multiple systems is the core foundation of smart radiotherapy. Following Service-Oriented Architecture style, a set of named operations (Eclipse Web Service API, EWSAPI) was developed for realizing network call of Eclipse.Approach.Under the guidance of Vertical Slice Architecture, EWSAPI was implemented in the C# language and based on ASP .Net Core 6.0. Each operation consists of three components: Request, Endpoint and Response. Depending on the function, the exchanged data for each operation, as input or output parameters, is the empty or a predefined JSON data. These operations were realized and enriched gradually, layer by layer, with reference to the clinical business classification. The business logic of each operation was developed and maintained independently. In situations where Eclipse Scripting API(ESAPI) was required, constraints of ESAPI were followed.Main results.Selected features of Eclipse TPS were encapsulated as standard web services, which can be invocated by other software through network. Several processes for data quality control and planning were encapsulated into interfaces, thereby extending the functionality of Eclipse. Currently, EWSAPI already covers testing of service interface, quality control of radiotherapy data, automation tasks for plan designing and DICOM RT files' transmission. All the interfaces support asynchronous invocation. A separate Eclipse context will be created for each invocation, and is released in the end.Significance.EWSAPI which is a set of standard web services for calling Eclipse features through network is flexible and extensible. It is an efficient way to integration of Eclipse and other systems and will be gradually enriched with the deepening of clinical applications.
The phonocardiogram (PCG) is a crucial tool for the early detection, continuous monitoring, accurate diagnosis, and efficient management of cardiovascular diseases. It has the potential to revolutionize cardiovascular care and improve patient outcomes. The PhysioNet/CinC Challenge 2016 database, a large and influential resource, encourages contributions to accurate heart sound state classification (normal versus abnormal), achieving promising benchmark performance (accuracy: 99.80%; sensitivity: 99.70%; specificity: 99.10%; and score: 99.40%). This study reviews recent advances in analytical techniques applied to this database, and 104 publications on PCG signal analysis are retrieved. These techniques encompass heart sound preprocessing, signal segmentation, feature extraction, and heart sound state classification. Specifically, this study summarizes methods such as signal filtering and denoising; heart sound segmentation using hidden Markov models and machine learning; feature extraction in the time, frequency, and time-frequency domains; and state-of-the-art heart sound state recognition techniques. Additionally, it discusses electrocardiogram (ECG) feature extraction and joint PCG and ECG heart sound state recognition. Despite significant technical progress, challenges remain in large-scale high-quality data collection, model interpretability, and generalizability. Future directions include multi-modal signal fusion, standardization and validation, automated interpretation for decision support, real-time monitoring, and longitudinal data analysis. Continued exploration and innovation in heart sound signal analysis are essential for advancing cardiac care, improving patient outcomes, and enhancing user trust and acceptance.