A decrease in Minimum Foot Clearance (MFC), which represents the minimum vertical distance of the foot from the ground surface during the swing phase of the gait cycle, is one of the primary contributors to tripping-related falls among people with stroke. Accurate prediction of upcoming MFC values, understanding their potential range, and evaluating whether future values will fall within the variability of actual MFC values are crucial for early fall risk assessment. This work proposes a new Transformer model for predicting multistep MFC values in stroke survivors collected from the affected lower-limb during walking on a treadmill utilizing a two-head self-attention mechanism. We introduce a data-driven conditional post-normalization projection approach to enhance the performance of multistep prediction for MFC values. In addition, we introduce a statistical moment-matching loss function during training to account for significant data variability, such as that observed in individuals with stroke. We compare the performance of our model with two other deep learning models. Our findings indicate that the Transformer model training is faster and achieves an average Mean Absolute Error (MAE) of approximately 0.0035, Maximum Absolute Error (MaxAE) of 0.0085 and a Root Mean Square Error (RMSE) of 0.0043, using the leave-one-out cross-validation (LOOCV) method. The highest prediction within the upper bound (PWUB) of around 89% is achieved with self-attention LSTM model. The main contribution of our work is to identify when an individual stroke survivor is at increased risk of tripping-related falls due to a lower MFC value in their affected lower limb.
The proliferation of high-velocity big data streams from contemporary technologies, such as the internet of things, social media platforms, wireless sensor networks, blockchain systems, etc., has intensified the need for scalable methods to model and analyze evolving big data representations. Advanced evolutionary representation learning techniques, employing deep learning, graph neural networks, and transformer architectures, generate time-varying feature vectors for diverse entities, including users, sensor nodes, and cryptocurrency wallets. Unsupervised learning over these evolving representations enables the discovery of latent structure, emerging patterns, and anomalies, which are critical for characterizing concept drift. Existing evolutionary clustering approaches, however, often assume fixed data points and a constant number of clusters across timestamps, or lack the scalability required for large-scale, real-world applications. Furthermore, most methods are unable to account for essential dynamics, such as cluster splitting and merging, which limits their ability to capture complex drift behaviors. To overcome these limitations, this paper presents evolVAT, a fast and scalable evolutionary clustering algorithm that incrementally updates clustering results by leveraging previously inferred structure. The proposed method accommodates multiple data-point transitions between consecutive snapshots and supports the addition and removal of entities over time. Experimental evaluations on a broad set of synthetic and real-world datasets demonstrate the effectiveness, robustness, and adaptability of evolVAT in diverse application domains.
Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data. However, SSL performance depends not only on model architecture but also on whether the self-supervised objective preserves the information required by the downstream clinical task. This review presents a task-oriented synthesis of SSL methods for medical imaging, focusing on how the design of the self-supervised objective interacts with imaging modality, label availability, and downstream performance. We analyze 78 studies published from 2017 to 2025 and organize them into four paradigms: contrastive, non-contrastive and predictive, generative and reconstruction-based, and hybrid learning. Rather than cataloging methods chronologically, we examine how these paradigms support classification, segmentation, detection, reconstruction, and regression. The evidence suggests that effectiveness is governed by the match among objective, modality, and downstream task rather than by any single strategy. Contrastive objectives favor global discriminative representations suited to classification but may underrepresent localized pathology, whereas spatial-prediction, masked-modeling, and reconstruction objectives better preserve anatomical structure for segmentation and dense prediction. Critically, misaligned objectives can cause negative transfer through shortcut learning on acquisition signatures or augmentation that erases diagnostic signal rather than merely weaker gains. SSL is most beneficial in low-label regimes, but its effectiveness depends on modality-aware augmentation, pathology-preserving corruption, and clinically meaningful evaluation. We conclude with practical design guidelines and open challenges for clinically aligned SSL.
Fetal compromise is a significant global health issue that can result in severe long-term disability and mortality. Current fetal monitoring is predominantly performed using cardiotocography (CTG) machines in healthcare settings, limiting clinical oversight to routine hospital visits or during labor. CTG monitoring typically relies on intermittent visual interpretation, sometimes resulting in inconsistent clinical judgement and delayed intervention. Artificial intelligence (AI) has been proposed to assist in CTG interpretation, but conventional AI models require substantial computational resources, making them impractical for wearable, battery-powered devices designed for continuous monitoring. In this study, we propose a quantized Edge AI model for detecting fetal compromise (pH < 7.05) in a resource-constrained setting. Utilizing the low-power MAX78002 AI microcontroller, we demonstrate real-time fetal compromise detection with a deep learning model optimized for power efficiency. Our final model, trained using knowledge distillation combined with quantization-aware training on an internal dataset of 9,887 CTG recordings and tested on 552 CTG recordings from the public CTU-UHB dataset, achieves an AUC of 0.81 while consuming only 2 mJ of energy per 60 minutes of inference data. This represents comparable performance to a GPU-based model on the same dataset, while achieving a 94% reduction in energy consumption. This work demonstrates that deep learning models for fetal compromise detection can be quantized and deployed on low-power Edge AI hardware, paving the way for wearable devices capable of continuous, real-time monitoring.
There are currently no measures to accurately predict the onset of labor at term. Currently, the onset of labor is anticipated based on the estimated due date (EDD), which is derived from the day of the last menstrual period or ultrasound-based anatomical information. However, the EDD is not intended to identify physiological factors which may result in the early onset of labor. Therefore, there is a need to identify potential biomarkers that are associated with the onset of labor to accurately predict the timing of delivery. In this exploratory study, we investigated the associations between maternal RR interval (mRRI), maternal heart rate variability (mHRV) features, and the onset of labor. A total of 37 participants were analyzed, including 25 with Electrohysterogram (EHG)-derived signals (age: 28 ± 5.9 years; gestational age (GA): 34 ± 2.7 weeks) and 12 with non-invasive electrocardiogram (NIFECG)-derived signals (age: 32 ± 4.5 years; GA: 38 ± 1.5 weeks). The association of mHRV with the onset of labor was quantified by calculating correlations with time to delivery, defined as the difference between GA at recording and GA at delivery. Correlation analysis revealed that several standard mHRV indices showed strong associations (r > 0.5) with time to delivery.
Background Diagnosis of occult atrial fibrillation (AF) is difficult as it is often asymptomatic, leading to under detection. Current diagnostic tests have variable limitations in feasibility and accuracy. Machine learning is gaining greater traction for clinical decision making and may help facilitate the detection of undiagnosed AF when applied to magnetic resonance imaging (MRI). We hypothesise that machine learning algorithm increases the accurate classification of MRIs of stroke patients into those due to AF vs large artery atherosclerosis. Methods Stroke aetiology for each patient was determined by a review of medical records and investigations. Patients with either AF or large artery atherosclerosis were included. Patients were randomly divided into the training and validation groups (4:1). A 3D convolutional neural network (ConvNeXt) was developed to train and validate the algorithm. After training, the models were evaluated using common metrics for binary classification. Results A total of 235 patients were analysed (97 with AF, 138 without AF). The mean age of the sample was 71.1 (SD 14.2) and 35% percent were female. The best discriminative performance was obtained in the 5th fold of cross-validation (AUC-ROC 0.88) and the overall model performance was 0.81. The best performing metrics were precision (0.84) and the F1-score (0.77). Conclusion Our machine learning algorithm has reasonable classification power in categorizing stroke patients into those with and without underlying AF. Testing in external validation data sets are critical to confirm these results.
Atrial fibrillation (AF) is a significant risk factor for ischemic stroke recurrence, yet its diagnosis remains challenging through short-term heart monitoring due to its often paroxysmal and silent nature. Despite its diagnostic superiority, prolonged cardiac monitoring is typically impractical and not cost-effective for widespread implementation. We propose a novel AF risk stratification framework using a multimodal deep learning approach that integrates diffusion-weighted imaging (DWI) of the brain with clinical patient data. Our methodology combines convolutional neural networks (CNNs) for image analysis and gradient-boosted decision trees (GBDT) for clinical data, leveraging an innovative fusion strategy and an auxiliary loss function based on infarct location. The proposed approach achieves an area under the receiver operating characteristic (AUROC) of 89.18%, outperforming unimodal counterparts. This work contributes to the field by enabling AF risk stratification from brain DWI, utilizing weak supervision, and introducing a novel early and late-stage data fusion approach. Our method easily integrates with existing workflows and can identify high-risk individuals requiring intensive cardiac monitoring.
Continuous monitoring of fetal heart rate (FHR) and uterine contractions (UC), otherwise known as cardiotocography (CTG), is often used to assess the risk of fetal compromise during labor. However, interpreting CTG recordings visually is challenging for clinicians, given the complexity of CTG patterns, leading to poor sensitivity. Efforts to address this issue have focused on data-driven deep-learning methods to detect fetal compromise automatically. However, their progress is impeded by limited CTG training datasets and the absence of a standardized evaluation workflow, hindering algorithm comparisons. In this study, we use a private CTG dataset of 9,887 CTG recordings with pH measurements and 552 CTG recordings from the open-access CTU-UHB dataset to conduct a cross-database evaluation of six deep-learning models for fetal compromise detection. We explore the impact of input selection of FHR and UC signals, signal pre-processing, downsampling frequency, and the influence of removing intermediate pH samples from the training dataset. Our findings reveal that using only FHR and pre-processing FHR with artefact removal and interpolation provides a significant improvement to classification performance for some model architectures while excluding intermediate pH samples did not significantly improve performance for any model. From our comparison of the six models, ResNet exhibited the strongest fetal compromise classification performance across both databases at a downsampling rate of 1Hz. Finally, class activation maps from highly contributing signal regions in the ResNet model aligned with clinical knowledge of compromised FHR patterns, highlighting the model’s interpretability. These insights may serve as a standardized reference for developing and comparing future works in this domain. Clinical and Translational Impact: This study provides a standardized workflow for comparing deep-learning methods for CTG classification. Ensuring new methods show generalizability and interpretability will improve their robustness and applicability in clinical settings.
Cardiotocography (CTG) is essential for monitoring high-risk pregnancies, yet perinatal asphyxia prediction accuracy remains limited to 50–55%. Regions of artifacts (missing valid signals)-including signal processing aberrations-possibly contribute to this limitation, highlighted by 40% of FDA reports on intrapartum stillbirths. This cohort study applied causal inference to two digitized CTG databases, analyzing 36,792 labor episodes (>36 weeks) at a tertiary Australian hospital (2010–2021) and externally validating on a Czech dataset (n = 552).High rates of missing valid signals (>30% fetal heart rate signal dropout or >1% maternal-fetal heart rate coincidence) was independently associated with asphyxia (aOR 1.47, 95% CI 1.19–1.81); dropout >30% showing stronger link (aOR 1.58, 95% CI 1.13–2.20 Australian dataset; aOR 2.30, 95% CI 1.08–4.91 Czech dataset). Risk of asphyxia increased with higher dropout (>37.45%, aOR 2.21 Australian dataset; >34.01%, aOR 4.08 Czech dataset). Integrating measures of missing valid signals into fetal monitoring algorithms may improve decision-making and neonatal outcomes.
Industry 5.0 integrates advanced technologies like Automated Guided Vehicles (AGVs) and Augmented/Virtual Reality (AR/VR) with human expertise, requiring ultra-reliable communication for safe and efficient manufacturing. Network resource allocation in this context is challenging, demanding efficient support for diverse applications while meeting stringent performance targets, including 99.9999% availability. This study presents a novel application-aware resource allocation scheme for an Industry 5.0 system connected to a 5G network. Our approach dynamically adapts to industrial application states, bridging network optimization and real-time factory operations. The solution comprises (1) a learning-based framework for safety and productivity-conscious allocation policies, (2) a heuristic real-time resource allocation policy addressing the computational scalability problem of learning-based methods, and (3) statistical analysis for bandwidth requirement estimation. Simulation results show our method achieves 99.9999% availability while reducing bandwidth usage by nearly 50% compared to traditional methods. This work contributes to more efficient and scalable Industry 5.0 systems, potentially doubling the number of supported industrial components within the same network infrastructure.
Stroke is one of the leading causes of disability among the elderly population and is a significant public health problem worldwide. The main impact of stroke is functional disabilities due to motor impairment after stroke. Advances in modern medicine and technology have significantly improved diagnosis and treatment; however, most post-stroke care is based on the effectiveness of rehabilitation. Stroke rehabilitation depends on two main components: (i) training (or therapy) to restore the patient to pre-stroke mobility and (ii) assessing motor functionality of affected patients performing activities to track motor recovery. This article highlights how combining wearable devices and machine learning (ML) produces new pathways for effective stroke rehabilitation. While wearable devices help capture patient movements at much finer time resolutions, ML allows us to build predictive models from wearable data to assist clinicians in diagnosis and treatments. Specifically, we expand on how wearable devices and ML can improve monitoring quality in training intervention, assessment, and remote monitoring. In addition, we provide our main findings from the literature, research challenges, and future directions in post-stroke therapies using wearable devices and ML.
Stroke rehabilitation interventions require multiple training sessions and repeated assessments to evaluate the improvements from training. Biofeedback-based treadmill training often involves 10 or more sessions to determine its effectiveness. The training and assessment process incurs time, labor, and cost to determine whether the training produces positive outcomes. Predicting the effectiveness of gait training based on baseline minimum foot clearance (MFC) data would be highly beneficial, potentially saving resources, costs, and patient time. This work proposes novel features using the Short-term Fourier Transform (STFT)-based magnitude spectrum of MFC data to predict the effectiveness of biofeedback training. This approach enables tracking non-stationary dynamics and capturing stride-to-stride MFC value fluctuations, providing a compact representation for efficient processing compared to time-domain analysis alone. The proposed STFT-based features outperform existing wavelet, histogram, and Poincaré-based features with a maximum accuracy of 95%, F1 score of 96%, sensitivity of 93.33% and specificity of 100%. The proposed features are also statistically significant (p<0.001) compared to the descriptive statistical features extracted from the MFC series and the tone and entropy features extracted from the MFC percentage index series. The study found that short-term spectral components and the windowed mean value (DC value) possess predictive capabilities regarding the success of biofeedback training. The higher spectral amplitude and lower variance in the lower frequency zone indicate lower chances of improvement, while the lower spectral amplitude and higher variance indicate higher chances of improvement.
Self-supervised learning has achieved state-of-the-art performance in various tasks and applications. In computer vision, self-supervised learning often employs contrastive learning and masked image modeling, each with its limitations: contrastive learning heavily relies on strong data augmentation and large batch sizes, etc., while masked image modeling struggles to capture high-level semantics and discrimination. In this work, we introduce MAsked Contrastive Representation Learning (MACRL), a novel framework that integrates both paradigms through an asymmetric siamese network design. The online and momentum branches of the network receive asymmetric data augmentation operations and extract features through their encoders. The decoder in the online branch reconstructs the original image, while the projectors in both branches compute the contrastive loss. The online branch and the momentum branch are updated through gradient backpropagation and exponential moving average, respectively. MACRL jointly optimizes the reconstruction and the contrastive objectives to encourage representations with enhanced discrimination and semantics. Experimental results show that MACRL achieves competitive performance in downstream vision tasks, including image classification and semantic segmentation. Moreover, it demonstrates consistent performance across both large-scale and small-scale datasets.
In recent years, spatio-temporal graph neural networks (GNNs) have successfully been used to improve traffic prediction by modeling intricate spatio-temporal dependencies in irregular traffic networks. However, these approaches may not capture the intrinsic properties of traffic data and can suffer from overfitting due to their local nature. This paper introduces the Implicit Sensing Self-Supervised learning model (ISSS), which leverages a multi-pretext task framework for traffic flow prediction. By transforming data into an alternative feature space, ISSS effectively captures both specific and general representations through self-supervised tasks, including contrastive learning and spatial jigsaw puzzles. This enhancement promotes a deeper understanding of traffic features, improved regularization, and more accurate representations. Comparative experiments on six datasets demonstrate the effectiveness of ISSS in learning general and discriminative features in both supervised and unsupervised modes. ISSS outperforms existing models, demonstrating its capabilities in improving traffic flow predictions while addressing challenges associated with local operations and overfitting. Comprehensive evaluations across various traffic prediction datasets, have established the validity of the proposed approach. Unsupervised learning scenarios have shown the improvements in RMSE for the METR-LA and PEMSBAY datasets of 0.39 and 0.35 for location-dependent and location-independent tasks, respectively. In supervised learning scenarios, for the same datasets, the improvements were 1.16 for location-dependent tasks and 0.55 for location-independent tasks.
You have accessJournal of UrologyPenile & Testicular Cancer II (MP61)1 May 2024MP61-02 SNAP DIAGNOSIS: PILOT STUDY OF AI-POWERED SMARTPHONE APPLICATION FOR PENILE CANCER DETECTION FROM THE COMFORT OF YOUR HOME Jianliang Liu, Jonathan S. O'Brien, Nandakishor Desai, Marimuthu Palaniswami, and Nathan Lawrentschuk Jianliang LiuJianliang Liu , Jonathan S. O'BrienJonathan S. O'Brien , Nandakishor DesaiNandakishor Desai , Marimuthu PalaniswamiMarimuthu Palaniswami , and Nathan LawrentschukNathan Lawrentschuk View All Author Informationhttps://doi.org/10.1097/01.JU.0001009536.58867.87.02AboutPDF ToolsAdd to favoritesDownload CitationsTrack CitationsPermissionsReprints ShareFacebookLinked InTwitterEmail Abstract INTRODUCTION AND OBJECTIVE: Penile squamous cell carcinoma (PSCC) is a rare disease with devastating psychosocial consequences. Early recognition and treatment are paramount to preserving function and long-term survival. However, many men present with advanced disease due to a lack of awareness, social stigma, and limited access to culturally appropriate care. Machine learning combined with smartphone photography has demonstrated growing utility by outperforming clinicians in diagnosing skin cancer. This project aims to evaluate the accuracy of an artificial intelligence algorithm for stratifying penile lesions. METHODS: A Google image search was performed for high-quality colour images of penile lesions in peer-reviewed English language articles with a formal diagnosis discussed within the article. The search terms used were "penile cancer", "penile squamous cell carcinoma (SCC)", "penile lesion", "benign penile lesion", "penile carcinoma in situ (CIS)", "penile neoplasm in situ (PeIN)". Images that fit inclusion criteria were downloaded as JPEG files and categorized as benign, pre-malignant, and penile SCC. Two penile cancer subspecialist urologists independently reviewed all images to confirm categorization. A deep learning algorithm was created to extract and automatically segment images into pixels. Contours of pixel edges were used to generate a "mask" representation of the lesion from normal skin. Features such as lesion elevation, erythema, ulceration, redness, and irregularity were extracted and compared. A hold-out validation methodology was performed for training and internally assessing accuracy. RESULTS: One hundred thirty-eight images from 83 articles were included – 67 invasive PSCC, 44 carcinoma in situ (CIS), and 27 benign. Subspecialist urologist agreement on image categorization was 96%. Ten rounds of algorithm training were performed on 98 randomly assigned images, equating to 980 experiments. A total of 40 images were randomly assigned to the test subset, sequestered from the algorithm, and used for internal accuracy validation. The algorithm demonstrated an overall triage accuracy of 87.5%. CONCLUSIONS: We present the first study to identify the role of artificial intelligence in accurately categorizing penile lesions. Penile lesions beneath non-retractile foreskin were identified as a limitation for image-based PSCC detection. Further image incorporation and validation with clinical images from electronic medical records are underway. An opportunity exists to refine artificial intelligence as an education, triage, and referral optimization tool via a smartphone application. Source of Funding: Grant has been awarded by the Epworth Medical Foundation, and the Australian Chinese Medical Association of Victoria for this project. No conflict of interest to disclose © 2024 by American Urological Association Education and Research, Inc.FiguresReferencesRelatedDetails Volume 211Issue 5SMay 2024Page: e1010 Advertisement Copyright & Permissions© 2024 by American Urological Association Education and Research, Inc.Metrics Author Information Jianliang Liu More articles by this author Jonathan S. O'Brien More articles by this author Nandakishor Desai More articles by this author Marimuthu Palaniswami More articles by this author Nathan Lawrentschuk More articles by this author Expand All Advertisement PDF downloadLoading ...
Self-supervised learning has achieved remarkable performance in computer vision, utilizing two key paradigms: contrastive learning and masked image modeling. Contrastive learning focuses on global representations by learning similarities and dissimilarities from different views of the inputs. On the other hand, masked image modeling learns from a pixel-level reconstruction objective and has shown improved performance compared to contrastive learning. However, masked image modeling lacks global semantics due to its pixel-level objective. To this end, we propose MOMA, a novel self-supervised distillation framework that employs a contrastive learning teacher to enhance the global representation of the masked image modeling student. Specifically, the teacher provides masks for the student, encouraging reconstructions that favor better global semantics. The feature alignment between the teacher and the student further enhances the global features in masked image modeling. Experimental results demonstrate that the proposed MOMA outperforms other masked image modeling methods and achieves competitive performance compared to other self-supervised baselines.
Background/Objective: Penile cancer is aggressive and rapidly progressive. Early recognition is paramount for overall survival. However, many men delay presentation due to a lack of awareness and social stigma. This pilot study aims to develop a convolutional neural network (CNN) model to differentiate penile cancer from precancerous and benign penile lesions. Methods: The CNN was developed using 136 penile lesion images sourced from peer-reviewed open access publications. These images included 65 penile squamous cell carcinoma (SCC), 44 precancerous lesions, and 27 benign lesions. The dataset was partitioned using a stratified split into training (64%), validation (16%), and test (20%) sets. The model was evaluated using ten trials of 10-fold internal cross-validation to ensure robust performance assessment. Results: When distinguishing between benign penile lesions and penile SCC, the CNN achieved an Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.94, with a sensitivity of 0.82, specificity of 0.87, positive predictive value of 0.95, and negative predictive value of 0.72. The CNN showed reduced discriminative capability in differentiating precancerous lesions from penile SCC, with an AUROC of 0.74, sensitivity of 0.75, specificity of 0.65, PPV of 0.45, and NPV of 0.88. Conclusion: These findings demonstrate the potential of artificial intelligence in identifying penile SCC. Limitations of this study include the small sample size and reliance on photographs from publications. Further refinement and validation of the CNN using real-life data are needed.
OBJECTIVES:To assess artificial intelligence (AI) ability to evaluate intraprostatic prostate cancer (PCa) on prostate-specific membrane antigen positron emission tomography (PSMA PET) scans prior to active treatment (radiotherapy or prostatectomy). MATERIALS AND METHODS:This systematic review was registered on the International Prospective Register of Systematic Reviews (PROSPERO identifier: CRD42023438706). A search was performed on Medline, Embase, Web of Science, and Engineering Village with the following terms: 'artificial intelligence', 'prostate cancer', and 'PSMA PET'. All articles published up to February 2024 were considered. Studies were included if patients underwent PSMA PET scan to evaluate intraprostatic lesions prior to active treatment. The two authors independently evaluated titles, abstracts, and full text. The Prediction model Risk Of Bias Assessment Tool (PROBAST) was used. RESULTS:Our search yield 948 articles, of which 14 were eligible for inclusion. Eight studies met the primary endpoint of differentiating high-grade PCa. Differentiating between International Society of Urological Pathology (ISUP) Grade Group (GG) ≥3 PCa had an accuracy between 0.671 to 0.992, sensitivity of 0.91, specificity of 0.35. Differentiating ISUP GG ≥4 PCa had an accuracy between 0.83 and 0.88, sensitivity was 0.89, specificity was 0.87. AI could identify non-PSMA-avid lesions with an accuracy of 0.87, specificity of 0.85, and specificity of 0.89. Three studies demonstrated ability of AI to detect extraprostatic extensions with an area under curve between 0.70 and 0.77. Lastly, AI can automate segmentation of intraprostatic lesion and measurement of gross tumour volume. CONCLUSION:Although the current state of AI differentiating high-grade PCa is promising, it remains experimental and not ready for routine clinical application. Benefits of using AI to assess intraprostatic lesions on PSMA PET scans include: local staging, identifying otherwise radiologically occult lesions, standardisation and expedite reporting of PSMA PET scans. Larger, prospective, multicentre studies are needed.
Standard clinical practice to assess fetal well-being during labour utilises monitoring of the fetal heart rate (FHR) using cardiotocography. However, visual evaluation of FHR signals can result in subjective interpretations leading to inter and intra-observer disagreement. Therefore, recent studies have proposed deep-learning-based methods to interpret FHR signals and detect fetal compromise. These methods have typically focused on evaluating fixed-length FHR segments at the conclusion of labour, leaving little time for clinicians to intervene. In this study, we propose a novel FHR evaluation method using an input length invariant deep learning model (FHR-LINet) to progressively evaluate FHR as labour progresses and achieve rapid detection of fetal compromise. Using our FHR-LINet model, we obtained approximately 25% reduction in the time taken to detect fetal compromise compared to the state-of-the-art multimodal convolutional neural network while achieving 27.5%, 45.0%, 56.5% and 65.0% mean true positive rate at 5%, 10%, 15% and 20% false positive rate respectively. A diagnostic system based on our approach could potentially enable earlier intervention for fetal compromise and improve clinical outcomes.
Self-supervised learning, specifically masked image modeling, has achieved significant success, surpassing earlier contrastive learning methods. However, the robustness of these methods against adversarial attacks, which subtly manipulate inputs to mislead models, remains largely unexplored. This study investigates the adversarial robustness of self-supervised learning methods, exposing their vulnerabilities to various adversarial attacks. We introduce Adversarial Masked Autoencoders (AMAE), a novel framework designed to enforce adversarial robustness during the masked image modeling process. Through extensive experiments on four classification benchmarks involving eight different adversarial attacks, we demonstrate that AMAE consistently outperforms seven state-of-the-art baseline self-supervised learning methods in terms of adversarial robustness.
Yee Wei Law合作论文数Department of Electrical and Electronic Engineering, The University of Melbourne31