Dynamic graph learning methods typically capture local structural information and short-range temporal dependencies at each time step. In this work, we introduce a dynamic graph learning architecture that generates time-step embeddings capturing both local structural context and progression-trajectory patterns for each node across an entire longitudinal sequence. The framework clusters fused embeddings that integrate (i) the global temporal trajectory of each node and (ii) its local spatial context at every graph snapshot to discover meaningful temporal patterns in longitudinal datasets. The proposed model was evaluated in the context of Parkinson’s disease (PD) progression using six years of longitudinal cerebrospinal fluid profiles from 24 patients. Visit-based graphs were constructed by representing patients as nodes enriched with peptide-abundance features, and by connecting patients with similar features profiles. A graph convolutional network captures visit-specific spatial relationships, while a sequential model learns global temporal representations. A fusion module integrates both sources of information to produce enriched node embeddings that reflect inter- and intra-patient molecular dynamics. Clustering of learned embeddings reveals four distinct stages of PD’s progression, supported by strong validity indices (Davies–Bouldin: 0.169; Calinski–Harabasz: 1264.24).Significant differences in motor severity (UPDRS 2 and UPDRS 3; p < 0.05) were observed in the groups, while the non-motor scores showed a more diffuse pattern (p = 0.11). Compared with existing features representation approaches, these findings demonstrate the potential of the proposed dynamic graph learning for data-driven disease staging and offer a generalizable framework to uncover latent temporal patterns in longitudinal datasets.
Early and accurate identification of Parkinson’s disease (PD) is essential for enabling timely treatment and effective disease management. In this study, we propose a deep learning approach to automate PD detection using convolutional neural networks trained on images derived from Pentagon copying tasks performed by patients and healthy controls. These drawings were collected using a digital pen and tablet. A total of 82 drawings (56 PD patients, 26 controls) were converted into three visual modalities: time, pressure, and pen angle relative to the X–Y axes. Given the small dataset size, we implemented several data augmentation strategies to increase training diversity and balance class distributions, including traditional geometric transformations, and synthetic augmentation using generative adversarial networks (GANs) and deep convolutional GANs (DCGANs).Among all experiments, the highest classification performance was achieved using representations derived from pen pressure, combined with traditional augmentation techniques, reaching 80.14% accuracy and a Kappa coefficient of 0.57. This modality consistently outperformed both time and angle-based representations. While GAN and DCGAN models produced visually varied images, in small clinical datasets, they required extensive training epochs to generate sufficient sample diversity, limiting their current practicality in this context. These findings demonstrate the potential of combining convolutional neural networks with drawing-based representations and augmentation methods as a non-invasive approach for PD screening in controlled settings. Importantly, strict subject-wise cross-validation was enforced to eliminate data leakage. The results support the use of pressure-based representations with traditional augmentation as a promising non-invasive approach for PD detection in binary PD-versus-healthy-control settings.
Understanding the trajectory of Huntington's disease (HD) is critical for patient stratification and the development of targeted interventions. Traditionally, studies relied on age-CAG models to estimate disease onset and progression, based on the well-established relationship between CAG repeat length and age at onset. However, additional genetic, environmental, and clinical factors can cause substantial variability. Recent machine learning approaches integrate clinical, imaging, and molecular data for more precise prediction of disease progression. Following PRISMA guidelines, we systematically reviewed studies on HD onset and progression. Using Web of Science, PubMed, and IEEE Xplore, 20 studies published between 2003 and 2024 met the inclusion criteria. We analyzed the machine learning approaches and input features used, assessed methodological quality, and evaluated risk of bias using the PROBAST tool. Overall, machine learning models, particularly support vector machines and ensemble approaches, consistently outperformed traditional age-CAG models. Several studies predicted conversion from premanifest to manifest HD within 5-10 years with high accuracy (88-98%). Beyond predicting onset, machine learning models have also been used to model dis-ease progression using clinical scores assessing motor, cognitive, and functional impairment. Performance was higher in studies incorporating structural and functional MRI biomarkers, and improved further with longitudinal clinical integration, enabling pre-diction of decline years before symptoms onset. Overall, machine learning shows strong potential to improve prognostic modeling in HD, especially through multimodal and longitudinal data. However, common methodological weaknesses and bias highlight the need for larger, externally validated studies using objective biomarkers.
Abstract Huntington’s disease (HD) presents a heterogeneous neurodegenerative course, with motor, cognitive, and functional symptoms progressing differently across individuals. This atypical progression complicates the definition of discrete disease stages, hindering understanding of disease trajectories, timely patient care, and therapy development. Consequently, current clinical staging systems rely heavily on clinician-defined, domain-specific criteria and fixed clinical measurement boundaries for stage assignment, reducing objectivity and often leading to overlapping clinical measurements across stages. While machine learning methods can help, existing approaches cannot fully capture complex temporal relationships within and across patients. We propose URL-STFN, a dynamic graph-based representation learning model that encodes both inter- and intra-patient temporal patterns from longitudinal clinical measures. We then evaluate disease stages formed through clustering and stability analysis of URL-STFN latent representations, and compare them with representations obtained from conventional embedding approaches. We further benchmark these clustering-based stages against states derived from conventional temporal models, including DHMM. We hypothesize that clustering URL-STFN latent representations enables identification of HD stages with reduced overlap in clinical measurements. The proposed framework is evaluated using 1,477 clinical visits from the Enroll-HD dataset, a large longitudinal cohort with repeated clinical assessments. For staging, we used 44 clinical measurements spanning motor, cognitive, and functional domains. URL-STFN identifies clinically meaningful HD stages consistent with established disease progression while reducing overlap in clinical feature values compared with DHMM-derived and clinical staging approaches. These findings highlight the potential of a dynamic graph-based representation learning and clustering framework to support more objective, data-driven, and precise HD staging. Highlights A novel dynamic graph captures inter- and intra-patient HD progression latent. Clustering captured patterns identified clinically meaningful HD disease stages. Discovered stages showed reduced clinical feature overlap across clusters. Dynamic graph-based staging outperformed conventional embeddings and DHMM staging. Clustering multidomain latent features revealed new clusters beyond Shoulson–Fahn.
Huntington's disease (HD) is a progressive brain disorder that gradually affects movement, cognitive function, and behavior. Identifying the stage of the disease accurately and consistently is important for understanding its course, grouping patients, personalized care, and discovering treatment. Existing clinical staging frameworks rely primarily on predefined clinical measurement thresholds and clinical expert decisions, yet these discrete cut-offs may obscure meaningful intra-stage variability and remain vulnerable to inter-rater differences, especially in motor and functional assessments. To address these limitations, we developed an unsupervised machine learning framework based on dynamic graph representation learning to capture temporal relationships within and across patients from longitudinal clinical measurements. Using the learned representations, we applied K-means++ clustering to identify well-separated groups. We then iteratively increased the number of clusters (k), using stability analysis to assess robustness and reveal additional meaningful clusters beyond the initial optimal solution. We applied the framework to 302 individuals from the Enroll-HD cohort (1,477 visits, 44 clinical variables per visit; 80
The use and integration of Artificial Intelligence (AI) and machine learning technologies into biosystems engineering create unprecedented opportunities for modelling, optimisation, and decision support across agriculture, livestock, food systems, environmental management, and related domains. However, the increasing complexity and often opacity of these methods is raising concerns regarding scientific rigour, reproducibility, transparency, generalisation and ethical responsibility. This letter establishes a set of good practice principles for authors submitting AI-driven research to Biosystems Engineering journal. The guidelines outline essential requirements for data quality, documentation of methodologies, experimental protocols, model selection, evaluation, interpretability, and reproduction of results. They emphasise the importance of open datasets and code availability, appropriate validation strategies, meaningful novelty, and clear evidence of relevance to the Biosystems Engineering scope.
Amyotrophic Lateral Sclerosis (ALS) is a degenerative disease of motor neurons that leads to muscle wasting, paralysis, and death, with an average life expectancy of 2-5 years. Approximately 10-15% of ALS cases are familial (fALS), typically linked to, but not always caused by identifiable inherited genetic mutations. The remaining 85-90% are considered sporadic ALS (sALS), which typically occurs without a clear family history. It is thought to result from a combination of genetic and non-genetic risk factors. ALS imposes heavy physical, psychological, and financial burdens on patients and caregivers. Early diagnosis is critical but remains challenging due to clinical variability and overlapping symptoms with other motor neuron disorders. Current diagnostic methods, including genetic testing and neurophysiological techniques, face limitations in reproducibility and accessibility, while machine learning offers potential by detecting patterns that traditional methods overlook. This study applies machine learning to characterise disease status in C9orf72-ALS patients and evaluate how pathological biomarkers relate to disease mechanisms. A tabular dataset from post-mortem brain tissue of 10 C9orf72-ALS patients and 10 controls was used to train models and benchmark results against Rifai et al. (2022). Models included random forest, support vector machine, xgboost, logistic regression, artificial neural networks, and ensembles, validated using 3-fold and 5-group cross-validation. The best model result was of random forest with 3-fold cross-validation for Iba1, achieving 88% sensitivity (p = 0.0011) and 83% specificity (p = 0.0004). However, as 3-fold cross-validation is less robust, we expect more reliable and stable results from 5-fold grouped cross-validation and, in future, repeated cross validation approaches. Machine learning offers insights into ALS, with implications for potential patient stratification and case identification. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study did not receive any funding. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study used ONLY anonymised post-mortem brain tissue data obtained from openly available source located at: https://figshare.com/projects/Random\_forest\_modelling\_and\_neuropathological\_features\_of\_a\_large\_cohort\_of\_C9orf72-ALS\_all\_raw\_data_/128222. Data Ethics Ethical approvals for data and tissue collection were obtained from the East of Scotland Research Ethics Service (16/ES/0084), in accordance with the Human Tissue (Scotland) Act (2006). The use of post-mortem samples was reviewed and approved by the Edinburgh Brain Bank Ethics Committee and the Academic and Clinical Central Office for Research and Development (AMREC). Clinical data were collected through the Scottish Motor Neurone Disease Register (SMNDR) and the Care Audit Research and Evaluation for Motor Neurone Disease (CARE-MND) platform, with ethical approval from Scotland A Research Ethics Committee (10/MRE00/78 and 15/SS/0216). All patients provided informed consent for the use of their data for research while they were alive, and clinical and cognitive data such as ECAS and ALSFRS were collected. The owners of the data granted permission to use this dataset for secondary processing. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes This study is a secondary analysis of previously published, ethically approved post-mortem brain tissue data obtained from publicly available figshare repositories linked in the original publication (cite: Rifai, O.M., Longden, J., O'Shaughnessy, J., Sewell, M.D., Pate, J., McDade, K., Daniels, M.J., Abrahams, S., Chandran, S., McColl, B.W. and Sibley, C.R., 2022. Random forest modelling demonstrates microglial and protein misfolding features to be key phenotypic markers in C9orf72‐ALS. The Journal of pathology, 258(4), pp.366-381.). The SD numbers of cases from the Edinburgh Brain Bank included in the study are available upon request. [https://figshare.com/projects/Random\_forest\_modelling\_and\_neuropathological\_features\_of\_a\_large\_cohort\_of\_C9orf72-ALS\_all\_raw\_data_/128222][1] [https://figshare.com/articles/dataset/Random\_forest\_results/17145905][2] [1]: https://figshare.com/projects/Random_forest_modelling_and_neuropathological_features_of_a_large_cohort_of_C9orf72-ALS_all_raw_data_/128222 [2]: https://figshare.com/articles/dataset/Random_forest_results/17145905
Amyotrophic lateral sclerosis (ALS) is a fatal neurological disease marked by motor deterioration and cognitive decline. Early diagnosis is challenging due to the complexity of sporadic ALS and the lack of a defined risk population. In this study, we developed Miniset-DenseSENet, a convolutional neural network combining DenseNet121 with a Squeeze-and-Excitation attention mechanism, using 190 autopsy brain images from the Gregory Laboratory at the University of Aberdeen. The model distinguishes controls, ALS patients with no cognitive impairment, and ALS patients with cognitive impairment (ALS-frontotemporal dementia) with 97.37% accuracy, addressing a significant challenge in overlapping neurodegenerative disorders involving TDP-43 proteinopathy. Miniset-DenseSENet outperformed other transfer learning models, achieving a sensitivity of 1 and specificity of 0.95. These findings suggest that integrating transfer learning and attention mechanisms into neuroimaging can enhance diagnostic accuracy, enabling earlier ALS detection and improving patient stratification. This model has the potential to guide clinical decisions and support personalied therapeutic strategies.
This study explores the potential of advanced, context-aware machine learning algorithms, such as autoencoders, to represent longitudinal cerebrospinal fluid proteomic data, enabling the objective discovery of two patient strata with significance.
Label-free autofluorescence lifetime is a unique feature of the inherent fluorescence signals emitted by natural fluorophores in biological samples. Fluorescence lifetime imaging microscopy (FLIM) can capture these signals enabling comprehensive analyses of biological samples. Despite the fundamental importance and wide application of FLIM in biomedical and clinical sciences, existing methods for analysing FLIM images often struggle to provide rapid and precise interpretations without reliable references, such as histology images, which are usually unavailable alongside FLIM images. To address this issue, we propose a deep learning (DL)-based approach for generating virtual Hematoxylin and Eosin (H&E) staining. By combining an advanced DL model with a contemporary image quality metric, we can generate clinical-grade virtual H&E-stained images from label-free FLIM images acquired on unstained tissue samples. Our experiments also show that the inclusion of lifetime information, an extra dimension beyond intensity, results in more accurate reconstructions of virtual staining when compared to using intensity-only images. This advancement allows for the instant and accurate interpretation of FLIM images at the cellular level without the complexities associated with co-registering FLIM and histology images. Consequently, we are able to identify distinct lifetime signatures of seven different cell types commonly found in the tumour microenvironment, opening up new opportunities towards biomarker-free tissue histology using FLIM across multiple cancer types.
It is challenging to train a generalizable deep learning classifier with limited training images. Existing few-shot learning approaches try to improve classification performance largely by transferring prior knowledge from upstream large-sample tasks to the current small-sample task. Besides upstream image datasets, prior knowledge may also be obtained from signals of other modalities. In this study, we propose a novel learning framework that can utilize prior knowledge from audio signals to help train an image classifier. In the framework, a pre-trained and fixed audio encoder can transform the audio signal of each class label into a class-specific audio prototype. By attracting image representations to the corresponding audio prototypes during training of the image classifier, within-class image representations become more clustered, while image representations become further apart if they are from different classes. To the best of our knowledge, this is the first work that utilizes audio-based prior knowledge to help train an image classifier with limited training images. The proposed learning framework is compatible with existing learning approaches, making it flexible enough to be combined with existing approaches. Extensive empirical evaluations on both natural and medical image datasets demonstrate that the proposed learning framework significantly outperforms existing methods in image classification with limited training images, thus establishing a new state of the art. The source code will be released publicly.
Many medical imaging modalities have benefited from recent advances in Machine Learning (ML), specifically in deep learning, such as neural networks. Computers can be trained to investigate and enhance medical imaging methods without using valuable human resources. In recent years, Fluorescence Lifetime Imaging (FLIm) has received increasing attention from the ML community. FLIm goes beyond conventional spectral imaging, providing additional lifetime information, and could lead to optical histopathology supporting real-time diagnostics. However, most current studies do not use the full potential of machine/deep learning models. As a developing image modality, FLIm data are not easily obtainable, which, coupled with an absence of standardisation, is pushing back the research to develop models which could advance automated diagnosis and help promote FLIm. In this paper, we describe recent developments that improve FLIm image quality, specifically time-domain systems, and we summarise sensing, signal-to-noise analysis and the advances in registration and low-level tracking. We review the two main applications of ML for FLIm: lifetime estimation and image analysis through classification and segmentation. We suggest a course of action to improve the quality of ML studies applied to FLIm. Our final goal is to promote FLIm and attract more ML practitioners to explore the potential of lifetime imaging.
Diagnosing Amyotrophic Lateral Sclerosis (ALS) remains a hand challenge due to its inherent heterogeneity. Notably, the occurrence of TDP-43 cytoplasmic aggregation in approximately 95% of ALS cases has emerged as a potential indicative hallmark. In order to develop deep learning models capable of distinguishing TDP-43 proteinopathic samples from their healthy counterparts, a comprehensive understanding of the sample set becomes imperative, particularly when the sample size is limited. The samples in question encompassed images obtained via an immunofluorescence procedure, employing super high-resolution microscopy coupled with meticulous processing. A feature-extracted dataset was created to collect meaningful features from every sample to approach three different classification problems (TDP-43 Pathology, TDP-43 Pathology Grades and ALS) based on the number of red and pink pixels, signifying cytoplasmic and nuclear TDP-43 presence. A series of diverse statistical approaches were undertaken. However, definitive outcomes remained elusive, although it was suggested that a classification based on the presence of TDP-43 proteinopathy was better than the one based on the presence of ALS for training the model. The dataset was reduced by eliminating the problematic samples through curation. Analyses were repeated using t-student tests and ANOVA, and visualisation of patient inter-variability was performed using hierarchical clustering. The TDP-43 pathology classification results showed significant differences in the number of red and pink pixels, the total amount of protein and the cytoplasmic and nuclear proportions between healthy and pathological samples between groups. These findings suggested that images classified according to the presence of TDP-43 proteinopathy are more suitable for training deep learning models.
Tremor is an involuntary, rhythmic, oscillatory movement of a body part. It is a common symptom in neurological disorders and significantly impacts daily life. Traditional clinical diagnosis of tremor includes a physical examination of the patient and a review of their medical history. It often fails to assess tremor accurately as the prevalence of tremor changes throughout the day. Therefore, continuous free-living monitoring of tremor symptoms with wearable technology can provide more comprehensive, objective and ecologically valid information compared to traditional clinical diagnoses. In this paper, a miniaturised flexible wearable patch-type device is presented to continuously assess tremor with an accelerometer and a gyroscope, whose recorded data is sent wirelessly to a mobile phone. The proposed patch is 30 mm x 20mm x 8mm in size and only 2 g, which makes it convenient for long periods of monitoring with wearables. Its miniaturised size facilitates an ergonomical attachment to any part of the body while offering an impressive measurement range, i.e., acceleration reaches up to +/- 8 g and angular velocity measures up to +/- 2000 degrees per second (dps). This study evaluates the efficacy of three Machine Learning (ML) models, Linear Discriminant Analysis (LDA), Logistic Regression, and AdaBoost, in classifying tremor patterns using data collected from six healthy volunteers and simulated with tremor occurrences by the proposed device. Preliminary results indicate LDA achieved the highest accuracy (68.30%) and F1-Score (66.46%), suggesting its potential effectiveness in tremor pattern recognition, though limited by the small sample size and simulation method.
Label-free virtual Haematoxylin and Eosin (H&E) staining has the potential to generate realistic histological images for rapid clinical diagnosis, eliminating the need for the conventional, costly, and time-consuming tissue staining procedure. Although various deep learning techniques have proven effective for this purpose, there has been limited attention given to how different loss functions influence the fidelity of synthesis. In this study, we focus on assessing the qualitative and quantitative impact of four widely applied loss functions designed for high-fidelity image synthesis in the context of virtual H&E staining using single-channel label-free autofluorescence images. Qualitative analysis involves a visual comparison between true and synthetic images, while quantitative analysis utilises several well-known image similarity metrics to measure the distance between real and virtual images. Our experimental results demonstrate the feasibility of extra regularisation terms with different weights on synthesis H&E images. Both visual inspection and quantitative outcomes align well with each other, but both should be facilitated to reach a conclusive decision for the virtual staining with optimal quality.
Autofluorescence lifetime images reveal unique characteristics of endogenous fluorescence in biological samples. Comprehensive understanding and clinical diagnosis rely on co-registration with the gold standard, histology images, which is extremely challenging due to the difference of both images. Here, we show an unsupervised image-to-image translation network that significantly improves the success of the co-registration using a conventional optimisation-based regression network, applicable to autofluorescence lifetime images at different emission wavelengths. A preliminary blind comparison by experienced researchers shows the superiority of our method on co-registration. The results also indicate that the approach is applicable to various image formats, like fluorescence intensity images. With the registration, stitching outcomes illustrate the distinct differences of the spectral lifetime across an unstained tissue, enabling macro-level rapid visual identification of lung cancer and cellular-level characterisation of cell variants and common types. The approach could be effortlessly extended to lifetime images beyond this range and other staining technologies.
In this paper, we introduce our unique dataset of fluorescence lifetime imaging endo/microscopy (FLIM), containing over 100,000 different FLIM images collected from 18 pairs of cancer/non-cancer human lung tissues of 18 patients by our custom fibre-based FLIM system. The aim of providing this dataset is that more researchers from relevant fields can push forward this particular area of research. Afterwards, we describe the best practice of image post-processing suitable per the dataset. In addition, we propose a novel hierarchically aggregated multi-scale architecture to improve the binary classification performance of classic CNNs. The proposed model integrates the advantages of multi-scale feature extraction at different levels, where layer-wise global information is aggregated with branch-wise local information. We integrate the proposal, namely ResNetZ, into ResNet, and appraise it on the FLIM dataset. Since ResNetZ can be configured with a shortcut connection and the aggregations by Addition or Concatenation, we first evaluate the impact of different configurations on the performance. We thoroughly examine various ResNetZ variants to demonstrate the superiority. We also compare our model with a feature-level multi-scale model to illustrate the advantages and disadvantages of multi-scale architectures at different levels.
Multi-scale architectures at a granular level are characterised by separating input features into groups and applying multi-scale feature extractions to the split input features, and thus the correlations among the input features as global information are no longer retained. Moreover, they usually require more input features due to the separation, and therefore, more complexity is introduced. To retain the global information while utilising the advantages of feature-level hierarchical multi-scale architectures, we propose a multi-scale aggregated-dilation architecture (MSAD) to perform hierarchical fusion of features at a layer level, with the integration of dilated convolutions to overcome these issues. To evaluate the model, we integrate it into ResNet, and apply it to a unique dataset, containing over 60,000 fluorescence lifetime endomicroscopic images (FLIM) collected on ex-vivo lung normal/cancerous tissues from 14 patients, by a custom fibre-based FLIM system. To evaluate the performance of our proposal, we use accuracy, precision, recall, and AUC. We first compare our MSAD model with eight networks achieving a superiority over 6%. To illustrate the advantages and disadvantages of multi-scale architectures at layer and feature-level, we thoroughly compare our MSAD model with the state-of-the-art feature-level multiscale network, namely Res2Net, in terms of parameters, scales, and effective convolutions.