ABSTRACT Introduction The population of older adults in India is projected to exceed 230 million within the next decade. With increasing life expectancy, maintaining cognitive and neurobiological health becomes essential for ensuring a good quality of life. This study is a comprehensive multimodal investigation of aging and cognition in a healthy Indian cohort. The aim is to characterize cognitive, neurobiological, physiological, psychological, and molecular trajectories and their associations across the adult lifespan in India. Methods and Analysis Data will be collected cross-sectionally from healthy young (19–35 years, n = 100), middle-aged (40–55 years, n = 100), and older adults (60–75 years, n = 100) with regional nativity and residence in Telangana or Andhra Pradesh. Individuals with cognitive impairment or any chronic health condition will be excluded. Assessments will include standardized cognitive tests, mental health questionnaires, brain imaging, heart rate variability, and serum samples for standard blood tests and epigenetic analyses. Hair cortisol and serum inflammatory markers will also be collected to profile stress and systemic inflammation, respectively. Data from this study will help to profile healthy aging in an Indian cohort, thereby addressing critical gaps in clinical research and health policy. Ethics and Dissemination This study was approved by the Institutional Review Board at IIIT Hyderabad (IIITH-IRB-PRO-2023-11). All participants will give written consent. Results from this study will be presented at conferences and published in peer-reviewed journals. Data will also be curated for an interactive web-portal to enable researchers, clinicians, and policy makers to understand the data in meaningful ways. Strengths and Limitations Most studies in aging focus on risk factors and progression of degenerative conditions. What healthy aging really means is not well understood. This study will recruit healthy Indian adults free from any chronic health problem and with regional nativity and residence to minimize geo-climatic effects and genetic heterogeneity. Hair cortisol and serum-inflammation levels may provide unique insights in healthy aging. The comprehensive data from this study will provide a profile of healthy aging in an Indian cohort that is highly vulnerable to metabolic disorders. It will enable clinicians, researchers, and policy makers to plan and promote healthy aging. The restricted native sample population and sample size may not capture the divergent trajectories in aging across other demographic populations within and outside India. Though cross-sectional data will not capture the changes that longitudinal data can, it will offer an opportunity to compare psychosocial and physiological measures across different age groups.
Chest X-ray (CXR) segmentation is an important step in computer-aided diagnosis, yet deploying large foundation models in clinical settings remains challenging due to computational constraints. We propose AdaLoRA-QAT, a two-stage fine-tuning framework that combines adaptive low-rank encoder adaptation with full quantization-aware training. Adaptive rank allocation improves parameter efficiency, while selective mixed-precision INT8 quantization preserves structural fidelity crucial for clinical reliability. Evaluated across large-scale CXR datasets, AdaLoRA-QAT achieves 95.6
Understanding how humans and artificial intelligence systems process complex narrative videos is a fundamental challenge at the intersection of neuroscience and machine learning. This study investigates how the temporal context length of video clips (3–12 s clips) and the narrative-task prompting shape brain-model alignment during naturalistic movie watching. Using fMRI recordings from participants viewing full-length movies, we examine how brain regions sensitive to narrative context dynamically represent information over varying timescales and how these neural patterns align with model-derived features. We find that increasing clip duration substantially improves brain alignment for multimodal large language models (MLLMs), whereas unimodal video models show little to no gain. Further, shorter temporal windows align with perceptual and early language regions, while longer windows preferentially align higher-order integrative regions, mirrored by a layer-to-cortex hierarchy in MLLMs. Finally, narrative-task prompts (multi-scene summary, narrative summary, character motivation, and event boundary detection) elicit task-specific, region-dependent brain alignment patterns and context-dependent shifts in clip-level tuning in higher-order regions. Together, our results position long-form narrative movies as a principled testbed for probing biologically relevant temporal integration and interpretable representations in long-context MLLMs.
Recent work has shown that scaling large language models (LLMs) improves their alignment with human brain activity, yet it remains unclear what drives these gains or which representational properties are responsible. Although larger models often yield better task performance and brain alignment, they are increasingly difficult to analyze mechanistically. This raises a fundamental question: \emph{what is the minimal model capacity required to capture brain-relevant representations?} To address this question, we systematically investigate how constraining model scale and numerical precision affects brain alignment. We compare full-precision LLMs, small language models (SLMs), and compressed variants (quantized and pruned) by predicting fMRI responses during naturalistic language comprehension. Across model families up to 14B parameters, we find that 3B SLMs achieve brain predictivity indistinguishable from larger LLMs, whereas 1B models degrade substantially, particularly in semantic language regions. Brain alignment is remarkably robust to compression: most quantization and pruning methods preserve neural predictivity, with GPTQ as a consistent exception. Linguistic probing reveals a dissociation between task performance and brain predictivity: compression degrades discourse, syntax, and morphology, yet brain predictivity remains largely unchanged. Overall, brain alignment saturates at modest model scales and is resilient to compression, challenging common assumptions about neural scaling and motivating compact models for brain-aligned language modeling.
Sleep architecture and integrity significantly influence neural recovery and cognitive restoration. These are particularly relevant in ischemic stroke survivors where sleep-disordered breathing (SDB) is a common comorbidity. To address the lack of stroke-specific sleep data, we present the Polysomnography Dataset for Sleep Analysis in Ischemic Stroke Patients (iSLEEPS), the first Asian and one of the largest stroke-specific sleep databases. Data collection was carried out between September-2018 and December-2021 at NIMHANS, India. iSLEEPS comprises 100 overnight PSG recordings with comprehensive expert annotations. Each recording includes sleep stages manually scored at 30-second epochs, detailed respiratory events, periodic limb movements, oxygen desaturation episodes, and clinical metrics, as per AASM (2017) guidelines. Our cohort demonstrates a high prevalence of SDB, enabling the investigation of stroke-sleep pathophysiology interactions. To illustrate dataset utility, we implemented automated sleep stage classification using deep learning methods. The Long Short-Term Memory model achieved the highest accuracy (74.70%), followed by Transformer (67.44%) and Convolutional Neural Network (61.65%). This dataset addresses crucial gap in stroke sleep research, supporting comprehensive analysis of post-stroke sleep disturbances.
Accurate sleep staging is essential for diagnosing OSA and hypopnea in stroke patients. Although PSG is reliable, it is costly, labor-intensive, and manually scored. While deep learning enables automated EEG-based sleep staging in healthy subjects, our analysis shows poor generalization to clinical populations with disrupted sleep. Using Grad-CAM interpretations, we systematically demonstrate this limitation. We introduce iSLEEPS, a newly clinically annotated ischemic stroke dataset (to be publicly released), and evaluate a SE-ResNet plus bidirectional LSTM model for single-channel EEG sleep staging. As expected, cross-domain performance between healthy and diseased subjects is poor. Attention visualizations, supported by clinical expert feedback, show the model focuses on physiologically uninformative EEG regions in patient data. Statistical and computational analyses further confirm significant sleep architecture differences between healthy and ischemic stroke cohorts, highlighting the need for subject-aware or disease-specific models with clinical validation before deployment. A summary of the paper and the code is available at https://himalayansaswatabose.github.io/iSLEEPS_Explainability.github.io/
Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. Most brain-encoding studies focus on aligning artificial models with brain activity during language comprehension or passive visual processing, while interactive brain-alignment studies have to date been largely limited to reinforcement-learning (RL) agents and theory-based models. To address this gap, we study brain alignment of representative models from two foundation-model families, namely vision-language models (VLMs) and large-action models (LAMs), using fMRI recordings from participants playing naturalistic Atari-style video games. Specifically, we examine how action-focused and reasoning-focused prompts shape model's internal representations and align with fMRI brain activity. First, we find that both VLMs and LAMs exhibit significantly exhibit voxel-wise encoding performance than RL baselines, with the advantage holding even under matched feature dimensionality. Second, prompt-driven gains scale with the cortical processing hierarchy: the largest improvements appear in frontal-parietal and motor-planning regions, while early visual cortex gains roughly half as much. Third, variance partitioning reveals a qualitatively different representational organization: VLM is prompt-symmetric (12.5
Accurate segmentation of medical images is challenging due to unclear lesion boundaries and mask variability. We introduce Segmentation Schödinger Bridge (SSB), the first application of Schödinger Bridge for ambiguous medical image segmentation, modelling joint image-mask dynamics to enhance performance. SSB preserves structural integrity, delineates unclear boundaries without additional guidance, and maintains diversity using a novel loss function. We further propose the Diversity Divergence Index (D_DDI) to quantify inter-rater variability, capturing both diversity and consensus. SSB achieves state-of-the-art performance on LIDC-IDRI, COCA, and RACER (in-house) datasets.
Recent voxel-wise multimodal brain encoding studies have shown that multimodal large language models (MLLMs) exhibit a higher degree of brain alignment compared to unimodal models. More recently, instruction-tuned multimodal (IT) models have been shown to generate task-specific representations that align strongly with brain activity, yet most prior evaluations focus on unimodal stimuli or non-instruction-tuned models under multimodal stimuli. We still lack a clear understanding of whether instruction-tuning is associated with IT-MLLMs organizing their representations around functional task demands or if they simply reflect surface semantics. To address this, we estimate brain alignment by predicting fMRI responses recorded during naturalistic movie watching (video with audio) from MLLM representations. Using instruction-specific embeddings from six video and two audio IT-MLLMs, across 13 video task instructions, we find that instruction-tuned video MLLMs show higher brain alignment than in-context learning (ICL) multimodal models ( 9
The intricate link between brain functional connectivity (FC) and structural connectivity (SC) is explored through models performing diffusion on SC to derive FC, using varied methodologies from single to multiple graph diffusion kernels. However, existing studies have not correlated diffusion scales with specific brain regions of interest (RoIs), limiting the applicability of graph diffusion. We propose a novel approach using graph diffusion wavelets to learn the appropriate diffusion scale for each RoI to accurately estimate the SC-FC mapping. Using the open Human Connectome Project dataset, we achieve an average Pearson's correlation value of 0.833, surpassing the state-of-the-art methods for the prediction of FC. It is important to note that the proposed architecture is entirely linear, computationally efficient, and notably demonstrates the power-law distribution of diffusion scales. Our results show that the bilateral frontal pole, by virtue of it having large diffusion scale, forms a large community structure. The finding is in line with the current literature on the role of the frontal pole in resting-state networks. Overall, the results underscore the potential of graph diffusion wavelet framework for understanding how the brain structure leads to FC.
The phenomenon of intentional binding pertains to the perceived connection between a voluntary action and its anticipated result. When an individual intends an outcome, it appears to subjectively extend in time due to a pre-activation of the intended result, particularly evident at shorter action-outcome delays. However, there is a concern that the operationalisation of intention might have led to a mixed interpretation of the outcome expansion attributed to the pre-activation of intention, given the sensitivity of time perception and intentional binding to external cues that could accelerate the realisation of expectations. To investigate the expansion dynamics of an intended outcome, we employed a modified version of the temporal bisection task in two experiments. Experiment 1 considered the action-outcome delay as a within-subject factor, while experiment 2 treated it as a between-subject factor. The results revealed that the temporal expansion of an intended outcome was only evident under the longer action-outcome delay condition. We attribute this observation to working memory demands and attentional allocation due to temporal relevancy and not due to pre-activation. The discrepancy in effects across studies is explained by operationalising different components of the intentional binding effect, guided by the cue integration theory. Moreover, we discussed speculative ideas regarding the involvement of specific intentions based on the proximal intent distal intent (PIDI) theory and whether causality plays a role in temporal binding. Our study contributes to the understanding of how intention influences time perception and sheds light on how various methodological factors, cues, and delays can impact the dynamics of temporal expansion associated with an intended outcome.
Brain stroke has become a significant burden on global health and thus we need remedies and prevention strategies to overcome this challenge. For this, the immediate identification of stroke and risk stratification is the primary task for clinicians. To aid expert clinicians, automated segmentation models are crucial. In this work, we consider the publicly available dataset ATLAS v2.0 to benchmark various end-to-end supervised U-Net style models. Specifically, we have benchmarked models on both 2D and 3D brain images and evaluated them using standard metrics. We have achieved the highest Dice score of 0.583 on the 2D transformer-based model and 0.504 on the 3D residual U-Net respectively. We have conducted the Wilcoxon test for 3D models to correlate the relationship between predicted and actual stroke volume. For reproducibility, the code and model weights are made publicly available: https://github.com/prantik-pdeb/BeSt-LeS.
Autism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by varied social cognitive challenges and repetitive behavioral patterns. Identifying reliable brain imaging-based biomarkers for ASD has been a persistent challenge due to the spectrum’s diverse symptomatology. Existing baselines in the field have made significant strides in this direction, yet there remains room for improvement in both performance and interpretability. We propose HyperGALE, which builds upon the hypergraph by incorporating learned hyperedges and gated attention mechanisms. This approach has led to substantial improvements in the model’s ability to interpret complex brain graph data, offering deeper insights into ASD biomarker characterization. Evaluated on the extensive ABIDE II dataset, HyperGALE not only improves interpretability but also demonstrates statistically significant enhancements in key performance metrics compared to both previous baselines and the foundational hypergraph model. The advancement HyperGALE brings to ASD research highlights the potential of sophisticated graph-based techniques in neurodevelopmental studies. The source code and implementation instructions are available at GitHub.
Categorical data classification and clustering are essential to many fields, including pattern recognition, data mining, knowledge discovery, and machine learning. It is crucial to understand how to provide categorical input to machine learning algorithms because the majority of them are developed for numerical inputs. Any machine learning algorithm’s performance depends not only on the model and hyper parameters but also on the way the data is prepared for and fed into the model. With high cardinality features and categorical predictor variables, machine learning approaches that handle categorical data frequently encounter challenges. In this scenario, we investigate categorical encoding techniques, develop a taxonomy of encoding techniques, and use the machine learning algorithms KNN, MLP, SVM, and ELM and make a comparative study. We suggest a method of hybrid encoding using CNN that gives better performance results. In this paper, we propose the machine learning strategy employing multi-dimensional encodings with CNN, which outperforms other machine learning approaches on majority datasets. It is clear that CNN’s performance in learning categorical features has been much enhanced by the use of multi-dimensional encoded input vectors. In our suggested method for supervised and unsupervised encoding, we have used four distinct encodings. These results encourage us to further study the outcomes of different hybrid combinations of encodings for categorical feature classification.
We introduce an innovative approach to automated sleep stage classification using EOG signals, addressing the discomfort and impracticality associated with EEG data acquisition. In addition, it is important to note that this approach is untapped in the field, highlighting its potential for novel insights and contributions. Our proposed SE-Resnet-Transformer model provides an accurate classification of five distinct sleep stages from raw EOG signal. Extensive validation on publically available databases (SleepEDF-20, SleepEDF-78, and SHHS) reveals noteworthy performance, with macro-F1 scores of 74.72, 70.63, and 69.26, respectively. Our model excels in identifying REM sleep, a crucial aspect of sleep disorder investigations. We also provide insight into the internal mechanisms of our model using techniques such as 1D-GradCAM and t-SNE plots. Our method improves the accessibility of sleep stage classification while decreasing the need for EEG modalities. This development will have promising implications for healthcare and the incorporation of wearable technology into sleep studies, thereby advancing the field's potential for enhanced diagnostics and patient comfort.
The coronavirus disease (COVID-19) caused over 170 million illnesses and over 3 million deaths worldwide. Researchers from all over the world have been using a variety of machine-learning techniques to classify the DNA sequences of the SARS-CoV-2 virus. To classify SARS-CoV-2 (Covid-19) whole genome sequences with other viruses including Dengue, Ebola, Influenza, and also Coronaviruses that may infect humans and are members of the family of Coronaviridae like SARS-CoV, MERS-CoV, we have proposed a multi-dimensional frequency encoding scheme. In the proposed method each genome is converted into four frequency-encoded sequences for each nucleotide. Then each encoded sequence is partitioned into quartiles and fed to the CNN. The findings of the binary classification of SARS-CoV-2 with other viruses and multi-class classification of six different viruses using five well-known machine learning algorithms and the proposed approach using CNN were reported in this study. Furthermore, we applied machine learning methods for the multi-class classification of eight different SARS-CoV-2 variants. The proposed approach performed well among all machine-learning techniques for classifying genome sequences of viruses and SARS-CoV-2 variants with accuracy of 99% and 98% respectively. In addition, we have also compared the performance of our proposed approach with the existing methods from the literature.
Over the last decade, there has been growing interest in learning the mapping from structural connectivity (SC) to functional connectivity (FC) of the brain. The spontaneous fluctuations of the brain activity during the resting-state as captured by functional MRI (rsfMRI) contain rich non-stationary dynamics over a relatively fixed structural connectome. Among the modeling approaches, graph diffusion-based methods with single and multiple diffusion kernels approximating static or dynamic functional connectivity have shown promise in predicting the FC given the SC. However, these methods are computationally expensive, not scalable, and fail to capture the complex dynamics underlying the whole process. Recently, deep learning methods such as GraphHeat networks and graph diffusion have been shown to handle complex relational structures while preserving global information. In this paper, we propose a novel attention-based fusion of multiple GraphHeat networks (A-GHN) for mapping SC-FC. A-GHN enables us to model multiple heat kernel diffusion over the brain graph for approximating the complex Reaction Diffusion phenomenon. We argue that the proposed deep learning method overcomes the scalability and computational inefficiency issues but can still learn the SC-FC mapping successfully. Training and testing were done using the rsfMRI data of 1058 participants from the human connectome project (HCP), and the results establish the viability of the proposed model. On HCP data, we achieve a high Pearson correlation of 0.788 (Desikan-Killiany atlas with 87 regions) and 0.773 (AAL atlas with 86 regions). Furthermore, experiments demonstrate that A-GHN outperforms the existing methods in learning the complex nature of the structure-function relation of the human brain.
Automated Sleep stage classification using raw single channel EEG is a critical tool for sleep quality assessment and disorder diagnosis. However, modelling the complexity and variability inherent in this signal is a challenging task, limiting their practicality and effectiveness in clinical settings. To mitigate these challenges, this study presents an end-to-end deep learning (DL) model which integrates squeeze and excitation blocks within the residual network to extract features and stacked Bi-LSTM to understand complex temporal dependencies. A distinctive aspect of this study is the adaptation of GradCam for sleep staging, marking the first instance of an explainable DL model in this domain with alignment of its decision-making with sleep expert's insights. We evaluated our model on the publically available datasets (SleepEDF-20, SleepEDF-78, and SHHS), achieving Macro-F1 scores of 82.5, 78.9, and 81.9, respectively. Additionally, a novel training efficiency enhancement strategy was implemented by increasing stride size, leading to 8x faster training times with minimal impact on performance. Comparative analyses underscore our model outperforms all existing baselines, indicating its potential for clinical usage.
One of the challenges in stroke rehabilitation is to identify bio-markers that correlate with amelioratory changes in the recovery of brain function that can be verified by a clinician. For strokes related to upper extremity, clinicians use Fugl Meyer Assessment - Upper Extremity (FMA-UE) score for verification. We hypothesize that even before clinical measures of function recovery (FMA-UE) show facilitatory changes, structural and functional changes in the brain connectome indicate early changes in stroke patients. Toward establishing such early biomarkers of rehabilitation related changes in the brain plasticity, we propose to use graph theoretic measures on the structural (SC) and functional connectivity (FC) matrices. We used longitudinal multi-modal neuroimaging data acquired within 1 to 6 months of the onset of stroke and after 3 months of rehabilitation for 15 acute ischemic stroke subjects with deficit in motor function of the upper extremity. We compared structural and functional network properties of the brain between the baseline (pre-) and post-rehabilitation stages. While significant changes are observed in the nodal properties, global network properties did not reach statistical significance. However, when correlation patterns across pre-, post-rehabilitation, and healthy controls are investigated, some interesting patterns emerge that are not captured by statistical analysis as well as the clinical assessment scores. Global functional network properties of patients postrehabilitation resemble that of a healthy control group, than that of pre-rehabilitation (Baseline). We also observed that greater the change in normalized percent change in FMA-UE scores, closer are the correlation maps of global graph metrics of FC to that of the healthy controls. We did not observe any such resemblance in correlation patterns in structural connectivity (SC) of pre- and post-rehabilitation with that of the healthy control group. Overall, the results point out the suitability of using graph metrics for characterizing early markers of functional change in the brain regions.