Understanding how differences in brain structure relate to differences in cognition across the lifespan is essential for addressing age-related cognitive decline. Since age is strongly associated with both brain structure and cognition, predictive models often risk simply capturing age effects. To mitigate this risk, deconfounding is typically applied to remove the effects of age. Here, beyond treating age as a confound, we treat it as a moderator by estimating brain-cognition associations separately across age groups. This captures age-stratified changes in how brain structure and cognitive performance are statistically connected. For this view to hold, variations in brain structure linked to differences in cognitive performance in older subjects (eg related to disease) would differ from those in younger subjects. Using structural brain imaging data from the UK Biobank we found an asymmetry in generalisability: models trained on younger subjects successfully predicted cognition in older subjects, but models trained on older subjects failed to generalize to younger individuals. These findings reveal a trade-off between model specificity and generalisability, suggesting the optimal approach-whether age-specific or pooled-depends on the research or clinical goal for the target population.
Abstract The UK Biobank (UKB) Brain Imaging cohort contains data from almost 100,000 subjects and has yielded invaluable understanding of the links between the brain and health outcomes and lifestyles. Much of the understanding of these links has come from exploring the association between Imaging Derived Phenotypes (IDPs) and other variables that are unrelated to brain imaging, so called non-Imaging Derived Phenotypes (nIDPs). When performing analysis of this kind, it is very important to control for well known confounding factors such as age, sex and socio-economic status, as well as confounds which are related to the imaging protocol itself. In previous work, we created a pipeline for constructing imaging confounds for use in statistical inference via a standard multivariate linear regression approach (Alfaro-Almagro et. al. 2021). However, this approach is problematic when the number of confounds exceeds the number of subjects, and is severely underpowered when the number of number of subjects is not much larger than the number of confounds. In this work, we perform a simulation study to evaluate 13 modelling approaches to account for confounds when their number is similar to or exceeds the number of subjects. Based on the simulation results, we recommend a ridge regression based permutation test for low sample sizes ( n ≤ 50), a version of de-sparsified LASSO for intermediate sample sizes (50 < n ≤ 500), and multivariate linear regression aided by Principal Component Analysis (PCA) for larger sample sizes ( n > 500). We also demonstrate the use of our recommended methodology on a real data example of finding associations between Alzheimer’s Disease (AD) and IDPs.
This paper presents a framework for modelling the topography of whole-brain connectivity in resting-state functional MRI. The aim is to disentangle functional segregation, which manifests as abrupt changes in connectivity, from so-called gradients, i.e., smooth variations in connectivity across the brain. Our core assumption is that functional segregation leads to low-rank structure in the dense (point-to-point) connectome, whereas connectivity gradients imply a sparse and non-low-rank structure in the dense connectome. Our method thus decomposes the connectome into low-rank and sparse components, enabling the integration of local-nonlinear and global-linear embedding strategies. We show that this hybrid model approximates the empirical dense connectome more effectively than purely low-rank or purely gradient approaches. We also find that connectivity gradients derived from this model exhibit strong correspondence with task-based topographic maps. We hope that this approach can provide insight into the organisational principles of brain regions where gradients remain poorly characterised.
Low-rank matrix decompositions can uncover patterns and structure in data and have a number of different applications across many disciplines. Extensions to "joint" low-rank decompositions have been proposed to link datasets from different modalities. While these methods enable the discovery of common patterns across modalities, they require that all the multimodal data share one or more dimensions. We propose a new analysis method, PathFinder, that enables co-analysis of datasets that do not necessarily all share a dimension. The key insight is that as long as pairs or subgroups of matrices do share some dimension, and that there are one or more paths that link across the data matrices, a global joint decomposition can be sought out. This enables the joint estimation of common patterns across different modalities, species, or scales, where a one-to-one mapping across all data along some dimension is not necessarily available. We show that PathFinder is a general umbrella under which many matrix decomposition methods fall as special cases. It can be used to discover common patterns across disparate datasets and to make predictions for missing data or modalities.
Abstract Dynamic functional connectivity (dFC) models have become increasingly popular over the past decade for characterising time-varying interactions between brain regions. However, assessing and comparing dFC models remains challenging. Here, we introduce bi-cross-validation as a general framework for evaluating dFC models and selecting key hyperparameters, such as the number of states. By jointly partitioning the data across subjects and brain regions, bi-cross-validation enables out-of-sample evaluation without re-estimating latent states on the same data used for testing, thereby avoiding circularity. Using simulated data with known ground-truth dynamics, we show that bi-cross-validation favours models that accurately capture the underlying state structure. Applying the framework to real resting-state fMRI data, we demonstrate that bi-cross-validation naturally balances goodness-of-fit against model complexity, with performance improving and then declining as model complexity increases. Finally, we use bi-cross-validation to directly compare static and dynamic FC models, showing that dynamic models underperform static models at low spatial dimensionality, but outperform static models at sufficiently high dimensionality. Together, these results establish bi-cross-validation as a principled tool for dFC model selection, evaluation, and comparison.
Background:Multi-system impacts of long COVID remain unknown. We aimed to compare multi-system deficits between people with long COVID and controls. Methods:We conducted a case-control study and recruited participants from two UK population-based cohorts: the Avon Longitudinal Study of Parents and Children (ALSPAC) and TwinsUK. Participants provided samples for SARS-CoV-2 serology between 2020 and 2021 and were asked about duration of COVID-19 symptoms between July and December 2021. Cases had long COVID (evidence of COVID-19 infection and persistent symptoms ≥4 weeks post infection); controls comprised 3 groups: acute COVID-19 only (symptoms reported for <4 weeks) and serological evidence of infection; self-reported long COVID-like symptoms but without wild-type SARS-CoV-2 virus antibodies; no symptoms or history of COVID-19 infection. People who were severely unwell or pregnant were excluded. Participants attended a clinic follow-up visit between 2021 and 2023 and underwent multi-system MRI, (cardiac, brain, lung, kidney), measurement of blood pressure and autonomic function, spirometry, renal function, exercise tests, strength and physical capability. Severity of deficit was then scored for each system as 0 (none) to 3 (severe). Primary outcome was a single composite multi-domain score summing each of nine domains: autonomic, brain, exercise capacity, heart, lungs, physical, renal, strength and vascular, with a maximum score of 27. Findings:In total, 349 participants, 141 with long COVID (40%) and 208 (60%) controls were recruited. Overall deficit score in cases was 0.22 (95% CI -0.44, 0.88) units greater than controls, adjusted for age, sex, ethnicity, cohort membership and relatedness. This estimate was little changed (0.32 (-0.34, 0.98)) when additionally adjusted for educational status, index of multiple deprivation, physical activity, smoking and co-morbidity. Restricting cases to those reporting symptoms including fatigue (n = 46) increased the excess deficit score to 0.81 (-0.19, 1.81) units in the minimally adjusted model. A difference was only observed in the vascular domain, largely attributable to elevated blood pressure, showing a 1.76 (1.04, 2.97) multivariable adjusted odds ratio excess in cases, and 3.04 (1.36, 6.80) when restricted to cases with fatigue. Interpretation:There was no evidence of marked residual subclinical deficits in most systems in people with long-COVID, although there was evidence of persistent deficits in the vascular system, largely related to elevated blood pressure. Mechanisms unrelated to organ dysfunction may contribute to symptoms. Blood pressure measurement and control should be included in clinical follow-up. Funding:Jointly funded by the National Institute for Health and Care Research and UK Research and Innovation (CONVALESCENCE, COV-LT-0009, MC_PC_20051).
Functional magnetic resonance imaging (fMRI) of awake macaque monkeys holds promise for advancing our understanding of primate brain organization, including humans. However, estimating functional connectivity in awake animals is challenging due to the limited duration of imaging sessions and the relatively low sensitivity to neural activity. To overcome these challenges, we developed a 24-channel 3T receive radiofrequency (RF) coil optimized for parallel imaging of awake macaques. This enabled the acquisition of cross-plane and in-plane accelerated ferumoxytol-weighted resting-state fMRI. The Human Connectome Project-style data processing pipelines were adapted to address the unique preprocessing demands of cerebral blood volume-weighted (CBVw) imaging, including motion correction, functional-to-structural image co-registration, and training a multi-run independent component analysis-based X-noiseifier (ICA-FIX) classifier for removal of structured artifacts. Our CBVw fMRI approach resulted in an elevated contrast-to-noise ratio compared with blood oxygenation level-dependent (BOLD) imaging in anesthetized macaques. However, structured imaging artifacts still contributed more variance to the functional timeseries than neural activity. By applying the ICA-FIX classifier, we achieved highly reproducible parcellated functional connectivity at the single-subject level, with test-retest matrix correlations and subject identification accuracy comparable with those observed in humans. At the group level, we identified dense functional networks with spatial features homologous to those observed in humans. The developed RF receive coil, image acquisition protocols, and data analysis pipelines are publicly available, providing the broader scientific community with tools to leverage these advances for further research.
Lesion network mapping (LNM) and related techniques have been used in over 200 studies, primarily to test whether anatomically distributed lesions that cause the same symptom fall within a common brain network. A recent article1 challenges the specificity and validity of this technique, suggesting that lesion network maps primarily reflect intrinsic properties of the normative connectome rather than lesion-symptom relationships. However, the data and procedures in van den Heuvel et al. do not reflect those used in most LNM studies. Further, the main conclusions were based on similarity between maps, but similarity does not imply the absence of meaningful differences. In contrast, LNM provides evidence for meaningful differences using specificity testing. Exemplary analyses of 1090 lesion locations from 34 prior LNM studies do not support van den Heuvel's concerns and confirm the lesion-deficit specificity of LNM. While we encourage further methodological investigation, the analyses of van den Heuvel et al. do not invalidate prior LNM findings or future applications.
Individual differences in the volumes of brain structures are often linked to various conditions, including Alzheimer's disease, schizophrenia, and overall brain health. However, it remains unclear to what extent these differences reflect individual levels present from young adulthood or diverging aging trajectories from later ages. In this study, we analyze the aging dynamics of the volumes of six brain structures based on magnetic resonance imaging (MRI) scans from a large cross-cohort longitudinal sample of cognitively healthy adults (n = 8,311 with 18,520 MRIs, ages from 18 to 97 years). From general assumptions about structural brain dynamics and measurement noise, a stochastic dynamical model was fitted to the data to estimate both the variability and persistence of structural changes across adulthood. Using this model, we calculated how much of the variance of volumetric differences between individuals can be attributed to stable levels from young adulthood versus systematic changes at older ages, as well as the theoretical sensitivity of longitudinal studies to detect individual differences in change. The findings were as follows: (1) Before age 60 years, inter-individual differences in neuroanatomical volumes almost exclusively reflect stable differences between individuals, while the influence from systematic differences in rate-of-change increases thereafter: up to 50% of the variation being due to differences in change at 80 years. In contrast, ventricular volume reflects differences in change from early adulthood. (2) Current brain-age models are unlikely to be sensitive to detect differences in aging trajectories. (3) Imaging studies have low reliability in detecting inter-individual brain changes before age 60 years. After 60 years, the study reliability increases sharply with longer intervals between scans and more modestly with additional intermediate observations. In conclusion, our results reinforce the view that it is critical to distinguish stable early adulthood levels from systematic differences in change when studying adult brain aging.
We present PANDORA (Population Archive of Neuroimaging Data Organized for Rapid Analysis), a huge brain imaging data archive and analysis resource for UK Biobank neuroimaging data. PANDORA UKBv1 contains 81,939 subjects’ voxel-level images, created by the core UKB brain image processing pipeline. PANDORA also includes highly efficient supervoxel versions of the data – much smaller and faster to work with than the full voxelwise representation while losing virtually no signal or spatial detail and providing denoising. For each of 98 sub-modalities (outputs from processing functional, structural, and diffusion modalities, e.g., fractional anisotropy from dMRI), images are collated into one massive subjectsXvoxels matrix, stored as a single convenient HDF5 file, alongside the compact subjectsXsupervoxels representations (1K and 10K). Also included are robust preprocessing, a curated set of imaging confounds, and a tool for easy voxelwise cross-subject regression against variables such as genetics or lifestyle factors. The regression tool performs model estimation in the encoded supervoxel space while computing mass-univariate statistics in full-resolution space via supervoxel-to-voxel mappings. PANDORA is easy and quick to use on the UK Biobank RAP (Research Analysis Platform). Downloading a local copy of a single sub-modality from the central data store takes a few minutes, and then running a regression across 80K subjects and around 1M voxels takes between a few seconds and an hour, depending on the modality, the number of subjects, and the size of the regression design matrix. Validation shows that 10K-supervoxels closely match full-resolution spatial maps, whereas 1K-supervoxels are extremely compute-efficient and boost power for less localized effects. Benchmarks indicate substantive speed and memory gains. Across five experiments (common substance use, cumulative trauma, anxiety vs. depression, early-vs. late-diagnosis autism polygenic scores, and an EPHA3 SNP), results reproduce known patterns and reveal new associations.
UK Biobank is the world's most comprehensive longitudinal population-based data and biosample resource. Twenty years after UK Biobank was first established, the incidence of dementia among participants is rising and is set to increase rapidly over the next 5-10 years, creating a distinct opportunity for studies of dementia risk and onset. In addition to extensive clinical phenotyping of >500,000 volunteers from across the UK at recruitment and at follow-up time points, UK Biobank includes data from serial lifestyle questionnaires, cognitive testing, multimodal imaging, accelerometry, genomics and other omics that are linked to individual health, cancer and death records. In this Perspective, we discuss how the use of UK Biobank data has enabled the discovery of new interactions between systemic and brain health and illustrate how these data can be used to characterize and identify risk factors, support mechanistic hypotheses and identify new biomarkers that predict the onset and course of dementia and related disorders. We also consider future developments of UK Biobank, including the UK Biobank Brain Health Study, which will build on and leverage the increasing incidence of dementias to advance understanding of these conditions.
The original Human Connectome Project multimodal cortical parcellation (HCP_MMP1.0) used MRI-derived local features and long-distance functional connectivity measures to define a multimodal parcellation at the group level, accompanied by an automated areal classifier, to create subject-specific mappings of the human cerebral cortex. These mappings, referred to as individual (cortex) parcellations, aim to capture individual variability in areal organization by learning from both structural and functional data. However, a strict supervised learning approach using the group parcellation as labels would have no incentive to learn individual differences that registration is unable to reconcile (e.g., atypical 55b topologies). Furthermore, there are many types of resting state network (RSN) feature maps, and it is unclear which type would most accurately or effectively classify areas, or even what should be the primary criteria for evaluating classification performance. Here, we introduce an A real R ecognition E nsemble with N ested A pproach (ARENA) classifier that learns from uncertain labels, using a novel application of weakly supervised learning to this type of problem. Additionally, in comparing multiple candidate RSN decompositions, temporal ICA and PROFUMO maps outperformed the original spatial ICA-based approach based on objective criteria. With these refinements, the ensemble classifier achieved a reliable individual variability score of 8380, an average areal detection rate of 97.8%, and test-retest reproducibility of 73.3%, outperforming a retrained version of the original Multi-layer Perceptron (MLP) model (whose reliable individual variability score was 3128, average areal detection rate was 97.2%, and test-retest reproducibility was 71.6% on the same dataset). Furthermore, the ARENA classifier demonstrated stronger generalization for all three measures when applied to task fMRI data that were not part of the training dataset. Using the refined classifier and leveraging all 1071 HCP-Young Adult subjects, we identified new types of atypical organization of language-related area 55b. Here we provide the fully data-driven HCP_MMP1.0_1071_MPM (Maximum Probability Map) group parcellation and a summary of area 55b organization in both hemispheres. Our automated individual parcellation pipeline powered by the novel ARENA classifier is now integrated into the HCP pipelines, offering a user-friendly tool for the neuroimaging community.
Neuroimaging faces a reproducibility crisis, where studies on small, heterogeneous datasets produce unreliable brain-wide associations and AI models that fail to generalize. To address this, we introduce GenBrain, a generative foundation model pretrained on approximately 1.2 million 3D scans from over 44,000 individuals across 34 imaging modalities to learn a population prior of brain structure and function. Crucially, GenBrain enables rapid, data-efficient adaptation, allowing any targeted study to generate biologically valid synthetic cohorts, conditioned on demographics, disease status, or other modalities, to augment statistical power and enhance generalizability. We demonstrate GenBrain’s transformative utility across 81 independent datasets spanning diverse populations, protocols, and clinical conditions. For image-level tasks, it achieves state-of-the-art performance in image enhancement and cross-modality synthesis while preserving subject-specific neurobiology. In population neuroscience, synthetic cohorts from GenBrain stabilize effect-size estimates and significantly improve the reproducibility of brain-wide association studies. For clinical AI, disease-specific fine-tuning of GenBrain substantially boosts the cross-site generalizability of prediction models. Finally, we prove its direct translational value when adapted to unseen modality and scarce clinical stroke data. GenBrain significantly improves predictions of acute stroke severity and chronic aphasia, demonstrating actionable utility under extreme data scarcity. By empowering small-scale studies with large-scale population priors, GenBrain provides a unified framework for more reproducible and clinically generalizable neuroimaging analysis.
Differences in the volumes of brain structures between individuals are often linked to various conditions, including Alzheimer's disease, schizophrenia, and overall brain health. However, it remains unclear to what extent these differences reflect individual levels present at young adulthood or diverging aging trajectories at later ages. In this study, we analyze the aging dynamics of the volume of six brain structures based on MRI scans from a large cross-cohort longitudinal sample of cognitively healthy adults (n = 8,311 with 18,520 MRIs, ages from 18 to 97 years). From general assumptions about structural brain dynamics and measurement noise, a stochastic dynamical model was fit to the data to estimate both the variability and persistence of structural changes across adulthood. Using this model, we calculated how much of the variance in individual volumetric differences can be attributed to stable levels from young adulthood versus systematic changes at older ages, as well as the theoretical sensitivity of longitudinal studies to detect individual differences in changes. The findings were as follows: 1) Before age 60 years, inter-individual differences in neuroanatomical volumes almost exclusively reflect stable differences between individuals, while the influence from systematic differences in rate-of-change increases thereafter; up to 40 % of the variation being due to differences in change at 80 years. In contrast, ventricular volume reflects differences in change from early adulthood. 2) Current brain-age models are unlikely to be sensitive to detect differences in aging trajectories. 3) Imaging studies have a low reliability to detect inter-individual brain change before age 60. After 60 years, the study reliability increases sharply with longer intervals between scans and more modestly with additional intermediate observations. In conclusion, it is critical to distinguish between stable levels from early adulthood and systematic differences in change when studying adult brain aging.
The potential value of large scale datasets is constrained by the ubiquitous problem of missing data, arising in either a structured or unstructured fashion. When imputation methods are proposed for large scale data, one limitation is the simplicity of existing evaluation methods. Specifically, most evaluations create synthetic data with only a simple, unstructured missing data mechanism which does not resemble the missing data patterns found in real data. For example, in the UK Biobank missing data tends to appear in blocks, because non-participation in one of the sub-studies leads to missingness for all sub-study variables. We propose a tool for generating mixed type missing data mimicking key properties of a given real large scale epidemiological data set with both structured and unstructured missingness while accounting for informative missingness. The process involves identifying sub-studies using hierarchical clustering of missingness patterns and modelling the dependence of inter-variable correlation and co-missingness patterns. On the UK Biobank brain imaging cohort, we identify several large blocks of missing data. We demonstrate the use of our tool for evaluating several imputation methods, showing modest accuracy of imputation overall, with iterative imputation having the best performance. We compare our evaluations based on synthetic data to an exemplar study which includes variable selection on a single real imputed dataset, finding only small differences between the imputation methods though with iterative imputation leading to the most informative selection of variables. We have created a framework for simulating large scale data with that captures the complexities of the inter-variable dependence as well as structured and unstructured informative missingness. Evaluations using this framework highlight the immense challenge of data imputation in this setting and the need for improved missing data methods.
Information processing in the brain spans from localised sensorimotor processes to higher-level cognition that integrates across multiple regions. Interactions between and within these subsystems enable multiscale information processing. Despite this multiscale characteristic, functional brain connectivity is often either estimated based on 10-30 distributed modes or parcellations with 100-1000 localised parcels, both missing across-scale functional interactions. We present Multiscale Probabilistic Functional Modes (mPFMs), a new mapping which comprises modes over various scales of granularity, thus enabling direct estimation of functional connectivity within- and across-scales. Crucially, mPFMs emerged from data-driven multilevel Bayesian modelling of large functional MRI (fMRI) populations. We demonstrate that mPFMs capture both distributed brain modes and their co-existing subcomponents. In addition to validating mPFMs using simulations and real data, we show that mPFMs can predict ~900 personalised traits from UK Biobank more accurately than current standard techniques. Therefore, mPFMs can offer a paradigm shift in functional connectivity modelling and yield enhanced fMRI biomarkers for traits and diseases.
"Brain age delta" is the difference between age estimated from brain imaging data and actual age. Positive delta in adults is normally interpreted as implying that an individual is aging (or has aged) faster than the population norm, an indicator of unhealthy aging. Unfortunately, from cross-sectional (single timepoint) imaging data, it is impossible to know whether a single individual's positive delta reflects a state of faster ongoing aging, or an unvarying trait (in other words, a "historical baseline effect" in the context of the population being studied). However, for a cross-sectional dataset comprising many individuals, one could attempt to disambiguate varying aging rates from fixed baseline effects. We present a method for doing this, and show that for the common approach of estimating a single delta per subject, baseline effects are likely to dominate. If instead one estimates multiple biologically distinct modes of brain aging, we find that some modes do reflect aging rates varying strongly across subjects. We demonstrate this, and verify our modelling, using longitudinal (two timepoint) data from 4,400 participants in UK Biobank. In addition, whereas previous work found incompatibility between cross-sectional and longitudinal brain aging, we show that careful data processing does show consistency between cross-sectional and longitudinal results.
Purpose The pathogenesis of the long-lasting symptoms which can follow an infection with the SARS-CoV-2 virus (‘long covid’) is not fully understood. The ‘COroNaVirus post-Acute Long-term EffectS: Constructing an evidENCE base’ (CONVALESCENCE) study was established as part of the Longitudinal Health and Wellbeing COVID-19 UK National Core Study. We performed a deep phenotyping case-control study nested within two cohorts (the Avon Longitudinal Study of Parents and Children and TwinsUK) as part of CONVALESCENCE.Participants From September 2021 to May 2023, 349 participants attended the CONVALESCENCE deep phenotyping clinic at University College London. Four categories of participants were recruited: cases of long covid (long covid(+)/SARS-CoV-2(+)), alongside three control groups: those with neither long covid symptoms nor evidence of prior COVID-19 (long covid(-)/SARS-CoV-2(-); control group 1), those who self-reported COVID-19 and had evidence of SARS-CoV-2 infection, but did not report long covid (long covid(-)/SARS-CoV-2(+); control group 2) and those who self-reported persistent symptoms attributable to COVID-19 but no evidence of SARS-CoV-2 infection (long covid(+)/SARS-CoV-2(-); control group 3). Remote wearable measurements were performed up until February 2024.Findings to date This cohort profile describes the baseline characteristics of the CONVALESCENCE cohort. Of the 349 participants, 141 (53±15 years old; 21 (15%) men) were cases, 89 (55±16 years old; 11 (12%) men) were in control group 1, 75 (49±15 years old; 25 (33%) men) were in control group 2 and 44 (55±16 years old; 9 (21%) men) were in control group 3.Future plans The study aims to use a multiorgan score calculated as the cumulative total for each of nine domains (ie, lung, vascular, heart, kidney, brain, autonomic function, muscle strength, exercise capacity and physical performance). The availability of data preceding acute COVID-19 infection in cohorts may help identify the consequences of infection independent of pre-existing subclinical disease and also provide evidence of determinants that influence the development of long covid.
Introduction SARS-CoV-2 disease (COVID-19) has had an enormous health and economic impact globally. Although primarily a respiratory illness, multi-organ involvement is common in COVID-19, with evidence of vascular-mediated damage in the heart, liver, kidneys and brain in a substantial proportion of patients following moderate-to-severe infection. The pathophysiology and long-term clinical implications of multi-organ injury remain to be fully elucidated. Age, gender, ethnicity, frailty and deprivation are key determinants of infection severity, and both morbidity and mortality appear higher in patients with underlying comorbidities such as ischaemic heart disease, hypertension and diabetes. Our aim is to gain mechanistic insights into the pathophysiology of multiorgan dysfunction in people with COVID-19 and maximise the impact of national COVID-19 studies with a comparison group of COVID-negative controls.Methods and analysis COmorbidities and Sociodemographic factors on Multiorgan Injury following COVID-19 (COSMIC) is a prospective, multicentre UK study which will recruit 200 subjects without clinical evidence of prior COVID-19 and perform extensive phenotyping with multiorgan imaging, biobank serum storage, functional assessment and patient reported outcome measures, providing a robust control population to facilitate current work and serve as an invaluable bioresource for future observational studies.Ethics and dissemination Approved by the National Research Ethics Service Committee East Midlands (REC reference 19/EM/0295). Results will be disseminated via peer-reviewed journals and scientific meetings.Trial registration number COSMIC is registered as an extension of C-MORE (Capturing Multi-ORgan Effects of COVID-19) on ClinicalTrials.gov (NCT04510025).