The Rasch model is not mathematically limited to psychometrics or human respondents; but should also be applicable to more technical agents. This work, as a bridging exercise exploiting this agnostic aspect, explores how psychometric tools can support a major activity in metrology across the disciplines: namely, Interlaboratory Comparisons (ILCs), where one or more common objects are circulated for measurement amongst several laboratories, as a regular tool for assessing performance. ILCs are less developed for ordinal and nominal data and have been rarely studied to date in the human sciences. The present clinical case study examines whether the performance of different surgical interventions—specifically based on carotid artery stenosis outcomes—can be interpreted in terms of metrological ILCs using Rasch psychometric modelling as an alternative to earlier Generalised Linear Mixed Modelling meta-analyses. Particular attention is given to statistical weighting and the definition of a reference value for these qualitative and categorical ILCs. Task difficulty assessed according to Rasch’s principle of specific objectivity provides the ILC reference value for ordinal data. Beyond surgical applications of the present case study, other potential applications include evaluating the performance of in vitro medical devices, large language models, and automated machine-readability for semantic interoperability.
Dementia is a major challenge in our aging society. Apart from its health economic impact, cognitive decline may affect the quality of life (QoL) and affective states of individuals. From a measurement perspective, it is of high importance to use quality-assured measurements to derive valid and reliable results. This paper investigates to what extent improved memory metrics can be exploited to investigate the strength of correlation between memory ability, depressive symptoms, and well-being. It also aims to shed light on earlier heterogeneous results through the application of quality-assured measurements. In this paper, we analyzed associations between memory ability, measured by the new and more accurate NeuroMET Memory Metric (NMM), and depressive symptoms and wellbeing of the same individuals measured using the Geriatric Depression Scale (GDS) and the WHOQoL-BREF of older individuals (n=332, ranging from healthy controls to patients with dementia due to suspected Alzheimer’s Disease (AD), with a mean age of 72 years). While decreased memory ability was moderately associated with more depressive symptoms (Pearson R = -0.36), it was only negligibly associated with worse well-being (associations with WHOQoL-BREF and subdomains ranging between R = 0.08 and 0.12), which is partly in line with previous studies. A group-wise analysis revealed that these effects were stronger for individuals with subjective cognitive decline (SCD). Our study leveraged measurements scales and models with improved accuracy.
The aim of this paper is two-fold: to investigate development of a Construct Specification Equation (CSE) for UL task difficulty, , and a CSE for person UL ability, , in support of the validity of these two constructs. Measurements of UL task difficulty, , and person UL ability, were derived from applying the Rasch model on the Tetraplegia Upper Limb Activity Questionnaire (TUAQ). The formulations of CSEs as explanations of the two constructs were done using Principal Component Regression (PCR). The CSE for UL task difficulty, , was to a large degree explained by the number of joints involved and the CSE for person UL ability, , was dominated by grasp-related variables. Pearson coefficients of 0.94 and 0.73 were obtained between UL task difficulty and UL person ability from the CSE, respectively, when correlated with each empirical measure. The present work has both explored and extended the methodology for using more qualitative explanatory variables. Specifically, for UL measurements for people with tetraplegia a good CSE for task difficulty, , supports the validity of TUAQ when measuring person UL ability. Additionally, the CSE formulated for person ability, , can be used both for validation purposes as well as a clinical tool.
Functional neuroimaging studies suggest a dynamic trajectory in Alzheimer’s disease (AD), with early hyperconnectivity followed by hypoconnectivity due to pathologic progression. This study aimed to investigate this hypothesis using neurochemical markers derived from high-field magnetic resonance spectroscopy (MRS). We analyzed data from 126 older adults enrolled in the NeuroMET studies, spanning from normal aging to dementia due to suspected AD. Using ultra-high-field 7 Tesla MRS, we quantified levels of the neurotransmitters glutamate (excitatory) and GABA (inhibitory) in the precentral cortex. Plasma p -Tau181 was used as a proxy for AD pathology, and memory was assessed using the NeuroMET Memory Metric (NMM). A p -Tau181 threshold of >2.08 pg/mL was applied as an estimated marker of amyloid positivity (Aß+). Linear mixed-effects models, adjusted for age at baseline, were used to evaluate associations and interaction effects. Glutamate levels showed a non-linear association with age, decreasing up to around 70 years and increasing thereafter, while GABA levels remained stable (Figure 1A–C). Before reaching the suggested pathological conversion threshold for amyloid positivity, both glutamate and GABA showed a slight, non-significant increase in individuals with higher p -Tau181 levels (Figure 1D–F). Notably, amyloid status moderated the relationship between memory ability and glutamate levels (Figure 1G–I). Among Aß-negative individuals, lower memory ability was associated with elevated glutamate, potentially reflecting a very early compensatory response. Our findings support the hypothesis of a non-linear neurochemical trajectory in aging and early AD pathology. The observed age-dependent fluctuations in glutamate, together with subtle shifts in response to rising plasma p -Tau181, may reflect early pathologic or compensatory mechanisms. The interaction between glutamate and memory ability, particularly in supposedly Aß-negative individuals, suggests that excitatory neurotransmission could transiently support cognitive function in the face of emerging pathology. Notably, the threshold of plasma p -Tau181 used to define amyloid positivity may identify individuals in progressed stages, potentially missing earlier windows of therapeutic opportunity. As potential treatments may have markedly different effects depending on the stage of disease progression, investigating the trajectory of neuronal connectivity and activity is critical.
BACKGROUND:Associations between longitudinal changes of plasma biomarkers and cerebral magnetic resonance (MR)-derived measurements in Alzheimer's disease (AD) remain unclear. METHODS:In a study population (n = 127) of healthy older adults and patients within the AD continuum, we examined associations between longitudinal plasma amyloid beta 42/40 ratio, tau phosphorylated at threonine 181 (p-tau181), glial fibrillary acidic protein (GFAP), neurofilament light chain (NfL), and 7T structural and functional MR imaging and spectroscopy using linear mixed models. RESULTS:Increases in both p-tau181 and GFAP showed the strongest associations to 7T MR-derived measurements, particularly with decreasing parietal cortical thickness, decreasing connectivity of the salience network, and increasing neuroinflammation as determined by MR spectroscopy (MRS) myo-inositol. DISCUSSION:Both plasma p-tau181 and GFAP appear to reflect disease progression, as indicated by 7T MR-derived brain changes which are not limited to areas known to be affected by tau pathology and neuroinflammation measured by MRS myo-inositol, respectively. HIGHLIGHTS:This study leverages high-resolution 7T magnetic resonance (MR) imaging and MR spectroscopy (MRS) for Alzheimer's disease (AD) plasma biomarker insights. Tau phosphorylated at threonine 181 (p-tau181) and glial fibrillary acidic protein (GFAP) showed the largest changes over time, particularly in the AD group. p-tau181 and GFAP are robust in reflecting 7T MR-based changes in AD. The strongest associations were for frontal/parietal MR changes and MRS neuroinflammation.
The justification for making a measurement can be sought in asking what decisions are based on measurement, such as in assessing the compliance of a quality characteristic of an entity in relation to a specification limit, SL. The relative performance of testing devices and classification algorithms used in assessing compliance is often evaluated using the venerable and ever popular receiver operating characteristic (ROC). However, the ROC tool has potentially all the limitations of classic test theory (CTT) such as the non-linearity, effects of ordinality and confounding task difficulty and instrument ability. These limitations, inherent and often unacknowledged when using the ROC tool, are tackled here for the first time with a modernised approach combining measurement system analysis (MSA) and item response theory (IRT), using data from pregnancy testing as an example. The new method of assessing device ability from separate Rasch IRT regressions for each axis of ROC curves is found to perform significantly better, with correlation coefficients with traditional area-under-curve metrics of at least 0.92 which exceeds that of linearised ROC plots, such as Linacre's, and is recommended to replace other approaches for device assessment. The resulting improved measurement quality of each ROC curve achieved with this original approach should enable more reliable decision-making in conformity assessment in many scenarios, including machine learning, where its use as a metric for assessing classification algorithms has become almost indispensable.
Measuring a person’s cognitive abilities, such as memory and learning, is central in many medical conditions to reliably diagnose, treat and monitor disease progression. Common tests typically include tasks of recalling sequences of blocks, digits or words. Recalling a word list is affected by so-called serial position effects (SPE), meaning that words at the beginning or end of the list are more likely to be recalled. In our earlier work, as part of including ordinal and nominal properties in metrology, compensation for ordinality in the raw test scores has been performed with psychometric Rasch measurement theory. Thereafter, SPE have been successfully explained with construct specification equations (CSE) dominated by information theoretical entropy as candidate reference measurement procedures. Here, we present how previous German results for explaining memory difficulty in the immediate recalling (IR, trial 1) task of the Rey’s Auditory Verbal Learning Test (RAVLT) can be replicated with a Swedish cohort (the Gothenburg Mild Cognitive Impairment study, n = 251). This CSE replicability for RAVLT demonstrates comparability across the two cohorts in a kind of inter-laboratory study. Moreover, RAVLT includes repeated trials and learning through practice is expected. How memory task difficulty changes over the eight trials in RAVLT is studied: SPE are not so prominent for the delayed recalling sequences and there is an overall reduction in the task difficulty CSE intercept with trial number, interpreted as an effect of learning. To conclude, the methodology and evidence provided here can be clinically used not only to measure a person’s memory ability but also his or her learning ability, as well as understanding the relationship between learning ability and other cognitive domains.
Currently, there are no widely accepted causal models of how common biomarkers linked to dementia present as cognitive symptoms. In the current study, we present a methodology and preliminary findings to ascertain relations between sets of biomarkers and cognitive symptoms where memory ability was quantified by the newly developed NeuroMET Memory Metric (NMM). The sample was from the EMPIR NeuroMET cohort (n = 213 individual assessments; healthy controls [33%], subjective cognitive decline [31%], mild cognitive impairment [15%], Alzheimer’s disease [20%]). Person memory ability was measured with the NMM and the legacy test Corsi block tests (CBT), MRI and MRS biomarkers were acquired with a 7T whole-body Magnetom MRI system and plasma biomarker levels of GFAP and pTau were measured with Simoa. To study how person memory ability can be explained, a so-called construct specification equation (CSE) was developed using principal component regression, where person memory ability was found to depend on a linear combination of a set of explanatory co-variables. Memory ability, ϴi, was found to depend on several biomarkers expressed by the CSE and the most promising CSE yielded was: ϴi = 7(3) + 2(1) x Cortical thickness + 1.7(5) x Hippocampus left – 0.0045(8) GFAP – 0.03(10) x pTau181 – 0.127(2) x Myoinositol – 0.8(4) x Amygdala right + 0.035 x Age. A Pearson correlation coefficient of 0.77 of memory ability based on the CSE prediction against the NMM measured memory ability was achieved. In contrast, all univariate correlations between individual biomarkers and memory ability were lower (highest for Cortical Thickness, R = 0.59) and correlations to memory ability measured with CBT was also lower (R = 0.46). This study is the first to combine a metrologically validated metric derived from set of legacy memory tests with biomarker data in an attempt to better understand predictive links through CSEs. Reflecting the CSE methodology and superior accuracy of the NMM we can get more accurate and clearer understanding of how changes in biomarkers relate to changes in cognitive symptoms can contribute to the advancement in targeting novel drugs and interventions.
There are different views in the literature about the number and inter-relationships of cognitive domains (such as memory and executive function) and a lack of understanding of the cognitive processes underlying these domains. In previous publications, we demonstrated a methodology for formulating and testing cognitive constructs for visuo-spatial and verbal recall tasks, particularly for working memory task difficulty where entropy is found to play a major role. In the present paper, we applied those insights to a new set of such memory tasks, namely, backward recalling block tapping and digit sequences. Once again, we saw clear and strong entropy-based construct specification equations (CSEs) for task difficulty. In fact, the entropy contributions in the CSEs for the different tasks were of similar magnitudes (within the measurement uncertainties), which may indicate a shared factor in what is being measured with both forward and backward sequences, as well as visuo-spatial and verbal memory recalling tasks more generally. On the other hand, the analyses of dimensionality and the larger measurement uncertainties in the CSEs for the backward sequences suggest that caution is needed when attempting to unify a single unidimensional construct based on forward and backward sequences with visuo-spatial and verbal memory tasks.
The ability to measure, track over time, and compare memory ability for people with neurodegeneration is important. However, currently, full comparability of memory test data is limited by a lack of quality assurance of memory measurements. At AAIC 2021, we presented a preliminary item bank to assess memory, composed by selecting items from legacy tests according to metrological principles through use of the Rasch model and with item equivalence based on entropy. Here, we demonstrate direct comparability of measurements generated from different tests, with the new NeuroMET Memory Metric comprising 87 selected items for task difficulty from: the Corsi Block Test, Digit Span Test, Rey’s Auditory Verbal Learning Test, Word Learning List from the CERAD test battery and the Mini Mental State Examination. Data were collected from the European EMPIR NeuroMET and the SmartAge studies recruited at Charité Hospital (Healthy controls n = 92; Subjective cognitive decline n = 160; Mild Cognitive Impairment n = 50; and Alzheimer’s Disease n = 58; age range 55-87). The Rasch analysis showed well-targeted items for all participants’ abilities; good fit to the measurement model, with 83 items (95%) having fit residuals within the expected range and satisfactory unidimensionality, and item reliability of 0.96. The full item bank comprising 87 short-term memory items gave a person reliability of 0.85. Subsequently, a conversion table was created linking the raw scores from the legacy tests to the common NeuroMET Memory Metric and to individual legacy tests. Legacy memory tests have previously proven useful in clinical practice and research, and will continue to be used, but have to date been metrologically limited. The provided conversion table, linking these legacy memory tests to a metrologically assured scale, viz., the NeuroMET memory metric, remedies this deficiency. The NeuroMET memory metric will be included in the first ever prototype of a metrological validated app used to deliver memory tests. Clinicians and researchers will be able to select sets of items to produce data, via a scoring algorithm for transforming patient responses to measures, in a common frame of reference.
Both construct specification equations (CSEs) and entropy can be used to provide a specific, causal, and rigorously mathematical conceptualization of item attributes in order to provide fit-for-purpose measurements of person abilities. This has been previously demonstrated for memory measurements. It can also be reasonably expected to be applicable to other kinds of measures of human abilities and task difficulty in health care, but further exploration is needed about how to incorporate qualitative explanatory variables in the CSE formulation. In this paper we report two case studies exploring the possibilities of advancing CSE and entropy to include human functional balance measurements. In case study I, physiotherapists have formulated a CSE for balance task difficulty by principal component regression of empirical balance task difficulty values from Berg's Balance Scale transformed using the Rasch model. In case study II, four balance tasks of increasing difficulty due to diminishing bases of support and vision were briefly investigated in relation to entropy as a measure of the amount of information and order as well as physical thermodynamics. The pilot study has explored both methodological and conceptual possibilities and concerns to be considered in further work. The results should not be considered as fully comprehensive or absolute, but rather open up for further discussion and investigations to advance measurements of person balance ability in clinical practice, research, and trials.
Better metrics of cognition can be formed by carefully combining selected items from legacy short-term memory tests so as to enhance coherence in item design while not jeopardizing validity. In this paper, we report on how Rasch Measurement Theory and Construct specification equations (CSE) have been brought together when composing the NeuroMET Memory Metric (NMM). The NMM is guided by: i) entropy-based equivalence criteria; ii) a comprehensive understanding of the construct purported to be measured; and iii) how a collection of items works together. CSEs play a major role in ensuring the metrological legitimacy of the NMM in a way analogous to certified reference materials in more established areas of metrology. The resulting NMM for short-term memory recall has up to a five-fold reduction in measurement uncertainties for memory ability compared with an individual legacy test, and the entropy-based CSEs should enable more efficient and valid assessment.
Blood-based biomarkers (BBM) have shown promising potential in diagnosis and prognosis of patients affected by Alzheimer’s disease (AD) pathology. This observational study analyzed the cross-sectional relation between BBMs and disease relevant outcomes (e.g., cognition, imaging modalities). The study sample comprised individuals with subjective cognitive decline (N = 35), mild cognitive impairment (N = 30), dementia due to suspected AD (N = 27) and healthy controls (N = 35). Data for the following measures were assessed: 1) AD-related BBMs measured in plasma on Simoa: amyloid beta ratio 42/40 (Ab42/40), tau phosphorylated at threonine-181 (p-Tau 181), glial fibrilic acid protein (GFAP), and neurofilament light chain (NfL), 2) several cognitive parameters, 3) structural volumes, functional connectivity and brain metabolite concentrations measured by 7T magnetic resonance imaging (MRI) and spectroscopy (MRS), and 4) concentrations of general blood count. Linear mixed models were used to assess associations between the AD-related BBMs (1) and the other measures (2 to 4). While none of the associations between Ab42/40 and the AD-related outcome measures reached significance (p > 0.05), higher concentrations of p-Tau 181, GFAP and NfL were associated with lower values for MMSE, memory and executive function and parietal cortical thickness, as well as higher concentrations of MRS Myo-inositol and blood creatinine (figure 1). Additionally, GFAP and NfL were associated with lower MRS N-Acetylaspartic acid (NAA), and GFAP was associated with smaller hippocampus volume. Beta-coefficients and p-values for all associations are provided in table 1. Unlike Ab42/40, abnormal concentrations of p-Tau 181, GFAP and NfL showed strong and robust associations to other AD-related outcome measures making them a better target for screening, diagnosis and possibly prognosis for individuals with AD.
Accurate assessment of memory ability for persons on the continuum of Alzheimer’s disease (AD) is vital for early diagnosis, monitoring of disease progression and evaluation of new therapies. However, currently available neuropsychological tests suffer from a lack of standardization and metrological quality assurance. Improved metrics of memory can be created by carefully combining selected items from legacy short-term memory tests, whilst at the same time retaining validity, and reducing patient burden. In psychometrics, this is known as “ crosswalks ” to link items empirically. The aim of this paper is to link items from different types of memory tests. Memory test data were collected from the European EMPIR NeuroMET and the SmartAge studies recruited at Charité Hospital (Healthy controls n = 92; Subjective cognitive decline n = 160; Mild cognitive impairment n = 50; and AD n = 58; age range 55–87). A bank of items (n = 57) was developed based on legacy short-term memory items (i.e., Corsi Block Test, Digit Span Test, Rey’s Auditory Verbal Learning Test, Word Learning Lists from the CERAD test battery and Mini Mental State Examination; MMSE). The NeuroMET Memory Metric (NMM) is a composite metric that comprises 57 dichotomous items (right/wrong). We previously reported on a preliminary item bank to assess memory based on immediate recall, and have now demonstrated direct comparability of measurements generated from the different legacy tests. We created crosswalks between the NMM and the legacy tests and between the NMM and the full MMSE using Rasch analysis (RUMM2030) and produced two conversion tables. Measurement uncertainties for estimates of person memory ability with the NMM across the full span were smaller than all individual legacy tests, which demonstrates the added value of the NMM. Comparisons with one (MMSE) of the legacy tests showed however higher measurement uncertainties of the NMM for people with a very low memory ability (raw score ≤ 19). The conversion tables developed through crosswalks in this paper provide clinicians and researchers with a practical tool to: (i) compensate for ordinality in raw scores, (ii) ensure traceability to make reliable and valid comparisons when measuring person ability, and (iii) enable comparability between test results from different legacy tests.
AbstractThe quality-assurance of measurement in person-centered care (PCC) – is introduced firstly by “bookending” the topic in the overall context of the quality assurance of the care itself. At the start the chapter we ask: What are the end-user objects and constructs of PCC – for instance as specified by the profession and in legislation? At the end: What decisions about PCC objects and constructs can be made and how reliable are they? Examples and illustrations from PCC have included (i) neuropsychological cases (dealt with in more detail in the accompanying chapter by Melin and Pendrill (Person centered outcome metrology. Springer, 2022)) and (ii) patient participation. In the two central sections of the chapter, assuring the quality of measurement in PCC has obliged consideration of how traditional metrological concepts – particularly metrological references for comparability via traceability and reliable estimates of uncertainty – need to be extended. In providing an overview of the benefits of combining Rasch measurement theory and quality assurance, the unique properties of Rasch Measurement Theory are exploited to the full. Replacing the instrument at the heart of a traditional measurement system with a human being provides a truly “person-centered” model of the metrology. This in turn enables a viable procedure to establish metrological references in fields such as PCC in the form of “recipes” analogous to certified reference materials or procedures in analytical chemistry and materials science. It also informs the measurement uncertainties which determine the final decisions about PCC taken at the end of the chapter.
Metrological methods for word learning list tests can be developed with an information theoretical approach extending earlier simple syntax studies. A classic Brillouin entropy expression is applied to the analysis of the Rey's Auditory Verbal Learning Test RAVLT (immediate recall), where more ordered tasks-with less entropy-are easier to perform. The findings from three case studies are described, including 225 assessments of the NeuroMET2 cohort of persons spanning a cognitive spectrum from healthy older adults to patients with dementia. In the first study, ordinality in the raw scores is compensated for, and item and person attributes are separated with the Rasch model. In the second, the RAVLT IR task difficulty, including serial position effects (SPE), particularly Primacy and Recency, is adequately explained (Pearson's correlation R=0.80) with construct specification equations (CSE). The third study suggests multidimensionality is introduced by SPE, as revealed through goodness-of-fit statistics of the Rasch analyses. Loading factors common to two kinds of principal component analyses (PCA) for CSE formulation and goodness-of-fit logistic regressions are identified. More consistent ways of defining and analysing memory task difficulties, including SPE, can maintain the unique metrological properties of the Rasch model and improve the estimates and understanding of a person's memory abilities on the path towards better-targeted and more fit-for-purpose diagnostics.
Myo-inositol (MI) is a presumed marker for glial activation which can be quantified by non-invasive Magnetic Resonance Spectroscopy (MRS) and is suspected to be elevated in early Alzheimer’s disease (AD). In the framework of the NeuroMET2 project, we investigated the relationship between MI and other AD-relevant measures. Absolute concentrations of MI were measured by 7 tesla MRS in the posterior cingulate cortex (PCC)/precuneus of 26 cognitively healthy individuals (HC), 23 patients with subjective cognitive decline (SCD), 23 with mild cognitive impairment (MCI) and 24 with dementia due to suspected AD. Using linear models, MI’s association with memory ability (NeuroMET Memory Metric), volumes of the hippocampus and PCC/precuneus, and seed-based functional connectivity were investigated. Interaction terms involving the apolipoprotein E (APOE) ε4 allele were analyzed additionally. All models were adjusted to age and education and weighted based on Cramér-Lao lower bounds. Absolute MI concentrations were substantially elevated in AD patients (adjusted mean [95% CI]=8.5 mmol/l [7.8; 9.3]). However, there were no relevant differences between SCD (6.7 mmol/l [6.0; 7.3]) or MCI (7.5 mmol/l [6.8; 8.2]) and participants of the HC group (7.2 mmol/l [6.6, 7.9]). Across the whole cohort, higher levels of MI were associated with lower memory ability (ß [95% CI]=-0.31 [-0.44; -0.17], std. ß=-0.40, p<0.001), smaller hippocampus volume (ß=-59.89 [-131.40; 11.62], std. ß=-0.16, p=0.100) and lower seed-based functional connectivity (ß=-3948.41 [-6480.69; -1416.13], std. ß=-0.34, p=0.003). There was no relevant effect between MI and PCC/precuneus volume (ß=83.27 [-117.89; 284.43], std. ß=0.08, p=0.413). Finally, there was no substantial mediation effect of APOE ε4 carriership on the association between MI and neither memory ability (ß=0.09 [-0.18; 0.36], std. ß=0.37, p=0.519), hippocampus volume (ß=-55.60 [-223.48; 112.29], std. ß=-0.45, p=0.511), PCC/precuneus volume (ß=-110.68 [-546.22; 324.86], std. ß=-0.32, p=0.614) nor seed-based functional connectivity (ß=2009.46 [-4207.40; 8226.32], std. ß=0.55, p=0.520). Across a cohort with participants ranging from healthy to AD, we showed that glial activation, measured by MI concentration, was associated with key AD-related processes. The potential of MI as a non-invasive biomarker for increasing memory impairment in pre-dementia stages remains to be explored in longitudinal studies.
AbstractMemory ability, together with many other constructs related to disability and quality of life, is of growing interest in the social sciences, psychology and in health care examinations. This chapter will focus on two elements aiming at understanding, predicting, measuring and quality-assuring constructs with examples from memory measurements: (i) explicit methods for testing theories of the measurement mechanism and establishment of metrological standards and (ii) substantive theories explaining the constructs themselves. Building on entropy as a principal explanatory variable, analogous to its use in thermodynamics and information theory, we demonstrate how more fit-for-purpose and valid memory measurements can be enabled. Firstly, memory task difficulty, extracted from a Rasch psychometric analysis of memory measurements of experimental data such as from the European NeuroMET project, can be explained with a construct specification equation (CSE). Based on that understanding, the CSE can facilitate the establishment of objective and scalable units through the generation of novel certified reference “materials” for metrological traceability and comparability. These formulations of CSEs can also guide how best to compose new memory metrics, through a judicious choice of items from various legacy tests guided by entropy-based equivalence, which opens up opportunities for formulating new, less onerous but more sensitive and representative tests. Finally, we propose and demonstrate how to formulate CSEs for person ability, correlated statistically and clinically with sets of biomarkers, that can be a means of providing diagnostic information to enhance clinical decisions and targeted interventions.
The extent to which existing legacy memory tests are accurate and sensitive in detecting early cognitive decline is questionable. This reflects the wider picture in neurodegenerative conditions where clinical rating scales have not been developed for early-stage disease. Item banks delivered through computer adaptive tests can help. But, in order to be fit-for-purpose, this approach requires a metrological framework with recourse to units, traceability, and interoperability. Here we describe the initial research to build such an item bank, based on legacy tests. Memory tests (i.e., Corsi Block Test [CBT], Digit Span Test [DST], Rey’s Auditory Verbal Learning Test [RAVLT] and Word Learning List [WLL]) data were collected in the European EMPIR NeuroMET and the SMART cohorts recruited in Charité Hospital (Healthy control n 86; Subjective Cognitive Decline n 99; Mild Cognitive Impairment n 37; and Alzheimer’s Disease n 45). In order to align with metrological requirements, Rasch measurement theory in conjunction with construct specification equations were chosen to analyse the data. Based on the combined dataset, when analyzed on their own, the data from the CBT, DST, RAVLT and WLL revealed skewness, gaps, and large measurement uncertainties. The addition of items from each of the tests into a bank improved these psychometric issues, improving reliability from a minimum of 0.65 to 0.85. The metrological legitimization of the ‘NeuroMET Memory Metric’ (formed from the combination of items) was confirmed through construct specification equations, which provided a Pearson correlation coefficient for empirical values vs. predicted (zR) values of up to 0.98 for task difficulty. Our early promising findings provide a strong foundation for a metric of memory ability, which can be used as the basis for computer adaptive testing, better targeted measurement, and more accurate discrimination of early cognitive decline. Establishing units can also provide for the potential of developing crosswalks between a wider range of memory test items. Further work includes further longitudinal data collection to confirm item estimates, development of a digital platform, and a quality assurance program to establish traceability, and interoperability.