
In legal contexts, psychiatric and psychological assessment decisions can have substantial consequences, so mental-health symptoms reported by evaluees must be assessed for credibility rather than taken at face value. To this end, the Inventory of Problems–29 (IOP-29) is particularly relevant because it was specifically developed to assess the credibility of self-reported symptoms and has been investigated in numerous language versions. Because no study had yet addressed the Russian-language IOP-29, this study replicated Akca et al.’s (2023) research paradigm in a Kazakhstan community sample. Overall, 154 volunteers completed the IOP-29 three times, yielding 462 administrations under honest responding, random responding, and instructed feigning conditions; the latter involved feigned schizophrenia, depression, or PTSD. The False Disorder Probability Score (FDS) discriminated feigned from honest responding, with specificity = 0.92 and sensitivity = 0.81 at the standard cutoff (FDS ≥ 0.50). The Random Responding Scale (RRS) differentiated random from non-random responding at the mean-score level, but classification accuracy was mixed. These findings provide support for the validity of the Russian IOP-29 as a measure of the credibility of reported mental health symptoms, whereas the RRS requires further study before it can be used routinely in applied settings.
Invalid responding can bias clinical judgement, particularly in adult attention-deficit/hyperactivity disorder (ADHD), where rates of symptom overreporting and cognitive underperformance are elevated. Performance validity tests (PVTs) and symptom validity tests (SVTs) are commonly used to assess response credibility, yet their relevance for long-term clinical outcomes in ADHD remains unclear. The present study examined whether PVT and SVT scores obtained during diagnostic assessment are associated with long-term clinical outcomes in adults with ADHD. Data were drawn from an outpatient cohort of adults diagnosed with ADHD (N = 181). Participants completed standardised diagnostic assessment including two memory-based PVTs, one attention-based PVT, and two embedded SVTs. Long-term clinical outcomes, including psychotherapy use, pharmacotherapy, and diagnostic re-evaluation, were assessed via self-report approximately one year later. Associations were examined using correlational analyses. No significant associations were observed between PVT and SVT scores and long-term clinical outcomes. In contrast, baseline ADHD symptom ratings showed small and domain-specific associations with treatment-related outcomes. These findings suggest that validity indicators may be more relevant to proximal clinical processes than to the prediction of longer-term treatment outcomes in adult ADHD. Future research should examine how validity indicators influence decision-making and treatment processes in adult ADHD assessment, using longitudinal and multi-method designs with more sensitive outcome measures.
The Ontario Psychological Association (OPA) has released the 2025 Mild Traumatic Brain Injury (mTBI) Guideline, a comprehensive update to the original 2016 OPA mTBI/concussion Guideline. This Guideline will refer to mTBI and will place “concussion” after a forward slash to acknowledge the common usage of the term concussion in both clinical and public discussions when discussing mTBI—“mTBI/concussion.” However, despite concussion being widely recognized and used, this Guideline will lead with mTBI to align with diagnostic fidelity and clarity. Guidelines serve an essential role in clinical practice, providing evidence-based recommendations that help standardize care, promote best practices, and ensure that diagnostic and treatment decisions are informed by the latest scientific knowledge. In the rapidly evolving scientific field of mTBI/concussion, updated guidelines are particularly critical to maintaining the highest standard of care and preventing diagnostic errors. The present Guideline is designed specifically to support Neuropsychologists in the assessment and diagnosis of mTBI/concussion. While Clinical Psychologists play a key role in identifying possible injuries and evaluating and treating the psychological and cognitive symptoms associated with mTBI/concussion, only Neuropsychologists, in addition to medical doctors and nurse practitioners, can make a formal diagnosis of mTBI/concussion in Ontario. This guideline reflects the current regulatory and professional practice framework in Ontario and may not apply to other jurisdictions, where scopes of practice and diagnostic authority may differ. A detailed discussion of jurisdiction-specific regulatory frameworks is beyond the scope of the present Guideline. Foundational to this Guideline is the integration of the 2023 American Congress of Rehabilitation Medicine (ACRM) Diagnostic Criteria for mTBI (Silverberg et al., 2023), the current gold standard for the evaluation and communication of an mTBI/concussion diagnosis. The ACRM forms the diagnostic basis from which this Guideline aims to supplement clinical practice. The ACRM diagnostic criteria is as follows: • Criterion 1: The Facts of the Plausible Mechanism of Injury • Criterion 2: Clinical Signs (one or more) • Criterion 3: Acute Symptoms (two or more) • Criterion 4: Clinical Examination and Laboratory Findings (one or more) • Criterion 5: Neuroimaging Abnormality (if completed) • Criterion 6: Not Better Accounted for by Confounding Factors By aligning with this internationally recognized standard, the present Guideline seeks to ensure that the diagnosis of mTBI/concussion is rooted in firm, evidence-based science, while addressing mis-interpretive pitfalls that may compromise diagnostic accuracy and individual outcomes.
In an increasingly diverse world, assessors are likely to encounter examinees who are not proficient in the language of standard service delivery. For such occasions, some tests provide versions translated into multiple languages. Establishing the psychometric and cultural equivalence of scores on the original versus translated versions requires further empirical investigation. We report the results of a preliminary study on the effect of language-of-administration on three rapid-assessment instruments designed to measure emotional functioning in a small bilingual student sample. The PHQ-9, GAD-7 and the V-8 (visual analog scale) were administered both in English and Spanish to 47 Cuban university students, in counterbalanced order. The classification accuracy of the V-8 was calculated against the two legacy measures as criteria. There was no difference as a function of the language of administration. The Depression and the Anxiety scales of the V-8 were significant predictors of the PHQ-9 and GAD-7 during the English administration. On the Spanish version, the V-8 Depression was a stronger predictor of the PHQ-9; the V-8 Anxiety was a non-significant predictor of the GAD-7. Twice as many participants scored ≥ 10 on the PHQ-9 (31.9
Giromini et al. (2026) provide a rigorous framework for Symptom Validity Test (SVT) research. This commentary extends that framework to an interpretive problem in severe, longstanding posttraumatic stress disorder (PTSD). Elkana et al. (2026) used a design closely aligned with Giromini et al.‘s clinical-honest versus nonclinical-feigner paradigm. At the Israeli-recommended Structured Inventory of Malingered Symptomatology (SIMS) cutoff (> 20), sensitivity for instructed feigning was 85
This simulation study sought to investigate a response style marked by simultaneous over- and under-reporting (OR + UR) in a criminal forensic evaluation context. Using a known groups design, Whitman et al. (2023) previously examined the impact of simultaneous OR + UR in a sample of civil disability claimants, finding it to be associated with higher scores on symptom validity tests, a higher likelihood of producing scores indicative of noncredible responding on performance validity tests, higher scores on measures of psychopathology (except those related to externalizing dimensions, which were lower), and lower scores on performance-based cognitive tests when compared with claimants who engaged in over-reporting only (OR-only). The current study utilized a simulation design to investigate OR + UR relative to an OR-only response style and genuine responding using the MMPI-3 in a simulated criminal forensic context. College student participants completed a battery of measures covering a broad range of psychopathology following standard instructions (SI). They then completed the MMPI-3 under one of three instructional manipulations: SI, OR-only, or OR + UR. The OR + UR group produced higher OR scale scores than the SI group and higher L scores than the SI and OR-only groups. Classification accuracy estimates supported use of the MMPI-3 Validity Scales for identifying OR-only and OR + UR response styles, particularly with respect to specificity. Finally, both the OR + UR and the OR-only response manipulations led to attenuation of validity coefficients for the MMPI-3 substantive scales.
Background: Medical expert witness assessments (MEWAs) evaluate case validity on the basis of both psychometric and non-psychometric modalities. Although symptom validity tests (SVTs) and performance validity tests (PVTs) are widely used, many thresholds were developed in analogue and known-groups validation contexts, and their transferability to MEWAs remains uncertain. Prior research validated a Criteria-Based Validity Assessment (CVA) as a multimodal framework for assessing case plausibility, but it remained unclear whether CVA and psychometric thresholds perform similarly across different legal contexts, such as pension insurance versus accident insurance evaluations. Objective: This study aimed to investigate differences in CVA, SVT, and PVT results across pension and accident insurance cases. Methods: A total of 721 MEWAs (572 pension; 149 accident) were analyzed. CVA criteria were rated by trained raters, and response biases were assessed psychometrically using the SIMS and the ASTM. Two-component beta-binomial mixture models were applied to derive CVA plausibility thresholds. Results: Mixture modeling reliably identified a distinct bimodal distribution of conspicuous CVA criteria counts, interpreted as plausible and implausible subgroups in both pension and accident cases, with ≥4 conspicuous CVA criteria representing the optimal plausibility threshold. Pension claimants showed significantly higher SIMS scores and lower ASTM scores than accident claimants, independent of case plausibility. Conclusion: CVA can be used to assess validity information collected from multiple data sources (longitudinal and cross-sectional) and to assess the validity of health-related claims across different legal settings. CVA, however, was not a tool for measuring context-specific response biases. Cases with ≤2 conspicuous CVA criteria may be valid, cases with three conspicuous criteria require individual examination because feigning is possible but not certain, and ≥4 conspicuous criteria supported a conservative classification of CVA-defined case implausibility.
The assessment of symptom validity plays a major role in most forensic and non-forensic psychological evaluations, and symptom validity tests (SVTs) are commonly used instruments in these evaluations. As SVT research has expanded dramatically in recent years, the time has come to reflect on and further refine the methodological standards guiding this growing body of work. In this article, we review current practices, propose refinements and future directions, and invite further scholarly dialogue on which lines of SVT research are most urgently needed and how such studies should ideally be designed. First, we describe the current status of SVT research and highlight several key open questions. In particular, we note that, compared to the substantial progress made in the field of performance validity testing – where numerous well-validated measures are available and clear professional consensus has emerged regarding best practices for integrating results from multiple tests – far fewer fully validated SVTs currently exist, and much less is known about how to combine results from multiple SVTs. We then discuss the strengths and limitations of commonly used research methods for evaluating the psychometric properties of existing SVTs and underscore the need for clearer guidelines for criterion-group studies, for which we therefore offer some initial recommendations. More specifically, for criterion-group studies we suggest that (a) validity classifications should be based on multiple, rather than single, validity indicators; (b) in the absence of a universally accepted gold standard, SVTs with the strongest empirical foundation should be preferred as criterion variables; and (c) optimal cut scores for these criterion variables should draw on the most recent meta-analytic or research survey findings rather than relying solely on test manuals. Ultimately, our goal is to stimulate scholarly contributions and advance the development of a coherent methodological framework to guide future research on symptom validity testing, and we invite comments, critiques, and proposals to support this collaborative effort.
This study provides proof-of-concept for the solution to the invalid before clinically elevated paradox introduced by Schneider et al. (2026). As they predicted, non-credible responding was a significant confound for self-reported depression and anxiety. Identifying and removing invalid response sets from normative samples seems necessary to establish valid clinical cutoffs for determining the presence/absence and severity rating of psychiatric symptoms. Results suggest that the invalid before impaired/clinically elevated paradox may be (at least partly) an artifact of contaminated norms (i.e., failure to exclude non-credible response sets). Data were analyzed from 73 university students who volunteered for academic research and passed a free-standing symptom validity test (SVT). Participants were administered the PHQ-9, GAD-7 and the V-8, a visual analog scale of Depression and Anxiety. The classification accuracy of the V-8 was calculated using the PHQ-9 and GAD-7 as criterion measures. A V-8 Depression score ≥ 40 was specific (0.90 − 1.00) to severe depression. If this cutoff is used to redefine clinical elevation, it opens up a wide range of credible and clinically significant depression (40–69) before the SVT cutoff (≥ 70) deems the response set invalid. A V-8 Anxiety score ≥ 60 at Time 1 had comparable specificity (0.85-0.96), allowing for a range of credible and clinically significant anxiety (60–79) before crossing the SVT cutoff (≥ 80) invalidates the response set. At Time 2, the clinical cutoff had to be lowered to ≥ 50 (0.86-0.96 specificity) to make room for credible and clinically significant anxiety (50–64) before the SVT cutoff (≥ 65) is activated. This study provides proof-of-concept for the solution to the invalid before clinically elevated paradox introduced by Schneider et al. (2026). As they predicted, non-credible responding was a significant confound for self-reported depression and anxiety. Identifying and removing invalid response sets from normative samples seems necessary to establish valid clinical cutoffs for determining the presence/absence and severity rating of psychiatric symptoms. Results suggest that the invalid before impaired/clinically elevated paradox may be (at least partly) an artifact of contaminated norms (i.e., failure to exclude non-credible response sets).
To evaluate the diagnostic utility of the Structured Inventory of Malingered Symptomatology (SIMS) for differentiating instructed feigned PTSD, from a Ministry of Defense (MoD)-recognized PTSD group (Israel), in comparison with healthy controls, and to examine whether PTSD symptom severity and neuroticism contribute to SIMS misclassification. Participants (N = 112) were Israeli men and included a MoD-recognized PTSD group (n = 38), healthy controls (n = 34) and instructed feigning group (n = 40). Measures included the SIMS, Test of Memory Malingering (TOMM), PTSD Checklist for DSM-5 (PCL-5), and Big Five Inventory (BFI). Using the Israeli-recommended SIMS cutoff (> 20), sensitivity for detecting instructed feigning was 85
Accurate detection of noncredible symptom reporting is critical for procedural fairness, the integrity of forensic psychological assessment, and informed legal decision-making. The Symptom Validity Test–Thai (SVT-Th), originally developed in 2019 and comprising 57 items, warrants revision to strengthen its conceptual and psychometric foundations. In Study 1, we revised and evaluated the SVT-Th-Revised using a large field dataset and derived a provisional cutoff via a forensic simulation design. The original scales were mapped onto two detection strategies: the Unlikely Detection Strategy (UDS) and the Amplified Detection Strategy (ADS). Confirmatory factor analysis (n = 1,173) supported a hierarchical structure, yielding a 49-item version with excellent model fit. The model represents noncredible symptom reporting as a higher-order construct indexed by UDS and ADS, which may support profile-based interpretation in forensic and disability-related assessments (FDRA). In Study 2, we used a newly collected simulation sample (n = 233); ROC analysis indicated outstanding classification accuracy under simulation conditions (AUC = 0.99), with an optimal cutoff of ≥ 58 yielding 99.20
Premorbid intellectual function estimation is essential in neuropsychological evaluation to differentiate acquired neurocognitive decline from baseline functioning. Although confounding variables may impact performance-based premorbid estimates, little is known about the impact of invalid test performance on such estimates. This cross-sectional study examined the impact of performance invalidity on performance based vs. demographics-based premorbid estimation using the Test of Premorbid Functioning (TOPF). Using four freestanding performance validity tests (PVTs), 489 adults referred for comprehensive neuropsychological evaluation were divided into valid (≤ 1 failure) or invalid (≥ 2 failures) groups. Three premorbid estimates were then calculated for each patient from the TOPF (i.e., Reading-Only, Demographics-Only, and Combined). IQ estimates from all 3 TOPF models were significantly lower in the invalid group compared to the valid, with the largest effect size for the Reading-Only model and the smallest for the Demographics-Only model. Supplementary analyses revealed negligible within-group effects in the valid group and medium effects in the invalid group. TOPF prediction models including the performance-based irregular word reading score (Reading-Only and Combined Models) were significantly impacted by invalid performance. Although considered a hold test, irregular word reading scores are susceptible to invalid performance. The Demographics-Only model should be utilized in such cases.
Within-person neuroanatomical changes and neurocognitive decline in Mild Cognitive Impairment (MCI) and Alzheimer’s disease (AD) differ from those observed in healthy older adults. Given the overlapping clinical presentations of MCI and mild AD, along with their differences, tailored and consecutive neuropsychological assessments can be required to determine the financial capacity of those with MCI and mild AD. The current literature on financial capacity lacks a comprehensive conceptual framework, and assessment tools often possess inadequate psychometric properties, focusing primarily on cognitive dysfunctions to determine financial capacity. Moreover, few longitudinal studies have been published to date, and the diversity of various cultural groups is underrepresented. Hence, we aim to provide a review of financial capacity assessment in MCI and mild AD, considering neuroanatomical, neuropsychiatric, and neuropsychological correlates, while defining and differentiating financial capacity from competency to identify limitations and gaps in the literature. Furthermore, we highlight clinical and methodological issues: (A) in evaluating financial capacity; (B) in recognizing cognitive markers of risk for financial exploitation; and (C) in assessment methods, from screening to comprehensive neuropsychological assessment. We also provide a summary of assessment scales used and practical recommendations for practice and future research direction.
Factitious Disorder Imposed on Another (FDIA), historically known as Munchausen by Proxy (MBP), is a psychologically and physically injurious form of child maltreatment that comes before the courts, raising questions about harm, safety, legal responsibility, and risk management. In 2019, Sanders and Bursch introduced ACCEPTS, a management and treatment protocol for addressing the injurious circumstances brought about by FDIA. Given the legal complexities of addressing this injurious condition, we review diagnostic considerations, assessment strategies, and treatment approaches for FDIA. We also explore ACCEPTS within broader ethical, clinical, and legal frameworks. Much of the research regarding ACCEPTS notes that acknowledgement of abuse is needed for successful treatment. Therefore, particular attention is given to the role of caregiver acknowledgement of injurious conduct, especially in court-involved contexts, and the clinical, ethical, and legal complexities that may arise when acknowledgement is incorporated into treatment. This paper offers recommendations to support ethically-grounded, developmentally-informed, and clinically flexible use of ACCEPTS, alongside directions for future research and policy development.
Leonhard and Leonhard (2025) argue that Symptom Validity Tests (SVTs) and Performance Validity Tests (PVTs) constitute a form of junk science. This qualification stands in sharp contrast to the breadth and depth of the scientific work on validity tests. Leonhard and Leonhard treat these tests as if they were equivalent to polygraph evidence, a move that betrays a fundamental misunderstanding of the conceptual and empirical foundations of these instruments. More than a year ago, we invited Leonhard and Leonhard to provide case law examples —if only a few— demonstrating that SVTs and/or PVTs contributed to risky legal decisions. So far, they have not been willing or able to cite a single instance. We therefore reiterate our invitation: show us the cases.
A growing body of research has demonstrated variable performance on validity assessment measures among individuals from culturally and linguistically diverse backgrounds. This experimental feigning study examined the classification accuracy of the revised Test of Memory Malingering (TOMM 2) and the Dot Counting Test (DCT). Participants were recruited in Mexico City and included a community control group (n = 49), a community group instructed to feign psychosis (n = 33), and a clinical control group of individuals receiving inpatient and outpatient mental health services (n = 27). Results revealed significant differences in TOMM 2 performance, with the experimental feigning group scoring lower than the community and clinical control samples. While the TOMM 2 < 45 cutoff demonstrated adequate sensitivity (Trial 2 = 66.7
This exploratory study examined the psychometric performances of widely used validity measures in interpreter-mediated, cross-cultural forensic contexts, a critical but underexplored area. The Test of Memory Malingering (TOMM), Morel Emotional Numbing Test (MENT), Miller Forensic Assessment of Symptoms Test (M-FAST), and embedded indices from the PTSD Checklist for DSM-5 (PCL-5) were examined in this linguistically diverse forensic sample. Archival data from 59 consecutively evaluated claimants, assessed for posttraumatic stress disorder under the Defense Base Act, were analyzed. Assessments were conducted in native languages spanning Indo-European, Niger-Congo, and Dravidian language families, with certified interpreters using quality assurance procedures. Descriptive statistics, non-parametric group comparisons, and correlation analyses were used to explore test performances, demographic influences, and inter-test concordance. TOMM and MENT demonstrated strong internal consistency, high concordance, and minimal demographic effects. The M-FAST showed variable subscale performance and was influenced by geographic origin and language, suggesting potential cultural or linguistic effects. PCL-5 embedded validity indices exhibited strong internal consistency but a lower base rate of failure, raising questions about sensitivity. Concordance was highest within performance validity tests (77
This study was designed to examine the potential of the clinical scales of the V-5 (a visual analog scale) to serve as embedded symptom validity tests (SVTs) in a mixed clinical sample. Archival data were collected from a consecutive case sequence of 100 adult outpatients physician-referred for neuropsychological evaluation. The V-5 was administered twice: at the beginning (Time 1) and the end of the battery (Time 2). The classification accuracy of the V-5 as an SVT was calculated against psychometrically operationalized non-credible symptom report and invalid performance on cognitive testing. Using SVTs as criterion measures, ≥70 on the Depression scale had .41-.67 sensitivity at .89-.94 specificity. Different cutoffs were needed on the Anxiety scale at Time 1 (≥80; .44-.63 sensitivity at .88-.90 specificity) and Time 2 (≥65; .44-.57 sensitivity at .88-.93 specificity). A score ≥60 on the Pain scale had low sensitivity (.20-.41) but high specificity (.87-.94). Aggregating the number of failures on the V-5 within and across administrations consolidated specificity (.90-.99) at a proportional cost to sensitivity (.22-.57), correctly classifying between 84
The application of trauma-informed principles to forensic mental health assessment has recently gained scholarly attention (Goldenson et al., 2022; Goldenson, 2025b). Central to this framework is the expectation that forensic mental health professionals possess a nuanced understanding of the prevalence and psychological impact of trauma among evaluees, and that they apply this knowledge to adapt assessment procedures in ways that promote dignity, autonomy, and fairness while also enhancing the quality and accuracy of the data collected. A persistent challenge in both criminal and civil contexts is the potential for feigned posttraumatic stress disorder. One underexplored ethical tension involves balancing the commitment to psychological safety, a core tenet of trauma-informed practice, with the use of assessment strategies designed to detect noncredible responding. These strategies may be experienced by evaluees as confrontational, invalidating, or deceptive. This article explores this tension and examines how forensic evaluators can best uphold trauma-informed principles while conducting scientifically rigorous evaluations, including those involving suspected feigning. The importance of a multi-source, multi-method approach is emphasized. When data appear inconsistent, evaluators should engage in critical thinking and contextual analysis rather than defaulting to assumptions of deception.
Medical Expert Witness Assessments (MEWA) are international standard to generate information involving medical questions in litigations, such as the capacity to work. While this procedure is widely utilized and guidelines for assessing the validity of symptoms using psychometric tools and non-psychometric criteria have been developed, the scientific foundation of this multimodal Criteria-Based Validity Assessment (CVA) is weak. This study aims to provide empirical validation of CVA using psychometric Symptom (SVT) and Performance Validity Tests (PVT) as areference point. 466 MEWA conducted in the law of the German Statutory Pension Insurance (GPI), all uniformly having addressed the question of the capacity to work, were analyzed. Information about scores regarding the Structured Inventory of Malingered Symptomatology (SIMS), Amsterdam Short-Term Memory Test (ASTM) aswell as the seven CVA criteria were extracted. A logistic regression using CVA data to group the MEWA into plausible and implausible (over- and/or under reporting of symptoms) cases showed a significant association between implausible cases and SIMS scores (OR= 1. 067; 95