Objective: Novel lipid-lowering therapies are being introduced. Few studies exist of the real-world effectiveness of adenosine-tri-phosphate citrate lyase inhibition with bempedoic acid. Methods: This study audited bempedoic acid therapy in 216 consecutive patients from three hospital centres - a university hospital (n = 77) and two district general hospitals (n = 106 and 33). Cardiovascular disease (CVD) risk factors, prescription qualification criteria, efficacy and adverse effects were assessed. Results: The population was aged 65.9 +/- 11.0 years, 42% were male, 25% had type 2 diabetes, and 31% had familial hypercholesterolaemia. CVD was present in 19% and multibed vascular disease in 8%. Statin intolerance was reported in 92%. Bempedoic acid reduced total cholesterol by 1.58 +/- 1.44 mmol/L (20%), LDL-C by 1.37 +/- 1.31 mmol/L (27%), triglycerides by 0.22 mmol/L (2%) with an 0.06 mmol/L (1%) increase in HDL-C after 22 +/- 9 months follow-up. An LDL-C <2.5 mmol/L was achieved in 40% and <2 mmol/L in 20%. Efficacy (r2 = .33) was predicted by baseline LDL-C (beta = .54; p <.001). No significant changes were seen in transaminases, creatinine, creatine kinase, urate or HbA1c. Treatment was discontinued by 33% of patients and occurred due to myalgia (43%), lack of efficacy (16%) and gastrointestinal adverse effects (15%). No cases of gout were observed. In a logistic regression only the number of previous drug classes not tolerated (beta = 1.60; p = .009) was a contributing factor to discontinuation. Conclusion: This audit suggests that bempedoic acid therapy is effective but that adverse effects and discontinuation are common. This suggests nocebo effects might be generalizable to all lipid-lowering drug therapies in susceptible individuals.
Background: Retinopathy of Prematurity (ROP) is a condition confined to the premature retina. Hence, it is more commonly associated with extreme prematurity and low birth weight.
Urinary tract infection (UTI) is one of the common infections in childhood. Prompt diagnosis and treatment reduces the risk of complications. The choice of antibiotic to treat UTI varies from region to region. Rational use and appropriately chosen antibiotic reduces the emergence of resistant uropathogens. We investigated the resistance pattern of uropathogens for commonly used antibiotics to treat UTI locally. Data was collected between 2009 and 2019 on all infants and children under 16 years of age with culture proven UTI. Results were compared with previously published figures between 2002 and 2008. A total of 1002 samples were analysed (91/year). Male to female ratio was 1:4.6. About 94% of the samples grew E. coli. As before, high resistance rates were recorded to Amoxicillin and Trimethoprim (Z = −0.325: P = 0.7452; not significant). Overall, average resistance has decreased for Nitrofurantoin from 10% between 2002 and 2008 to 5.84% between 2009 and 2019 (Z = 3.002: P = 0.0027). On the other hand, Cefalexin resistance has increased from 7.4 to 14.56% between the two study periods (Z = −4.2: P = < 0.0002). Despite rising resistance rates, we recommend that Cefalexin should cautiously remain the antibiotic of choice for empirically treating uncomplicated urinary tract infections in secondary care pending urine culture. Nitrofurantoin should be reserved for treating non-coliform/atypical UTIs or multi-drug resistant UTIs. There is an ongoing need for clinicians in all geographic regions to continue to monitor antibiotic resistance rates every few years.
AIMS:Lysosomal β-glucocerebrosidase A (GBA) deficiency causes Gaucher disease (GD), a recessive disorder caused by bi-allelic mutations in GBA. The prevalence of GD is associated with ethnicity but largely unknown and potentially underestimated in many countries. GD may manifest with organomegaly, bone involvement, and neurological symptoms as well as abnormal laboratory biomarkers. This study attempted to screen for GD in patients using abnormal platelet, alkaline phosphatase (ALP), and ferritin results from laboratory databases. METHODS:Electronic laboratory databases were interrogated using a 2- to 4-year time interval to identify from clinical biochemistry records patients with a phenotype of reduced platelets (<150 × 109 /L) and either elevated ALP (>130 iu/L) or ferritin [>150 (female) or >250 µg/L (male)]. The mean value over the screening window was used to reduce variability in results. A dried blood spot sample was collected for the determination of GBA activity in patients meeting these criteria. If low GBA activity was found, then the concentration of the GD-specific biomarker glucosyl-sphingosine (lyso-GB1) was assayed, and the GBA gene sequenced. RESULTS:Samples were obtained from 1058 patients; 232 patients had low GBA activity triggering further analysis. No new cases of GD with homozygosity for pathogenic variants were identified, but 12 patients (1%) were identified to be carriers of a pathogenic variant in GBA. CONCLUSIONS:Pathology databases hold routine information that can be used to screen for patients with inherited errors of metabolism. However, biochemical screening using mean platelets, ALP, and ferritin has a low yield for unidentified cases of GD.
To collect and review data from consecutive patients admitted to Queen’s Hospital, Burton on Trent for treatment of Covid‐19 infection, with the aim of developing a predictive algorithm that can help identify those patients likely to survive.
The National Institute on Drug Abuse and Joint Institute for Biological Sciences at the Oak Ridge National Laboratory hosted a meeting attended by a diverse group of scientists with expertise in substance use disorders (SUDs), computational biology, and FAIR (Findability, Accessibility, Interoperability, and Reusability) data sharing. The meeting's objective was to discuss and evaluate better strategies to integrate genetic, epigenetic, and 'omics data across human and model organisms to achieve deeper mechanistic insight into SUDs. Specific topics were to (a) evaluate the current state of substance use genetics and genomics research and fundamental gaps, (b) identify opportunities and challenges of integration and sharing across species and data types, (c) identify current tools and resources for integration of genetic, epigenetic, and phenotypic data, (d) discuss steps and impediment related to data integration, and (e) outline future steps to support more effective collaboration-particularly between animal model research communities and human genetics and clinical research teams. This review summarizes key facets of this catalytic discussion with a focus on new opportunities and gaps in resources and knowledge on SUDs.
Disease diagnosis and treatment is challenging in part due to the misalignment of diagnostic categories with the underlying biology of disease. The evaluation of large-scale genomic experimental datasets is a compelling approach to refining the classification of biological concepts, such as disease. Well-established approaches, some of which rely on information theory or network analysis, quantitatively assess relationships among biological entities using gene annotations, structured vocabularies, and curated data sources. However, the gene annotations used in these evaluations are often sparse, potentially biased due to uneven study and representation in the literature, and constrained to the single species from which they were derived. In order to overcome these deficiencies inherent in the structure and sparsity of these annotated datasets, we developed a novel Network Enhanced Similarity Search (NESS) tool which takes advantage of multi-species networks of heterogeneous data to bridge sparsely populated datasets. NESS employs a random walk with restart algorithm across harmonized multi-species data, effectively compensating for sparsely populated and noisy genomic studies. We further demonstrate that it is highly resistant to spurious or sparse datasets and generates significantly better recapitulation of ground truth biological pathways than other similarity metrics alone. Furthermore, since NESS has been deployed as an embedded tool in the GeneWeaver environment, it can rapidly take advantage of curated multi-species networks to provide informative assertions of relatedness of any pair of biological entities or concepts, e.g., gene-gene, gene-disease, or phenotype-disease associations. NESS ultimately enables multi-species analysis applications to leverage model organism data to overcome the challenge of data sparsity in the study of human disease. Availability and Implementation Implementation available at https://geneweaver.org/ness . Source code freely available at https://github.com/treynr/ness . Author summary Finding consensus among large-scale genomic datasets is an ongoing challenge in the biomedical sciences. Harmonizing and analyzing such data is important because it allows researchers to mitigate the idiosyncrasies of experimental systems, alleviate study biases, and augment sparse datasets. Additionally, it allows researchers to utilize animal model studies and cross-species experiments to better understand biological function in health and disease. Here we provide a tool for integrating and analyzing heterogeneous functional genomics data using a graph-based model. We show how this type of analysis can be used to identify similar relationships among biological entities such as genes, processes, and disease through shared genomic associations. Our results indicate this approach is effective at reducing biases caused by sparse and noisy datasets. We show how this type of analysis can be used to aid the classification gene function and prioritization of genes involved in substance use disorders. In addition, our analysis reveals genes and biological pathways with shared association to multiple, co-occurring substance use disorders.
AIMS:Lysosomal α-galactosidase A deficiency (Fabry disease (FD)) was considered an X-linked recessive disorder but is now viewed as a variable penetrance dominant trait. The prevalence of FD is 1 in 40 000-117 000 but the ascertainment of late-onset cases and degree of female penetrance makes this unclear. Its prevalence in the general population, especially in patients with abnormal renal function is unclear. This study attempted to identify the prevalence of FD in patients with abnormal renal function results from laboratory databases.METHODS:Electronic laboratory databases were interrogated to identify from clinical biochemistry records patients with a phenotype of reduced estimated glomerular filtration rate categorised by age on one occasion or more over a 3-year time interval. Patients were recalled and a dried blood spot sample was collected for the determination of α-galactosidase A activity by fluorimetric enzyme assay in men and mass spectrometry assays of α-galactosidase A and lyso-globotriaosylceramide (lyso-GL-3) concentrations in women.RESULTS:Samples were obtained from 1084 patients identified with reduced renal function. No cases of FD were identified in 505 men. From 579 women, one subject with reduced α-galactosidase activity (1.5 µmol/L/h) and increased Lyso-GL-3 (5.5 ng/mL) was identified and shown to be heterozygous for a likely FD pathogenic variant (GLA c.898C>T; p.L300F; Leu300Phe). It was later confirmed that she was a relative of a known affected patient.CONCLUSIONS:Pathology databases hold routine information that can be used to identify patients with inherited errors of metabolism. Biochemical screening using reduced eGFR alone has a low yield for unidentified cases of Fabry Disease.
Genome-wide association studies and other discovery genetics methods provide a means to identify previously unknown biological mechanisms underlying behavioral disorders that may point to new therapeutic avenues, augment diagnostic tools, and yield a deeper understanding of the biology of psychiatric conditions. Recent advances in psychiatric genetics have been made possible through large-scale collaborative efforts. These studies have begun to unearth many novel genetic variants associated with psychiatric disorders and behavioral traits in human populations. Significant challenges remain in characterizing the resulting disease-associated genetic variants and prioritizing functional follow-up to make them useful for mechanistic understanding and development of therapeutics. Model organism research has generated extensive genomic data that can provide insight into the neurobiological mechanisms of variant action, but a cohesive effort must be made to establish which aspects of the biological modulation of behavioral traits are evolutionarily conserved across species. Scalable computing, new data integration strategies, and advanced analysis methods outlined in this review provide a framework to efficiently harness model organism data in support of clinically relevant psychiatric phenotypes.
Aims Adult-onset inherited errors of metabolism can be difficult to diagnose. Some cases of potentially treatable myopathy are caused by autosomal recessive acid α-1,4 glucosidase (acid maltase) deficiency (Pompé disease). This study investigated whether screening of asymptomatic patients with elevated creatine kinase (CK) could improve detection of Pompé disease. Methods Pathology databases in six hospitals were used to identify patients with elevated CK results (>2× upper limit of normal). Patients were recalled for measurement of acid α-1,4 glucosidase activity in dried blood spot samples. Results Samples were obtained from 812 patients with elevated CK. Low α-glucosidase activity was found in 13 patients (1.6%). Patients with neutropaenia (n=4) or who declined further testing (n=1) were excluded. Confirmation plasma specimens were obtained from eight individuals (1%) for a white cell lysosomal enzyme panel, and three (0.4%) were confirmed to have low α-1,4-glucosidase activity. One patient was identified as a heterozygous carrier of an acid α-1,4 glucosidase c.-32–13 G>T mutation. Screening also identified one patient who was found to have undiagnosed Fabry disease and one patient with McArdle’s disease. One patient later presented with Pompé’s after an acute illness. Including the latent case, the frequency of cases at 0.12% was lower than the 2.5% found in studies of patients with raised CK from neurology clinics (p<0.001). Conclusions Screening pathology databases for elevated CK may identify patients with inherited metabolic errors affecting muscle metabolism. However, the frequency of Pompé’s disease identified from laboratory populations was less than that in patients referred for neurological investigation.
Background Understanding mechanisms underlying specific chemotherapeutic responses in subtypes of cancer may improve identification of treatment strategies most likely to benefit particular patients. For example, triple-negative breast cancer (TNBC) patients have variable response to the chemotherapeutic agent cisplatin. Understanding the basis of treatment response in cancer subtypes will lead to more informed decisions about selection of treatment strategies. Methods In this study we used an integrative functional genomics approach to investigate the molecular mechanisms underlying known cisplatin-response differences among subtypes of TNBC. To identify changes in gene expression that could explain mechanisms of resistance, we examined 102 evolutionarily conserved cisplatin-associated genes, evaluating their differential expression in the cisplatin-sensitive, basal-like 1 (BL1) and basal-like 2 (BL2) subtypes, and the two cisplatin-resistant, luminal androgen receptor (LAR) and mesenchymal (M) subtypes of TNBC. Results We found 20 genes that were differentially expressed in at least one subtype. Fifteen of the 20 genes are associated with cell death and are distributed among all TNBC subtypes. The less cisplatin-responsive LAR and M TNBC subtypes show different regulation of 13 genes compared to the more sensitive BL1 and BL2 subtypes. These 13 genes identify a variety of cisplatin-resistance mechanisms including increased transport and detoxification of cisplatin, and mis-regulation of the epithelial to mesenchymal transition. Conclusions We identified gene signatures in resistant TNBC subtypes indicative of mechanisms of cisplatin. Our results indicate that response to cisplatin in TNBC has a complex foundation based on impact of treatment on distinct cellular pathways. We find that examination of expression data in the context of heterogeneous data such as drug-gene interactions leads to a better understanding of mechanisms at work in cancer therapy response.
'What I cannot create (and control), I do not understand' (Richard Feynman; modified Bertolero & Bassett, 20191) As anyone with an interest in the works of JK Rowling knows nifflers like shiny treasure and go to extreme lengths to find it. Cardiovascular disease (CVD) physicians are similar in their wish to find deposits of atherosclerosis but are far less accomplished at it. Atherosclerosis is a cryptic disease starting in the vascular wall and only later manifesting within the artery lumen with long-term consequences in the form of plaque rupture or erosion (type 1 lesions) but also vasospasm through secondary endothelial dysfunction (type 2 disease).2 Detecting atherosclerosis is possible using imaging either thorough the detection of early lesions on ultrasound or in the vessel wall (intima-media thickness) and late-stage calcified plaques (coronary artery calcium).3 The most sophisticated approach is to image atheroma in the wall either in large arteries by magnetic resonance imaging or in coronary arteries by intravascular ultrasound on angiography. Further developments now include three-dimensional imaging techniques applying computerised image reconstruction techniques. However, all of these direct approaches are limited in their application by the expense and size of the machinery required and the logistics of managing patient flows to central sites. Instead a cheaper and easier approach is pursued by all health systems. The availability of large epidemiological databases and cohort studies now extending in some cases to up to three generations (Framingham)4 means that high-risk individuals can be identified easily from common parameters. These studies maintain assay standardisation which may not apply to electronic health records (EHRs) linked to standard laboratory assays which evolve with time.5 Landmark analyses starting in 1987 identified certain key CVD risk factors and remarkably quickly these were standardised as age, gender, smoking, blood pressure, diabetes and cholesterol (later divided into total and high-density lipoprotein (HDL) cholesterol).4, 6, 7 A multitude of additional CVD risk factors have since been described but all of these added little to the basic predictive model which is mostly driven by age, gender and ethnicity.8 Risk factor counting and set intervention levels were the basis of defining high-risk patients for intervention. These still persist in modern guidelines, for example, stage 2 hypertension or total cholesterol > 7.5mmol/L and more usefully the concept of two CVD risk factors predicting lifetime risk from age 55.9 Yet these crude cut-offs had the significant limitation that they only identified a small fraction of patients at risk of CVD—that is, high specificity but limited sensitivity. The next development in the 1990s was the beginning of the use of mathematical models based on logistic regression analyses of epidemiological datasets once semiconductor-based scientific notation calculators became available. These could be simplified into paper-based systems or mechanical tools for routine clinical use.10, 11 Now that substantial computing capacity is available through cell phones or internet-based systems these are now universally recommended for assessment of patients with a risk of CVD. The desire to increase convenience has now led to the wish to simplify the process further by abolishing the most logistically difficult (and expensive) aspect which comprises the cholesterol blood tests. In fact, the Framingham risk engine can be easily reformatted by substituting body mass index for lipids but surprisingly this has not achieved great popularity for initial risk stratification despite its simplicity.12 The main quest in CVD risk estimation, however, has been to improve sensitivity and specificity. The main methods used have been to use larger more representative datasets based either on aggregating epidemiological cohort studies (eg US atherosclerotic CVD score- ASCVD13) or national EHRs (eg QRISK in the UK14). The best predictive performance of epidemiological datasets is an average area under curve (AUC) for receiver operator characteristic (ROC) curves (ie C-statistic) of approximately 0.70-0.75.11 Adding imaging data from coronary artery calcium increases this to 0.79 with less benefit from the far more convenient ultrasound techniques or biomarkers such as high-sensitivity troponin measurements.15, 16 The next great hope is to exploit the developments in electronic databases and advances in computing. Models to date have relied on deterministic processes guided by humans yet the suspicion has remained that information may have been lost by these decisions so other statistical data interpretation techniques are being explored. Machine learning and Bayesian analysis are the current trendy concepts but many others exist. Neural nets, the best known form of machine learning, were first described in 1943 but it has taken 70 years to make them practical as they require large scale computing to make them practical.17, 18 The techniques of neural net analysis rely on large scale data inputs, intermediate layer (or layers) of nodes linked back to the data and forward to the outputs—in this case CVD events (Figure 1). Nodes are set randomly and then iterate and adjust input weights to optimise the prediction of the outputs.17, 18 Finally, as in classic epidemiological models the outputs are validated in another dataset. In contrast to classic calculators, neural nets are multilayer of which many aspects are obscured but if collapsed down to a single layer these can be isolated and described in classic terms. Whether this concept represents an electronic obscurial more than just a black box is the subject of debate. Until now the commonest application of neural nets in medicine has been in the analysis of images as these were data rich and the most problematic for classical methods.19 The problem in CVD for risk prediction has been the availability of large EHR datasets. This is now changing with the rapid computerisation of health systems. In this issue of International Journal of Clinical Practice, Quesada and colleagues describe multi-model analyses of an EHR comprising 38.527 patients with a 5-year follow-up and a likely 5%-10% CVD event rate as is typical in cohorts of this type.20 In their analysis quadratic discriminant analysis and Naïve Bayes ranked above (area under curve [AUC] 0.70) neural nets and classical logistic regression model -derived calculators such as the European Systematic COronary Risk Evaluation (SCORE; CVD mortality alone21) or the US Framingham study-related REGICOR score for CVD events (AUC = 0.63).22 Ten of 15 computer models were better than the classical methods but not by much. This is common in studies which attempt to improve the standard AUC of 0.65-0.75 found for classical CVD risk calculators in populations that match their original derivation and validation cohorts. This study lacked comparisons with recalibrated Spanish cohorts as opposed to generic models so it is unclear how much extra predictive capacity was actually added. Other studies have compared computerised models including neural nets with logistic regression models. One study using 689 patients from India, but using a validation population of 5209 US patients from the Framingham study, pre-specified classical risk factors and a quantum neural net approach suggested that this model was superior to the classical FRS.23 This is not surprising as the CVD risk factor weighting is different in Indian Asians from US populations. A similar criticism would apply to the Korean National Health and Nutrition Evaluation study (KNHANES-6) using 4244 EHR records with complete pre-specified six CVD risk factor data and a deep belief network (DBN) analysis and a restricted Boltzman Hopfield network that optimised to six nodes in one layer.24 The statistical DBN gave an AUC of 0.79 compared with 0.72 for logistic regression. This study did not assess their performance against classical or modified (ie recalibrated) CVD risk calculators. These have been investigated in Korean populations where in a study of 200 010 patients the ASCVD equation has an AUC of 0.73-0.75 but calibration errors with an excess 57%-74% in men and a deficit of 28% in women but was useful in enabling a Korean-specific CVD score to be derived.25 The neural net analysis of the Multi-Ethnic Study of Atherosclerosis (MESA) cohort of 6814 patients followed up for 12 years and 735 variables derived from biochemistry, questionnaires and imaging was used by random survival forest analysis to derive top 20 predictors for individual CVD outcomes.15 In this study, nine models were tested including Cox and LASSO-Cox models, and Aikake information criterion applied to regression analysis as well as random survival forest analysis. Predictably age was the most important predictor of mortality. Coronary artery calcium was the best predictor of coronary heart disease or CVD with glucose and carotid ultrasound for stroke. In contrast to usual expectations, troponin was the strongest predictive of heart failure while NTproB-type natriuretic peptide was the best predictor of CVD. A UK study used data from 378 256 primary care patients in the Clinical Practice Research Database (CPRD) and 24 970 recorded CVD events (6.6%) to compare various computational methods of CVD risk prediction.26 This study compared the US ASCVD score (not interestingly UK QRISK) with machine learning models. The standard ASCVD model had an AUC of 0.73, with the random forest model 0.75, logistic regression, gradient boosting or neural networks 0.76. The neural network algorithm predicted 4998 of 7404 cases (sensitivity 68%, positive predictive value (PPV) 18%) and 53 458 of 75 585 non-cases (specificity 72%, negative predictive value (NPV) 96%), predicting 8% more patients who developed CVD compared with the established ASCVD baseline model which predicted 53 106 non-cases from 75 585 non-cases, resulting in a specificity of 70% and NPV of 95%. As is true of all CVD risk prediction models because of their structure of containing many unaffected patients the greatest power is to rule out disease (negative predictive value). The small addition to risk prediction, which is better described in the form of net (or total) reclassification indices (NRI), is not unusual in this type of analyses.27 More recently a comparative study was conducted in a cohort of 109 490 individual using aggregated and longitudinal features from EHR involving analysis of historical and prospective phases.28 The models tested included logistic regression, random forests, gradient boosting trees, convolutional neural networks (CNN) and recurrent neural networks with long short-term memory (LSTM) units. A further analysis of 10 612 patients used late-fusion approach to incorporate genetic risk score data. The ASCVD equation achieved a typical ROC AUC of 0.73, while machine learning models using only classical CVD risk factors doing no better. Incorporation of EHR features mostly relating to the length of the EHR and variances in biochemical analytes achieved an AUC of 0.77-0.78. By adding temporal features, logistic regression (LR), gradient boosting trees (GBT) and deep learning models improved the AUC to 0.78-0.79. Both GBT and convolutional neural networks (CNN) achieved an AUC of 0.79 (ie 7.9% improvement from baseline). Most of the studies reviewed in this article use ROC curves which present graphically the trade-off between the true positive rate (TP) (sensitivity) and false positive (FP) (1-specificity) rate for a predictive model using different probability thresholds. In contrast Precision-Recall curves (PRC) and their graphical outputs summarise the trade-off between the true positive (TP) rate and the positive predictive value (PPV; precision) for a predictive model using different probability thresholds.29 Mathematically ROC curves are appropriate when the observations are balanced between each class, whereas PRCs are appropriate for unbalanced datasets as is commonly the case for epidemiological cohort datasets being used to predict events as only a minority develop CVD. In this study Area under PRC (AUPRC) analysis showed that machine learning using temporal features improved predictions founded on baseline data (0.25-0.29 vs 0.19, a 33%-44% improvement) more clearly than that for ROC curves. The top features in all machine learning models include some conventional CVD risk factors such as age, blood pressure (BP) and total cholesterol, as well as several new features not included in standard CVD risk calculators such as body mass index (BMI),30 creatinine,31 glucose.32 However, all of these have been previously identified in the Framingham study or have been used other CVD scoring systems (eg QRISK).12, 14 Among drug therapies use of anti-platelet agents was also predictive. Distribution data for laboratory values (eg fasting lipid values) and physical measurements (eg BMI and blood pressure) contributed more than median values to the models. In the models incorporating longitudinal data such as logistic regression selected biochemical data distribution in two separate sampling periods while random forests selected BMI. The effect of variation in CVD risk factors such as blood pressure, cholesterol, glucose and body mass index (BMI) has previously been linked to risk of CVD in classic epidemiological studies.33 This has been validated for blood pressure and is included in one currently nationally approved CVD risk calculator (QRISK-3).14 It also exists for glucose and cholesterol but this data has not been included in any guideline approved CVD risk calculator to date. Gradient boosting tree (GBT) analysis preferred historic diagnostic codes such as heart valve disorders, lipid disorders and hypertension over other features. One problem of large scale EHRs is the quality of data recording so this historical data may reflect single anomalous values being entered as diagnostic codes (ie a proxy for variance) or the lack of original untreated values in the EHR. Similar considerations apply to anti-platelet therapies such aspirin-clopidogrel acting as proxies for unrecorded diagnoses of significant CVD (or peripheral arterial disease) or in the case of aspirin alone—clinical suspicion of high-risk status. Genetic risk scores (GRS) are easily derived given the increasing ease of obtaining large scale genome variation data. Many studies are now investigating the utility of adding GRS to classical CVD risk factors in risk prediction.34 A multiplicity of scores have been investigated using limited panels and whole genome data applied to cohort data sets of up to 300 000 patients but whether any of these are superior to imaging remains unclear.34 GBT using classical CVD risk factors gave similar results to standard methods. Adding longitudinal EHR features to GBT increased AUC to 0.71 vs 0.70; AUPRC of 0.43 vs 0.40 and the genetic risk score (GRS) improved the AUROC and AUPRC by 2% and 9%. The GRS data included known CVD risk factor genes such as melanoma inhibitory activity protein 3 (MIA3; 2 loci) also known as Transport and Golgi organisation protein 1 (TANGO1) involved in chylomicron and very low-density lipoprotein transport, and lipoprotein (a) (LPA;2 loci) as well as chemokine C-X-C motif chemokine 12 (CXCL12) (stromal cell derived factor-1) involved in inflammation and a check point gene cyclin-dependent kinase Inhibitor 2A (CDKN2A) involved in angiogenesis. As in the field of CVD risk scores standardisation of inputs and data transparency are becoming essential to allow comparison of different strategies for the purposes of quality appraisal for evidence-based guidelines. The variety and quality of data set reporting, analytical and statistical approaches, provision of absolute as opposed to relative effect sizes and lack of specificity and sensitivity data at set points remain common problems.35, 36 Such approaches are now standard for epidemiological cohorts (CONSORT statement) and diagnostic assays (STARD).35, 36 Reporting standards have been introduced for single-nucleotide polymorphism (SNP) association data and genome wide association studies to provide greater clarity for journal referees and editors assessing these studies and for readers to understand them and conduct validation studies. The increasing popularity and complexity of mathematical models applied to CVD and other endpoint data means that similar provisions need to be applied to these studies as well.37 The electronic nature of modern scientific literature means model derivation structures and data can easily be added as appendices or contributed to public scientific data repositories. A number of publications and review articles have begun to request certain details of mathematical models in addition to data and ideally model transparency and ideally availability. A suggested scheme based on data presented in studies reviewed in this field is presented in Table 1. 1. Formal presentation of research questions 2. Data selection Public databases vs electronic health record databases vs registry data 3. Hardware selection 4. Data preparation 5. Feature selection This should not be necessary. Multi-dimensional datasets may require strategies such as vector embedding to enable features to be passed to other directed learning models 6. Data splitting Design and justify the proportion of training, validation and testing in the dataset (ie 70/10/20 or 80/10/10 or 60/20/20) and ideally provide comparison data 7. Modelling selection 8. Technical details for model Specification of technical terms to communicate with data scientists or programmers and allow understanding of the process of model development (learning rate selection, tuning hyperparameter, batch dropout and normalisation, regularisation strategies, loss function selection and network optimisation). Methods used in model structure- logistic regression, Cox regression, random forest or gradient boosting models, neural networks 9. Evaluation of model discrimination and calibration Precision recall curve (PRC; unbalanced data) or receiver operator curve (ROC; balanced data) analysis of data with presentation of C-statistics, Brier scores from probabilistic outcomes Presentation of NPV, PPV, sensitivity and specificity at specific set points Comparison with standard statistical approaches (ie multi-variable regression), goodness-of-fit, calibration plots or the decision-curve analysis 10. Clinical Validation Comparison with expert opinion or published data of other current clinical strategies 11. Publication and transparency Sharing of codes with journal (ie online supplements) or public space (ie Github, bioRxiv). Directed learning methodologies should be clearly explained in data appendices. Consider strategies for computational anonymisation It will take consensus conferences between investigators, journal editors, computer modelling specialists, evidence assessment groups and ideally academic funding agencies to finalise and agree the final set of quality metrics. These aurors will then pronounce on the quality of the work submitted. After all you need to remove the magic is to identify the underlying nature of the work—or to truly see Grindelwald. Otherwise fantastic means mythical as opposed to wonderful. The authors thank Dr Scooter Morris of the Pharmaceutical Chemistry group in the School of Pharmacy at the University of California, San Francisco, USA for his helpful comments on this manuscript. None.
Genomic data interpretation often requires analyses that move from a gene-by-gene focus to a focus on sets of genes that are associated with biological phenomena such as molecular processes, phenotypes, diseases, drug interactions or environmental conditions. Unique challenges exist in the curation of gene sets beyond the challenges in curation of individual genes. Here we highlight a literature curation workflow whereby gene sets are curated from peer-reviewed published data into GeneWeaver (GW), a data repository and analysis platform. We describe the system features that allow for a flexible yet precise curation procedure. We illustrate the value of curation by gene sets through analysis of independently curated sets that relate to the integrated stress response, showing that sets curated from independent sources all share significant Jaccard similarity. A suite of reproducible analysis tools is provided in GW as services to carry out interactive functional investigation of user-submitted gene sets within the context of over 150 000 gene sets constructed from publicly available resources and published gene lists. A curation interface supports the ability of users to design and maintain curation workflows of gene sets, including assigning, reviewing and releasing gene sets within a curation project context.
Purpose of review Extensive work has gone into understanding the genetics of cardiovascular disease (CVD) and implicating genes involved in hyperlipidaemia. Translation into routine practise involves using genetic risk scores (GRS) to identify high-risk individuals in the general population. Some of these risk scores are beginning to disentangle the complex nature of CVD and inherited dyslipidaemias. Recent findings GRS of varying complexity have been used to identify high-risk groups of patients with polygenic CVD including some individuals with risk equivalent to monogenic disease. In phenotypic familial hypercholesterolaemia a six or 12 gene lipid GRS may identify polygenic cases that comprise up to 50% of cases. In high triglyceride syndromes including even cases of familial chylomicronaemia syndrome more than 80% of cases are polygenic and not even associated with rare variants. In both familial hypercholesterolaemia and familial chylomicronaemia syndrome individuals with polygenic disease have a lower risk than those with monogenic disease. Summary GRS show promise in identifying individuals with high risks of CVD. They have a close relationship with imaging markers. It is unclear whether GRS, imaging or both will be used to identify individuals at high risk of future events.
Objective To investigate recent (2011–2015) research productivity in clinical biochemistry and compare it with a previous audit (1994–1998). Design A retrospective audit of peer-reviewed academic papers published in Medline listed journals. Setting UK chemical pathology/clinical biochemistry laboratories and other clinical scientific staff working in departments of pathology. Participants Medically qualified chemical pathologists and clinical scientists. Main outcome measures Publications were identified from electronic databases for individuals and sites. Analyses were conducted for individuals, sites and regional educational groups. Results Clinical scientific staff numbers fell by 3.9% and medical staff by 17.4% from 1998 to 2015. Publication rates declined as publication count centiles rose between 1998 and 2015 (e.g. n = 5; 67th→84th centile; p < 0.001). A reduction in productivity was seen in medically qualified staff but less from clinical scientists. Regional staffing was 77 ± 37 (range 30–150) with university hospital laboratory staff accounting for 58 ± 19% (range 30–92%). Medically qualified staff comprised 20 ± 4% of staff with lowest numbers in some London regions. Publication rates varied widely with a median of 155 papers per region (range 98–1035) and 2.82 (1.21–8.62) papers/individual. The skew was attenuated, increasing the publication rate to 6.0 ± 2.73 papers (range 2.29–11.76)/individual after correction for the number of university hospital sites per region and was not related to numbers of trainees. High publication rates were associated with the presence of one highly research-active individual. Their activity correlated over their careers from recruitment to today (r2 = 0.45; p = 0.05). The productivity rates of recent cohorts of trainees are inferior to previous cohorts. Conclusions Research remains a minority interest in clinical biochemistry. A small and decreasing proportion of individuals publish 90% of the work. A reduction was seen in clinical scientist and especially medical research productivity. No correlation of training activity with research productivity was seen implying weak links with translational medicine.
Alirocumab is a fully human monoclonal antibody to proprotein convertase subtilisin/kexin type 9 (PCSK9) and has been previously shown, in the phase III ODYSSEY clinical trial program, to provide significant lowering of low-density lipoprotein cholesterol (LDL-C) and reduction in risk of major adverse cardiovascular events. However, real-world evidence to date is limited.
International Journal of Clinical PracticeVolume 73, Issue 3 e13311 PERSPECTIVE Primum non nocere: Demand management in pathology and preventing harm Anthony S. Wierzbicki, Corresponding Author Anthony S. Wierzbicki Anthony.Wierzbicki@kcl.ac.uk orcid.org/0000-0003-2756-372X Dept Metabolic Medicine/Chemical Pathology, Guy’s & St Thomas’ Hospitals, London, UK Correspondence Email: Anthony.Wierzbicki@kcl.ac.ukSearch for more papers by this authorTimothy M. Reynolds, Timothy M. Reynolds orcid.org/0000-0002-9729-4775 Dept Metabolic Medicine/Chemical Pathology, Queen’s Hospital, Burton-on-Trent, Staffordshire, UKSearch for more papers by this author Anthony S. Wierzbicki, Corresponding Author Anthony S. Wierzbicki Anthony.Wierzbicki@kcl.ac.uk orcid.org/0000-0003-2756-372X Dept Metabolic Medicine/Chemical Pathology, Guy’s & St Thomas’ Hospitals, London, UK Correspondence Email: Anthony.Wierzbicki@kcl.ac.ukSearch for more papers by this authorTimothy M. Reynolds, Timothy M. Reynolds orcid.org/0000-0002-9729-4775 Dept Metabolic Medicine/Chemical Pathology, Queen’s Hospital, Burton-on-Trent, Staffordshire, UKSearch for more papers by this author First published: 11 January 2019 https://doi.org/10.1111/ijcp.13311Citations: 1Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinkedInRedditWechat Citing Literature Volume73, Issue3March 2019e13311 RelatedInformation