Background: Literature regarding opioid use disorder (OUD) is often difficult for nonscientific communities to access. The OUD database on RefBin categorizes scientific findings and may facilitate access to information regarding OUD.Objectives: To evaluate if the RefBin OUD database improves access to information about OUD for policymakers and medical students.Methods: 31 medical students and 13 individual policymakers completed this study. Using a cross-over method, participants answered questions about OUD. Speed, accuracy, confidence, and satisfaction metrics were collected and compared between searches that used RefBin vs other resources chosen by participants.Results: At baseline, medical students reported being comfortable with scientific literature and familiar with OUD. Policymakers reported low comfort levels with scientific literature and variable familiarity with OUD. Within the medical student sample, the odds of answering correctly were 2.43 times higher for RefBin searches than for searches using resources other than RefBin (non-RefBin searches) (p = .005; 95% CI: (1.31, 4.51)). For policymakers, the odds of answering correctly were 3.65 times higher for RefBin vs non-RefBin searches (p = .0496; 95% CI: [1.002, 13.279]). Medical students reported feeling confident in their results 50.7% of the time when using RefBin, compared to 28.3% with non-RefBin searches (p = .006).Conclusion: When compared with searching using non-RefBin sources, searches performed using RefBin resulted in improved accuracy and efficiency for both medical students and policymakers. This demonstrates the potential utility of the RefBin OUD database in improving access to reliable information about OUD.
Although allele frequency data for most HLA loci provide strong evidence for balancing selection at the allele level, the DPB1 locus is a notable exception, with allele frequencies compatible with neutral evolution (genetic drift) or directional selection in most populations. This discrepancy is especially interesting as evidence for balancing selection has been seen at the nucleotide and amino acid (AA) sequence levels for DPB1. We describe methods used to examine the global distribution of DPB1 alleles and their constituent AA sequences. These methods allow investigation of the influence of natural selection in shaping DPβ diversity in a hierarchical fashion for DPB1 alleles, all polymorphic DPB1 exon 2-encoded AA positions, as well as all pairs and trios of these AA positions. In addition, we describe how asymmetric linkage disequilibrium for all DPB1 exon 2-encoded AA pairs can be used to complement other methods. Application of these methods provides strong evidence for the operation of balancing selection on AA positions 56, 85-87, 36, 55 and 84 (listed in decreasing order of the strength of selection), but no evidence for balancing selection on DPB1 alleles.
The DPB1 locus is notable among the classical HLA loci in that allele frequencies at this locus are consistent with genetic drift, whereas the frequencies of specific DPP amino acids are consistent with the action of balancing selection. We investigated the influence of natural selection in shaping the diversity of three functional categories of DPB1 diversity defined by specific amino acid motifs, DPB1 T-cell epitopes, DPB1 supertypes and DP1-DP4 serologic categories (SCs), via Ewens-Watterson (EW) selective neutrality and asymmetric Linkage Disequilibrium (ALD) analyses in a worldwide sample of 136 populations. These EW analyses provide strong evidence for the operation of balancing selection on DP SCs, but no evidence for balancing selection on T- cell epitopes or supertypes. We further investigated the global distribution of SCs. Each SC is common in a different region of the world, with the DP1 SC most common in Southeast Asia and Oceania, the DP2 SC in North and South America, the DP3 SC in South America, and the DP4 SC in Europe. The DP2 SC is present in all populations, while 14% of populations are missing at least one DP1, DP3, or DP4 SC. We observed consistent DPA1SDP SC haplotype associations across 10 populations from five global regions, and found that asymmetric linkage disequilibrium (LD) between the DPB1 locus and the four most-common DPA1 alleles (DPA1"01:03, "02:01, "02:02 and "03:01) is determined by variation at DPP AA positions 85-87. These positions are in LD with both DP alpha positions 31 and 50. We conclude from these EW analyses that natural selection is primarily operating to maintain population-level diversity of DP SCs, rather than DPB1 alleles or other functional categories of DPB1 diversity.
Python for Population Genomics (PyPop) is a software package that processes genotype and allele data and performs large-scale population genetic analyses on highly polymorphic multi-locus genotype data. In particular, PyPop tests data conformity to Hardy-Weinberg equilibrium expectations, performs Ewens-Watterson tests for selection, estimates haplotype frequencies, measures linkage disequilibrium, and tests significance. Standardized means of performing these tests is key for contemporary studies of evolutionary biology and population genetics, and these tests are central to genetic studies of disease association as well. Here, we present PyPop 1.0.0, a new major release of the package, which implements new features using the more robust infrastructure of GitHub, and is distributed via the industry-standard Python Package Index. New features include implementation of the asymmetric linkage disequilibrium measures and, of particular interest to the immunogenetics research communities, support for modern nomenclature, including colon-delimited allele names, and improvements to meta-analysis features for aggregating outputs for multiple populations.Code available at: https://zenodo.org/records/10080668 and https://github.com/alexlancaster/pypop
Human leukocyte antigen (HLA) class I and II loci are essential elements of innate and acquired immunity. Their functions include antigen presentation to T cells leading to cellular and humoral immune responses, and modulation of NK cells. Their exceptional influence on disease outcome has now been made clear by genome-wide association studies. The exons encoding the peptide-binding groove have been the main focus for determining HLA effects on disease susceptibility/pathogenesis. However, HLA expression levels have also been implicated in disease outcome, adding another dimension to the extreme diversity of HLA that impacts variability in immune responses across individuals. To estimate HLA expression, immunogenetic studies traditionally rely on quantitative PCR (qPCR). Adoption of alternative high-throughput technologies such as RNA-seq has been hampered by technical issues due to the extreme polymorphism at HLA genes. Recently, however, multiple bioinformatic methods have been developed to accurately estimate HLA expression from RNA-seq data. This opens an exciting opportunity to quantify HLA expression in large datasets but also brings questions on whether RNA-seq results are comparable to those by qPCR. In this study, we analyze three classes of expression data for HLA class I genes for a matched set of individuals: (a) RNA-seq, (b) qPCR, and (c) cell surface HLA-C expression. We observed a moderate correlation between expression estimates from qPCR and RNA-seq for HLA-A, -B, and -C (0.2 ≤ rho ≤ 0.53). We discuss technical and biological factors which need to be accounted for when comparing quantifications for different molecular phenotypes or using different techniques.
Screening, Brief Intervention, and Referral to Treatment (SBIRT) is an important secondary prevention strategy to address substance use and depression risk beginning in youth and continuing across the lifespan. Ten healthcare settings in Virginia implemented the SBIRT model between 2017 and 2020. A total of 65,315 participants ages 18 and older were universally screened to determine the severity of their substance use and depression and offered a risk-informed intervention. 12.7 percent of individuals endorsed some level of risky substance use and 4.5 percent screened positive for depression overall (11.1 percent in the outpatient setting). 10 percent of all brief intervention recipients were enrolled for follow-up screening 6 months later. Younger adults had significantly greater prevalence of risky drug use and depression compared to older age groups while middle-age adults displayed higher prevalence of moderate to severe alcohol risk highlighting the need for early intervention among younger adults. Significant reductions were observed in risky alcohol use (52.2%), as well as illicit drug use (44.7%) and depression (63.0%).
The Human Leukocyte Antigen (HLA) loci are extremely well documented targets of balancing selection, yet few studies have explored how selection affects population differentiation at these loci. In the present study we investigate genetic differentiation at HLA genes by comparing differentiation at microsatellites distributed genomewide to those in the MHC region. Our study uses a sample of 494 individuals from 30 human populations, 28 of which are Native Americans, all of whom were typed for genomewide and MHC region microsatellites. We find greater differentiation in the MHC than in the remainder of the genome (FST-MHC = 0.130 and FST-Genomic = 0.087), and use a permutation approach to show that this difference is statistically significant, and not accounted for by confounding factors. This finding lies in the opposite direction to the expectation that balancing selection reduces population differentiation. We interpret our findings as evidence that selection favors different sets of alleles in distinct localities, leading to increased differentiation. Thus, balancing selection at HLA genes simultaneously increases intra-population polymorphism and inter-population differentiation in Native Americans.
The American continent was the last to be occupied by modern humans, and native populations bear the marks of recent expansions, bottlenecks, natural selection, and population substructure. Here we investigate how this demographic history has shaped genetic variation at the strongly selected HLA loci. In order to disentangle the relative contributions of selection and demography process, we assembled a dataset with genome-wide microsatellites and HLA-A, -B, -C, and -DRB1 typing data for a set of 424 Native American individuals. We find that demographic history explains a sizeable fraction of HLA variation, both within and among populations. A striking feature of HLA variation in the Americas is the existence of alleles which are present in the continent but either absent or very rare elsewhere in the world. We show that this feature is consistent with demographic history (i.e., the combination of changes in population size associated with bottlenecks and subsequent population expansions). However, signatures of selection at HLA loci are still visible, with significant evidence selection at deeper timescales for most loci and populations, as well as population differentiation at HLA loci exceeding that seen at neutral markers.
Linkage disequilibrium (LD) is the nonrandom association of alleles at two or more loci. The standard measures of LD strength are described in detail, including extensions to multiple alleles. When there are different numbers of alleles at two loci, there are asymmetries that are not captured by standard measures. A complementary pair of conditional asymmetric LD (ALD) measures more accurately describes the correlation among locus pairs in this situation. Understanding and incorporating the LD structure of a genetic region into analyses is crucial for detecting disease predisposing variants, as well as for understanding the evolutionary history of a genetic region.
Standard measures of linkage disequilibrium (LD) provide an incomplete description of the correlation between two loci. Recently, Thomson and Single (2014) described a new asymmetric pair of LD measures (ALD) that give a more complete description of LD. The ALD measures are symmetric and equivalent to the correlation coefficient r when both loci are bi-allelic. When the numbers of alleles at the two loci differ, the ALD measures capture this asymmetry and provide additional detail about the LD structure. In disease association studies the ALD measures are useful for identifying additional disease genes in a genetic region, by conditioning on known effects. In evolutionary genetic studies ALD measures provide insight into selection acting on individual amino acids of specific genes, or other loci in high LD (see Thomson and Single (2014) for these examples). Here we describe new software for computing and visualizing ALD. We demonstrate the utility of this software using haplotype frequency data from the National Marrow Donor Program (NMDP). This enhances our understanding of LD patterns in the NMDP data by quantifying the degree to which LD is asymmetric and also quantifies this effect for individual alleles.
For multiallelic loci, standard measures of linkage disequilibrium provide an incomplete description of the correlation of variation at two loci, especially when there are different numbers of alleles at the two loci. We have developed a complementary pair of conditional asymmetric linkage disequilibrium (ALD) measures. Since these measures do not assume symmetry, they more accurately describe the correlation between two loci and can identify heterogeneity in genetic variation not captured by other symmetric measures. For biallelic loci the ALD are symmetric and equivalent to the correlation coefficient r. The ALD measures are particularly relevant for disease-association studies to identify cases in which an analysis can be stratified by one of more loci. A stratified analysis can aid in detecting primary disease-predisposing genes and additional disease genes in a genetic region. The ALD measures are also informative for detecting selection acting independently on loci in high linkage disequilibrium or on specific amino acids within genes. For SNP data, the ALD statistics provide a measure of linkage disequilibrium on the same scale for comparisons among SNPs, among SNPs and more polymorphic loci, among haplotype blocks of SNPs, and for fine mapping of disease genes. The ALD measures, combined with haplotype-specific homozygosity, will be increasingly useful as next-generation sequencing methods identify additional allelic variation throughout the genome.
BACKGROUND: Several previous studies have reported conflicting data on recent trends in use of initial total mastectomy (TM); the factors that contribute to TM variation are not entirely clear. Using a multi-institution database, we analyzed how practice, patient, and tumor characteristics contributed to variation in TM for invasive breast cancer.STUDY DESIGN: We collected detailed clinical and pathologic data about breast cancer diagnosis, initial, and subsequent breast cancer operations performed on all female patients from 4 participating institutions from 2003 to 2008. We limited this analysis to 2,384 incident cases of invasive breast cancer, stages I to III, and excluded patients with clinical indications for mastectomy. Predictors of initial TM were identified with univariate analyses and random effects multivariable logistic regression models.RESULTS: Initial TM was performed on 397 (16.7%) eligible patients. Use of preoperative MRI more than doubled the rate of TM (odds ratio [OR] = 2.44; 95% CI, 1.58-3.77; p < 0.0001). Increasing tumor size, high nuclear grade, and age were also associated with increased rates of initial TM. Differences by age and ethnicity were observed, and significant variation in the frequency of TM was seen at the individual surgeon level (p < 0.001). Our results were similar when restricted to tumors <20 mm.CONCLUSIONS: We identified factors associated with initial TM, including preoperative MRI and individual surgeon, that contribute to the current debate about variation in use of TM for the management of breast cancer. Additional evaluation of patient understanding of surgical options and outcomes in breast cancer and the impact of the surgeon provider is warranted. ((c) 2013 by the American College of Surgeons)
BACKGROUND:Treatment with neoadjuvant chemotherapy (NAC) has made it possible for some women to be successfully treated with breast conservation therapy (BCT ) who were initially considered ineligible. Factors related to current practice patterns of NAC use are important to understand particularly as the surgical treatment of invasive breast cancer has changed. The goal of this study was to determine variations in neoadjuvant chemotherapy use in a large multi-center national database of patients with breast cancer.METHODS:We evaluated NAC use in patients with initially operable invasive breast cancer and potential impact on breast conservation rates. Records of 2871 women ages 18-years and older diagnosed with 2907 invasive breast cancers from January 2003 to December 2008 at four institutions across the United States were examined using the Breast Cancer Surgical Outcomes (BRCASO) database. Main outcome measures included NAC use and association with pre-operatively identified clinical factors, surgical approach (partial mastectomy [PM] or total mastectomy [TM]), and BCT failure (initial PM followed by subsequent TM).RESULTS:Overall, NAC utilization was 3.8%l. Factors associated with NAC use included younger age, pre-operatively known positive nodal status, and increasing clinical tumor size. NAC use and BCT failure rates increased with clinical tumor size, and there was significant variation in NAC use across institutions. Initial TM frequency approached initial PM frequency for tumors >30-40 mm; BCT failure rate was 22.7% for tumors >40 mm. Only 2.7% of patients undergoing initial PM and 7.2% undergoing initial TM received NAC.CONCLUSIONS:NAC use in this study was infrequent and varied among institutions. Infrequent NAC use in patients suggests that NAC may be underutilized in eligible patients desiring breast conservation.
The human leucocyte antigen (HLA) system shows extensive variation in the number and function of loci and the number of alleles present at any one locus. Allele distribution has been analysed in many populations through the course of several decades, and the implementation of molecular typing has significantly increased the level of diversity revealing that many serotypes have multiple functional variants. While the degree of diversity in many populations is equivalent and may result from functional polymorphism(s) in peptide presentation, homogeneous and heterogeneous populations present contrasting numbers of alleles and lineages at the loci with high-density expression products. In spite of these differences, the homozygosity levels are comparable in almost all of them. The balanced distribution of HLA alleles is consistent with overdominant selection. The genetic distances between outbred populations correlate with their geographical locations; the formal genetic distance measurements are larger than expected between inbred populations in the same region. The latter present many unique alleles grouped in a few lineages consistent with limited founder polymorphism in which any novel allele may have been positively selected to enlarge the communal peptide-binding repertoire of a given population. On the other hand, it has been observed that some alleles are found in multiple populations with distinctive haplotypic associations suggesting that convergent evolution events may have taken place as well. It appears that the HLA system has been under strong selection, probably owing to its fundamental role in varying immune responses. Therefore, allelic diversity in HLA should be analysed in conjunction with other genetic markers to accurately track the migrations of modern humans.
In this chapter, we outline some basic principles for the consistent management of immunogenetic data. These include the preparation of a single master data file that can serve as the basis for all subsequent analyses, a focus on the quality and homogeneity of the data to be analyzed, the documentation of the coding systems used to represent the data, and the application of nomenclature standards specific for each immunogenetic system being evaluated. The data management principles discussed here are intended to provide a foundation for the data analysis methods detailed in Chaps. 13 and 14 . The relationship between the data management and analysis methods covered in these three chapters is illustrated in Fig. 3.The application of these data management principles is a first step toward consistent and reproducible data analyses. While it may take extra time and effort to apply them, we feel that it is better to take this approach than to assume that low data quality can be compensated for by large sample sizes.In addition to their relevance for analytical reproducibility, it is important to consider these data management principles from an ethical perspective. The reliability of the data collected and generated as part of a research study should be as important a component of the ethical review of a research application as the security of those data. Finally, in addition to ensuring the integrity of the data from collection to publication, the application of these data management principles will provide a means to foster research integrity and to improve the potential for collaborative data sharing.
Heather Feigelson1, Adedayo Onitilo2, Ted James3, Erin Aiello Bowles4, Richard Single3, Tom Barney5, Jessica Engel2 and Laurence McCahill6 1Kaiser Permanente Colorado 2Marshfield Clinic 3University of Vermont 4Group Health Cooperative 5VanAndel Research Institute 6Lacks Cancer Center
CONTEXT:Health care reform calls for increasing physician accountability and transparency of outcomes. Partial mastectomy is the most commonly performed procedure for invasive breast cancer and often requires reexcision. Variability in reexcision might be reflective of the quality of care.OBJECTIVE:To assess hospital and surgeon-specific variation in reexcision rates following partial mastectomy.DESIGN, SETTING, AND PATIENTS:An observational study of breast surgery performed between 2003 and 2008 intended to evaluate variability in breast cancer surgical care outcomes and evaluate potential quality measures of breast cancer surgery. Women with invasive breast cancer undergoing partial mastectomy from 4 institutions were studied (1 university hospital [University of Vermont] and 3 large health plans [Kaiser Permanente Colorado, Group Health, and Marshfield Clinic]). Data were obtained from electronic medical records and chart abstraction of surgical, pathology, radiology, and outpatient records, including detailed surgical margin status. Logistic regression including surgeon-level random effects was used to identify predictors of reexcision.MAIN OUTCOME MEASURE:Incidence of reexcision.RESULTS:A total of 2206 women with 2220 invasive breast cancers underwent partial mastectomy and 509 patients (22.9%; 95% CI, 21.2%-24.7%) underwent reexcision (454 patients [89.2%; 95% CI, 86.5%-91.9%] had 1 reexcision, 48 [9.4%; 95% CI, 6.9%-12.0%] had 2 reexcisions, and 7 [1.4%; 95% CI, 0.4%-2.4%] had 3 reexcisions). Among all patients undergoing initial partial mastectomy, total mastectomy was performed in 190 patients (8.5%; 95% CI, 7.2%-9.5%). Reexcision rates for margin status following initial surgery were 85.9% (95% CI, 82.0%-89.8%) for initial positive margins, 47.9% (95% CI, 42.0%-53.9%) for less than 1.0 mm margins, 20.2% (95% CI, 15.3%-25.0%) for 1.0 to 1.9 mm margins, and 6.3% (95% CI, 3.2%-9.3%) for 2.0 to 2.9 mm margins. For patients with negative margins, reexcision rates varied widely among surgeons (range, 0%-70%; P = .003) and institutions (range, 1.7%-20.9%; P < .001). Reexcision rates were not associated with surgeon procedure volume after adjusting for case mix (P = .92).CONCLUSION:Substantial surgeon and institutional variation were observed in reexcision following partial mastectomy in women with invasive breast cancer.
In this chapter, we describe analyses commonly applied to immunogenetic population data, along with software tools that are currently available to perform those analyses. Where possible, we focus on tools that have been developed specifically for the analysis of highly polymorphic immunogenetic data. These analytical methods serve both as a means to examine the appropriateness of a dataset for testing a specific hypothesis, as well as a means of testing hypotheses. Rather than treat this chapter as a protocol for analyzing any population dataset, each researcher and analyst should first consider their data, the possible analyses, and any available tools in light of the hypothesis being tested. The extent to which the data and analyses are appropriate to each other should be determined before any analyses are performed.
BACKGROUND:Common measures of surgical quality are 30-day morbidity and mortality, which poorly describe breast cancer surgical quality with extremely low morbidity and mortality rates. Several national quality programs have collected additional surgical quality measures; however, program participation is voluntary and results may not be generalizable to all surgeons. We developed the Breast Cancer Surgical Outcomes (BRCASO) database to capture meaningful breast cancer surgical quality measures among a non-voluntary sample, and study variation in these measures across providers, facilities, and health plans. This paper describes our study protocol, data collection methods, and summarizes the strengths and limitations of these data.METHODS:We included 4524 women ≥18 years diagnosed with breast cancer between 2003-2008. All women with initial breast cancer surgery performed by a surgeon employed at the University of Vermont or three Cancer Research Network (CRN) health plans were eligible for inclusion. From the CRN institutions, we collected electronic administrative data including tumor registry information, Current Procedure Terminology codes for breast cancer surgeries, surgeons, surgical facilities, and patient demographics. We supplemented electronic data with medical record abstraction to collect additional pathology and surgery detail. All data were manually abstracted at the University of Vermont.RESULTS:The CRN institutions pre-filled 30% (22 out of 72) of elements using electronic data. The remaining elements, including detailed pathology margin status and breast and lymph node surgeries, required chart abstraction. The mean age was 61 years (range 20-98 years); 70% of women were diagnosed with invasive ductal carcinoma, 20% with ductal carcinoma in situ, and 10% with invasive lobular carcinoma.CONCLUSIONS:The BRCASO database is one of the largest, multi-site research resources of meaningful breast cancer surgical quality data in the United States. Assembling data from electronic administrative databases and manual chart review balanced efficiency with high-quality, unbiased data collection. Using the BRCASO database, we will evaluate surgical quality measures including mastectomy rates, positive margin rates, and partial mastectomy re-excision rates among a diverse, non-voluntary population of patients, providers, and facilities.