The rise of interdisciplinary and no-boundary engagement has created a need to train the next generation of No-Boundary Thinking (NBT) scholars and practitioners. So it is essential that students be provided with NBT experiences in the classroom and through group-based research experiences. Our no-boundary community has offered a first generation of classes to provide an environment where students can engage in no-boundary projects and exercises, and reflect upon the nature of this type of thinking and problem solving. The following five classes were first offered in fall 2015 through spring 2018 at four institutions for undergraduate and graduate students. The experience has been enriching for both students and faculty. In all cases the courses have been well received by the students and institutions, and most instructors plan to continue to provide the classes as permanent offerings. We describe the early offerings of each class.
Ideal healthcare should provide prevention and treatment strategies in the context of individual variability. The promise of genomics and big data for understanding the complex disease etiology and development of treatment strategies for translating research findings in a laboratory setting to the bedside requires a paradigm shift in how we conduct biomedical research. The take-home message from the Human Genome Sequencing Project is the need for a bold vision, even in the absence of a clear path. The No-Boundary Thinking (NBT) approach that advocates a scientific dialogue among individuals with varying expertise in a “discipline-free” manner at the problem definition stage is a pragmatic approach to leverage big data for precision medicine. Genomics big data as it pertains to understanding the molecular function of genes and proteins is discussed in this chapter. We also discuss the challenges in the adoption of NBT to genomics research.
In this era of Big Data and AI, expertise in multiple aspects of data, computing, and the domains of application is needed. This calls for teams of experts with different training and perspectives. Because data analysis can have serious ethical implications, it is important that these teams are well and deeply integrated. No-Boundary Thinking (NBT) teams can provide support for team formation and maintenance, thereby attending to the many dimensions of the ethics of data and analysis. In this NBT workshop session, we discuss the ethical concerns that arise from the use of data and AI, and the implications for team building; and provide and brainstorm suggestions for ethical data enabled science and AI.
Team building can be challenging when participants are from the same discipline or sub-discipline, but needs special attention when participants use a different vocabulary and have different cultural views on what constitutes viable problems and solutions. Essential to No Boundary Thinking (NBT) teams is proper formulation of the problem to be solved, and a basic tenant is that the NBT team must come together with diverse perspectives to decide the problem before solutions can be considered. Given that participants come with different views on problem formulation and solution, it is important to consider a robust process for team formation and maintenance. This takes extra effort and time, but scholars studying teams of experts with diverse training have found that they are better positioned to be successful in solving even deep and difficult problems especially if they have learned to work well with each other. At this workshop we will discuss principles that scholars who have worked in NBT teams have discovered as effective. We will then engage with the workshop participants to consider discuss these principles and brainstorm to consider other approaches.
In this era of Big Data and AI, expertise in multiple aspects of data, computing, and the domains of application is needed. This calls for teams of experts with different training and perspectives. Because data analysis can have serious ethical implications, it is important that these teams are well and deeply integrated. No-Boundary Thinking (NBT) teams can provide support for team formation and maintenance, thereby attending to the many dimensions of the ethics of data and analysis. In this NBT workshop session, we discuss the ethical concerns that arise from the use of data and AI, and the implications for team building; and provide and brainstorm suggestions for ethical data enabled science and AI.
An organism’s transcriptome is the set of all transcripts within a cell at a certain time. We often analyze the transcriptome by quantifying gene expression and performing subsequent analyses such as a differential expression or a network analysis. Such analysis helps us in understanding and interpreting the functional elements of the genome. Many challenges limit the accuracy and ability to map all the RNA-Seq correctly into its genome sequence. Some of these challenges are exemplified when mapping sequences fall at exon junctions, sequences containing polymorphisms, multiple insertions or deletions, and reads falling partially or wholly within introns. One of the most significant problems is the loss of data occurring from the inability to map sequences when they align to multiple genomic locations, sometimes called ambiguous sequence mappings. In this paper, we present a novel method to increase the accuracy of gene expression estimation by relying on a statistical approach to increase the accuracy of mapping the ambiguous reads to their proper locations within the genome. This approach allows us to better identify significantly expressed genomic locations so we can accurately map ambiguous reads to their most likely accurate genomic locations and to define more precisely which genes are expressed throughout the genome. Due to its statical nature the approach can be easily combined with other existing mapping tools and mechanisms as well.
Bovine respiratory disease (BRD), the leading disease complex in beef cattle production systems, remains highly elusive regarding diagnostics and disease prediction. Previous research has employed cellular and molecular techniques to describe hematological and gene expression variation that coincides with BRD development. Here, we utilized weighted gene co-expression network analysis (WGCNA) to leverage total gene expression patterns from cattle at arrival and generate hematological and clinical trait associations to describe mechanisms that may predict BRD development. Gene expression counts of previously published RNA-Seq data from 23 cattle (2017; n = 11 Healthy, n = 12 BRD) were used to construct gene co-expression modules and correlation patterns with complete blood count (CBC) and clinical datasets. Modules were further evaluated for cross-populational preservation of expression with RNA-Seq data from 24 cattle in an independent population (2019; n = 12 Healthy, n = 12 BRD). Genes within well-preserved modules were subject to functional enrichment analysis for significant Gene Ontology terms and pathways. Genes which possessed high module membership and association with BRD development, regardless of module preservation ("hub genes"), were utilized for protein-protein physical interaction network and clustering analyses. Five well-preserved modules of co-expressed genes were identified. One module ("steelblue"), involved in alpha-beta T-cell complexes and Th2-type immunity, possessed significant correlation with increased erythrocytes, platelets, and BRD development. One module ("purple"), involved in mitochondrial metabolism and rRNA maturation, possessed significant correlation with increased eosinophils, fecal egg count per gram, and weight gain over time. Fifty-two interacting hub genes, stratified into 11 clusters, may possess transient function involved in BRD development not previously described in literature. This study identifies co-expressed genes and coordinated mechanisms associated with BRD, which necessitates further investigation in BRD-prediction research.
Background Transcriptomics has identified at-arrival differentially expressed genes associated with bovine respiratory disease (BRD) development; however, their use as prediction molecules necessitates further evaluation. Therefore, we aimed to selectively analyze and corroborate at-arrival mRNA expression from multiple independent populations of beef cattle. In a nested case-control study, we evaluated the expression of 56 mRNA molecules from at-arrival blood samples of 234 cattle across seven populations via NanoString nCounter gene expression profiling. Analysis of mRNA was performed with nSolver Advanced Analysis software ( p < 0.05), comparing cattle groups based on the diagnosis of clinical BRD within 28 days of facility arrival ( n = 115 Healthy; n = 119 BRD); BRD was further stratified for severity based on frequency of treatment and/or mortality (Treated_1, n = 89; Treated_2+, n = 30). Gene expression homogeneity of variance, receiver operator characteristic (ROC) curve, and decision tree analyses were performed between severity cohorts. Results Increased expression of mRNAs involved in specialized pro-resolving mediator synthesis ( ALOX15 , HPGD ), leukocyte differentiation ( LOC100297044 , GCSAML , KLF17 ), and antimicrobial peptide production ( CATHL3 , GZMB , LTF ) were identified in Healthy cattle. BRD cattle possessed increased expression of CFB, and mRNA related to granulocytic processes ( DSG1 , LRG1 , MCF2L ) and type-I interferon activity ( HERC6, IFI6, ISG15, MX1 ). Healthy and Treated_1 cattle were similar in terms of gene expression, while Treated_2+ cattle were the most distinct. ROC cutoffs were used to generate an at-arrival treatment decision tree, which classified 90% of Treated_2+ individuals. Conclusions Increased expression of complement factor B, pro-inflammatory, and type I interferon-associated mRNA hallmark the at-arrival expression patterns of cattle that develop severe clinical BRD. Here, we corroborate at-arrival mRNA markers identified in previous transcriptome studies and generate a prediction model to be evaluated in future studies. Further research is necessary to evaluate these expression patterns in a prospective manner.
The rapid developments in high-throughput sequencing technologies have allowed researchers to analyze the full genomic sequence of organisms faster and cheaper than ever before. An important application of such advancements is to identify the impact of single nucleotide polymorphisms (SNPs) on the phenotypes and genotypes of the same species by discovering the factors that affect the occurrence of SNPs. The focus of this study is to determine whether climate factors such as the main climate, the precipitation, and the temperature affecting a certain geographical area might be associated with specific variations in certain ecotypes of the plant Arabidopsisthaliana. To test our hypothesis we analyzed 18 genes that encode Forkhead-Associated domain-containing proteins. They were extracted from 80 genomic sequences gathered from within 8 Eurasian regions. We used k-means clustering to separate the plants into distinct groups and evaluated the clusters using an innovative scoring system based upon the Köppen-Geiger climate classification system. The methods we used allow the selection of candidate clusters most likely to contain samples with similar polymorphisms. These clusters show that there is a correlation between genomic variations and the geographic distribution of those ecotypes.
Today Artificial Intelligence (AI) supports difficult decisions about policy, health, and our personal lives. The AI algorithms we develop and deploy to make sense of information, are informed by data, and based on models that capture and use pertinent details of the population or phenomenon being analyzed. For any application area, more importantly in precision medicine which directly impacts human lives, the data upon which algorithms are run must be procured, cleaned, and organized well to assure reliable and interpretable results, and to assure that they do not perpetrate or amplify human prejudices. This must be done without violating basic assumptions of the algorithms in use. Algorithmic results need to be clearly communicated to stakeholders and domain experts to enable sound conclusions. Our position is that AI holds great promise for supporting precision medicine, but we need to move forward with great care, with consideration for possible ethical implications. We make the case that a no-boundary or convergent approach is essential to support sound and ethical decisions. No-boundary thinking supports problem definition and solving with teams of experts possessing diverse perspectives. When dealing with AI and the data needed to use AI, there is a spectrum of activities that needs the attention of a no-boundary team. This is necessary if we are to draw viable conclusions and develop actions and policies based on the AI, the data, and the scientific foundations of the domain in question.
Bovine respiratory disease (BRD) is a multifactorial disease involving complex host immune interactions shaped by pathogenic agents and environmental factors. Advancements in RNA sequencing and associated analytical methods are improving our understanding of host response related to BRD pathophysiology. Supervised machine learning (ML) approaches present one such method for analyzing new and previously published transcriptome data to identify novel disease-associated genes and mechanisms. Our objective was to apply ML models to lung and immunological tissue datasets acquired from previous clinical BRD experiments to identify genes that classify disease with high accuracy. Raw mRNA sequencing reads from 151 bovine datasets (n = 123 BRD, n = 28 control) were downloaded from NCBI-GEO. Quality filtered reads were assembled in a HISAT2/Stringtie2 pipeline. Raw gene counts for ML analysis were normalized, transformed, and analyzed with MLSeq, utilizing six ML models. Cross-validation parameters (fivefold, repeated 10 times) were applied to 70% of the compiled datasets for ML model training and parameter tuning; optimized ML models were tested with the remaining 30%. Downstream analysis of significant genes identified by the top ML models, based on classification accuracy for each etiological association, was performed within WebGestalt and Reactome (FDR ≤ 0.05). Nearest shrunken centroid and Poisson linear discriminant analysis with power transformation models identified 154 and 195 significant genes for IBR and BRSV, respectively; from these genes, the two ML models discriminated IBR and BRSV with 100% accuracy compared to sham controls. Significant genes classified by the top ML models in IBR (154) and BRSV (195), but not BVDV (74), were related to type I interferon production and IL-8 secretion, specifically in lymphoid tissue and not homogenized lung tissue. Genes identified in Mannheimia haemolytica infections (97) were involved in activating classical and alternative pathways of complement. Novel findings, including expression of genes related to reduced mitochondrial oxygenation and ATP synthesis in consolidated lung tissue, were discovered. Genes identified in each analysis represent distinct genomic events relevant to understanding and predicting clinical BRD. Our analysis demonstrates the utility of ML with published datasets for discovering functional information to support the prediction and understanding of clinical BRD.
Bovine respiratory disease (BRD) remains the leading infectious disease in post-weaned beef cattle. The objective of this investigation was to contrast the at-arrival blood transcriptomes from cattle derived from two distinct populations that developed BRD in the 28 days following arrival versus cattle that did not. Forty-eight blood samples from two populations were selected for mRNA sequencing based on even distribution of development (n = 24) or lack of (n = 24) clinical BRD within 28 days following arrival; cattle which developed BRD were further stratified into BRD severity cohorts based on frequency of antimicrobial treatment: treated once (treated_1) or treated twice or more and/or died (treated_2+). Sequenced reads (~ 50 M/sample, 150 bp paired-end) were aligned to the ARS-UCD1.2 bovine genome assembly. One hundred and thirty-two unique differentially expressed genes (DEGs) were identified between groups stratified by disease severity (healthy, n = 24; treated_1, n = 13; treated_2+, n = 11) with edgeR (FDR ≤ 0.05). Differentially expressed genes in treated_1 relative to both healthy and treated_2+ were predicted to increase neutrophil activation, cellular cornification/keratinization, and antimicrobial peptide production. Differentially expressed genes in treated_2+ relative to both healthy and treated_1 were predicted to increase alternative complement activation, decrease leukocyte activity, and increase nitric oxide production. Receiver operating characteristic (ROC) curves generated from expression data for six DEGs identified in our current and previous studies ( MARCO, CFB, MCF2L, ALOX15, LOC100335828 (aka CD200R1 ) , and SLC18A2 ) demonstrated good-to-excellent (AUC: 0.800–0.899; ≥ 0.900) predictability for classifying disease occurrence and severity. This investigation identifies candidate biomarkers and functional mechanisms in at arrival blood that predicted development and severity of BRD.
Background Despite decades of extensive research, bovine respiratory disease (BRD) remains the most devastating disease in beef cattle production. Establishing a clinical diagnosis often relies upon visual detection of non-specific signs, leading to low diagnostic accuracy. Thus, post-weaned beef cattle are often metaphylactically administered antimicrobials at facility arrival, which poses concerns regarding antimicrobial stewardship and resistance. Additionally, there is a lack of high-quality research that addresses the gene-by-environment interactions that underlie why some cattle that develop BRD die while others survive. Therefore, it is necessary to decipher the underlying host genomic factors associated with BRD mortality versus survival to help determine BRD risk and severity. Using transcriptomic analysis of at-arrival whole blood samples from cattle that died of BRD, as compared to those that developed signs of BRD but lived (n = 3 DEAD, n = 3 ALIVE), we identified differentially expressed genes (DEGs) and associated pathways in cattle that died of BRD. Additionally, we evaluated unmapped reads, which are often overlooked within transcriptomic experiments. Results 69 DEGs (FDR<0.10) were identified between ALIVE and DEAD cohorts. Several DEGs possess immunological and proinflammatory function and associations with TLR4 and IL6. Biological processes, pathways, and disease phenotype associations related to type-I interferon production and antiviral defense were enriched in DEAD cattle at arrival. Unmapped reads aligned primarily to various ungulate assemblies, but failed to align to viral assemblies. Conclusion This study further revealed increased proinflammatory immunological mechanisms in cattle that develop BRD. DEGs upregulated in DEAD cattle were predominantly involved in innate immune pathways typically associated with antiviral defense, although no viral genes were identified within unmapped reads. Our findings provide genomic targets for further analysis in cattle at highest risk of BRD, suggesting that mechanisms related to type I interferons and antiviral defense may be indicative of viral respiratory disease at arrival and contribute to eventual BRD mortality.
Mutations that provide environment-dependent selective advantages drive adaptive divergence among species. Many phenotypic differences among related species are more likely to result from gene expression divergence rather than from non-synonymous mutations. In this regard, cis-regulatory mutations play an important part in generating functionally significant variation. Some proposed mechanisms that explore the role of cis-regulatory mutations in gene expression divergence involve microsatellites. Microsatellites exhibit high mutation rates achieved through symmetric or asymmetric mutation processes and are abundant in both coding and non-coding regions in positions that could influence gene function and products. Here we tested the hypothesis that microsatellites contribute to gene expression divergence among species with 50 individuals from five closely related Helianthus species using an RNA-seq approach. Differential expression analyses of the transcriptomes revealed that genes containing microsatellites in non-coding regions (UTRs and introns) are more likely to be differentially expressed among species when compared to genes with microsatellites in the coding regions and transcripts lacking microsatellites. We detected a greater proportion of shared microsatellites in 5′UTRs and coding regions compared to 3′UTRs and non-coding transcripts among Helianthus spp. Furthermore, allele frequency differences measured by pairwise FST at single nucleotide polymorphisms (SNPs), indicate greater genetic divergence in transcripts containing microsatellites compared to those lacking microsatellites. A gene ontology (GO) analysis revealed that microsatellite-containing differentially expressed genes are significantly enriched for GO terms associated with regulation of transcription and transcription factor activity. Collectively, our study provides compelling evidence to support the role of microsatellites in gene expression divergence.
Bovine respiratory disease (BRD) is a multifactorial disease complex and the leading infectious disease in post-weaned beef cattle. Clinical manifestations of BRD are recognized in beef calves within a high-risk setting, commonly associated with weaning, shipping, and novel feeding and housing environments. However, the understanding of complex host immune interactions and genomic mechanisms involved in BRD susceptibility remain elusive. Utilizing high-throughput RNA-sequencing, we contrasted the at-arrival blood transcriptomes of 6 beef cattle that ultimately developed BRD against 5 beef cattle that remained healthy within the same herd, differentiating BRD diagnosis from production metadata and treatment records. We identified 135 differentially expressed genes (DEGs) using the differential gene expression tools edgeR and DESeq2. Thirty-six of the DEGs shared between these two analysis platforms were prioritized for investigation of their relevance to infectious disease resistance using WebGestalt, STRING, and Reactome. Biological processes related to inflammatory response, immunological defense, lipoxin metabolism, and macrophage function were identified. Production of specialized pro-resolvin mediators (SPMs) and endogenous metabolism of angiotensinogen were increased in animals that resisted BRD. Protein-protein interaction modeling of gene products with significantly higher expression in cattle that naturally acquire BRD identified molecular processes involving microbial killing. Accordingly, identification of DEGs in whole blood at arrival revealed a clear distinction between calves that went on to develop BRD and those that resisted BRD. These results provide novel insight into host immune factors that are present at the time of arrival that confer protection from BRD.
Diagnosis and management of bovine respiratory disease (BRD), the most significant disease complex in post-weaned beef cattle, relies predominantly upon nonspecific clinical signs. Research is necessary to discern the underlying biological processes associated with BRD acquisition and severity. We hypothesize that at-arrival whole blood transcriptomes of stocker cattle will distinguish host immunologic responses that influence BRD severity, defined by BRD-attributed mortality.
BACKGROUND:A key use of high throughput sequencing technology is the sequencing and assembly of full genome sequences. These genome assemblies are commonly assessed using statistics relating to contiguity of the assembly. Measures of contiguity are not strongly correlated with information about the biological completion or correctness of the assembly, and a commonly reported metric, N50, can be misleading. Over the years, multiple research groups have rejected the overuse of N50 and sought to develop more informative metrics.RESULTS:This paper presents a review of problems that arise from relying solely on contiguity as a measure of genome assembly quality as well as current alternative methods. Alternative methods are compared on the basis of how informative they are about the biological quality of the assembly and how easy they are to use. A comprehensive method for using multiple metrics of measuring assembly quality is presented.CONCLUSIONS:This study aims to report on the status of assembly assessment methods and compare them, as well as to offer a comprehensive method that incorporates multiple facets of quality assessment. Weaknesses and strengths of varying methods are presented and explained, with recommendations based on speed of analysis and user friendliness.
Microsatellites are common in genomes of most eukaryotic species. Due to their high mutability, an adaptive role for microsatellites has been considered. However, little is known concerning the contribution of microsatellites towards phenotypic variation. We used populations of the common sunflower (Helianthus annuus) at two latitudes to quantify the effect of microsatellite allele length on phenotype at the level of gene expression. We conducted a common garden experiment with seed collected from sunflower populations in Kansas and Oklahoma followed by an RNA-Seq experiment on 95 individuals. The effect of microsatellite allele length on gene expression was assessed across 3,325 microsatellites that could be consistently scored. Our study revealed 479 microsatellites at which allele length significantly correlates with gene expression (eSTRs). When irregular allele sizes not conforming to the motif length were removed, the number of eSTRs rose to 2,379. The percentage of variation in gene expression explained by eSTRs ranged from 1%-86% when controlling for population and allele-by-population interaction effects at the 479 eSTRs. Of these eSTRs, 70.4% are in untranslated regions (UTRs). A gene ontology (GO) analysis revealed that eSTRs are significantly enriched for GO terms associated with cis- and trans-regulatory processes. Our findings suggest that a substantial number of transcribed microsatellites can influence gene expression.
Intestinal microbiota contributes health of living organisms. Florfenicol is an approved antimicrobial (AM) prescribed for several bacterial fish diseases. The present study investigated the extent to which florfenicol modulates intestinal microbial populations of channel catfish (Ictalurus punctatus). Florfenicol was administered orally to catfish at a standard therapeutic dose (10-15 mg/kg of body weight for 10 days), and the intestinal contents were collected and the 16S rRNA was subjected to Illumina sequencing. Alpha diversity analysis indicated that florfenicol significantly decreased microbiota richness and diversity. Beta diversity reflected a clear separation had occurred between the florfenicol-fed and control groups. Results indicated a significant increase in the abundance of phylum Proteobacteria (98.97% vs. 79.35% of the population) and decrease in phyla Firmicutes and Bacteroidetes (0.72% and 0.29% vs. 18.04% and 2.53%, respectively) in the florfenicol-fed fish in comparison with the control fish. At the genus level, unclassified Enterobacteriaceae and Escherichia populations increased in the florfenicol group. In contrast, Plesiomonas, Aeromonas, Lactococcus, Clostridium sensu stricto, Romboutsia, Klebsiella, Turicibacter and Lactobacillus decreased in florfenicol-fed fish in comparison with the control fish, indicating that intestinal microbiota of catfish was substantially modulated by florfenicol administration. Knowledge of changes in gut microbiota during medicated feed administration is important to improve fish performance and disease management and could enable the development of alternative therapeutic strategies.
PiRNAs are a particular type of small non-coding RNA. They are distinct from miRNA in size as well as other characteristics, such as the lack of sequence conservation and increased complexity when compared to their miRNA counterparts. PiRNA is considered the largest class of sRNA that is expressed especially in the animal cells. piRNAs are derived from long single-stranded RNAs, which are transcribed from genomic clusters, in contrast to other small silencing RNAs. It has been speculated that one locus could generate more than one piRNA. PiRNA corresponding to repetitive elements is fewer in mammals than in other species like Drosophila and Danio rerio, which signifies that piRNA might have possessed or gained some additional functionality in mammals. While the functionality of piRNAs may not be fully understood, they are believed to be involved in gene silencing. In this paper, we will examine a novel approach to identify potential piRNA clusters based on genes downstream and upstream location and order.
Xiuzhen Huang合作论文数Department of Computer Science,;Arkansas State University.5