RNA covalent modifications (RCMs) influence RNA stability and translation efficiency, and they thus play critical roles in eukaryotic growth and development. However, their role in regulating plant performance under abiotic stress remains largely unexplored. Here, we integrated multi-omics data in 6 Sorghum bicolor accessions under water-limiting conditions in the field to explore the relationship between RCMs and drought response. Within a stress- and photosynthesis-associated gene co-expression module, we identified SbDUS2, a member of a family of enzymes conserved across eukaryotes, that catalyzes the reduction of uracil to dihydrouridine (DHU) on RNA molecules. DHU-modified transcripts in this module were enriched for photosynthetic functions and showed strong correlation with photosynthetic traits. To elucidate the function of this RCM, we characterized loss-of-function dus2 mutants in Arabidopsis thaliana. Under control conditions, these DHU-deficient mutants exhibited impaired germination and delayed development. Furthermore, under water-limiting or heat conditions, these mutants showed significantly reduced net CO2 assimilation and survival. Using multiple transcriptome-wide RNA stability assays, we demonstrated that transcripts associated with lower DHU levels in a dus2 background generally exhibited increased stability compared to Col-0 controls. Particularly, lack of DUS2 led to the hyperstability of photosynthesis-related transcripts, impeding their turnover and likely preventing proper photosynthetic acclimation during stress. We propose a model where DHU acts as a critical post-transcriptional regulator marking mRNAs for rapid turnover under stress, highlighting an overlooked regulatory layer contributing to plant resilience.
As sequencing technologies advance and costs decline, there has been a surge in the application of RNA sequencing (RNA-seq) to understand the effects of gene expression regulation on specific biological processes. In addition to the typical uses of RNA-seq for transcriptomics, gene annotation, novel gene discovery, and network analysis, these data can enable a deeper understanding of cellular processes through the identification of RNA modifications (epitranscriptome) and long non-coding RNAs (lncRNAs). To expedite discovery, we developed a portable, centralized computational pipeline for the high-throughput annotation of modified ribonucleotides and long non-coding ribonucleic acids (HAMRLNC). HAMRLNC differs from existing methods by integrating three workflows for transcript abundance quantification, RNA modification inference, and lncRNA annotation using the same RNA-seq pre-processing and mapping steps. This facilitates reproducibility across multiple analyses and allows researchers to perform post hoc analyses of archived sequencing data. In addition, we include novel analysis features to enable downstream visualization of annotated modified RNAs. HAMRLNC generates over a dozen well-defined and labeled figures as output, including gene ontology heatmaps, modification enrichment landscapes, and modification clustering statistics.
As sequencing technologies advance and costs decline, there has been a surge in the application of RNA sequencing (RNA-seq) in understanding biological processes. In addition to the typical uses of RNA-seq for transcriptomics, gene annotation, novel gene discovery, and network analysis, these data can enable a deeper understanding of cellular processes through the identification of RNA modifications (epitranscriptome) and long non-coding RNAs (lncRNAs). To expedite discovery, we developed a portable, centralized computational pipeline for High-throughput Annotation of Modified Ribonucleotides and Long Non-Coding ribonucleic acids (HAMRLNC). HAMRLNC differs from existing methods by incorporating three workflows for quantifying transcript abundance, inferring RNA modifications, and lncRNA annotation using the same RNA- seq pre-processing and mapping steps. This facilitates reproducibility across multiple analyses and allows researchers to perform post-hoc analyses of archived sequencing data. In addition, we include novel analysis features to enable downstream visualization of annotated modified RNAs. HAMRLNC generates over a dozen well-defined and labeled figures as output, including gene ontology heatmaps, modification enrichment landscape, and modification clustering statistics. Availability and Implementation HAMRLNC is an open-source software, and the source code is available at . The pipeline can be installed and used through a docker container (). HAMRLNC is also available as an app in the CyVerse Discovery Environment . Supplementary information Supplementary data are available at BioRxiv online. ### Competing Interest Statement The authors have declared no competing interest. NSF, MCB-2427729, IOS-2023310
Covalent RNA modifications (RCMs) are post-transcriptional changes to the chemical composition of RNA. RCMs influence mRNA stability, regulate transcription and translation efficiency, and play critical roles in the growth and development of eukaryotes. However, their role in plants remains poorly understood, particularly as they relate to performance in the field under stress conditions. In this study, we grew a panel of six diverse sorghum (S. bicolor) accessions in the field during the summer in central Arizona and examined their physiological and molecular responses to drought and heat stress over time. We then explored the molecular features that contributed to plant performance under stress. To do so, we combined genomic, transcriptomic, epitranscriptomic, physiological, and metabolomic data in a systems-level approach. Co-expression network analyses uncovered two modules of interest, one controlled primarily by a single stress-responsive transcription factor, SbCDF3, and the other by an RCM, dihydrouridine. While the CDF3 module largely contained a set of stress response and photosynthesis-associated genes that were positively correlated with plant performance, the dihydrouridine-associated module was largely comprised of photosynthesis and metabolism genes, including SbPPDK1, which is integral to C4 photosynthesis in the grasses. In addition, the transcript encoding the enzyme responsible for this RCM, dihydrouridine synthase (SbDUS2), was also present in this module, and its abundance was positively correlated with photosynthetic traits and SbPPDK1 abundance. Given that this highly conserved RCM has never been characterized in plants, we examined loss of function mutants for the DUS2 enzyme in Arabidopsis, demonstrating decreases in plant growth and performance under heat stress in this background. Our work highlights both a key transcription factor, CDF3, for breeding in the Poaceae. In addition, for the first time in plants, we reveal a role for the RCM dihydrouridine in modifying conserved, core metabolic and photosynthesis-associated transcripts in plants. ### Competing Interest Statement The authors have declared no competing interest.
RNA Covalent Modifications (RCMs) are post-transcriptional chemical alterations that influence RNA stability and translation efficiency, thus play critical roles in eukaryotic growth and development. However, their role in regulating plant performance under abiotic stress remain largely unexplored. Here, we integrated multi-omics data in six Sorghum bicolor accessions under water-limiting conditions in the field to explore the relationship between RCMs and drought response. Within a stress and photosynthesis-associated gene co-expression module, we identified SbDUS2, a member of family of enzymes, conserved across eukaryotes, which catalyzes the reduction of uracil to dihydrouridine (DHU) on RNA molecules. DHU-modified transcripts in this module were enriched for photosynthetic functions and showed strong correlation with photosynthetic traits. To elucidate the function of this RCM, we characterized loss of function dus2 mutants in the genetic model, Arabidopsis thaliana. Under control conditions, these DHU-deficient mutants exhibited impaired germination and delayed development. Furthermore, when exposed to heat or water-limiting conditions, these mutants showed significantly reduced net CO2 assimilation and survival. Using multiple transcriptome-wide RNA stability assays, we demonstrated that transcripts associated with lower DHU level in a dus2 background generally exhibited increased stability compared to Col-0 controls. Particularly, lack of DUS2 led to the hyperstability of photosynthesis-related transcripts, impeding their turnover and likely preventing proper photosynthetic acclimation during stress. We propose a model based on these data where DHU acts as a critical post-transcriptional regulator marking mRNAs for rapid turnover under stress, highlighting an overlooked regulatory layer contributing to plant resilience.
CyVerse, the largest publicly-funded open-source research cyberinfrastructure for life sciences, has played a crucial role in advancing data-driven research since the 2010s. As the technology landscape evolved with the emergence of cloud computing platforms, machine learning and artificial intelligence (AI) applications, CyVerse has enabled access by providing interfaces, Software as a Service (SaaS), and cloud-native Infrastructure as Code (IaC) to leverage new technologies. CyVerse services enable researchers to integrate institutional and private computational resources, custom software, perform analyses, and publish data in accordance with open science principles. Over the past 13 years, CyVerse has registered more than 124,000 verified accounts from 160 countries and was used for over 1,600 peer-reviewed publications. Since 2011, 45,000 students and researchers have been trained to use CyVerse. The platform has been replicated and deployed in three countries outside the US, with additional private deployments on commercial clouds for US government agencies and multinational corporations. In this manuscript, we present a strategic blueprint for creating and managing SaaS cyberinfrastructure and IaC as free and open-source software.
Abstract Charcoal rot of sorghum (CRS) is a significant disease affecting sorghum crops, with limited genetic resistance available. The causative agent, Macrophomina phaseolina (Tassi) Goid, is a highly destructive fungal pathogen that targets over 500 plant species globally, including essential staple crops. Utilizing field image data for precise detection and quantification of CRS could greatly assist in the prompt identification and management of affected fields and thereby reduce yield losses. The objective of this work was to implement various machine learning algorithms to evaluate their ability to accurately detect and quantify CRS in red‐green‐blue images of sorghum plants exhibiting symptoms of infection. EfficientNet‐B3 and a fully convolutional network emerged as the top‐performing models for image classification and segmentation tasks, respectively. Among the classification models evaluated, EfficientNet‐B3 demonstrated superior performance, achieving an accuracy of 86.97%, a recall rate of 0.71, and an F1 score of 0.73. Of the segmentation models tested, FCN proved to be the most effective, exhibiting a validation accuracy of 97.76%, a recall rate of 0.68, and an F1 score of 0.66. As the size of the image patches increased, both models’ validation scores increased linearly, and their inference time decreased exponentially. This trend could be attributed to larger patches containing more information, improving model performance, and fewer patches reducing the computational load, thus decreasing inference time. The models, in addition to being immediately useful for breeders and growers of sorghum, advance the domain of automated plant phenotyping and may serve as a foundation for drone‐based or other automated field phenotyping efforts. Additionally, the models presented herein can be accessed through a web‐based application where users can easily analyze their own images.
Dramatic improvements in measuring genetic variation across agriculturally relevant populations (genomics) must be matched by improvements in identifying and measuring relevant trait variation in such populations across many environments (phenomics). Identifying the most critical opportunities and challenges in genome to phenome (G2P) research is the focus of this paper. Previously (Genome Biol, 23(1):1–11, 2022), we laid out how Agricultural Genome to Phenome Initiative (AG2PI) will coordinate activities with USA federal government agencies expand public–private partnerships, and engage with external stakeholders to achieve a shared vision of future the AG2PI. Acting on this latter step, AG2PI organized the “Thinking Big: Visualizing the Future of AG2PI” two-day workshop held September 9–10, 2022, in Ames, Iowa, co-hosted with the United State Department of Agriculture’s National Institute of Food and Agriculture (USDA NIFA). During the meeting, attendees were asked to use their experience and curiosity to review the current status of agricultural genome to phenome (AG2P) work and envision the future of the AG2P field. The topic summaries composing this paper are distilled from two 1.5-h small group discussions. Challenges and solutions identified across multiple topics at the workshop were explored. We end our discussion with a vision for the future of agricultural progress, identifying two areas of innovation needed: (1) innovate in genetic improvement methods development and evaluation and (2) innovate in agricultural research processes to solve societal problems. To address these needs, we then provide six specific goals that we recommend be implemented immediately in support of advancing AG2P research.
In this study, we introduce PlantSegNet, a novel neural network model for instance segmentation of nearby objects with similar geometric structures. Our work addresses the challenges of instance segmentation of plant point clouds, including the difficulty of annotating and labeling point clouds, the loss of local structural information in neural network components, and the generation of large numbers of incorrect small clusters due to poor choices of the loss function. One of the key contributions of our approach is a digital twin of sorghum, i.e., a procedural sorghum model, which was used to generate point clouds of sorghum fields. This allowed us to create a large-scale, annotated, synthetic dataset of sorghum plants that we used to train our PlantSegNet model. We demonstrated the effectiveness of our method in segmenting instances of sorghum leaves grown in outdoor field settings. To the best of our knowledge, this is the first study to address this specific instance segmentation problem for plants grown in such a setting. We compared our proposed method with other state-of-the-art methods for indoor settings, including SGPN and TreePartNet, on both synthetic and real data. Our results show that PlantSegNet outperforms these methods regarding accuracy, robustness, and efficiency.
Unmanned Aerial Vehicles (UAVs), i.e., drones, are expected to be widely used in various applications, such as parcel delivery and passenger transport, with the benefits of mitigating traffic congestion and reducing carbon emissions. In this paper, we study a UAV path planning problem under uncertain weather conditions, and design a data-driven dynamic decision support system for multiple types of UAVs. To this end, we categorize all relevant costs into three types, namely, economic, environmental, and social costs, and formulate a nonlinear two-stage stochastic programming model to establish optimal paths for UAV missions under weather uncertainty. We then discretize the nonlinear model and propose a tight linear approximation for the discretized problem to allow for a near real-time implementation. To quantify weather uncertainty, we propose a weather scenario generation algorithm to map ensemble-based weather forecast information to airspace blockage maps. With comprehensive computational studies through simulations, we show that our proposed stochastic approach can lower operating costs by an average of around 6%, where the savings increase as weather conditions become more severe and complex. We also find that, for missions operated by small UAVs, it is not sufficient to determine a path solely based on economic cost minimization, but it should rather be through total cost minimization, which involves environmental and social costs. Considering only the economic cost in the optimization may lead to much higher non-economic costs. However, for missions operated by large UAVs, it is sufficient to determine paths through economic cost optimization, as including environmental and social costs in the optimization process does not result in solutions that are much different from those obtained by considering only the economic costs. For both small and large UAVs, a path established solely through environmental or social cost minimization may not be economically sustainable, as doing so would imply very high economic costs.
Video-based applications form one of the most popular applications on the Internet that is continually evolving. There is a need to develop novel network services that enable reliable video transmission over network paths with dynamic cross-traffic, as well as services that utilize programmable data planes enabled by Protocol-independent Packet Processors (P4). In this paper, we describe experiences in developing network services for reliable video transmission using resources from the FABRIC network instrument, which supports high-performance edge/cloud as well as programmable networking infrastructure. Specifically, we deploy a programmable network using local Ethernet (Layer 2) sites, as well as geographically distributed wide-area network (LAN extension) sites in FABRIC, in order to experiment with visual cloud computing application use cases. Our experiment results provide insights into benefits of data plane programmability (i.e., port forwarding) on improving video streaming quality and compare CPU vs. GPU processing times while completing object detection pipeline processing to obtain visual situational awareness.
Many Internet of Things (IoT) applications require compute resources that cannot be provided by the devices themselves. On the other hand, processing of the data generated by IoT devices and sensors often has to be performed in real- or near real-time, i.e., with stringent latency requirements in constrained environments (e.g., intermittent network connectivity and limited power envelopes). Examples of such scenarios are autonomous vehicles in the form of cars and drones where the processing and analysis of observational data (e.g., video feeds) need to be performed expeditiously to allow for safe operation of the vehicles and to deliver the results in a timely fashion to the stakeholders of the mission. To support the compute and timeliness requirements of such applications, it is essential to include suitable edge resources to process these workflows, and to develop an end-to-end system that can route the vehicles dynamically and process and deliver mission-critical data and analyzed results. In this paper, we develop and evaluate a dynamic scheduling approach that considers complex tradeoffs between real-time constraints, network availability, and latency sensitivity of the mission. We devise an optimized route planning and data transmission schedule for drone flights. The scheduling algorithm is encapsulated in a novel end-to-end architecture (FlyPaw) and an associated adaptive drone mission control system, which enables deployment and management of an integrated cyberphysical system (CPS) – from real drone testbed to base stations to edge-to-cloud resources. The planning algorithm takes into account measured network communication characteristics, estimated uncertainties of future data link connectivity, and data timeliness requirements of the mission to prioritize candidate decision tree solutions based on a risk metric derived from Sharpe's ratio. Our results show that for given task sets, Net Time to Retrieve, our metric describing the time required to perform end-to-end collection and downstream processing of data, can be significantly reduced compared to other naive approaches. The theoretical improvement provided by our algorithm over other naive approaches is dependent on several factors — task locations, network connectivity, processing times and available resources, and is bounded by the duration of the drone flight.
As phenomics data volume and dimensionality increase due to advancements in sensor technology, there is an urgent need to develop and implement scalable data processing pipelines. Current phenomics data processing pipelines lack modularity, extensibility, and processing distribution across sensor modalities and phenotyping platforms. To address these challenges, we developed PhytoOracle (PO), a suite of modular, scalable pipelines for processing large volumes of field phenomics RGB, thermal, PSII chlorophyll fluorescence 2D images, and 3D point clouds. PhytoOracle aims to (i) improve data processing efficiency; (ii) provide an extensible, reproducible computing framework; and (iii) enable data fusion of multi-modal phenomics data. PhytoOracle integrates open-source distributed computing frameworks for parallel processing on high-performance computing, cloud, and local computing environments. Each pipeline component is available as a standalone container, providing transferability, extensibility, and reproducibility. The PO pipeline extracts and associates individual plant traits across sensor modalities and collection time points, representing a unique multi-system approach to addressing the genotype-phenotype gap. To date, PO supports lettuce and sorghum phenotypic trait extraction, with a goal of widening the range of supported species in the future. At the maximum number of cores tested in this study (1,024 cores), PO processing times were: 235 minutes for 9,270 RGB images (140.7 GB), 235 minutes for 9,270 thermal images (5.4 GB), and 13 minutes for 39,678 PSII images (86.2 GB). These processing times represent end-to-end processing, from raw data to fully processed numerical phenotypic trait data. Repeatability values of 0.39-0.95 (bounding area), 0.81-0.95 (axis-aligned bounding volume), 0.79-0.94 (oriented bounding volume), 0.83-0.95 (plant height), and 0.81-0.95 (number of points) were observed in Field Scanalyzer data. We also show the ability of PO to process drone data with a repeatability of 0.55-0.95 (bounding area).
To meet the challenges of an increasing global population and multiple environmental stressors on agricultural production, it is essential to better understand how genotype (G) and environment (E) influence phenotype for traits of economic importance in agriculture. To accomplish this requires interdisciplinary teams consisting of researchers from crop and livestock sciences, genetics, genomics, computational and data sciences, and engineering. Towards this end, in 2020, Congress established the Agricultural Genome to Phenome Initiative (AG2PI). An important objective of this presentation is to encourage strong participation of animal science researchers and industry in ongoing AG2PI activities. The purpose of the initial AG2PI project, awarded to Iowa State University and the Universities of Arizona, Idaho and Nebraska, was to assemble and prepare a community to conduct AG2P research across the plant and animal kingdoms. The AG2PI project aims to: a) identify research gaps and opportunities, b) foster community solutions to these challenges, and c) rapidly disseminate findings for community use. Training workshops, field days, conferences, community surveys, and seed grants have all contributed to these goals. Opinions and ideas from the crop and livestock genetics communities are key to advancing AG2P science and related work. AG2PI has used online community surveys and in-person conferences to gather these and to facilitate transdisciplinary communications. In September 2022, AG2PI held a conference titled “Thinking Big: Visualizing the Future of AG2PI”, in which attendees participated in focused small group discussions to identify which agricultural G2P-related topics are most critical for future R&D funding and which may be the most difficult to achieve. These discussions are summarized in a concept paper for the community, shared with USDA-NIFA, and submitted for publication. A follow-up conference was held June 14-15, 2023, in Kansas City, where people working in both the animal and plant genetics research communities and industry stakeholders attended. Outcomes from this conference will be summarized in this presentation, and requests for animal scientists to provide further input will be made. The future of agricultural genomics will be shaped by both the participants in AG2PI and the solutions developed by the AG2PI community. Funding for this work comes from USDA NIFA awards 2022-70412-38454, 2021-70412-35233, and 2020-70412-32615.
Over the past few years, due to the boom of advances in image processing, edge computing, and wireless networking, unpiloted aerial vehicles, often referred to as drones, have become an important enabler to support a wide variety of scientific applications, ranging from environmental monitoring, disaster response, and wildfire monitoring to the survey of archaeological sites. In this article, we present the FlyNet platform, which extends an existing workflow management system to support and manage scientific workflows. FlyNet enables automated resource allocation, workflow instrumentation, and network service support to support researchers in their goal to analyze data for new scientific discoveries. In addition, FlyNet provides network services management to support quality of service for efficient data transport between edge devices, edge servers, and the cloud.
The USDA-NIFA is developing a vision for agricultural genomics to phenomics research (AG2P) through two programs. In 2017, the Functional Annotation of Animal Genomes (FAANG) request for proposals was released, which has provided over $9 Million in funding to groups over the past five years to create public resources for the study of regulatory and other functional elements in agriculturally important species. This effort has been extended by other countries who have contributed another $20 Million+ (equivalent). These initial resource-building efforts are providing extensive new epigenetic data that are used to better understand the function of animal genomes and to enhance animal genetic improvement. However, such research demonstrates new infrastructure is needed across the agricultural enterprise to predict phenotypes from genome information. In 2020, NIFA inaugurated the Agricultural Genome to Phenome Initiative program (AG2PI), and an important objective of this presentation is to encourage strong participation of animal genomic researchers and industry in the AG2PI project. The purpose of AG2PI is to assemble and prepare a community to conduct AG2P research across the plant and animal kingdoms. The focus of the AG2PI is: a) identifying research gaps and opportunities, b) fostering community solutions to these challenges, and c) rapidly disseminating findings. Training workshops, field days, conferences, community surveys, and seed grants have all contributed to these goals. Opinions and ideas from the crop, livestock and aquaculture genetics communities are key to advancing G2P science and related work. The AG2PI uses online community surveys and in-person conferences to gather these and to facilitate transdisciplinary communications. In September 2022, AG2PI held a conference entitled “Thinking Big: Visualizing the Future of AG2PI”, in which attendees focused small group discussions to identify which agricultural G2P-related topics are most critical for future R&D funding and which may be the most difficult to achieve. These discussions have been summarized in a concept paper for the community and shared with USDA-NIFA. A follow-up conference is planned for June 2023 in Kansas City, which we hope many members from animal genetics research and industry will attend. The future of agricultural genomics will be shaped by both the participants in AG2PI and the solutions developed by the AG2PI community. USDA NIFA awards 2022-70412-38454, 2021-70412-35233, and 2020-70412-32615.
Visual Cloud Computing (VCC) applications provide highly efficient solutions in video data processing pipelines on edge/cloud infrastructures. These applications and their infrastructures demand end-to-end monitoring and fine-grained application traffic control to meet user quality of experience requirements. In this paper, we propose a novel network services management methodology for VCC applications by leveraging the advantages of programmable data planes enabled by Protocol-independent Packet Processors (P4). Specifically, we define a custom fixed-length application header and use it to improve the performance of a video streaming application through congestion avoidance using Multi-Hop Route Inspection (MRI), a variant of In-band Network Telemetry (INT), and switch port forwarding (tunneling) capabilities. For evaluation experiments, we use P4 per-packet telemetry metadata for routing paths, ingress/egress timestamps, queue occupancy in a given node, and egress port link utilization in a VCC testbed on the NSF-supported FABRIC infrastructure. Our experiment results demonstrate performance improvement obtained with our methodology in terms of both packet loss and throughput metrics.