Study objective: Develop a new genome browser using modern web technologies that is fast, easy to use, and achieves a high density of information. Methods: The new BioCyc genome browser uses Javascript plus an HTML5 canvas to provide fast, real-time zooming and panning from a chromosome-level view of a genome to a sequence-level view. Results: The BioCyc genome browser communicates the positions of functional elements within 20,000 genomes to scientists, and supports inference of new genome features from large datasets. Example functional elements include genes, transcription start sites, transcription-factor binding sites, and terminators. The BioCyc genome browser offers high-speed zooming across a scale factor of 1500 because its graphics are generated in the user's web browser rather than in a remote server. Its zooming is highly intuitive because it is performed using a mouse wheel or trackpad that enables both fast and precise zooming. As the user increases the magnification level, semantic zooming displays new features such as promoters, transcription-factor binding sites, and nucleotide and amino-acid sequences. Most other genome browsers use tracks to present genome features such as coding regions and transcription-factor binding sites. In contrast, the BioCyc genome browser generates diagrams that are more information dense than tracks-based diagrams, and in which positional relationships are easier to discern visually. The BioCyc diagrams integrate a standard array of genome features into compact, information-rich displays that can wrap across multiple lines. These diagrams are so compact that this genome browser can typically present a much longer genome region in the same screen size as can other browsers, providing scientists with more genome context. For example, in one example on a wide monitor the BioCyc browser depicted 2400 bases, whereas JBrowse depicted 125 bases. The graphical genome features incorporated in the BioCyc browser are: operons, protein-coding genes, RNA-coding genes, pseudogenes, transcription start sites, transcription-factor binding sites, terminators, attenuators, mRNA binding sites, ribosome binding sites, extragenic sites, and insertion elements. The BioCyc browser does provide a tracks mechanism for depicting large-scale datasets against the genome. In tracks mode, the browser no longer wraps its diagrams across multiple lines. The BioCyc browser provides a comparative mode in which multiple chromosomes are displayed aligned at orthologous genes. When the user zooms or pans the diagram, all chromosomes move in synchrony. When the user zooms to the sequence level, a sequence selection tool enables selection of arbitrary nucleotide and amino-acid sequences. Conclusions: The new BioCyc genome browser provides fast and intuitive operation; its innovative graphical layout enables more genome context to be present at a given time than other browsers. Example URLs for the new BioCyc genome browser, available by approximately December 11, 2023: Escherichia coli K-12 https://biocyc.org/genbro/genbro.shtml?orgid=ECOLI&replicon=COLI-K12 Comparative mode https://biocyc.org/genbro/ortho.shtml?lead-orgid=GCF_000963815&lead-genes=ABUW_RS08185&orgids=GCF_000963815,GCF_004797155,GCF_002734145,GCF_000008525 This work was funded by grant NSF2109898.
ABSTRACT Are two adjacent genes in the same operon? What are the order and spacing between several transcription factor binding sites? Genome browsers are software data visualization and exploration tools that enable biologists to answer questions such as these. In this paper, we report on a major update to our browser, Genome Explorer, that provides nearly instantaneous scaling and traversing of a genome, enabling users to quickly and easily zoom into an area of interest. The user can rapidly move between scales that depict the entire genome, individual genes, and the sequence; Genome Explorer presents the most relevant detail and context for each scale. By downloading the data for the entire genome to the user’s web browser and dynamically generating visualizations locally, we enable fine control of zoom and pan functions and real-time redrawing of the visualization, resulting in smoother and more intuitive exploration of a genome than is possible with other browsers. Further, genome features are presented together, in-line, using familiar graphical depictions. In contrast, many other browsers depict genome features using data tracks, which have low information density and can visually obscure the relative positions of features. Genome Explorer diagrams have a high information density that provides larger amounts of genome context and sequence information to be presented in a given-sized monitor than for tracks-based browsers. Genome Explorer provides optional data tracks for the analysis of large-scale data sets and a unique comparative mode that aligns genomes at orthologous genes with synchronized zooming. IMPORTANCE Genome browsers provide graphical depictions of genome information to speed the uptake of complex genome data by scientists. They provide search operations to help scientists find information and zoom operations to enable scientists to view genome features at different resolutions. We introduce the Genome Explorer browser, which provides extremely fast zooming and panning of genome visualizations and displays with high information density.
Spoken dialog systems, lacking the means to address the complex phenomena of spontaneous speech and conversational dynamics, force users into a constrained mode of dialog that resembles text-based interaction more closely than spoken conversation. Turn-taking is simplified and discourse-related information is lost, as discourse markers are largely ignored and prosodic information is not captured or utilized. We hypothesize that incorporating a few of these key conversational phenomena at specific points in a dialog will reduce cognitive load in spoken human-computer interaction and expand the potential application areas of dialog systems to tasks requiring more complex interactions. In this paper, we describe our approach to adding conversational intelligence to dialog systems and our work to date validating the hypothesis that adding conversational intelligence to existing dialog systems will significantly reduce users’ cognitive load.
Recent years have seen rapid advances in intelligent technology to support online learning, but these have primarily targeted formal educational contexts such as classrooms and e-courses. In contrast, the predominant form of adult learning in the workplace is informal and self-directed. Learners selfassess competency, set goals, find relevant learning resources, and initiate learning activities covering many topics at different depths at different points in time. Our approach to supporting selfregulated learning is embodied in PERLS, a mobile personal assistant application that serves as a virtual mentor for informal learning. A key component of PERLS is its recommendation system, designed to adaptively co-construct a path towards desired learning outcomes with the learner. Recommendations are made largely on the basis of value propositions, each a persuasive explanation for taking a particular learning action to advance along a particular learning path. In this paper, we present a process model of self-regulated learning used in PERLS and our approach to generating explained recommendations and using them to co-construct learning paths.
The Omics Dashboard is a software tool for interactive exploration and analysis of gene-expression datasets. The Omics Dashboard is organized as a hierarchy of cellular systems. At the highest level of the hierarchy the Dashboard contains graphical panels depicting systems such as biosynthesis, energy metabolism, regulation and central dogma. Each of those panels contains a series of X-Y plots depicting expression levels of subsystems of that panel, e. g. subsystems within the central dogma panel include transcription, translation and protein maturation and folding. The Dashboard presents a visual read-out of the expression status of cellular systems to facilitate a rapid top-down user survey of how all cellular systems are responding to a given stimulus, and to enable the user to quickly view the responses of genes within specific systems of interest. Although the Dashboard is complementary to traditional statistical methods for analysis of gene-expression data, we show how it can detect changes in gene expression that statistical techniques may overlook. We present the capabilities of the Dashboard using two case studies: the analysis of lipid production for the marine alga Thalassiosira pseudonana, and an investigation of a shift from anaerobic to aerobic growth for the bacterium Escherichia coli.
Pathway Tools is a bioinformatics software environment with a broad set of capabilities. The software provides genome-informatics tools such as a genome browser, sequence alignments, a genome-variant analyzer and comparative-genomics operations. It offers metabolic-informatics tools, such as metabolic reconstruction, quantitative metabolic modeling, prediction of reaction atom mappings and metabolic route search. Pathway Tools also provides regulatory-informatics tools, such as the ability to represent and visualize a wide range of regulatory interactions. This article outlines the advances in Pathway Tools in the past 5 years. Major additions include components for metabolic modeling, metabolic route search, computation of atom mappings and estimation of compound Gibbs free energies of formation; addition of editors for signaling pathways, for genome sequences and for cellular architecture; storage of gene essentiality data and phenotype data; display of multiple alignments, and of signaling and electron-transport pathways; and development of Python and web-services application programming interfaces. Scientists around the world have created more than 9800 Pathway/Genome Databases by using Pathway Tools, many of which are curated databases for important model organisms.
Inquire Biology is a prototype of a new kind of intelligent textbook that answers students questions, engages their interest, and improves their understanding. Inquire uses knowledge representation of the conceptual knowledge from the textbook and uses inference procedures to answer questions. Students ask questions by typing free-form natural language queries or by selecting passages of text. The system then attempts to answer the question and also generates suggested questions related to the query or selection. The questions supported by the system were chosen to be educationally useful, for example: what is the structure of X?; compare X and Y?; how does X relate to Y? In user studies, students found this question-answering capability to be useful while reading and while doing problem solving. In a controlled experiment, community college students using Inquire Biology outperformed students using either a hard copy or conventional E-book version of the same textbook. While additional research is needed to fully develop Inquire, the prototype demonstrates the promise of applying knowledge representation and question-answering to electronic textbooks.
Cognitive simulation of analogical processing can be used to answer comparison questions such as: What are the similarities and/or differences between A and B, for concepts A and B in a knowledge base (KB). Previous attempts to use a general-purpose analogical reasoner to answer such questions revealed three major problems: (a) the system presented too much information in the answer, and the salient similarity or difference was not highlighted; (b) analogical inference found some incorrect differences; and (c) some expected similarities were not found. The cause of these problems was primarily a lack of a well-curated KB and, and secondarily, algorithmic deficiencies. In this paper, relying on a well-curated biology KB, we present a specific implementation of comparison questions inspired by a general model of analogical reasoning. We present numerous examples of answers produced by the system and empirical data on answer quality to illustrate that we have addressed many of the problems of the previous system.
When designing the natural language question asking interface for a formal knowledge base, managing and scoping the user expectations regarding what questions the system can answer is a key challenge. Allowing users to type ask arbitrary English questions will likely result in user frustration, because the system may be unable to answer many questions even if it correctly understands the natural language phrasing. We present a technique for responding to natural language questions, by suggesting a series of questions that the system can actually answer. We also show that the suggested questions are useful in a variety of ways in an intelligent textbook to improve student learning.
Textbooks are increasingly moving into the digital realm, which presents an opportunity for them to evolve from providing the reader with a static, linear experience, into an interactive application that can adapt to a student as well as to specific learning goals. As a step in this direction, we present Inquire: Biology, an electronic textbook that provides question-answering capability.
The Adept Task Learning system is an end-user programming environment that combines programming by demonstration and direct manipulation to support customization by nonprogrammers. Previously, Adept enforced a rigid procedure-authoring workflow consisting of demonstration followed by editing. However, a series of system evaluations with end users revealed a desire for more feedback during learning and more flexibility in authoring. We present a new approach that interleaves incremental learning from demonstration and assisted editing to provide users with a more flexible procedure-authoring experience. The approach relies on maintaining a "soup" of alternative hypotheses during learning, propagating user edits through the soup, and suggesting repairs as needed. We discuss the learning and reasoning techniques that support the new approach and identify the unique interaction design challenges they raise, concluding with an evaluation plan to resolve the design challenges and complete the improved system.
EcoCyc (http://EcoCyc.org) is a comprehensive model organism database for Escherichia coli K-12 MG1655. From the scientific literature, EcoCyc captures the functions of individual E. coli gene products; their regulation at the transcriptional, post-transcriptional and protein level; and their organization into operons, complexes and pathways. EcoCyc users can search and browse the information in multiple ways. Recent improvements to the EcoCyc Web interface include combined gene/protein pages and a Regulation Summary Diagram displaying a graphical overview of all known regulatory inputs to gene expression and protein activity. The graphical representation of signal transduction pathways has been updated, and the cellular and regulatory overviews were enhanced with new functionality. A specialized undergraduate teaching resource using EcoCyc is being developed.
PathwayTools is a production-quality software environment for creating a type of model-organism database called a Pathway/Genome Database (PGDB). A PGDB such as EcoCyc integrates the evolving understanding of the genes, proteins, metabolic network and regulatory network of an organism. This article provides an overview of Pathway Tools capabilities. The software performs multiple computational inferences including prediction of metabolic pathways, prediction of metabolic pathway hole fillers and prediction of operons. It enables interactive editing of PGDBs by DB curators. It supports web publishing of PGDBs, and provides a large number of query and visualization tools. The software also supports comparative analyses of PGDBs, and provides several systems biology analyses of PGDBs including reachability analysis of metabolic networks, and interactive tracing of metabolites through a metabolic network. More than 800 PGDBs have been created using PathwayTools by scientists around the world, many of which are curated DBs for important model organisms. Those PGDBs can be exchanged using a peerto-peer DB sharing system called the PGDB Registry.
In the winter 2004 issue of AI Magazine, we reported Vulcan Inc.'s first step toward creating a question‐answering system called Digital Aristotle. The goal of that first step was to assess the state of the art in applied knowledge representation and reasoning (KRR) by asking AI experts to represent 70 pages from the advanced placement (AP) chemistry syllabus and to deliver knowledge‐based systems capable of answering questions from that syllabus. This article reports the next step toward realizing a Digital Aristotle: we present the design and evaluation results for a system called AURA, which enables domain experts in physics, chemistry, and biology to author a knowledge base and that then allows a different set of users to ask novel questions against that knowledge base. These results represent a substantial advance over what we reported in 2004, both in the breadth of covered subjects and in the provision of sophisticated technologies in knowledge representation and reasoning, natural language processing, and question answering to domain experts and novice users.
End-user development (EUD), the practice of users creating, modifying, or extending programs for personal use, is a valuable but often challenging task for nonprogrammers. From the beginning, EUD systems have shown that recommendations can improve the user experience. However, these usability improvements are limited by a reliance on handcrafted rules and heuristics to generate reasonable and useful suggestions. When the number of possible recommendations is large or the available context is too limited for traditional reasoning techniques, recommender technologies present a promising solution. In this paper, we provide an overview of the state of the art in end-user development, focusing on the different kinds of recommendations made to users. We identify four classes of suggestion that could most directly benefit from existing recommendation techniques. Along the way we explore straightforward applications of recommender algorithms as well as a few difficult but high-value recommendation problems in EUD. We discuss the ways that EUD systems have been evaluated in the past and suggest the modifications necessary to evaluate recommenders within the EUD context. We highlight EUD research as one area that can facilitate the transition of recommender system evaluation from algorithmic performance evaluation to a more user-centered approach. We conclude by restating our findings as a new set of research challenges for the recommender systems community.
As Web services become more diverse and powerful, end user programming (EUP) systems for the Web become increasingly compelling. However, many user workflows do not exist exclusively online. To support these workflows completely, EUP systems must allow the user to program across multiple applications in different domains. To this end, we created Integrated Task Learning (ITL), a system that integrates several learning components to learn end user workflows as user-editable executable procedures. In this chapter, we illustrate a motivating cross-domain task and describe the various learning techniques that support learning such a task with ITL. These techniques include dataflow reasoning to learn procedures from demonstration, symbolic analysis and compositional search to support procedure editing, and machine learning to infer new semantic types. Then, we describe the central engineering concept that ITL uses to facilitate cross-domain learning: pluggable domain models, which are independently generated type and action models over different application domains that can be combined to support cross-domain procedure learning. Finally, we briefly discuss some open questions that cross-domain EUP systems will need to address in the future.
Numerous techniques exist to help users automate repetitive tasks; however, none of these methods fully support end-user creation, use, and modification of the learned tasks. We present an integrated task learning system (ITL) that learns executable procedures based on user demonstration and instruction, constituting a first step toward a broader solution for procedure management. We discuss our deployment of ITL into a collaborative command-and-control system. In this complex domain, ITL's performance with end users doing real tasks indicates that providing multiple, integrated learning techniques both extends functionality and improves user experience. Our experience in integrat-ing this system also provides key insights for future designs of domain-independent task learning systems, specifically in supporting users' ability to understand and edit lengthy procedures.
When creating algorithms or systems that are supposed to be used by people, we should be able to adopt a "binocular" view of users' interaction with intelligent systems: a view that regards the design of interaction and the design of intelligent algorithms as interrelated parts of a single design problem. This special issue offers a coherent set of articles on two levels of generality that illustrate the binocular view and help readers to adopt it.
Peter D Karp合作论文数Artificial Intelligence Center, SRI International8
Bruce Porter合作论文数Department of Computer Science The University of Texas at Austin4