
Deep neural networks (DNNs) achieve state-of-the-art performance in many areas, including computer vision, system configuration, and question-answering. However, DNNs are expensive to develop, both in intellectual effort (e.g., devising new architectures) and computational costs (e.g., training). Re-using DNNs is a promising direction to amortize costs within a company and across the computing industry. As with any new technology, however, there are many challenges in re-using DNNs. These challenges include both missing technical capabilities and missing engineering practices. This vision paper describes challenges in current approaches to DNN re-use. We summarize studies of re-use failures across the spectrum of re-use techniques, including conceptual (e.g., re-using based on a research paper), adaptation (e.g., re-using by building on an existing implementation), and deployment (e.g., direct re-use on a new device). We outline possible advances that would improve each kind of re-use.
History, including the history of computing, is always in danger of distortion by our very human love for legends. Legends are fun, but they can be inaccurate in both positive and negative directions, making heroes of some individuals and villains of others. This paper attempts a more fact-based and technical look at one of the most emotion-prone chapters in computing history, the creation of the Atanasoff-Berry Computer (ABC) in the late 1930s and its influence on the computers that came later.
The proliferation of scientific applications, the Internet of Things (IoT), social media, and e-commerce has led to an exponential growth in data generation, consequently fueling the demand for large-scale data analytics systems. Consequently, the transfer of data over the Internet has escalated to an unprecedented scale, surpassing the zettabyte mark [37]. As data generation rates continue to surge, the carbon footprint associated with data movement has emerged as a pressing concern, particularly for High-Performance Computing (HPC) and Cloud data centers. Projections indicate that by 2030, information and communication technologies will account for 8%–21% of global electricity usage [6]. HPC and Cloud data centers and communication networks contribute to 69% of the overall power consumption within the IT sector [4]. Within this segment, data transfers alone consume over a hundred terawatt-hours of energy annually, amounting to a staggering $20 billion USD [6]. The environmental implications are equally monumental, with information and communication technologies projected to be responsible for a staggering 14% of carbon emissions by 2040 [6]. These trends have spurred significant efforts to reduce energy consumption in hardware and software systems, as well as networking devices.
Today's rapid progress in AI and science is largely fueled by the availability of ever larger and more powerful compute systems. The “classic” HPC systems targeted at executing complex workflows and simulations have recently crossed the exaflop boundary in terms of their double-precision floating point performance. At the same time, new systems targeted at training large AI-models use alternative number representations and are already pushing the limits well beyond the ten exaflop mark. To continue scaling the performance of large HPC systems, system architects need to address several barriers including the slowdown of Moore's law, energy density limitations, production yield challenges and practical limits to overall power consumption in the 10s-of- MW range. All recent #1 HPC systems are already relying on specialized, heterogenous components to offset the slowdown. As specialization continues and advances, the heterogeneity will evolve from today's CPU-GPU combinations into a broad set of more specialized accelerators, but also entirely new computing paradigms, e.g., Quantum computing are emerging. With the continued scaling of total system size and the compute density within a single node, the intra- and inter-node communication requirements increase accordingly. Today, most available interconnect fabrics that support symmetric multiprocessing (SMP) and/or asymmetric variants of cache-coherent communication are based on proprietary implementations, which prevent the assembly of heterogeneous high-performance systems from components from more than a single vendor. Hence, system architects are looking at ways assemble innovative high-performance heterogenous systems using open standards. Under continued cost constraints, better utilization is desired to match hardware configuration to software usage needs. The ability to compose virtual compute nodes from a set of disaggregated components is a natural way of approaching the problem. The first challenge of composability that is currently being tackled is memory disaggregation. A vision of higher utilization and resource sharing is appealing, but low latency and high bandwidth need to be maintained. All these trends and limitations demand a fresh look at the architectures enabling an ecosystem from which domain-specific high-performance computing systems can be assembled. In this presentation, we discuss the motivation and requirements for a new node level and rack scale architecture as well as the need for an open standards-based, composable, high-performance interconnect fabric. The architecture of this system needs to be accompanied by an open and interoperable software stack as well as a fine-grained control plane. The control plane enables and supports composability under tight security and performance constraints. While composability originated to increase the efficiency of heterogenous computer systems, more recently, it has been proposed as a means for heterogeneous components to share a common memory pool, reduce data traffic, and increase the speed of cooperation among the system components. At the lowest level, it is critical for CPU-and accelerator cores to have their own memory hierarchies (L1, L2, LLC). Shared memory pools could be very effective mechanism for coordinating a workflow across the heterogenous system components. Workflows can then evolve from a file-based sharing method to a shared memory model utilizing a high-speed fabric, e.g., CXL. Composability also addresses the sustainability issues of large computing systems by allowing for upgrades of individual parts in the system. The use of heterogenous components must expand from the current rack-level or board-level integration down to chiplet-based modules, and even System-on-Chip (SoC), depending on the scale and demands of a workflow. A standards-based coherent interconnect fabric is a key element that will allow innovations from different heterogeneous components to be mixed beyond the limitations of any single vendor and is also a key ingredient for an industry growth play. For board-level connections outside of the SMP fabric, the evolving CXL standards are a good match for this role as they support traditional I/O connect plus a scalable memory extension. CXL over the emerging UCIe connection standard offers the possibility to extend this value proposition to a chiplet-based ecosystem, where tight integration into the SMP fabric is not required. For extended reach, CXL over an optics standard would provide for even larger scale composable systems. The composable elements in a compute fabric need a distributed control structure for initialization, resource management, and workflow control. Open standards such as OFMF will play a critical role in the overall system management. The emergence of confidential computing as a paradigm for reducing the trusted computing base (TCB) of a computation is also essential for HPC and cloud. Open standards are required to enable confidential computing's trusted execution environments (TEEs) across heterogeneous elements. Security features will be required to address supply chain attacks, secure and trusted boot, authentication and attestation of each component, enable secure and confidential communication between the various components in the heterogenous system.
Binary decision diagrams (BDDs) have been a huge success story in hardware and software verification and are increasingly applied to a wide range of combinatorial problems. While BDDs can encode boolean-valued functions of boolean-valued variables, many BDD variants have been proposed, not just to improve their efficiency, but to manage multivalued domains (a straightforward extension), multivalued ranges (using several competitive alternatives), and two-dimensional data (relations and matrices instead of sets or vectors). Orthogonally to these extensions, much effort has been spent on variable order heuristics, an essential aspect that can affect memory and time requirements by up to an exponential factor. We survey some of these exciting results and discuss some fruitful research directions for further work.
Radiology has been transformed in recent decades from a specialty that relied on subjective interpretation of qualitative assessment of medical images to one that is now informed by objective analyses of medical images as quantitative constructs through advanced computational methods. This paper discusses medical images as mathematical data structures, from the pipeline of physics and engineering through which medical images are acquired to the unique attributes of images obtained from patients in a clinical setting to post-acquisition image processing that leads to computer-aided diagnosis (now dominated by artificial intelligence systems). The magnitude of data contained within images from a single imaging study and the volume of imaging studies stored in a clinical archive or required for research clearly warrant the label “big data.”
The main goal of this plenary panel at IEEE SERVICES 2023 is to review John Vincent Atanasoff's revolutionary invention of electronic digital computing. It will serve as a tribute to this pioneer's exceptional accomplishments and as a celebration of Atanasoff's 120 th birthday. Back in 1939, the first proof-of-concept prototype of electronic digital computer became operational. The plenary panel discussion will recognize the contributions of John Vincent Atanasoff for the invention and early development of electronic digital computing and computers that changed the world.
Currently, AI and, in particular, deep learning play a major role in science, from data analytics and simulation surrogates to policy and system decisions. This role is likely to increase as ideas from early adopters spread across all academic fields. One can group the structure of “AI for Science” into a few patterns, where one needs to explore examples of each pattern, possibly leading to Foundation models for each or maybe all of them combined. We suggest examining each pattern and supporting it with high-performance, easy-to-use environments for end-to-end systems, including data engineering. This support would cover parallelism, storage and data movement, security, and the user interface. We discuss the relationship between patterns, Foundation models, and benchmarks.
In this paper, we present an overview of the evolution of autonomous streaming-based services. Our hypothesis presented in Section 1. is that the technologies for autonomous assisting in urbanized or structured environments are in transition from fixed or portable and wearable devices to those in which the main function is supported by independent and possibly self-initiated and controlled movement of the devices. Furthermore, as was the case with fixed devices, the moving devices initially function in isolation and separation. But as their penetration and technologies progress, service models of groups of multiple devices are also being implemented. This transition is outlined and categorized in Section 2. We propose a multidimensional and layered taxonomy for the basic case of single mobile autonomous systems for 1d-, 2d-, and 3d-movement models. Our classification covers both the functional and technological aspects of these systems. In Section 3. we consider the parameters of three exemplary platforms as use cases for the three main modes of movement i.e. surface walking, surface wheeled movement, and aerial (or fluid) propelling. Finally, in the Conclusion we address the applicability of our parametrization for the purposes of standardization and interoperability as the main conditions for the service implementations based on a group of multiple devices.
In this work, we demonstrate how quantum control methods can be applied in cloud-based quantum computers. In particular, we demonstrate the use of composite pulses to perform robust population transfer, single-qubit gates, and reduction of leakage out of the computational subspace. The experiments have been performed on one of IBM's quantum computers, demonstrating excellent agreement with theoretical predictions.
Development and deployment of software as a medical device (SaMD) have increased in modern times. At the U.S. Food and Drug Administration, we ensure that such medical devices are safe and effective within their intended use cases. This paper reviews the FDA approval process for medical devices and describes several regulatory science efforts within the Division of Imaging, Diagnostics, and Software Reliability to understand and facilitate the regulatory review of computational medical devices.
It's been 10 years since the IEEE launched the Rebooting Computing Initiative (RCI) with the intention to re-examine all levels of how we compute. Back then, the term “post-Moore” was just coming into vogue. The RCI held several invitation-only summits and then opened the doors to others by launching the International Symposium on Rebooting Computing. Through the intervening years, many new ideas in how to rethink our computing levels of abstraction have been proposed. This paper examines some of those proposals, discusses their current status, and reviews the road ahead
Computational medicine has revolutionized healthcare by leveraging cutting-edge computing technologies to advance medical research and improve patient outcomes. For decades, high performance computing (HPC) has been a great ally to biological modeling but over the past few years HPC use for biomedical applications has been growing steadily. With the rise of exascale computing, promising to deliver unprecedented computing power and scalability, the potential of computational medicine is even greater. This paper discusses opportunities and some challenges exascale computing presents for advancing computational medicine. In addition, the paper highlights some applications addressing pressing healthcare delivery challenges of our time.
We design a lightweight overlay network, called Arigatoni, that is suitable to deploy the global computing paradigm over the Internet. Communications over the behavioral units of the model are performed by a simple communication protocol. Basic global computers can communicate by first registering to a brokering service and then by mutually asking and offering services, in a way that is reminiscent to Rapoport's "tit-for-tat" strategy of cooperation based on reciprocity. In the model, resources are encapsulated in the administrative domain in which they reside, and requests for resources located in another administrative domain traverse a broker-2-broker negotiation using classical PKI mechanisms. The model is suitable to fit with various global scenarios from classical P2P applications, like file sharing, or band-sharing, to more sophisticated grid applications, like remote and distributed big (and small) computations, to possible, futuristic real migrating computations. Indeed, our model fits some of the objectives suggested by the CoreGrid network of excellence, as described in Schwiegelshohn et al. (Schwiegelshohn et al., 2005)
Recent development in the aspects of low-level image processing for feature extraction, as well as in the standardization of description schemes for video description, provide means for the development of Intelligent Transportation Systems. ITS may support a wide range of abilities, including vehicle identification and reidentification, event detection and optimum resource management. This paper analyzes the structure operation principles and the structure of an integrated highway management system. Its operation is based on the combination of sensor networks, low-level image processing algorithms and high-level description schemes. Low-level algorithms perform the tasks of license plate recognition for vehicle identification, as well as feature extraction through change detection, for event detection purposes. High level schemes are based on the recently developed MPEG-7 Description Schemes and are utilized as an information processing framework, which formulates identification and event representation procedures.
John Vincent Atanasoff has been called the father of the computer and his machine, the Atanasoff-Berry Computer, is known as the first electronic digital computer. This paper examines these statements, not from the point of view of the usual computer scientist, but from that of the professional historian. What does it mean to be first? Is it important to be first? On what basis is the claim made? What other devices might be considered to share this honor? We provide an historian's view of these questions and illustrate them with examples from the early history of computing machines. We hope these examples indicate that a professional historian looks at these questions differently than most
Data replication can reduce access time and improve fault tolerance and load balancing. Typical requirements for a replica management system include an upper bound on replica Round Trip Time, scalability, reliability, self-management and selforganization, and ability to maintain consistency of mutable replicas. This article presents the design and a prototype implementation of a scalable, autonomous, service-oriented replica management framework for Globus Toolkit Version 4 using DKS. DKS is a structured peer-to-peer middleware. Grid nodes are integrated into a P2P network. The framework uses the ant metaphor and techniques of multi-agent systems for collaborative replica selection. We propose also a complimentary "background" service that collects access statistics and optimizes replica placement based on access pattern and replica lifetimes statistics. We have tested and profiled the prototype.
Determining the context of users and machines is an important topic in current computing research. An essential detail of a physical object's context is its location, which includes both the actual position as well as the semantics of the surroundings. This paper focuses on the specific problem of determining the position of objects and people within buildings. A low-cost approach is based on wireless LANs, which are now widely deployed. The paper presents a sophisticated probabilistic algorithm for indoor positioning using wireless LANs, but also discusses the problems that need to be solved to make indoor geolocation commonplace.
The authors present a natural deduction calculus for the computation tree logic, CTL, defined with the full set of classical and temporal logic operators. The system extends the natural deduction construction of the linear-time temporal logic. This opens the prospect to apply our technique as an automatic reasoning tool in a deliberative decision making framework across various applications in AI and computer science, where the branching-time setting is required
In this paper we consider the capabilities of nearest neighbor as a criterion of data association to ensure correct decision in associating measurement to target. We continue the work started in our previous paper to establish analytical approach end to derive explicit expressions for calculating correct data association probability for different levels of false alarm densities. Recurrent equation for correct association probability is derived. By means of this equation expressions for some particular values of FA densities are received. The correctness of the derivation is proved by the results of exhaustive Monte Carlo simulations.