In this paper, we present a method that uses a physics based virtual environment to evaluate the feasibility of neural network-based generated designs. Deep learning models rely on large training data sets that are used for training. These training data sets are typically validated by human designers that have a conceptual understanding of the problem being solved. However, the requirement of human training data severely constrains the size and availability of training data for computer generated models due to the manual process of either creating or labeling such data sets. Furthermore, there may be misclassification errors that result from human labeling. To mitigate these challenges, we present a physics-based simulation environment that helps users discover correlations between the form of a generated design and the physical constraints that relate to its function. We hypothesize that training data that includes machine validated designs from a physics-based virtual environment will increase the probability of generative models creating functionally-feasible design concepts. A case study involving a generative model that is trained on over 70,000 human 2D boat sketches is used to test the hypothesis. Knowledge gained from testing this hypothesis will provide human designers with insights into the importance of training data in the resulting design solutions generated by deep neural networks.
The authors of this work present a method that mines big media data streams from large Social Media Networks in order to discover novel correlations between objects appearing in images and electricity utilization patterns. The hypothesis of this work is that there exist correlations between what users take pictures of, and electricity utilization patterns. This work employs a Convolutional Neural Network to detect objects in 578,232 images gathered from over 15,000,000 tweets sent in the San Diego area. These objects were considered in the context of concurrent power use, on a monthly and hourly basis. The results reveal both positive and negative correlations between power use and specific objects, such as lamps (.053 hourly), dogs (−.011 hourly), horses (.422 monthly) and motorcycles (−.415, monthly).
Quantifying the ability of a digital design concept to perform a function currently requires the use of costly and intensive solutions such as computational fluid dynamics. To mitigate these challenges, the authors of this work propose a deep learning approach based on three-dimensional (3D) convolutions that predict functional quantities of digital design concepts. This work defines the term functional quantity to mean a quantitative measure of an artifact's ability to perform a function. Several research questions are derived from this work: (i) Are learned 3D convolutions able to accurately calculate these quantities, as measured by rank, magnitude, and accuracy? (ii) What do the latent features (that is, internal values in the model) discovered by this network mean? (iii) Does this work perform better than other deep learning approaches at calculating functional quantities? In the case study, a proposed network design is tested for its ability to predict several functions (sitting, storing liquid, emitting sound, displaying images, and providing conveyance) based on test form classes distinct from training class. This study evaluates several approaches to this problem based on a common architecture, with the best approach achieving F scores of >0.9 in three of the five functions identified. Testing trained models on novel input also yields accuracy as high as 98% for estimating rank of these functional quantities. This method is also employed to differentiate between decorative and functional headwear, which yields an 84.4% accuracy and 0.786 precision.
An important part of the engineering design process is prototyping, where designers build and test their designs. This process is typically iterative, time consuming, and manual in nature. For a given task, there are multiple objects that can be used, each with different time units associated with accomplishing the task. Current methods for reducing time spent during the prototyping process have focused primarily on optimizing designer to designer interactions, as opposed to designer to tool interactions. Advancements in commercially available sensing systems (e.g., the Kinect) and machine learning algorithms have opened the pathway toward real-time observation of designer's behavior in engineering workspaces during prototype construction. Toward this end, this work hypothesizes that an object O being used for task i is distinguishable from object O being used for task j, where i is the correct task and j is the incorrect task. The contributions of this work are: (i) the ability to recognize these objects in a free roaming engineering workshop environment and (ii) the ability to distinguish between the correct and incorrect use of objects used during a prototyping task. By distinguishing the difference between correct and incorrect uses, incorrect behavior (which often results in wasted time and materials) can be detected and quickly corrected. The method presented in this work learns as designers use objects, and infers the proper way to use them during prototyping. In order to demonstrate the effectiveness of the proposed method, a case study is presented in which participants in an engineering design workshop are asked to perform correct and incorrect tasks with a tool. The participants' movements are analyzed by an unsupervised clustering algorithm to determine if there is a statistical difference between tasks being performed correctly and incorrectly. Clusters which are a plurality incorrect are found to be significantly distinct for each node considered by the method, each with p ≪ 0.001.
The accuracy of RGB-D sensing has enabled many technical achievements in applications such as gamification, task recognition, as well as pedagogical applications. The ability of these sensors to track many body parts simultaneously has introduced a new data modality for analysis. By analyzing body language, this work can predict if a student will struggle in the future, and if an instructor should intervene. To accomplish this, a study is performed to determine how early (after how many seconds) does it become possible to determine if a student will struggle. A simple neural network is proposed which is used to jointly classify body language and predict task performance. By modeling the input as both instances and sequences, a peak F Score of 0.459 was obtained, after observing a student for just two seconds. Finally, an unsupervised method yielded a model which could determine if a student would struggle after just 1 second with 59.9% accuracy.
This work describes how automated data generation integrates in a big data pipeline. A lack of veracity in big data can cause models that are inaccurate, or biased by trends in the training data. This can lead to issues as a pipeline matures that are difficult to overcome. This work describes the use of a Generative Adversarial Network to generate sketch data, such as those that might be used in a human verification task. These generated sketches are verified as recognizable using a crowd-sourcing methodology, and finds that the generated sketches were correctly recognized 43.8% of the time, in contrast to human drawn sketches which were 87.7% accurate. This method is scalable and can be used to generate realistic data in many domains and bootstrap a dataset used for training a model prior to deployment.
The hypothesis of this paper is that topics, expressed through large-scale social media networks, approximate electricity utilization events (e.g., using high power consumption devices such as a dryer) with high accuracy. Traditionally, researchers have proposed the use of smart meters to model device-specific electricity utilization patterns. However, these techniques suffer from scalability and cost challenges. To mitigate these challenges, we propose a social media network-driven model that utilizes large-scale textual and geospatial data to approximate electricity utilization patterns, without the need for physical hardware systems (e.g., such as smart meters), hereby providing a readily scalable source of data. The methodology is validated by considering the problem of electricity use disaggregation, where energy consumption rates from a nine-month period in San Diego, coupled with 1.8 million tweets from the same location and time span, are utilized to automatically determine activities that require large or small amounts of electricity to accomplish. The system determines 200 topics on which to detect electricity-related events and finds 38 of these to be valid descriptors of energy utilization. In addition, a comparison with electricity consumption patterns published by domain experts in the energy sector shows that our methodology both reproduces the topics reported by experts, while discovering additional topics. Finally, the generalizability of our model is compared with a weather-based model, provided by the U.S. Department of Energy.
The authors of this work present a computer vision approach that discovers and classifies objects in a video stream, towards an automated system for managing End of Life (EOL) waste streams. Currently, the sorting stage of EOL waste management is an extremely manual and tedious process that increases the costs of EOL options and minimizes its attractiveness as a profitable enterprise solution. There have been a wide range of EOL methodologies proposed in the engineering design community that focus on determining the optimal EOL strategies of reuse, recycle, remanufacturing and resynthesis. However, many of these methodologies assume a product/component disassembly cost based on human labor, which hereby increases the cost of EOL waste management. For example, recent EOL options such as resynthesis, rely heavily on the optimal sorting and combining of components in a novel way to form new products. This process however, requires considerable manual labor that may make this option less attractive, given products with highly complex interactions and components. To mitigate these challenges, the authors propose a computer vision system that takes live video streams of incoming EOL waste and i) automatically identifies and classifies products/components of interest and ii) predicts the EOL process that will be needed for a given product/component that is classified. A case study involving an EOL waste stream video is used to demonstrate the predictive accuracy of the proposed methodology in identifying and classifying EOL objects.
Static analysis has been successfully used in many areas, from verifying mission-critical software to malware detection. Unfortunately, static analysis often produces false positives, which require significant manual effort to resolve. In this paper, we show how to overlay a probabilistic model, trained using domain knowledge, on top of static analysis results, in order to triage static analysis results. We apply this idea to analyzing mobile applications. Android application components can communicate with each other, both within single applications and between different applications. Unfortunately, techniques to statically infer Inter-Component Communication (ICC) yield many potential inter-component and inter-application links, most of which are false positives. At large scales, scrutinizing all potential links is simply not feasible. We therefore overlay a probabilistic model of ICC on top of static analysis results. Since computing the inter-component links is a prerequisite to inter-component analysis, we introduce a formalism for inferring ICC links based on set constraints. We design an efficient algorithm for performing link resolution. We compute all potential links in a corpus of 11,267 applications in 30 minutes and triage them using our probabilistic approach. We find that over 95.1% of all 636 million potential links are associated with probability values below 0.01 and are thus likely unfeasible links. Thus, it is possible to consider only a small subset of all links without significant loss of information. This work is the first significant step in making static inter-application analysis more tractable, even at large scales.
Many program analyses require statically inferring the possible values of composite types. However, current approaches either do not account for correlations between object fields or do so in an ad hoc manner. In this paper, we introduce the problem of composite constant propagation. We develop the first generic solver that infers all possible values of complex objects in an interprocedural, flow and context-sensitive manner, taking field correlations into account. Composite constant propagation problems are specified using COAL, a declarative language. We apply our COAL solver to the problem of inferring Android Inter-Component Communication (ICC) values, which is required to understand how the components of Android applications interact. Using COAL, we model ICC objects in Android more thoroughly than the state-of-the-art. We compute ICC values for 460 applications from the Play store. The ICC values we infer are substantially more precise than previous work. The analysis is efficient, taking slightly over two minutes per application on average. While this work can be used as the basis for many whole-program analyses of Android applications, the COAL solver can also be used to infer the values of composite objects in many other contexts.
With the rise in both personal and official smartphone use by the military, large scale application store analysis is an appealing idea, but until now there has been no real method of quickly accumulating a large library for analysis. In this paper we present a population study of permission and library use in Android apps, obtained from the Google Play Store. To accomplish this, we present a novel method for quickly reconnoitering a large database using a novel wordlist approach. Employing this method we compiled a library of over 700, 000 applications. By leveraging program analysis and data mining techniques, we analyzed how permissions and libraries are used in real world settings on a population-wide level. From this we were able to make several claims about the health and apparently direction of the Android ecosystem.