Members of the Western Michigan Transformative Interdisciplinary Human+AI Research Group have been engaged in two consecutive NSF-funded projects to promote AI readiness in diverse STEM disciplines. Putting equal emphasis on theory and practice, our goal is to instill knowledge and competency in safe, secure, and reliable AI across a wide range of learners from high school students through to university students and practitioners who wish to upskill. The second project that is currently underway has a specific focus on machine-assisted processing of massive data. This presentation focuses on the development of immersive learning experiences.
Information systems are increasingly using artificial intelligence (AI). However, AI can be tricked into misbehaving, showing bias, or committing abuse. The root causes of these errors and uncertainties can be hidden away while parallelizing AI algorithms on high-performance computing (HPC) infrastructure. The project outlined in this paper aims to use artificial intelligence from the ground up to generate teaching materials and curricula for student-teachers. Students embark on a journey of discovery, taking calculated risks in a learning environment. The main purpose of this document is to present the primary research results of the two-year pilot project. A secondary purpose of this paper is to disseminate information about this exciting endeavor to encourage like-minded educators and researchers to participate in this project.
Phishing is a criminal act in which a Phisher sends a well- counterfeit Webpage, using its Lexical, Host-Based, and Content (LHC) features, containing unseen security threats and stealthy attacks to unwary victims tickling them to disclose sensitive credentials such as financial data, address, etc. The Webpage will probably pass under anti-phishing techniques (APT) because they mainly focus on detecting and classifying Webpages as either Malicious or Benign, neglecting Webpage traffic behavior (TB). In this research, we propose the detection, prevention, and classification (DPC) of Webpages' (W) suspicious and malicious activities based on their TB model, namely DPC based on the B - WTB or L2 model, as the second line of defense against a classified Benign Webpage has passed under the APT line of defense undetected. The L2 model is encapsulated in a sandbox to avoid system failure and keep attacks from spreading around the network, which will classify L1 Webpages as Benign, Suspicious, or Malicious based on their TB when they attempt to access unauthorized resources. Using 10369 records from ISCX - URL2016 dataset, the L2 model achieves an accuracy of 90.07%, 91.85%, and 92.62%, using KNN, LR, and SVM machine learning algorithms. In addition, the implementation of the proposed L2 model shows a significant observation regarding classified Webpages' attempt to access restricted resources based on their maximum number of access violation attempts for each of the restricted resources and an accumulative number of access attempts over time for each violation access attempts on the restricted resources. The experimental results show the precision score, the recall score, and the F1 score for each model.
Artificial intelligence (AI) is increasingly being applied to disciplines beyond computer science (CS). Engineers, statisticians, business analysts, biologists, physicists, physicians, and pharmacists, are among the many non-CS professionals who leverage the power of AI algorithms and systems for solving their domain-specific problems. Although AI has been found useful for solving a wide range of previously unsolvable problems, there are important limitations associated with contemporary AI. It is therefore important to inform current and future AI users regarding both strengths and weaknesses of AI in its current form, as well as what AI will be like in the foreseeable future. In this paper, the authors describe a pedagogical approach toward educating AI users from a range of STEM disciplines so that they can best exploit what AI has to offer. Specifically, a balanced approach is taken to ensure that learners gain knowledge and skills in what AI can or cannot do for them. A growing suite of experiential learning modules, which complement existing educational resources, serve as a vehicle for getting STEM learners ready for a future workplace characterized by significant use of AI technologies. These learning modules promote active learning and can be applied in a traditional classroom setting, self-directed online study, or a mix of the two modes. The paper ends with a presentation of encouraging results of actual use of the experiential learning modules in a mixed mode setting across multiple quantitative disciplines. All project artifacts, including the developed experiential learning modules, recommended uses, and best practices, are freely available on the project website. Interested educators and researchers are welcome to use the available resources and/or contribute to the on-going research.
Artificial Intelligence (AI) has experienced a strong revival recently. From autonomous vehicles to smart factories, AI is increasingly being used in STEM fields beyond computer science (CS). While AI is regularly taught in CS curricula, treatment of AI varies in other STEM disciplines. This paper describes an on-going project aimed at elevating AI knowledge and skills of non-CS STEM students and professionals. It emphasizes trustworthiness as a key factor of effective AI usage for large scale data analysis. Specifically, ten initial modules, which take an experiential learning and problem-solving pedagogical approach, have been developed and are now being piloted in fall 2021. They have been designed to be used both in a traditional classroom setting and as self-guided learning aid. This paper reports on the findings to date and aims to disseminate this exciting venture broadly for like-minded researchers to consider.
Artificial intelligence (AI) is increasingly applied to IT systems. However, AI can be manipulated to perform undesirably, exhibit biases or abusive behaviors. When AI algorithms are parallelized on high-performance computing-based cyberinfrastructure (CI), such misbehaviors and uncertainty can multiply to obscure the root causes. Secure, safe, and reliable computing techniques can mitigate these problems. The project described in this paper aims to inform curriculum and develop materials to educate students who use AI from the outset, so that they will first become aware of the issues and secondly practical considerations will be integrated with theory in classes. Intensive, multi-faceted, modular, experiential learning units are designed to rapidly upgrade the skills of current and future CI users, so they can apply new skills to their tasks. The loosely coupled modules can be taken as standalone self-directed units or integrated into existing classes, starting with CS 1 and CS 2, which are taken by many non-CS STEM students. In a sandpit environment, learners take measured risks when guided on a journey of discovery. The primary purpose of this paper is to present key findings of the research following a 2-year pilot. A secondary purpose of the paper is to disseminate this exciting endeavor broadly, so that likeminded educators and researchers can consider participating in the project.
Machine commonsense reasoning (MCR) systems can significantly improve the way we interact with machines. MCR systems are therefore an important element in any human-centric applications. Recent advances in machine learning (ML) have enabled breakthroughs in MCR technologies. This paper aims to improve healthcare outcomes by making human-machine interactions more intuitive than before. It presents learning models developed for MCR. Specifically, it presents a critical analysis of state-of-the-art deep learning (DL) models for MCR. These include recurrent neural network (RNN), transfer learning (TL), and transformers. Transformers, in particular, have been found to be effective for a range of natural language processing (NLP) applications, including MCR. Based on the analysis, another contribution of this paper is to assemble useful MCR tools into an adaptable MCR toolbox. To ensure broad applicability, the toolbox can be customizable for different MCR applications. Our research focuses on two specific MCR applications: commonsense validation and commonsense explanation. The former concerns identifying statements that do not make commonsense. The latter aims at explaining the reason why a given statement does not make commonsense. The paper presents some preliminary results of applying elements of the assembled toolbox to the two MCR applications. These results indicate that it is possible to achieve near human performances using finely-tuned state-of-the-art DL methods for the two MCR applications.
Artificial intelligence (AI) is increasingly applied to IT systems. However, AI can be manipulated to perform undesirably, exhibit biases or abusive behaviors. When AI algorithms are parallelized on high-performance computing-based cyberinfrastructure (CI), such misbehaviors and uncertainty can multiply to obscure the root causes. Secure, safe, and reliable computing techniques can mitigate these problems. The project described in this paper aims to inform curriculum and develop materials to educate students who use AI from the outset, so that they will first become aware of the issues and secondly practical considerations will be integrated with theory in classes. Intensive, multi-faceted, modular, experiential learning units are designed to rapidly upgrade the skills of current and future CI users, so they can apply new skills to their tasks. The loosely coupled modules can be taken as standalone self-directed units or integrated into existing classes, starting with CS 1 and CS 2, which are taken by many non-CS STEM students. In a sandpit environment, learners take measured risks when guided on a journey of discovery. The primary purpose of this paper is to present key findings of the research following a 2-year pilot. A secondary purpose of the paper is to disseminate this exciting endeavor broadly, so that likeminded educators and researchers can consider participating in the project.
Computational natural language processing (NLP) is indispensable in a humanized ambience intelligence environment. NLP facilitates ambient intelligence by making machines understand, and be understood by, humans. This in turn makes machines behave more human-like than they typically are today. Technological advances in machine learning (ML), and especially deep learning (DL), have been a key enabler of NLP research. This paper begins with a survey of recent developments of ML/DL for NLP. It then identifies some of the most promising techniques reported in recent literature. These most promising techniques are then assembled into a reusable toolkit for computational NLP. The adaptable nature of the assembled toolkit allows it to be reused in a broad range of NLP applications. The paper then describes experimental evaluation of our implemented solutions for comparative analysis. Two specific NLP applications form the basis of comparative evaluation. The first involves identifying one of M English sentences that does not make sense. The second, which is harder than the first, involves choosing from among N sentences the one that best explains why a presented sentence is invalid. Human baseline accuracies for these applications are 99.1% and 97.8%, respectively. The observation that these results are somewhat less than perfect demonstrates that even humans can occasionally find these tasks difficult. It further underscores the difficulties involved in some of these computational NLP applications. Experiments conducted on benchmark data show that advanced ML/DL can achieve near-human performance in both computational NLP applications with accuracy scores of 96.1% and 93.7%, respectively.
Phishing is a criminal act in which a Phisher creates almost identical website connections exploiting URL Lexical characteristics to dupe unsuspecting users into exposing sensitive information such as financial data, address, and other personal information. Phishers recently sought to trick security experts by masking malicious URLs with obfuscation techniques to make them appear legitimate. This action leads us to conclude that URL lexical features analysis approaches are absolute procedures, and extract analysis techniques are required. This research is a first step toward designing and developing a decision-making system that uses a combination of URL Lexical, and Network Traffic features to detect and classify malicious URLs rather than relying solely on Lexical or URL Network Traffic features. To achieve our goal, we examined and assessed the usage of URL Lexical and Network Traffic features to detect malicious URLs. In the study, three methodologies are used: Complete Features, KMO test as a features selection method, and PCA as a dimensionality method, which are tested by LR, SVM, and KNN classification algorithms and evaluated by the Confusion Matrix Accuracy measure. Using Network Traffics features (ISCXURL dataset), the W/O approach: LR, SVM, and KNN has 92%, 94%, and 93% accuracy. The KMO approach: SVM has 91% accuracy. The PCA approach: LR and SVM have 92% and 94% accuracy, surpassing the use of Lexical features (UCI dataset). In contrast, using Lexical features (UCI dataset), the KMO approach: LR and KNN has 90% and 94% accuracy. The PCA approach: KNN has 95% accuracy, surpassing the use of Network Traffic features (ISCXURL dataset). As a result, we are confident in proceeding with the next step of designing and developing a decision-making application that detects and classifies malicious URLs utilizing URL Lexical and Network Traffic features.
Western Michigan University, together with public and private partners, have developed ten learning units that promote artificial intelligence (AI) competency across STEM disciplines. The learning units are modular, experiential, customizable, and fun to use. They have been developed for both traditional classroom and self-directed learning. The modules are loosely coupled, so that learners can choose different pathways and modes of usage to suit. With an emphasis on “learn by doing”, the modules are experiential and can be customized for different STEM disciplines. The primary aim of the proposed workshop is to provide participants with hands-on experience of the learning units for themselves and go through a guided journey of discovery that is fun and engaging. An integral part of this dissemination effort involves a “train the trainers” component, encourages participants to share their experiences among others in their workplaces, etc., thereby creating a multiplier effect.
Although artificial intelligence (AI) promises to deliver ever more user-friendly consumer applications, recent mishaps involving fake information and biased treatment serve as vivid reminders of the pitfalls of AI. AI can harbor latent biases and flaws that can cause harm in diverse and unexpected ways. Before AI becomes interwoven into human society, it is important to understand how and when AI can fail. This article presents a timely survey of AI-induced mishaps that relate to consumer applications. The article also offers suggestions on mitigating strategies to manage the undesirable side effects of using AI for consumer applications. It, therefore, serves a dual purpose of creating awareness of current issues and encouraging other researchers in the consumer technology community to build better AI consumer applications.
The C language is used to develop software that implements fundamental mechanisms used by higher level software to protect data. Yet C continues to be difficult for students to understand and use securely, and integer errors continue to create vulnerabilities. In fact, \em Integer Overflow or Wraparound is listed at position 11 in the 2020 CWE Top 25 Most Dangerous Software Weaknesses. This paper presents the Expression Evaluation (EE) visualization tool that helps students understand the type conversions that take place implicitly within a C program. This tool depicts step-wise the coercions that take place within the compilation of an expression with mixed integer type operands. This enables students to create unlimited examples to test their understanding. We present the results of our evaluation of EE in both a lower-level class and an upper-level class. We also present the results of an expanded evaluation of a complementary integer security education tool Integer Representation (IR) in these same classes. This represents evaluation of IR across a wider student audience; prior evaluations of the IR tool were within classes focused on low-level programming and security. Our evaluation results showed that students in an upper-level course improved their understanding in both IR and EE more significantly than students in a lower-level course. As shown by the data collected from both classes, our tools were easy to use and very effective.
Integer errors continue to create vulnerabilities. In fact, Integer Overflow or Wraparound is listed at position 11 in the 2020 CWE Top 25 Most Dangerous Software Weaknesses. This poster describes the Expression Evaluation (EE) visualization tool that helps students understand the type conversions that take place implicitly within a C program. This tool depicts step-wise the coercions that take place within the evaluation of a user specified expression with mixed integer type operands. The system enables students to create unlimited examples to test their understanding. The tool was evaluated in the classroom and shown to be easy to use and effective.
Data integrity is critical to the secure operation of a computer system. Applications need to know that the data that they access is trustworthy. Many current production-level integrity models are tightly coupled to a specific domain, (e.g., databases), or only apply after the fact (e.g., backups). In this paper we propose a recommendation-based trust model, called Admonita, for data integrity that is applicable to any structured data in a system and provides a measure of trust to applications on-the-fly. The proposed model is based on the Biba integrity model and utilizes the concept of an Integrity Verification Procedure (IVP) proposed by Clark-Wilson. Admonita incorporates subjective logic to maintain the trustworthiness of data and applications in a system. To prevent critical applications from losing trust, Admonita also incorporates the principle of weak tranquility to ensure that highly trusted applications can maintain their trust levels. We develop a simple algebra around these elements and describe how it can be used to calculate the trustworthiness of system entities. By applying subjective logic, we build a powerful, artificial and reasoning trust model for implementing data integrity.
Seemingly small coding errors can create significant vulnerabilities in C programs. This often occurs due to memory being overwritten in unexpected ways. If a student understands where program variables appear in the process address space, then she can understand the effect of writing beyond the memory allocated to a variable. With this understanding, she can tie her code to its effect within an executing process and is more likely to appreciate the significance of these seemingly harmless errors and to avoid them. We have developed a program analysis and visualization tool to help students understand the impact of common memory errors with the goal to help students avoid introducing these errors into their code. The visualization is through the Program Address Space (PAS) window within a larger system for analysis and visualization of security issues in C programs. The larger system is called SecureCvisual. In this paper, we describe our experience with teaching students fundamental concepts about process address spaces and the impact of buffer overflows using the PAS window. We also present the results from an evaluation of the tool. Our results indicate that students found the tool useful and that it enhanced the course in which it was used.
In many undergraduate programs, students primarily write code in Java or other scripting languages. Yet C and C++ are widely used when performance is important. Poor understanding of a C program's layout in memory and its execution leads to the introduction of security vulnerabilities. We present the SecureCvisual system, which is designed to help students learn to develop more secure and robust C programs. The system takes input from dynamic analysis using Pintool. The analysis produces a sequence of events that are processed by the visualizations. A student or instructor can step forward or backwards through an execution. Source code is displayed, and events are linked to a line of source code. A program address space visualization depicts the values of registers and the program address space. Buffer overflows and other memory errors are easily seen. An integer representation window identifies integer coercions that take place within an equation. The result of a conversion between integer types is also shown. A sensitive data visualization teaches students how to protect data so that it does not appear unencrypted on secondary storage. The tool is convenient for lecture. Multiple levels of detail and different perspectives on an execution make the tool useful in a variety of courses. This work has been supported by the National Science Foundation under grants DUE-1245310, DGE-1522883 and DGE-1523017.
Integer errors can introduce significant vulnerabilities into C programs. We have developed a program analysis and visualization tool to help students understand integer representation and type conversions with the goal to help students avoid introducing these errors into the code they develop. The visualization is through the Integer Representation (IR) window within a larger system for analysis and visualization of security issues in C programs. The system is called the Visualization and Analysis for C Code Security (VACCS) system. In this paper, we describe our experience with teaching fundamental aspects of integer security in a junior-level systems programming course, the IR window, and an evaluation of the tool. Our results indicate that students found the tool to be useful and that it enhanced the course in which it was used.
The pervasive nature of the Internet of Things has resulted in generating a huge amount of data about the lives of IoT users. This data includes Personally Identifiable Information (PII) that reflects people's behaviors, interests, lifestyles, and everyday routines. Protecting PII from privacy violations is a challenge since IoT data need to be handled by public networks, servers, and clouds, which are untrusted parties for data owners. In this paper, a solution called Policy Enforcement Fog Module (PEFM) is proposed for protecting sensitive IoT data whenever they are accessed throughout their entire lifecycle. PEFM uses the power of policy enforcement in the edge-fog infrastructures for protecting data accessed within users’ local domains. For data that need to be sent to remote domains, PEFM uses Active Data Bundle (ADB); an executable and self-protecting construct that can run on any visited host and enforces privacy policies automatically for data accessed by these hosts. To test the feasibility of PEFM in realistic IoT systems, a framework of using PEFM as a privacy control for Foscam home security system is simulated. The experimental results show that PEFM assures data privacy via data minimization due to selective data disclosures. Better privacy controls with minimal overhead can be achieved if most PEFM processes are executed by the local fog nodes. Migrating parts of PEFM processes to remote fog nodes or the cloud incurs more overhead than using strictly local fog nodes. This overhead is the cost for a higher level of privacy regarding lifecycle data protection.
Memory distance analysis, the number of unique memory references made between two accesses to the same memory location, is an effective method to measure data locality and predict memory behavior. Many existing methods on memory distance measurement and analysis consider sequential programs only. With the trend towards concurrent programming, it is necessary to study the impact of memory distance on the performance of concurrent programs. Unfortunately, accurate measurement of concurrent program memory distance is non-trivial. In fact, due to non-determinism, the reuse distance of memory references may differ with the same input set across multiple runs. Since memory distance measurement is fundamental to analysis, we propose a measuring approach that is based on randomized executions. Our approach provides a probabilistic guarantee of observing all possible interleavings without repeated executions. In order to evaluate our approach, we propose a second symbolic execution based approach that is more rigorous but much less scalable than the first approach. We have compared the two approaches on small programs and evaluated the first one on Parsec benchmark suite and a large industrial-size benchmark MySQL. Our experiments confirm that the randomized execution based approach is effective and practical.
Ching-Kuang Shene合作论文数Department of Computer Science, Michigan Technological University21
Ajay K. Gupta合作论文数Computer Science at Western Michigan University10
Yu Chin Cheng合作论文数Department of Economics
University of Washington2