
Machine learning (ML)-based Artificial Intelligence (AI) systems rely on training data to function, but the internal workings of how ML models learn from this data are often black-box. Influence analysis provides valuable insights into the model's behavior by evaluating the effect of individual training instances on the model's predictions. However, calculating the influence of each training data can be computationally expensive. In this paper, we propose a proxy model-based approach to influence analysis called Proxima. The main idea of our approach is to use a subset of training instances to create a proxy model that is simpler than the original model, and then use the proxy model and the subset of training instances to perform influence analysis. We evaluate Proxima on ML models trained using seven real-world datasets. We compare Proxima to two state-of-the-art influence analysis tools, i.e., FastIF and Scaling-Up. Our experimental results suggest that the proposed approach can successfully perform and speed up the influence analysis process, and, in most cases, perform better than FastIF and Scaling-Up.
Finding critical scenarios is essential in testing autonomous driving and automated driving functions. Such scenarios describe a sequence of interactions between the autonomous vehicle or the vehicle equipped with automated driving functions and the environment, i.e., other cars, pedestrians, and the current road conditions, which challenge the system we want to test. In this paper, we present a search-based testing solution utilizing genetic algorithms for test generation coupled with a traffic simulator. As a fitness function, we rely on the amount of emergency braking required to prevent crashes. In addition, we compare two types of hyperparameter tuning. One type uses combinations of hyperparameters obtained from previous papers. The other is based on a design of experiment method. We show that the genetic algorithm using the design of experiments method for hyperparameter tuning outperforms the other implementation in terms of criticality (i.e., the time of emergency braking) and diversity. Furthermore, we show that both genetic algorithm implementations are superior to pure random testing in the application context of autonomous and automated driving.
AI-based systems possess distinctive characteristics and introduce challenges in quality evaluation at the same time. Consequently, ensuring and validating AI software quality is of critical importance. In this paper, we present an effective AI software functional testing model to address this challenge. Specifically, we first present a comprehensive literature review of previous work, covering key facets of AI software testing processes. We then introduce a 3D classification model to systematically evaluate the image-based text extraction AI function, as well as test coverage criteria and complexity. To evaluate the performance of our proposed AI software quality test, we propose four evaluation metrics to cover different aspects. Finally, based on the proposed framework and defined metrics, a mobile Optical Character Recognition (OCR) case study is presented to demonstrate the framework's effectiveness and capability in assessing AI function quality.
Artificial Intelligence (AI) has rapidly evolved, offering transformative potential across various sectors, including healthcare, finance, education, and transportation. However, its deployment raises significant ethical, societal, and policy concerns that have necessitated careful consideration and action since its inception. This panel will discuss the multifaceted implications of AI, focusing on ethical dilemmas such as privacy, bias, transparency, and accountability, as well as the broader societal impacts on employment, social equity, and human interaction. Additionally, it examines the current state of AI policy and governance, highlighting the need for comprehensive frameworks that balance innovation with responsible AI or Machine Learning (ML) use. This panel will discuss these critical issues and aims to contribute to the current and future development of AI systems that are not only technologically advanced but also ethically sound and socially beneficial.
Real-time monitoring plays a vital role in software testing by facilitating the identification and diagnosis of issues during canary deployments or limited-scale testing phases. This approach serves as a method for testing and issue identification in the deployment phase. Furthermore, efficiently detecting anomalies in software instrumentation data is a crucial aspect of real-time monitoring. However, due to the volatility of software telemetry data, achieving accurate anomaly detection is quite challenging. In this paper, we present an efficient transformer-based framework designed for real-time anomaly detection. This approach called TEAD deeply explores the temporal correlations in time series data to uncover anomalous patterns, ensuring stability in predictions amidst fluctuating data. Empirical validation conducted at a prominent Internet company highlights the effectiveness of TEAD, emphasizing its ability to improve both efficiency and accuracy in real-time monitoring. Despite its significance, the method is constrained by its business-centric nature, warranting further cross-industry dissemination research. Overall, this study provides a novel solution for the real-time monitoring phase of software testing, validating the applicability and effectiveness of transformer-like models in this field.
The functionalities of AI-powered mobile apps or systems heavily depend on the given training dataset. The challenge, in this case, is that a learning system will change its behavior due to a slight change in the dataset. Current alternative approaches for evaluating these apps either focus on individual performance measurements such as accuracy, etc. Inspired by principles of the decision tree test method in software engineering, this paper provides a tutorial discussion on intelligent AI test modeling chat systems including basic concepts, validation process, testing scopes, approaches, and needs. The report is about an intelligent AI test modeling chatbot system that is built and implemented based on an innovative 3D AI test model for AI-powered functions in intelligent mobile apps to support model-based AI function testing, test data generation, auto test scripting and execution, and adequate test coverage analysis.
In order to make sure that recent code changes haven't negatively impacted the software's current functionalities, regression testing is a crucial maintenance step in the software development lifecycle. Time and resource constraints make it impractical to run every test case during regression testing, especially as software systems get larger and more complex. This problem is addressed by test case prioritization (TCP) techniques, which arrange test cases in a way that maximizes the probability of finding errors early on. Several of advanced machine learning (ML) techniques that have recently shown great promise in improving the efficacy and efficiency of the prioritization process when integrated with TCP. The methods, advantages, and drawbacks of these cutting-edge ML-based approaches for TCP are examined in this paper. I present a comparative analysis of their performance metrics and provide an implementation example for TCP using ensemble methods in machine learning. Practical considerations such as data requirements, model selection, and integration with current continuous integration/continuous deployment (CI/CD) pipelines are covered when implementing these advanced ML-based TCP methods in real-world scenarios. The results show that advanced ML based TCP is a useful tactic for contemporary regression testing procedures since it enhances fault detection rates and maximizes resource utilization.
Question-answering based text summarization can produce personalized and specific summaries; however, the primary challenge is the generation and selection of questions that users expect the summary to answer. Large language models (LLMs) provide an automatic method for generating these questions from the original text. By prompting the LLM to answer these selected questions based on the original text, high-quality summaries can be produced. In this paper, we experiment with an approach for question generation, selection, and text summarization using the LLM tool GPT4o. We also conduct a comparative study of existing summarization approaches and evaluation metrics to understand how to produce personalized and useful summaries. Based on the experiment results, we explain why question-answering based text summarization achieves better performance.
It has been shown that multimodal text-vision large language models (LLMs) are vulnerable to visual adversarial examples. An adversarial example is a delicately perturbed input intended to cause deep neural networks to behave incorrectly. Currently, adversarial attacks can be categorized into two types based on the knowledge required to perform an attack: black box and white box. Most existing black-box attacks only try to generate adversarial examples within a given l(2) or l(infinity) perturbation budget. We propose a novel black-box attack to address this research gap to generate l(1), l(2) adversarial examples with minimal perturbation. To this end, our attack solves a constrained optimization problem by converting it to an unconstrained one and then finds its solution using the genetic algorithm. Extensive experiments on three test benches have been conducted to evaluate its performance. To facilitate reproducing the results presented in this work, we make the code for the experiments publicly available at https://github.com/sjysjy1/MPBA.
Welcome to the IEEE International Congress on Intelligent and Service-Oriented Systems Engineering (CISOSE) 2024, taking place July 15-18th in Shanghai, China. It is our pleasure to welcome you to this premier event that provides unique global platform to bring together all the experts, researchers, industry professionals, and decision makers to discuss and exchange the issues, challenges, solutions and applications in intelligent and service-oriented systems. In the past 18 years, the IEEE CISOSE congress has successfully established, organized and delivered six IEEE international conferences annually in different regions and countries in the world, including Shanghai/China, San Francisco/USA, Oxford/UK, Bamberg/Germany, Athens/Greece. Up to today, IEEE CISOSE consists of six IEEE international conferences, including IEEE SOSE, IEE MobileCloud (known as IMC), IEEE BDS, IEEE DAPPS, IEEE AITEST, and IEEE JCC. These conferences bring the talents and excellent researchers working on different cutting-edge research subjects and applications.
Cyber-physical systems are integral to the infrastructure of global communication and transportation networks, which makes it crucial to detect faults, prevent cyber attacks, and ensure operational safety. Although machine learning techniques, including large language models (LLMs), have been explored for fault detection, the efficacy of open-source LLMs remains underexplored. In this work, we assess the capabilities of eight open-source LLMs in identifying faults in cyber-physical systems using a simulation dataset from monitoring an electrified vehicle's battery management system. By applying pretrained LLMs without fine-tuning and incorporating retrieval augmented generation (RAG) techniques alongside textual encoding methods, our study aims to explore the potential of open LLMs in fault detection. Our results show that open LLMs can effectively identify faults, with Mistral outperforming alternative models such as Mixtral, codellama, and Gemma in precision, recall, and F1-score metrics. Furthermore, our results highlight the importance of textual encoding strategies in enhancing the fault detection capabilities of LLMs, which possess a degree of explanatory power with respect to the detected anomalies. This work demonstrates the feasibility of using open LLMs for fault detection in cyber-physical systems and opens avenues for future research to enhance fault detection and fault localization.
Speaker verification systems, crucial for securing online banking and Internet of Things (IoT) devices, employ human voice biometrics to identify users. However, as these state-of-the-art systems rely on machine learning models, they are susceptible to adversarial attacks. Presently, there exists no defense system entirely immune to every form of adversarial attack. This work aims to develop a defense mechanism capable of effectively countering a broad spectrum of known adversarial threats. Our approach introduces an ensemble defense system named MEH-FEST-NA, which integrates the MEH-FEST and NA strategies. Experimental results reveal that our system achieves a maximum false negative rate 13.5% across 110 adversarial attacks. Consequently, while a speaker verification system with a singular defense strategy may reach a maximum refined F1 score of 0.16, our ensemble defense system significantly improves performance, achieving a score of 0.91. This underscores its enhanced effectiveness in protecting against adversarial attacks.
The past three years have seen a rapid growth of generative AI (GenAI) technology. Several multi-modal large language models have been developed such as ChatGPT, Gemini and Sora, etc. They have demonstrated an impressive capability of generating contents and performing a wide range of natural language processing tasks including reasoning, programming, even producing images and video clips from text instructions. It is widely perceived that the technology is rapidly advancing towards an artificial general intelligence and will fundamentally change the world by revolutionise the ways we work and live. The panel invited a few active researchers from the academics and practitioners from the industry. Each panellist will give a short position statement to present their visions on the following issues related to generative AI.
Generative Artificial Intelligence is becoming an integral and enduring part of our lives, growing more powerful with each passing day. This paper explores Large Language Models and their application in text generation, specifically examining their potential to assist software quality assurance engineers in their daily tasks. Our focus is on the generation of unit tests as a critical component of software development. The research question is simple: Can Generative AI generate comprehensive unit tests? We started with Python and a very simple use case, and if Gen AI is successful, we will continue with complex tasks. Current literature focuses on success, but we are interested in failures as well. How many test cases are missing?
The panel aims to foster a critical and insightful discussion on the sustainable development of AI, examining both the ethical and environmental implications of its widespread use and the strategies to ensure its responsible growth and integration into society. The following key topics will be addressed in this panel include: •Sustainable AI Infrastructure: Given the exponential expansion of data storage and the computational demands of AI, it's important to monitor and optimize energy consumption at various layers of the full AI stack. •Environmental Impact, Standard and Assessment: We'll discuss the relationship between AI and environmental impact including how AI impacts environment, how to measure and assess the impact, and how AI can help address environmental challenges. •Ethical Consideration of AI: The confluence of privacy, bias mitigation, and equitable decisionmaking in AI necessitates a robust ethical framework and technologies. The panel will discuss and share real-world practices.
Spectrum-based fault localization approaches utilize the statistical information about the execution of test cases to rank statements denoting the most specious ones leading to the failure of test cases. We propose a new approach for software fault localization called SpecNLP, which combines natural language processing (NLP) techniques with spectrum-based fault localization (SBFL). SpecNLP uses a pre-trained NLP model called CodeBERT to extract semantic features. These features together with spectrum information from test executions are fed into a multi-layer perceptron (MLP) to predict the start and end locations of bugs. The key innovation is the integration of SBFL execution test cases with CodeBERT code embeddings, enabling more accurate bug localization. SpecNLP outperforms previous ML and NLP methods on the Codeflaws benchmark. On the key Top-N metric, SpecNLP achieves 31.9% accuracy on Top-1 predictions versus 5.4% for SBFL techniques. The results demonstrate that SpecNLP outperforms previous methods on a benchmark and achieves higher accuracy in predicting fault locations.
Machine learning (ML)-based systems are becoming increasingly ubiquitous even in safety critical environments. The strength of ML systems, to solve complex problems with a stochastic model, leads to challenges in the testing domain. This motivates us to introduce a rigorous testing method for ML-models and their application environment akin to classical software testing, which is independent of the training process and considers the probabilistic nature of ML. The approach is based on the concept of the Probabilistically Extended ONtology (PEON). In brief, PEON is a an ontology modeling the designated Operational Design Domain (ODD), which is extended by assigning probability distributions to classes and their individual attributes, as well as probabilistic dependencies between these attributes. The relevant statistical key figures like accuracy depend not only on the ML-based model but also strongly on the statistics of the test data set, which we refer to by quality assurance (QA) data set, to emphasize its independence from the test data set in the training process. This implies that we have to consider the statistical properties of the QA data in order to evaluate an ML-based system. In this paper we present first experimental results comparing established test selection methods e.g. N-wise, with a new approach the PEON. Our findings strongly suggest, that the underlying statistical properties of the QA data significantly influence the test results of ML-based systems. In this respect, careful attention must be paid to the statistical independence and balance of the QA data. The PEON provides a good basis for the composition of QA data sets, which are not only independent of the development process but also statistically representative and balanced with respect to the modeled ODD.
Anomaly Detection (AD) is widely used in security applications such as intrusion detection, but its vulnerability to nondeterminism attacks has not been noticed, and its robustness against such attacks has not been studied. Nondeterminism, i.e., output variation on the same input dataset, is a common trait of AD implementations. We show that nondeterminism can be exploited by an attacker that tries to have a malicious input point (outlier) classified as benign input (inlier). In our threat model, the attacker has extremely limited capabilities - they can only retry the attack; they cannot influence the model, manipulate the AD/IDS implementation, or insert noise. We focus on three concrete, orthogonal attack scenarios: (1) a restart attack that exploits a simple re-run, (2) a resource attack that exploits the use of less computationally-expensive parameter settings, and (3) an inconsistency attack that exploits the differences between toolkits implementing the same algorithm. We quantify attack vulnerability in popular implementations of four AD algorithms - IF, RobCov, LOF, and OCSVM - and offer mitigation strategies. We show that in each scenario, despite attackers' limited capabilities, attacks have a high likelihood of success.