We describe a new benchmark, https://github.com/wblangdon/rand_malloc for GI via parameter tuning, and demonstrate it using heap performance data collected via our modified Valgrind DHAT from Python/C++ simulation tool gem5 with the GNU gcc compiler’s glibc malloc. As well as GI applications, rand_malloc might help develop, stress test, tune, and benchmark other dynamic memory managers.
HotBugs.jar is a novel benchmark targeting time-critical (a.k.a. hot) fixes. We propose an approach to analyze the taxonomy of the bugs in HotBugs.jar by extending PatchCat into HotCat, integrating hotfix metadata with multi-objective optimization. Using NSGA-II, we evolve bitmask-based feature subsets that balance accuracy, Normalized Mutual Information (NMI), and runtime. On 88 records across 17 categories, HotCat achieved 0.59 accuracy and 0.58 NMI in 129 s, with maximum accuracy of 0.63 in 132 s, demonstrating accuracy improvements without additional resource use, thus supporting sustainability. Future work will expand and augment the dataset, refine optimization objectives, and improve semantic categorization, robustness, and cluster balance.
Fuzz testing finds security issues and improves robustness, however it has only two implicit test oracles: timeout and crash. Information theory gives a third: entropy, which is generic, low cost and widely applicable. Fuzz3 is programming language agnostic, treating software as a black box, searching its input space based on entropy distributions. Applied to 6 open source software, it found 4 bugs, 2 of which has already been fixed since we reported it to the original developers.
We extend recent 256 SSE vector work to 512 AVX giving a four fold speedup. We use MAGPIE (Machine Automated General Performance Improvement via Evolution of software) to speedup a C++ linear genetic programming interpreter. Local search is provided with three alternative hand optimised codes, revision history and the Intel 512 bit AVX512VL documentation as C++ XML. Magpie is applied to the new Single Instruction Multiple Data (SIMD) parallel interpreter for Peter Nordin's linear genetic programming GPengine. Linux mprotect sandboxes whilst performance is given by perf instruction count. In both cases, in a matter of hours local search reliably sped up 114 or 310 lines of manually written parallel SIMD code for the Intel Advanced Vector Extensions (AVX) by 2 percent.
Compression, e.g. gzip, gives algorithmic information theory (Kolmogorov Complexity) based measures of string population diversity. To boost it we use the GI tool Magpie and select programs of average fitness that contribute most to variety, allowing evolution to automatically tailor triangle.c for production speed. We calculate C source code diversity via approximations to the Normalised Compression Distance on Multisets (NCDm) using both Cohen and Vitanyi’s O(n^2) approach and our own, O(n) method, finding the cheaper, O(n), is equally good.
Using the language independent genetic improvement tool MAGPIE (Machine Automated General Performance Improvement via Evolution of software) and logarithmic sampling, we measure the parameter fitness landscape when optimising the GNU glibc heap management of a million line C++ application, gem5. The malloc_info landscape is far smoother than is commonly assumed and savings of 11 % with no loss of runtime speed are readily obtained by both Magpie and CMA-ES.
Mobile applications can be very network-intensive. Mobile phone users are often on limited data plans, while network infrastructure has limited capacity. There's little work on optimizing network usage of mobile applications. The most popular approach has been prefetching and caching assets. However, past work has shown that developers can improve the network usage of Android applications by making changes to Java source code. We built upon this insight and investigated the effectiveness of automated, heuristic application of software patches, a technique known as Genetic Improvement (GI), to improve network usage. Genetic improvement has already shown effective at reducing the execution time and memory usage of Android applications. We thus adapt our existing GIDroid framework with a new mutation operator and develop a new profiler to identify network-intensive methods to target. Unfortunately, our approach is unable to find improvements. We conjecture this is due to the fact source code changes affecting network might be rare in the large patch search space. We thus advocate use of more intelligent search strategies in future work.
We introduce a novel automated testing technique that combines LLM and search-based fuzzing. We use ChatGPT to parameterise C programs. We compile the resultant code snippets, and feed compilable ones to SearchGEM5, our extension to AFL++ fuzzer with customised new mutation operators. We run thus created 4005 binaries through our system under test, gem5, increasing its existing test coverage by more than 1000 lines. We discover 244 instances where gem5 simulation of the binary differs from the binary’s expected behaviour.
Following the formal presentations, which included keynotes by Prof. Myra B. Cohen of Iowa State University and Dr. Sebastian Baltes of SAP as well as six papers (which are recorded in the pro- ceedings) there was a wide ranging discussion at the twelfth inter- national Genetic Improvement workshop, GI-2023 @ ICSE held on Saturday 20th May 2023 in Melbourne and online via Zoom. Topics included GI to improve testing, and remove unpleasant surprises in cloud computing costs, incorporating novelty search, large language models (LLM ANN) and GI benchmarks.
Using MAGPIE (Machine Automated General Performance Improvement via Evolution of software) we show genetic improvement GI can reduce the cache load of existing computer programs. Cache miss reduction is tested on two industrial open source C programs (Google's Open Location Code OLC and Uber's Hexagonal Hierarchical Spatial Index H3) and two C++ 2D photograph image processing tasks, counting pixels and OpenCV's SEEDS segmentation algorithm. Magpie's patches functionally generalise. In one case they reduce data misses on the highest performance L1 cache by 47%.
Large arithmetic expressions are dissipative: they lose information and are robust to perturbations. Lack of conservation gives resilience to fluctuations. The limited precision of floating point and the mixture of linear and nonlinear operations make such functions anti-fragile and give a largely stable locally flat plateau a rich fitness landscape. This slows long-term evolution of complex programs, suggesting a need for depthaware crossover and mutation operators in tree-based genetic programming. It also suggests that deeply nested computer program source code is error tolerant because disruptions tend to fail to propagate, and therefore the optimal placement of test oracles is as close to software defects as practical.
research-article Share on Tributes to Julian F. Miller (1955 - 2022) Authors: Wolfgang Banzhaf Michigan State University Michigan State UniversitySearch about this author , James A. Foster University of Idaho, Moscow, Idaho University of Idaho, Moscow, IdahoSearch about this author , Simon Harding Machine Intelligence Ltd Machine Intelligence LtdSearch about this author , Roman Kalkreuth TU Dortmund University, Dortmund, Germany TU Dortmund University, Dortmund, GermanySearch about this author , Gul Muhammad Khan University of Engineering and Technology, Peshawar, Pakistan University of Engineering and Technology, Peshawar, PakistanSearch about this author , W. B. Langdon University College London, London, UK University College London, London, UKSearch about this author , Riccardo Poli University of Essex, Colchester, UK University of Essex, Colchester, UKSearch about this author , Stephen Smith University of York, York, UK University of York, York, UKSearch about this author , Susan Stepney University of York, York, UK University of York, York, UKSearch about this author , Martin Trefzer University of York, York, UK University of York, York, UKSearch about this author , Dennis G. Wilson University of Toulouse; Toulouse Mind & Brain Institute, Toulouse, France University of Toulouse; Toulouse Mind & Brain Institute, Toulouse, FranceSearch about this author Authors Info & Claims ACM SIGEVOlutionVolume 15Issue 1April 2022 Article No.: 1pp 1–7https://doi.org/10.1145/3532942.3532943Published:20 April 2022Publication History 0citation26DownloadsMetricsTotal Citations0Total Downloads26Last 12 Months26Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Limited precision floating point computer implementations of large polynomial arithmetic expressions are nonlinear and dissipative. They are not reversible (irreversible, lack conservation), lose information, and so are robust to perturbations (anti-fragile) and resilient to fluctuations. This gives a largely stable locally flat evolutionary neutral fitness search landscape. Thus even with a large number of test cases, both large and small changes deep within software typically have no effect and are invisible externally. Shallow mutations are easier to detect but their RMS error need not be simple.
If a software execution is disrupted, witnessing the execution at a later point may see evidence of the disruption or not. If not, we say the disruption failed to propagate. One name for this phenomenon is software robustness but it appears in different contexts in software engineering with different names. Contexts include testing, security, reliability, and automated code improvement or repair. Names include coincidental correctness, correctness attraction, transient error reliability. As witnessed, it is a dynamic phenomenon but any explanation with predictive power must necessarily take a static view. As a dynamic/static phenomenon it is convenient to take a statistical view of it which we do by way of information theory. We theorise that for failed disruption propagation to occur, a necessary condition is that the code region where the disruption occurs is composed with or succeeded by a subsequent code region that suffers entropy loss over all executions. The higher is the entropy loss, the higher the likelihood that disruption in the first region fails to propagate to the downstream observation point. We survey different research silos that address this phenomenon and explain how the theory might be exploited in software engineering.
With side effect free terminals and functions it is possible to evaluate the fitness of genetic programming trees from their parents without creating them. This allows selection before forming the next generation. Thus avoiding unfit runt Genetic Algorithm individuals, which will themselves have no children. In highly diverse GA populations with strong selection, more than 50% of children need not be created. Even with two parent crossover, in converged populations, $$e^{-2}$$ = 13.5% can be saved. Eliminating bachelors and spinsters and extracting the smaller genetic material of each mating before crossover, reduces storage in an N multi-threaded implementation for a population M to $$\le $$ 0.63M+N, compared to the usual M+2N. Memory efficient crossover achieves 692 billion GP operations per second, 692 giga GPops, at runtime on a 16 core 3.8 GHz desktop.
We study both genotypic and phenotypic convergence in GP floating point continuous domain symbolic regression over thousands of generations. Subtree fitness variation across the population is measured and shown in many cases to fall. In an expanding region about the root node, both genetic opcodes and function evaluation values are identical or nearly identical. Bottom up (leaf to root) analysis shows both syntactic and semantic (including entropy) similarity expand from the outermost node. Despite large regions of zero variation, fitness continues to evolve and near zero crossover disruption suggests improved GP systems within existing memory use.
Information-theoretic analysis of large, evolved programs produced by running genetic programming for up to a million generations has shown even functions as smooth and well behaved as floating-point addition and multiplication lose entropy and consequently are robust and fail to propagate disruption to their outputs. This means that, while dependent upon fitness tests, many genetic changes deep within trees are silent. For evolution to proceed at a reasonable rate it must be possible to measure the impact of most code changes, yet in large trees, most crossover sites are distant from the root node. We suggest that to evolve very large, very complex programs, it will be necessary to adopt an open architecture where most mutation sites are within 10--100 levels of the organism's environment.
L. Spector合作论文数Cognitive Science
Hampshire College
Evolutionary Computation
Genetic Programming and Evolvable Machines
International Society for Genetic and Evolutionary Computation
School of Cognitive Science at Hampshire College2