Reverse engineering binary code is notoriously difficult and, especially, understanding a binary’s dynamic data structures. Existing data structure analyzers are limited wrt. program comprehension: they do not detect complex structures such as skip lists, or lists running through nodes of different types such as in the Linux kernel’s cyclic doubly-linked list. They also do not reveal complex parent-child relationships between structures. The tool DSI remedies these shortcomings but requires source code, where type information on heap nodes is available. We present DSIbin, a combination of DSI and the type excavator Howard for the inspection of C/C++ binaries. While a naive combination already improves upon related work, its precision is limited because Howard’s inferred types are often too coarse. To address this we auto-generate candidates of refined types based on speculative nested-struct detection and type merging; the plausibility of these hypotheses is then validated by DSI. We demonstrate via benchmarking that DSIbin detects data structures with high precision.
As the complexity of malware grows, so does the necessity of employing program structuring mechanisms during development. While control flow structuring is often obfuscated, the dynamic data structures employed by the program are typically untouched. We report on work in progress that exploits this weakness to identify dynamic data structures present in malware samples for the purposes of aiding reverse engineering and constructing malware signatures, which may be employed for malware classification. Using a prototype implementation, which combines the type recovery tool Howard and the identification tool Data Structure Investigator (DSI), we analyze data structures in Carberp and AgoBot malware. Identifying their data structures illustrates a challenging problem. To tackle this, we propose a new type recovery for binaries based on machine learning, which uses Howard's types to guide the search and DSI's memory abstraction for hypothesis evaluation.
Comprehension of C programs containing pointer-based dynamic data structures can be a challenging task. To tackle this challenge we present Data Structure Investigator (DSI), a new dynamic analysis for automated data structure identification that targets C source code. Our technique first applies a novel abstraction on the evolving memory structures observed at runtime to discover data structure building blocks. By analyzing the interconnections between building blocks we are then able to identify, e.g., binary trees, doubly-linked lists, skip lists, and relationships between these such as nesting. Since the true shape of a data structure may be temporarily obscured by manipulation operations, we ensure robustness by first discovering and then reinforcing evidence for data structure observations. We show the utility of our DSI prototype implementation by applying it to both synthetic and real world examples. DSI outputs summarizations of the identified data structures, which will benefit software developers when maintaining (legacy) code and inform other applications such as memory visualization and program verification.
C programs that manipulate list-based dynamic data structures remain a challenging target for static verification. In this paper we employ the dynamic analysis of dsOli to locate and identify data structure operations in a program, and then use this information to automatically annotate that program with assertions in separation logic. These annotations comprise candidate pre/post-conditions and loop invariants suitable to statically verify memory safety with the verification tool VeriFast. By using both textbook and real-world examples on our prototype implementation, we show that the generated assertions are often discharged automatically. Even when this is not the case, candidate invariants are of great help to the verification engineer, significantly reducing the manual verification effort.
Comprehension of C programs can be a difficult task, especially when they contain pointer-based dynamic data structures. This paper describes our tool dsOli which aims to simplify this problem by automatically locating and identifying data structure operations in C programs, such as inserting into a singly linked list. The approach is based on a dynamic analysis that seeks to identify functional units in a program by observing repetitive temporal patterns caused by multiple invocations of code fragments. The behaviour of these functional units is then classified by matching the associated heap states against templates describing common data structure operations. The analysis results are available to the user via XML output, and can also be viewed using an intuitive GUI which overlays the learnt information on the program source code.
We investigate whether dynamic data structures in pointer programs can be identified by analysing program executions only. This paper describes a first step towards solving this problem by applying machine learning and pattern recognition techniques to analyse executions of C programs. By searching for repeating temporal patterns in memory caused by multiple invocations of data-structure operations, we are able to first locate and then identify these operations. Applying a prototypic tool implementing our approach to pointer programs that employ, e.g., lists, queues and stacks, we show that the identified operations can accurately determine the data structures used.
A current limitation of the Esterel language for reactive-systems design is its lack of support for accessing databases. This talk presents the results of a summer student project which investigated a way of integrating databases and Esterel by providing an API for database use inside Esterel. A case study, involving a warehouse storage system built using Lego Mindstorms robotics kits, demonstrates the utility of the API. This system employs an Esterel-programmed robot whose task it is to collect various items from a customer's order and assemble them in one place. To do so, the robot accesses customer-order data and floor-plan data stored in a database.
CSIRO Sustainable Ecosystems is constructing a spatially explicit modelling system capable of exploring alternative land and water policy alternatives against plausible price, cost, and climate scenarios for the next 20 years. INSIGHT will be used to identify the likely impacts of land and water policy options on regional economies and structural adjustment. Flowcharts have been constructed for most of the major crop and pasture and associated economic models for commodities produced in the Lachlan River Catchment of New South Wales. This enabled the most important components and interrelationships within these models to be readily identified. The next step has been to construct models at the regional scale that contain the essential elements of the more-detailed point models. The paper describes the progress to date in describing these models, and how they have been integrated into a coordinated agricultural crop production evaluation system.
Increasing atmospheric concentrations of ‘greenhouse gases’ are expected to result in global climatic changes over the next decades. Means of evaluating and reducing greenhouse gas emissions are being sought. In this study an existing simulation model of a tropical savanna woodland grazing system was adapted to account for greenhouse gas emissions. This approach may be able to be used in identifying ways to assess and limit emissions from other rangeland, agricultural and natural ecosystems.
Amino acid activation by anhydride formation in model tetrahedral silicate and aluminate sites in clays and neutral phosphates have been studied by semi-empirical molecular orbital calculations. the results have been compared to previousab initio studies on the reactant species and were found to be in good agreement. The geometries of all species were totally optimized and heats of formation obtained. Relative heats of formation of the anhydrides indicate the extent of anhydride formation to be Al > Si > P which is the same order as the stability of hydrolysis. The relative efficacy of the anhydrides in promoting peptide bond formation has been evaluated using both thermodynamic and chemical reactivity criteria. Heats of reaction for model reactions were calculated from calculated enthalpies of formation of the products and reactants. The electrophilicity of the carbonyl carbon and the nucleophilicity of the oxygen were specifically used as indicators of chemical reactivity towards dipeptide formation by the activated amino acids. Our results indicate that if the reaction mechanism is dominated by the nucleophilic character of the oxygen, tetrahedral Al sites should be more active than Si, and if the electrophilic character dominates, the order would be reversed.
We have undertaken a complete kinetic analysis of the template-directed oligoguanylate synthesis originated in Orgel's laboratory (Inoue and Orgel, 1982). The reaction of guanosine 5′-phospho-2-methylimidazolide, 2-MelmpG, with ribooligoguanylates all 3′–5′ linked, designatedn3 withn=7−12, was studied in the presence/absence of the complementary template polycytidylic acid, poly(C). Conditions were chosen where poly(C) and 2-MelmpG are in large excess over the oligoguanylate. In the absence of the template at 37 °C the reaction leads to three isomeric oligomers that are elongated by one monomer unit. They are the 3′–5′ linked, (n+1)3, the 2′–5′ linked, (n+1)2, and the pyrophosphate product, (n+1) p , formed in an approximate ratio 1:2:5. In the presence of the template the reaction is 20-fold faster and yields productsn+1,n+2,n+3 etc. as long as 2-MelmpG is available. Most importantly the formation of the natural, 3′–5′ linked isomer, is enhanced selectively by 140-fold at 37 °C. Qualitative observations allow the conclusion that this enhancement is temperature dependent and increases with decreasing temperature. For example, at 1 °C only the 3′–5′ linked isomers were detected. Initial rates for the disappearance of then3 oligoguanylate were determined at 1, 23, and 37 °C. It was found that the pseudo-first order rate constant for oligoguanylate elongation was linearly proportional to the 2-MelmpG concentration. This implies that the reaction complex poly(C)·n3·2-MelmpG does not accumulate under the reaction conditions, a conclusion which is also supported by infrared data (Miles and Frazier, 1982). The implication of the above results with respect to chemical evolution is that lower temperatures, i.e., close to freezing, enhance the regioselectivity of these template-directed reactions and that one way to improve replication models may be sought in finding conditions that favor stable reaction complexes.
Replacement of a carboxyl function by fluorine, fluorodecarboxylation, is a new process that can be accomplished by the reaction of alkanoic acids with xenon difluoride. Primary, tertiary, and benzylic acids perform best in the reaction, which is conducted at room temperature in methylene chloride or chloroform solution. A reaction mechanism is proposed in which the acid is initially converted to a fluoroxenon ester, RCO2XeF. The esters of the primary and secondary acids react by nucleophilic displacement by fluoride, as evidenced by incorporation of 18F− and no reactions common to free radicals or carbocations. The esters of the tertiary and benzylic acids react by converting to free radicals that can be further oxidized to carbocations. Thus incorporation of 18F− and racemization are observed with α-methoxy-α-trifluoromethylphenylacetic acid. Hydroxyl and amino functions inhibit the reaction. Aromatic and vinylic acids do not react.
AbstractThe key intermediate (VI) in the synthesis of the title compound (VII) is prepared from the monoacetal (I) by treatment with (II) and double Horner‐Emmons reaction to give (IV), which is transformed to (VI) via the mixture of cis/trans isomers (V) [in the conversion (V)‐(VI), deprotection has to precede the introduction of the double bonds to avoid isomerization].
In order to understand how self-reproducing molecules could have originated on the primitive Earth or extraterrestrial bodies, it would be useful to find laboratory models of simple molecules which are able to carry out processes of catalysis and templating. Furthermore, it may be anticipated that systems in which several components are acting cooperatively to catalyze each other's synthesis will have different behavior with respect to natural selection than those of purely replicating systems. As the major focus of this work, laboratory models are devised to study the influence of short peptide catalysts on template reactions which produce oligonucleotides or additional peptides. Such catalysts could have been the earliest protoenzymes of selective advantage produced by replicating oligonucleotides. Since this is a complex problem, simpler systems are also studied which embody only one aspect at a time, such as peptide formation with and without a template, peptide catalysis of nontemplated peptide synthesis, and model reactions for replication of the type pioneered by Orgel.
Gerald Lüttgen合作论文数University of York;Computer Science 7