The uses of web search engines are very frequent and common worldwide over the internet by end users for different purposes. A web search engine takes the query request from the end user and executes that query on relational database used to store the information on behalf of that web search engine. Based on input queries the dynamic response is generated by search engine, in the form of HTML based pages. Such pages are supported with the web databases. Every web page generated contains many results to display for particular query, called as Search Result Records (SRRs). Sometimes it becomes troublesome to extract relevant data from diverse sources. The SRRs generated may contain data units that are relevant to one common semantic. These SRRs are further required to be assigned with proper labels. The manual methods for record extraction and labeling have a worse scalability. Thus automatic annotation based method is needed to improve the accuracy as well as scalability of web search engines. This paper presents an automatic annotation technique for web search results. The proposed approach first aligns the data units on a result page into different groups such that the data in the same group have the same semantic. Then, each group is annotated from different aspects and aggregates the different annotations to predict a final annotation label for it. The annotation wrapper generated for the search site is automatically constructed and can be used to annotate new result pages from the same web database. Experiments indicate that
Classification, which involves finding rules that partition a given data set into disjoint groups, is one class of data mining problems. Approaches proposed so far for mining classification rules for large databases are mainly decision tree based symbolic learning methods. The connectionist approach based on neural networks has been thought not well suited for data mining. One of the major reasons cited is that knowledge generated by neural networks is not explicitly represented in the form of rules suitable for verification or interpretation by humans. This paper examines this issue. With our newly developed algorithms, rules which are similar to, or more concise than those generated by the symbolic methods can be extracted from the neural networks. The data mining process using neural networks with the emphasis on rule extraction is described. Experimental results and comparison with previously published works are presented.
With the gradual deepening of Changqing Oilfield horizontal wells, horizontal wells see water phenomenon is gradually increasing, more than 80% of the current water level of high water wells horizontal wells accounted for 25.5% of the total number of open wells.High water level of wells governance means there are horizontal wells to find the mechanical plugging wells, corresponding water depth profile control, bi-directional oil wells plugging technology, by contrast, the current techniques of horizontal wells in Changqing Oilfield high water governance best suited for the corresponding injection wells be deep profile.By reason of the horizontal well see the water, see water type analysis, and ultimately optimize the construction program, citing CDST-01 blocking agent system and application of multi-level multi-round profile control technology to achieve effective governance of the high water level of wells.Field test seven wells, increasing the average single-well oil 4.5t, average moisture content decreased 68%, cumulative oil 3825t, cumulative precipitation 4585m3, a significant treatment effect, provide a reference for the next step Changqing Oilfield high water level governance and reference wells .
Abstract The Changqing field is one of the largest fields in China, yet it is characterized by low permeability, low initial pressure, and low reservoir quality. During the past few years, technological advancements within the oil industry, horizontal drilling, and hydraulic fracture stimulation have helped create a development boom in this field, with total production of barrel of oil equivalence (BOE) reaching 50 million tons in 2013. Located in the Ordos basin of China, the primary pay formation of the Changqing field is Triassic, which features a complicated sedimentary system. 72.8% of reservoirs in the Changqing field have permeability lower than 1.0 md. Although it was apparent that all of these wells would require fracture treatments to help achieve optimum production, it was difficult to identify a routine treatment design or fracturing strategy to extend to all of the reservoirs because the Changqing field has 19 oilfields and three gas fields stretching over five provinces. The sediment source of the pay formation varies in terms of direction and lithology, and all of the sedimentary systems vary from fan delta to braided stream delta systems. To maximize well production, it is necessary to analyze the differences among these fields and identify the most favorable fracture treatment design for each field. Microseismic mapping has proven to be a valuable technology for assessing and optimizing the fracture treatment by describing fracture geometry. This paper introduces a real-time microseismic mapping project conducted on a multistage fracture treatment in a specific tight sand reservoir of Changqing; during this project, a quantifying relationship between the fracture treatment parameters and stimulated reservoir volume (SRV) was realized, enabling stimulation optimization in this area. This paper details the microseismic mapping acquisition, hydraulic fracture symmetry, practical drainage pattern, evaluation of the fracture treatment parameters, and perforation strategy that impact fracture geometry and SRV as well as optimization in this particular tight sand reservoir. Introduction Located in northern middle China, the Ordos basin (Fig. 1) is the second largest depositional basin, covering an aerial extent of approximately 370,000 square kilometers, stretching across the Shaanxi, Gansu, Shanxi, Ningxia, and Inner Mongolia, with original oil in place (OOIP) of approximately 8.6 billion tons. The Ordos basin is a superimposed basin having experienced multiple tectonic movements. The stratigraphic fill of the Ordos basin is composed of numerous facies. These types vary between Carbonate Platform, Clastic sedimentary of Paleozoic, fluvial, delta, and lacustrine of Mesozoic; its total thickness is approximately 5000m. In the Ordos basin, hydrocarbon contribution formations of Triassic and Jurassic produce oil; whereas, Permo-Carboniferous and Ordovician produce gas. Across the entire board of the basin, oil reservoirs are distributed in the southern portion, while gas reservoirs are abundant in the northern part. Because of complex tectonics characteristics and low permeability, reaching its current total production of BOE of 50 million tons per year proved overwhelmingly difficult. Before the 1950s, it took 50 years to achieve production of 2,000 tons per year; it then took another 20 years to achieve production reaching 20,000 tons. Even up to the year 2000, production was only 7.5 million tons per year. In the last 10 years, as conventional sources of oil and gas have declined, operators have been driven to explore unconventional resources globally, and the Changqing field had made astonishing progress during this period.
The temporal data is widely utilized and studied. The researchers are interested in the storage management and indexing technique on this special type of data. XML is a newly developed data structure and exchange standard in recent years. Because of its flexible structure of the XML document and independent communication protocol, XML document is especially suitable for transmitting on the internet network and World Wide Web. Different from the traditional database management systems, the numerous characteristics of XML make the data easier to exchange and communicate. As a result, more data are stored in XML format. If the temporal data stored in XML structure, searching data must check for each data of document, which results in high cost. For temporal data, the traditional indexing techniques considered only single attribute in data clustering, making it faster to retrieve data by the indexed attribute. However, the efficiency declines dramatically when the search criteria is not indexed. In our research, we proposed a new data clustering method for improving comprehensive efficiency of data retrieval. By building the temporal data in the multidimensional matrix observes the relation among the data. Then, we try to cut matrix apart into a lot of blocks by considering all attributes and utilized the characteristics of greedy algorithm to improve the speed of the data clustering. By this approach, we can improve comprehensive efficiency of data retrieval.
PURPOSE Proton pencil beam scanning can provide precise and efficient treatment delivery for both conventional and intensity modulated proton therapy (IMPT) techniques. It is challenging to measure the 3-D dose distribution of a pencil beam (1 - 1.5 cm diameter) because of its small field size as well as the complexity of dose changes in various directions. This study is the first investigation in characterizing the dosimetry of a single proton pencil beam using PRESAGE radiochromic dosimeters and an optical CT scanner. METHODS In this study, cylindrical PRESAGE dosimeters and an optical CT scanner were used to implement dose distribution measurements for proton pencil beams at Massachusetts General Hospital. Two different dosimeter formulations, with different physical density and atomic number, were used for measurements for proton pencil beams, with energies between 93 MeV and 110 MeV. Optical density dose response was studied by irradiating the dosimeters to various known doses (up to 9 Gy). PDD distributions were studied for proton pencil beams, as well as unmodulated and modulated scattered beams. RESULTS PRESAGE has a linear dose response from 0 up to 8 Gy for all proton energies studied. The physical density of PRESAGE was used to scale the depth of PRESAGE measurement to depth of water. For pencil beams, normalized at surface, the dose at the Bragg peak is underestimated by 15-20% due to LET effects. The spread of a proton pencil beam as a function of depth can be observed. Measured PDDs were compared with that in water for both modulated (SOBP) and unmodulated scattered beams. Some LET effects are observed. CONCLUSION The findings from this study suggest that it is feasible to use PRESAGE dosimeter for proton pencil beam dosimetry study. Future studies will be focused on the responses of various PRESAGE formulations in proton pencil beams.
Purpose: Proton imaging is very sensitive to changes in water equivalent pathlength (WEPL), and can therefore differentiate soft tissues from target volumes. Most systems under development today use pencil beam with detector grids for path and energy determination. We explored a novel method for proton imaging by taking advantage of the time-resolved dose rate of a scattered proton beam produced by passive scattering systems with a cyclical range-modulator. Methods: We show that for such system with the modulator spinning at a constant speed, the WEPL can be determined along any ray by measuring the time-resolved dose rate patterns at the ray endpoints on an imaging plane. If we use a beam with sufficient energy to pass through the patient and measure the dose rate function at points behind the patient, we can obtain a 2D image consisting of WEPLs. In order to simulate this effect, we implemented time-resolved dose calculations in our in-house proton dose-calculation framework and applied these to lung cancer cases. An AP beam was used with a proton range of 24 g/cm2 and the dose rate patterns in a plane right below the posterior skin of the patient were calculated. These patterns were used to determine the WEPL from the anterior to the posterior surfaces as an image of the tumor and the surrounding tissues. Results: The lung tumor volumes are easily identified in the resultant radiograph of WEPL. In addition, further processing of the patterns based on range mixing can enhance areas where WEPL changed rapidly, especially for the tumor edge and chest wall. The estimated dose for imaging is less than 0.5 cGy. Conclusions: Our method is rapid (∼100 ms) and simultaneous over the whole field, it can image mobile tumors without the problem of interplay effect inherently challenging for methods based on pencil beams.
Web logs collected at proxy servers, referred to as proxy logs, contain rich information about Web user activities. These logs are becoming critical data sources for various Web applications such as Web log mining. However, a raw proxy log treated as a at sequence of individual Web requests does not reliably represent correct information about Web user behavior, owing to a lack of semantic structure. This problem has consequently impaired Web mining results from discovering meaningful knowledge eectiv ely. The W3C WCA working group conceptualizes the online behavior of a Web user in terms of user sessions and episodes. A user session is a delimited set of user clicks across one or more Web servers. An episode is a subset of related user clicks in a user session. In this paper, we investigate the problem of restoring semantic structure in a proxy log by classifying individual Web requests into semantically meaningful episodes so that it can serve as a more reliable input than its raw format for various knowledge discovery processes. Existing approaches to transaction identication as well as other log preprocessing techniques are found inadequate to identify episodes with clear semantics from proxy logs. This is because, in a proxy log, Web requests issued by dieren t Web users and made for dieren t navigation purposes are interleaved with each other whereas there are neither explicit identiers of individual Web users nor clear boundaries among user sessions and episodes. In this paper, we propose a semantics-based approach, cut-and-pick, in which a semantic stamp is used for distinguishing dieren t Web requests for dieren t semantic purposes. As
In this study, we propose a simple and novel data structure using hyper-links, H-struct, and a new mining algorithm, H-mine, which takes advantage of this data structure and dynamically adjusts links in the mining process. A distinct feature of this method is that it has a very limited and precisely predictable main memory cost and runs very quickly in memory-based settings. Moreover, it can be scaled up to very large databases using database partitioning. When the data set becomes dense, (conditional) FP-trees can be constructed dynamically as part of the mining process. Our study shows that H-mine has an excellent performance for various kinds of data, outperforms currently available algorithms in different settings, and is highly scalable to mining large databases. This study also proposes a new data mining methodology, space-preserving mining, which may have a major impact on the future development of efficient and scalable data mining methods.
Quantile computation has many applications including data mining and financial data analysis. It has been shown that an /spl epsi/-approximate summary can be maintained so that, given a quantile query (/spl phi/,/spl epsi/), the data item at rank /spl lceil//spl phi/N/spl rceil/ may be approximately obtained within the rank error precision /spl epsi/N over all N data items in a data stream or in a sliding window. However, scalable online processing of massive continuous quantile queries with different /spl phi/ and /spl epsi/ poses a new challenge because the summary is continuously updated with new arrivals of data items. In this paper, first we aim to dramatically reduce the number of distinct query results by grouping a set of different queries into a cluster so that they can be processed virtually as a single query while the precision requirements from users can be retained. Second, we aim to minimize the total query processing costs. Efficient algorithms are developed to minimize the total number of times for reprocessing clusters and to produce the minimum number of clusters, respectively. The techniques are extended to maintain near-optimal clustering when queries are registered and removed in an arbitrary fashion against whole data streams or sliding windows. In addition to theoretical analysis, our performance study indicates that the proposed techniques are indeed scalable with respect to the number of input queries as well as the number of items and the item arrival rate in a data stream.
Summarizing topological relations is fundamental to many spatial applications including spatial query optimization. In this article, we present several novel techniques to effectively construct cell density based spatial histograms for range (window) summarizations restricted to the four most important level-two topological relations: contains, contained, overlap, and disjoint. We first present a novel framework to construct a multiscale Euler histogram in 2D space with the guarantee of the exact summarization results for aligned windows in constant time. To minimize the storage space in such a multiscale Euler histogram, an approximate algorithm with the approximate ratio 19/12 is presented, while the problem is shown NP-hard generally. To conform to a limited storage space where a multiscale histogram may be allowed to have only k Euler histograms, an effective algorithm is presented to construct multiscale histograms to achieve high accuracy in approximately summarizing aligned windows. Then, we present a new approximate algorithm to query an Euler histogram that cannot guarantee the exact answers; it runs in constant time. We also investigate the problem of nonaligned windows and the problem of effectively partitioning the data space to support nonaligned window queries. Finally, we extend our techniques to 3D space. Our extensive experiments against both synthetic and real world datasets demonstrate that the approximate multiscale histogram techniques may improve the accuracy of the existing techniques by several orders of magnitude while retaining the cost efficiency, and the exact multiscale histogram technique requires only a storage space linearly proportional to the number of cells for many popular real datasets.
In this paper, we propose a data mining-based approach to public buffer management in distributed database systems where database buffers are organized into two areas: public and private. While the private buffer areas contain pages to be updated by particular users, the public buffer area contains pages shared among users from different sites. Different from traditional buffer management strategies where limited knowledge of user access patterns is used, the proposed approach discovers knowledge from page access sequences of user transactions and uses it to guide public buffer placement and replacement. The knowledge to be discovered and the discovery algorithms are discussed. The effectiveness of the proposed approach was investigated through a simulation study. The results indicate that with the help of the discovered knowledge, the public buffer hit ratio can be improved significantly.
Traditionally, building a classifier requires two sets of examples: positive examples and negative examples. This paper studies the problem of building a text classifier using positive examples (P) and unlabeled examples (U). The unlabeled examples are mixed with both positive and negative examples. Since no negative example is given explicitly, the task of building a reliable text classifier becomes far more challenging. Simply treating all of the unlabeled examples as negative examples and building a classifier thereafter is undoubtedly a poor approach to tackling this problem. Generally speaking, most of the studies solved this problem by a two-step heuristic: first, extract negative examples (N) from U. Second, build a classifier based on P and N. Surprisingly, most studies did not try to extract positive examples from U. Intuitively, enlarging P by P' (positive examples extracted from U) and building a classifier thereafter should enhance the effectiveness of the classifier. Throughout our study, we find that extracting P' is very difficult. A document in U that possesses the features exhibited in P does not necessarily mean that it is a positive example, and vice versa. The very large size of and very high diversity in U also contribute to the difficulties of extracting P'. In this paper, we propose a labeling heuristic called PNLH to tackle this problem. PNLH aims at extracting high quality positive examples and negative examples from U and can be used on top of any existing classifiers. Extensive experiments based on several benchmarks are conducted. The results indicated that PNLH is highly feasible, especially in the situation where |P| is extremely small.
Outlier detection techniques are widely used in many applications such as credit-card fraud detection, monitoring criminal activities in electronic commerce, etc. These applications attempt to identify outliers as noises, exceptions, or objects around the border. The existing density-based local outlier detection assigns the degree to which an object is an outlier in a numerical space. In this paper, we propose a novel mutual-reinforcement-based local outlier detection approach. Instead of detecting local outliers as noise, we attempt to identify local outliers in the center, where they are similar to some clusters of objects on one hand, and are unique on the other. Our technique can be used for bank investment to identify a unique body, similar to many good competitors, in which to invest. We attempt to detect local outliers in categorical, ordinal as well as numerical data. In categorical data, the challenge is that there are many similar but different ways to specify relationships among the data items. Our mutual-reinforcement-based approach is stable, with similar but different user-defined relationships. Our technique can reduce the burden for users to determine the relationships among data items, and find the explanations why the outliers are found. We conducted extensive experimental studies using real datasets.
Mining frequent itemsets from transactional data streams is challenging due to the nature of the exponential explosion of itemsets and the limit memory space required for mining frequent itemsets. Given a domain of I unique items, the possible number of itemsets can be up to 2I−1. When the length of data streams approaches to a very large number N, the possibility of an itemset to be frequent becomes larger and difficult to track with limited memory. The existing studies on finding frequent items from high speed data streams are false-positive oriented. That is, they control memory consumption in the counting processes by an error parameter ϵ, and allow items with support below the specified minimum support s but above s−ϵ counted as frequent ones. However, such false-positive oriented approaches cannot be effectively applied to frequent itemsets mining for two reasons. First, false-positive items found increase the number of false-positive frequent itemsets exponentially. Second, minimization of the number of false-positive items found, by using a small ϵ, will make memory consumption large. Therefore, such approaches may make the problem computationally intractable with bounded memory consumption. In this paper, we developed algorithms that can effectively mine frequent item(set)s from high speed transactional data streams with a bound of memory consumption. Our algorithms are based on Chernoff bound in which we use a running error parameter to prune item(set)s and use a reliability parameter to control memory. While our algorithms are false-negative oriented, that is, certain frequent itemsets may not appear in the results, the number of false-negative itemsets can be controlled by a predefined parameter so that desired recall rate of frequent itemsets can be guaranteed. Our extensive experimental studies show that the proposed algorithms have high accuracy, require less memory, and consume less CPU time. They significantly outperform the existing false-positive algorithms.
Decision support systems issue a large number of online analytical processing (OLAP) queries to access very large databases. A data warehouse needs to precompute or materialize some of such OLAP queries in order to improve the system throughput, since many coming queries can benefit greatly from these materialized views. Materialized view selection with resource constraint is one of the most important issues in the management of data warehouses. It addresses how to fully utilize the limited resource, disk space, or maintenance time to minimize the total query processing cost. This paper revisits the problem of materialized view selection under a disk-space constraint S. Many efficient greedy algorithms have been developed to address this problem. The quality of greedy solutions is guaranteed by a lower bound. However, it is observed that, when S is small, this lower bound can be very small and even be negative. In such cases, their solution quality will not be guaranteed well. In order to improve further the solution quality in such cases, a new competitive A* algorithm is proposed. It is shown that it is just the distinctive topological structure of the dependent lattice that makes the A* search a very competitive strategy for this problem. Both theoretical and experimental results show that the proposed algorithm is a powerful, efficient, and flexible approach to this problem
False-negative frequent items mining from a high speed transactional data stream is to find an approximate set of frequent items with respect to a minimum support threshold, s. It controls the possibility of missing frequent items using a reliability parameter δ. The importance of false-negative frequent items mining is that it can exclude false-positives and therefore significantly reduce the memory consumption for frequent itemsets mining. The key issue of false-negative frequent items mining is how to minimize the possibility of missing frequent items. In this paper, we propose a new false-negative frequent items mining algorithm, called Loss-Negative, for handling bursting in data streams. The new algorithm consumes the smallest memory in comparison with other false-negative and false-positive frequent items algorithms. We present theoretical bound of the new algorithm, and analyze the possibility of minimization of missing frequent items, in terms of two possibilities, namely, in-possibility and out-possibility. The former is about how a frequent item can possibly pass the first pruning. The latter is about how long a frequent item can stay in memory while no occurrences of the item comes in the following data stream for a certain period. The new proposed algorithm is superior to the existing false-negative frequent items mining algorithms in terms of the two possibilities. We demonstrate the effectiveness of the new algorithm in this paper.
David W. Cheung (张偉犖)合作论文数Department of Computer Science,University of Hong Kong7