Grid computing technology makes it possibly to use scientific workflow to be shared and reused among different users. However, a scientific workflow model usually needs to be tailored during reuse because of different problem contexts met by different users. These change logs provide additional information to help users customize scientific workflow. In this paper, we present a change sequence mining approach to automatically induce customized workflow based on different scientific experiment contexts. A docking case for drug design is described to validate the approach.
Users of social networking services can connect with each other by forming communities for online interaction. Yet as the number of communities hosted by such websites grows over time, users have even greater need for effective community recommendations in order to meet more users. In this paper, we investigate two algorithms from very different domains and evaluate their effectiveness for personalized community recommendation. First is association rule mining (ARM), which discovers associations between sets of communities that are shared across many users. Second is latent Dirichlet allocation (LDA), which models user-community co-occurrences using latent aspects. In comparing LDA with ARM, we are interested in discovering whether modeling low-rank latent structure is more effective for recommendations than directly mining rules from the observed data. We experiment on an Orkut data set consisting of 492,104 users and 118,002 communities. Our empirical comparisons using the top-k recommendations metric show that LDA performs consistently better than ARM for the community recommendation task when recommending a list of 4 or more communities. However, for recommendation lists of up to 3 communities, ARM is still a bit better. We analyze examples of the latent information learned by LDA to explain this finding. To efficiently handle the large-scale data set, we parallelize LDA on distributed computers and demonstrate our parallel implementation's scalability with varying numbers of machines.
This paper presents PLDA, our parallel implementation of Latent Dirichlet Allocation on MPI and MapReduce. PLDA smooths out storage and computation bottlenecks and provides fault recovery for lengthy distributed computations. We show that PLDA can be applied to large, real-world applications and achieves good scalability. We have released MPI-PLDA to open source at http://code.google.com/p/plda under the Apache License.
Grid technologies give the chance to share distributed computing resources in Internet. Many researches and projects have done to deploy and integrate legacy applications and programs into service-oriented grid architecture. However, the coordination of legacy applications is still scarce of good approaches. ChinaGrid support platform (CGSP) is the kernel grid middleware of ChinaGrid project. It enables deployment, invocation and coordination of legacy applications in grid. This paper introduces the architecture of CGSP, presents the legacy application execution mechanism through CGSP, come up with an abstract workflow specification named ChinaGrid Workflow Description Language (CGWDL) which can be transformed into BPEL, analyze the life cycle of coordinating legacy applications based on CGSP. An image processing application is used to validate the feasibility and efficiency of our solution.
Service process management system (SPMS) plays a very important role in todaypsilas business process management environment and automatic configuration is essential to ensure its scalability and stability. However, current simple and static configuration mechanism doesnpsilat work well when the load of SPMS changes greatly with the time. An automatic configuration algorithm for distributed service process engine based on fuzzy control is introduced. The algorithm is implemented in a SPMS prototype based on the JINI platform. The experiment shows this automatic configuration mechanism improves the ability of adaptability and the efficiency of SPMS.
Frequent itemset mining (FIM) is a useful tool for discovering frequently co-occurrent items. Since its inception, a number of significant FIM algorithms have been developed to speed up mining performance. Unfortunately, when the dataset size is huge, both the memory use and computational cost can still be prohibitively expensive. In this work, we propose to parallelize the FP-Growth algorithm (we call our parallel algorithm PFP) on distributed machines. PFP partitions computation in such a way that each machine executes an independent group of mining tasks. Such partitioning eliminates computational dependencies between machines, and thereby communication between them. Through empirical study on a large dataset of 802,939 Web pages and 1,021,107 tags, we demonstrate that PFP can achieve virtually linear speedup. Besides scalability, the empirical study demonstrates that PFP to be promising for supporting query recommendation for search engines.
Service process usually needs to be changed during reuse because of the change of context. These change mapping relations provide additional information to help users customize service processes. In this paper, we present a mining approach to mining process change sequences based on different context and finding the best sequence to tailor the base process.
Wireless mesh sensor network (WMSN) merges advantages of wireless mesh networks and wireless sensor networks, especially on scalability, robustness and balanced energy dissipation. Routing in WMSNs faces with more challenges than that in traditional sensor networks on account of multiple sink nodes and the mobility of nodes. This paper focuses on two challenging problems. Firstly, we propose a reliable architecture of WMSNs by deploying multiple mobile mesh nodes in each sensor network to collect sensed data, which improves the scalability and performance of WMSNs. Also, we design a routing protocol characteristic to WMSNs. The routing protocol aims at maximizing the lifetime of sensor networks by reducing total energy consumption of a sensor network, as well as balancing energy usage among sensor nodes.
Services composition is a way to integrate a group of individual grid services into one more powerful service so that the time and efforts to develop a new application can be reduced greatly. Workflow modeling and enactment is one of the most important methods for service composition. However, the existing approaches do not provide enough functionality to support workflow modeling and enactment for flexible services composition. We have developed a system named ECA (event-condition-action)-rule-based workflow management system (EWMS) for grid services composition. The system architecture is introduced in detail in the paper. An image processing application is used to validate the feasibility and efficiency of the system.
As Grid technology is expanding from scientific computing to business applications, transactional workflow management emerges as one of the most important services for Grids. The ShanghaiGrid project launched in Shanghai, China implemented a basic workflow service without the reliability support. In this paper , we propose a transactional Grid workflow service (GridTW), providing a reliable and automatic workflow management for the ShanghaiGrid as well as other Grid workflow applications. This paper focuses on how to manage Grid transactions and combine transaction management with the workflow service. We present an automatic compensation based coordination algorithm with which the GridTW guarantees the reliability of Grid workflows. An important feature of our algorithm is that it can adapt to the dynamics Grid applications by the event- and condition-driven mechanism, and allows users to select execution results from committed subtransactions.
In grid environment, workflow process can be seen as not only cooperative approach of grid services and resources, but also reusable and sharable knowledge to settle specific problem. The research of grid workflow process clustering can promote knowledge discovery and reuse in grid. In this paper, we put forward a grid workflow process design method using event-condition-action (ECA) rule, and propose a new process similarity measure approach. Then, we use a case to prove the feasibility of the approach and show how to revise present clustering algorithm with the similarity measure approach briefly.
Since many of the key techniques now being applied in building services and a lot of service-based applications were developed, service-oriented computing (SOC) has become the main trend of distributed computing research field. In this paper, we proposes a workflow-based middleware named ShanghaiGrid Support Platform (SGSP) for SOC in the Grid environment, which provides a grid toolkit for ShanghaiGrid application developers and specific grid constructors. The toolkit can access, invoke and compose diverse services conveniently. The architecture and function of the each component of toolkit have been introduced in the paper. We also use an image processing application to validate the feasibility and efficiency of SGSP.
The technology of grid workflow has attracted growing attention in the recent years. With the desire for knowledge retrieval and reuse, the need of a se- mantic grid workflow description framework emerged. In this paper, we put forward a new angle of view that use goal ontology derived from request engineering to describe the semantic of grid workflow. An extendable and configurable semantic framework is proposed to describe and query workflow with semantic template based on the goal ontology. A practical case is introduced to validate the feasibility and efficiency of the framework.
With the development of Grid Technologies, Many researches and projects have done to deploy and integrate legacy programs into service-oriented Grid architecture. However, the legacy programs orchestration is still scarce of good approaches. In this paper, we overviews the architecture and legacy program execution mechanism of CGSP, presents an abstract workflow specification named ChinaGrid Workflow Description Language (CGWDL) which can be transformed into BPEL, and depicts the life cycle of legacy programs orchestration based on CGSP. An image processing application is used to validate the feasibility and efficiency of our solution.
Radio frequency identification (RFID) provides a quick, flexible, and reliable electronic means to detect, identify, track, and manage a variety of items. It has the potential to significantly alter how processes occur and how companies operate. Currently, the development of RFID and Grid technologies has opened the door to many new RFID applications, many of which require transaction support. This paper proposes a transaction commit protocol which can ensure reliable execution of RFID applications in Grid environment, and a transaction compensating approach to undo committed sub-transactions. Reliable RFID applications can be realized by the proposed protocol effectively.
Intelligent city traffic for travelling navigation, traffic prediction and decision support needs to collect large-scale real-time data from numerous vehicles. As a small, economical yet reasonably efficient device, wireless sensors can conveniently serve for this purpose. In this paper, we investigate how to deploy wireless sensor networks in buses to gather traffic data for intelligent city traffic. The paper presents a self-organization mechanism and a routing protocol for the proposed sensor networks. Our work has three advantages: (1)adaptive network topology, which satisfies highly mobile city traffic environment, (2)directed data transmission, saving energy consumption of sensor nodes with limited power resource, and (3)longer lifetime because of fewer redundant network communication and balanced power usage of sensor nodes in a network.
ShanghaiGrid is the first metropolitan grid in China. The project aims to develop an environment to share distributed resources in Shanghai conveniently based on grid technology. Yet, resources are finite for ever, especially for some expensive and specific resources. Users must reserve these resources in advance before use them. Present solutions need users confirm the absolute start time of reservation before the reservation which is impossible in practice at most time. But they can often estimate the relative time to start reservation after a specific event, e.g. beginning or finishing a task. As a part of ShanghaiGrid, we develop an ECA rule-based workflow management system (EWMS) which can serve for arranging the relative start time of advance resource reservation as users' demand. We introduce the design architecture of EWMS in the paper, and give a modeling demo to show how to realize advance resource reservation through our system
In this paper we propose GRAMS, a resource monitoring and analysis system in Grid environment. GRAMS provides an infrastructure for conducting online monitoring and performance analysis of a variety of Grid resources including computational and network devices. Based on analysis on real-time event data as well as historical performance data, steering strategies are given for users or resource scheduler to control the resources. Besides, GRAMS also provides a set of management tools and services portals for user not only to access performance data but also to handle these resources.
Grid computing is becoming a mainstream technology for large-scale distributed resource sharing and system integration. One of the most important grid services is workflow management. Grid workflow applications are also emerging as one of the most interesting application classes for the grid. In this paper, we give an introduction to our agent-based grid workflow management system (AGWMS). AGWMS has a four-layer framework. It bases on the adapter middleware and uses a multi-agent platform to make the system more robust, flexible and intelligent. Artificial Intelligence (AI) planning technology is also utilized to generate the agent plan automatically. Our AGWMS is a novel one.
Composing autonomous Web services to achieve new function, is receiving more and more attention in recent years in several computer science communities. The workflow model is the most important part in Web services composition. However, the existing approaches have difficulties in the process because their complexity. In this paper, we have developed an ECA (Event-Condition-Action)-rule-based workflow management system (EWMS) for Web service composition, which can help users compose Web services conveniently. We also use an image processing application to validate the feasibility and efficiency of our system.
Daqiang Zhang (张大强)合作论文数同济大学软件学院2