Top-k Skyline查询结合了Top-k与Skyline的特性,可以在数据集中找到最好的点。但是,现有的算法在大数据环境下具有较高的时间开销。文中提出一种新的算法DFTS,其可以高效地在大数据集中进行Top-k Skyline查询。DFTS包括3个步骤:首先,利用度值评价函数对数据集进行排序,快速过滤掉大量的点,仅保留足够少的候选集;然后,对候选集进行Skyline查询计算,进一步排除掉Skyline集合外的点;最后,筛选出Top-k的数据点作为最终结果。通过这种方式,DFTS有效减少了算法的运行时间。从理论上证明了DFTS查询的最终结果符合Top-k Skyline查询的要求。基于大数据集的大量实验表明,DFTS具有比现有算法更好的性能。
Recent years there has been increasing concern about the rider demand responsive systems and the vehicular ad hoc networks. On one hand, centralised taxi platforms such as Uber and Didi Taxi are popular and changing our daily life; on the other hand, vehicles are equipped with more and more sensors and are capable to calculate, store, and communicate with other vehicles or road side units, forming vehicleto-vehicle or vehicle-to-infrastructure communications. However, little effort has been devoted to integrating these two fields. In this paper, we propose a distributed public vehicle (PV) system that integrates the rider demand responsive system and ad hoc vehicular technologies, where the concept of fog computing and vehicular sensing are adopted for the system design. The challenges lie in that the PV scheduling problem itself is NP-hard, and careful design of scheduling and cooperation schemes among nodes are needed as they are ubiquitously connected at the edge of networks. The proposed PV system adopts a heuristic request insertion algorithm and a cooperative strategy among vehicle nodes, fog nodes, and the cloud to dispatch requests and to schedule routes for PVs. Experimental studies on real-world data sets demonstrate that the proposed scheme achieves higher service ratio of requests and better efficiency than other transit methods. Furthermore, the distributed vehicular sensing is demonstrated to be capable of collecting feasible metadata for scheduling applications. To the best of our knowledge, this paper is the first report on the integration of fog nodes and vehicular sensing for the rider request responsive scheduling systems.
The training of big data professionals is the foundation of a new round of scientific and technological contest in the world.Colleges and universities assume the responsibility of training big data talents.As a typical “new engineering” major,the major of big data is still in the exploratory stage in the construction of curriculum system.Firstly,the difficulties in the construction of big data courses were analyzed,and then the big data course system built by Xiamen University was introduced,including introductory courses,advanced courses and training courses.Also the experience and methods of the course construction of the principles and applications of big data technology were introduced,which includes course orientation,training objectives,preparatory knowledge,knowledge partitioning between big data and cloud computing courses,course content and arrangement,teaching material,experimental environment construction,matching resources construction,online service platform,offline training and communication,and so on.
A new class of monitoring applications is emerging, in which multiple embedded devices are deployed to sense the physical world and a large amount of data is injected into the network. Yet, existing monitoring algorithms usually output a result set that is trivial for users and too expensive for the resource-constraint network. In this paper, we study the problem of iceberg join processing in wireless sensor networks. The iceberg join query only includes a small fraction of data in its result set, yet, still contains the most 'interesting' and useful data relationships and linkages of the sensing data. The proposed algorithm SRJA is output sensitive and adopts a progressive refinement strategy for the query processing. Our algorithm first constructs flexible synopses according to the characteristics of the joining data, and then progressively refines these synopses to identify tuples that can match and meet the iceberg threshold in the joining regions. It fully utilises the iceberg threshold to filter out tuples that do not contribute to the final result set at early stages, saving lots of transmissions. Extensive experiments indicate that our algorithm gains a reduction up to 25% of message transmissions compared with other schemes.
In this paper,a selective ensemble learning algorithm was proposed based on hierarchical selection and dynamic updating,which can optimize the parameters of classifier with multi-thread technique and select the sub sequence set of classifiers based on hierarchical selection and dynamical information.It can solve the problem in the past for choosing classifier to ensemble learning inefficiently.In addition,divide-and-conquer strategy is employed to reduce the time cost for ensemble voting.The big voting task can be divided recursively into small child task by dichotomy,then the tasks are executed in parallel and it would conquer the voting result.Experimental results show that the selective algorithm can outperform the traditional classification algorithms on F1-Measure and AUC.
Cost-sensitive classification is an important research topic in the classification problem.In order to improve the accuracy of MetaCost,which serves as a cost-sensitive classification algorithm,and reduce its misclassification cost,we propose a new cost-sensitive algorithm,called D-MetaCost,for multi-class problems.In D-MetaCost algorithm,we can calculate the accuracy of multiple models generated in the beginning of MetaCost algorithm,and select first few base classifiers with higher accuracy,then integrate them together with the new model of the last stage to obtainthe final classification model.Experimental results show that the proposed algorithm enjoys obvious improvements in accuracy and cost in comparison with the classical MetaCost algorithm.
The crowdfunding industry is growing rapidly worldwide and poses new challenges on how to understand investment behavior. Indeed, a key challenge in this area is how to measure the similarity of an investor and a company, or the interest of an investor in a company. Tremendous effort has been made in previous research regarding the single effective factor or homogeneous network model based on link prediction for investment behavior prediction. In this study, we build an investment behavior prediction model of meta-path-based heterogeneous network, which considers multiple entity and relation types associated with the investment behavior of a particular investor. Our investment behavior prediction model provides an effective similarity measure function for meta-path. To validate the proposed model, we perform experiments on real-world data from CrunchBase. Experimental results reveal that our investment behavior prediction model is indeed a useful indicator.
Robot-assisted cell microinjection, which is precise and can enable a high throughput, is attracting interest from researchers. Conventional probe-type cell microforce sensors have some real-time injection force measurement limitations, which prevent their integration in a cell microinjection robot. In this paper, a novel supported-beam based cell micro-force sensor with a piezoelectric polyvinylidine fluoride film used as the sensing element is described, which was designed to solve the real-time force-sensing problem during a robotic microinjection manipulation, and theoretical mechanical and electrical models of the sensor function are derived. Furthermore, an array based cell-holding device with a trapezoidal microstructure is micro-fabricated, which serves to improve the force sensing speed and cell manipulation rates. Tests confirmed that the sensor showed good repeatability and a linearity of 1.82%. Finally, robot-assisted zebrafish embryo microinjection experiments were conducted. These results demonstrated the effectiveness of the sensor working with the robotic cell manipulation system. Moreover, the sensing structure, theoretical model, and fabrication method established in this study are not scale dependent. Smaller cells, e.g., mouse oocytes, could also be manipulated with this approach.
Mobile sensing emerges as an important application for mobile networks. Smartphones equipped with sensors are used to monitor a diverse range of human activities. One key and challenging procedure of the mobile sensing applications is data gathering, where the sensed data from distributed mobile nodes are captured and uploaded to the cloud or base station for further processing. Yet the mobile sensing application, which usually periodically generates some sensed data, would definitely deteriorate the 3G quality because the network cannot cope with the high demand; and users would be charged at high prices by using the 3G channel, which makes the mobile sensing application infeasible. In this paper, we proposed a hybrid data gathering and offloading algorithm DGO for the mobile sensing applications. Besides the direct uploading through 3G or Wifi offloading, the sensed data could also be forwarded to other peer nodes through short range communications. Nodes collect meta-data such as remaining energy, contact regularity, and expected contact duration to calculate the upload/offload utility and upload priority for data segments. Based on these utility factors, each data segment could decide its own approach at a specific time for uploading. Experimental studies show that DGO is efficient in data gathering and data offloading in mobile sensing applications. Given the low accessibility of Wifi APs, DGO still gains about more than 30 % of data offloading compared with existing algorithms without much extra transmission overhead or delay.
Keyword search over relational databases makes it easier to retrieve information from structural data. One solution is to first represent the relational data as a graph, and then find the minimum Steiner tree containing all the keywords by traversing the graph. However, the existing work involves substantial costs even for those based on heuristic algorithms, as the minimum Steiner tree problem is proved to be an NP-hard problem. In order to reduce the response time for a single search to a low level, a progressive ant-colony-optimization-based algorithm, called PACOKS, is proposed here, which achieves the best answer in a step-by-step manner, through the cooperation of large amounts of searches over time, instead of in an one-step manner by a single search. Through this way, the high costs for finding the best answer, are shared among large amounts of searches, so that low cost and fast response time for a single search is achieved. Extensive experimental results based on our prototype show that our method can achieve better performance than those state-of-the-art methods.
Background: MicroRNAs play important roles in the progression of various diseases. Therefore, it is of vital importance to predict novel microRNA-disease associations for understanding disease mechanisms.Objective: As far as we see, there are generally three problems for the microRNA-disease association prediction. The first one is the lack of similarity among miRNAs. The second one is the presence of a few defined relationships between miRNAs and diseases. The insufficient number of available negative samples for studies on miRNA-disease associations is another troubling issue. We aimed to solve the three problems with the inductive matrix completion method.Method: In this paper, the inductive matrix completion method is exploited to overcome the three problems. We also contributed multiple feature sets to address problems related to insufficient miRNA-disease association data. The method could be applied to predict unknown microRNA-disease associations and new pathogenic miRNAs for well-characterized diseases.Results: Experiments can prove the performance of our inductive matrix completion method. The experiment is compared with several current methods through cross-validation. Our result reveals the superiority of our method to other approaches.Conclusion: We can conclude that the inductive matrix completion method is more suitable than transductive one, for the prediction of microRNA-disease associations.
String similarity join is widely used in many fields, e.g. data cleaning, web search, pattern recognition and DNA sequence matching. During the recent years, many similarity join methods have been proposed, for example Pass-Join, Ed-Join, Trie-Join, and so on, among which the Pass-Join algorithm based on edit distance can achieve much better overall performance than the others. But Pass-Join can not effectively filter those candidate pairs which are partially similar. Here a novel algorithm called GFSF is proposed, which introduces two additional filtering steps based on character frequency vector. Through this way, the number of pairs which are only partially similar are greatly reduced, thus greatly reducing the total time of string similarity join process. The experimental results show that the overall performance of the proposed method is better than Pass-Join.
Classification with imbalanced class distributions is a major problem in machine learning. Researchers have given considerable attention to the applications in many real-world scenarios. Although several works have utilized the area under the receiver operating characteristic (ROC) curve to select potentially optimal classifiers in imbalanced classifications, limited studies have been devoted to finding the classification threshold for testing or unknown datasets. In general, the classification threshold is simply set to 0.5, which is usually unsuitable for an imbalanced classification. In this study, we analyze the drawbacks of using ROC as the sole measure of imbalance in data classification problems. In addition, a novel framework for finding the best classification threshold is proposed. Experiments with SCOP v.1.53 data reveal that, with the default threshold set to 0.5, our proposed framework demonstrated a 20.63% improvement in terms of F-score compared with that of more commonly used methods. The findings suggest that the proposed framework is both effective and efficient. A web server and software tools are available via http://datamining.xmu.edu.cn/prht/ or http://prht.sinaapp.com/.
With the assumption that each appointment patient up in clinic arrives on time,this paper investigates the appointment scheduling problem with no-show.An integer programming model is established with the numbers of appointment patient and probability distribution of remaining patients at each slot as optimization variables.The objective function includes the benefits of serving patients,the patient waiting time costs,and system overtime expenses.By relaxing the coupled constraints about probability distribution of remaining patients at each slot,this paper proposes a Lagrangian relaxation method to solve the model with a dynamic programming algorithm solving the relaxation problem and a sub-gradient method solving the dual problem.Numerical experiments show that the algorithm can find the optimal solution for small scale problems,and that the best solution from the proposed algorithm is better than the one in the related literature for large scale problems.
提出了一种新的基于B-树的闪存数据库索引——CF-HNLBI索引.使用链表组织缓冲区中的更新信息,减少了缓冲区遍历时间,通过链表结构减少冗余信息,提高了缓冲区利用率.将缓冲区分为冷区和热区,并采用基于更新信息频度的替换算法,有效地减少了闪存写操作次数.实验结果表明,CF-HNLBI索引比其他已有索引具有更好的性能.
Data gathering is a key operator for applications in wireless sensor networks; yet it is also a challenging problem in mobile sensor networks when considering that all nodes are mobile and the communications among them are opportunistic. This paper proposes an efficient data gathering scheme called ADG that adopts speedy mobile elements as the mobile data collector and takes advantage of the movement patterns of the network. ADG first extracts the network meta-data at initial epochs, and calculates a set of proxy nodes based on the meta-data. Data gathering is then mapped into the Proxy node Time Slot Allocation (PTSA) problem that schedules the time slots and orders, according to which the data collector could gather the maximal amount of data within a limited period. Finally, the collector follows the schedule and picks up the sensed data from the proxy nodes through one hop of message transmissions. ADG learns the period when nodes are relatively stationary, so that the collector is able to pick up the data from them during the limited data gathering period. Moreover, proxy nodes and data gathering points could also be timely updated so that the collector could adapt to the change of node movements. Extensive experimental results show that the proposed scheme outperforms other data gathering schemes on the cost of message transmissions and the data gathering rate, especially under the constraint of limited data gathering period.
MapReduce是一个并行分布式计算模型,已经被广泛应用于处理两个或多个大型表的连接操作.现有的基于MapReduce的多表连接算法,在处理链式连接时,不能处理多个大表的连接,或者需要顺序运行较多的MapReduce任务,效率较低.为此提出了一种基于MapReduce的多表连接算法——PipelineJoin,高效地实现任意多个大表的链式连接.PipelineJoin采用流水线模型和调度器来实现MapReduce任务的流水线式执行,从而有效提高多表连接的效率,同时可以较好地克服链式多表连接算法的缺陷.最后,在不同规模的数据集上进行了大量实验,实验结果表明Pipeline Join算法与原有链式多表连接算法相比,可以有效减少连接所需的时间.
MicroRNAs constitute an important class of noncoding, single-stranded, ~22 nucleotide long RNA molecules encoded by endogenous genes. They play an important role in regulating gene transcription and the regulation of normal development. MicroRNAs can be associated with disease; however, only a few microRNA-disease associations have been confirmed by traditional experimental approaches. We introduce two methods to predict microRNA-disease association. The first method, KATZ, focuses on integrating the social network analysis method with machine learning and is based on networks derived from known microRNA-disease associations, disease-disease associations, and microRNA-microRNA associations. The other method, CATAPULT, is a supervised machine learning method. We applied the two methods to 242 known microRNA-disease associations and evaluated their performance using leave-one-out cross-validation and 3-fold cross-validation. Experiments proved that our methods outperformed the state-of-the-art methods.
Keyword search over relational database has been researched a lot. By using simple keywords to search over relational data, ordinary users are not required to learn the difficult structural query language, thus resulting in better user friendliness. Search effectiveness is an important consideration for those solutions to this problem. The available methods adopt static ranking mechanism to ensure that the most relevant answer will be presented first to users. However, they are not able to dynamically optimize the search results according to the time-changing user interest. Here an ant-colony-optimizaton-based algorithm, called ACOKS, is proposed to deal with keyword search problem, in which node-temperature-based optimization is used to achieve dynamic search result optimization by following the track of user behavior. Extensive experimental results show that our methods can achieve better performance than the state-of-the-art methods.