
The Social Web makes visible the ebb and flow of popular interest in topics both newsworthy ("GulfSpill") and trivial ("Lolcat"). Understanding this emergent behavior is a fundamental goal for Social Web research. Key problems include discovering emergent topics from online text sources, modeling burst activity, and predicting the future trajectory of a given topic. Past work has addressed such problems individually for specific applications, but has lacked a generalizable framework for performing both classification and prediction of topic usage. Our approach is to model a topic as a temporally ordered sequence of derived feature states and capture characteristic changes in the topic trend. These sequences are drawn from a dynamic segmentation of frequency data based on change point analysis. We employ Partitioning Around Medoids clustering on these segments to produce signatures which highlight characteristic patterns of usage growth and decay. We demonstrate how this signature model can be used to define distinctive classes of topics in multiple online contexts, including tagging systems and web-based information retrieval. Additionally, we show how the model can predict the general trajectory of interest in a particular topic.
Recently, a variety of indexing techniques have been proposed for optimizing keyword search on graph. However, graph indexing has very high space and time complexities, and thus these single-machine in-memory indices are usually not affordable for massive graphs. In this paper, we propose a novel distributed disk-based index, which organizes the local topology information in the graph to track and prune matched vertices that will not participate in the top-k answers to a specified query before search with heuristics. The distributed index can be constructed in a MapReduce manner. Moreover, a parallel search algorithm is also developed. It runs multiple asynchronous search instances that incrementally enumerate the current best local answers and then produces the global top-k answers from them. Lastly, we perform experiments on both synthetic and real graphs with various configurations. The results show that our approach can improve search efficiency on massive graphs significantly with affordable indexing overheads.
Automatically extracting the headline of online web articles has many applications in web mining and information retrieval. In this paper, we developed a content-based and domain-and language-independent approach, TitleFinder , for unsupervised extraction of the headline of web articles. TitleFinder starts by using a heuristic to select a candidate headline. In a second step the contents of each text fragment in the HTML file are compared to the candidate headline. We implemented four types of similarity for this comparison: two variations of the cosine similarity based on tf and tf-idf weighting schemata, an overlap scoring similarity and an aggregated metric combining the scores of the previous three similarities. Our method achieves high performance in terms of effectiveness and efficiency and outperforms approaches operating on structural and visual features on a test set consisting of 11,218 news web pages from 15 different domains.
Contextual factors can greatly influence users' decisions in selecting items, such as songs when listening to music. The goal of a context-aware recommender system is to adapt its recommendations not just to the general preferences of users, but also to the context in which users are seeking those recommendations. In the domain of music recommendation, the explicit contextual factors and their values might not be known to the system, a priori. Moreover, the contextual state of a user can be dynamic and change during an interaction with the system. In this paper, we present a hybrid context-aware recommender system which infers contextual information from the sequence of songs listened to or specified by a user and uses this information to produce context-aware recommendations. Our system mines popular tags for songs from social media Web sites and uses a topic modeling approach to learn latent topics representing various contexts. We then model each song as a set of latent topics capturing the general characteristics of that song. This representation is used to track and detect changes in user's choice of music, as reflected in a playlist of song sequence, and adjust the recommendations to better meet the current context of the user. Using our approach, the contextual information can be integrated with any traditional recommendation algorithm to produce context-aware recommendations. For our system, we designed and evaluated two hybrid methods. The first hybrid combines collaborative filtering and content-based recommendation techniques, and the second hybrid additionally incorporates information about pairwise song associations. Our evaluation results show that both the hybrid approach and the contextualization can enhance the performance of baseline music recommendation method.
The paper presents a tractable subclass of DTDs, called DC-DTDs, for XPath satisfiability with sibling axes. A DC-DTD is a DTD such that each content model is in the form of a concatenation of single tag names and Kleene-starred regular expressions. DC-DTDs are a proper subclass of covering DTDs proposed by Montazerian et al., and a proper superclass of disjunction-free DTDs. In this paper, it is shown that tractability by covering DTDs is fragile against sibling axes. Then, tractability of XPath satisfiability with sibling axes under DC-DTDs is demonstrated. Finally, as a limitation of the tractability of DC-DTDs, it is shown that upward axes appearing in qualifiers bring intractability under even disjunction-free DTDs.
Modern intelligence analysis often involves a complex, iterative, highly branched sequence of information gathering and processing steps. Analysts can benefit greatly from Mind Snaps, semantic bookmarks that would allow them to return to a particular point in the analysis and recreate the complete context they had at that time. This paper addresses some basic issues related to creating and maintaining Mind Snaps. One issue is how frequently we need to take a Mind Snap. Our experiment shows that 10 to 30 analyst events offer 85 percent to 95 percent precision in the ability to distinguish analysts working on different tasks. This translates into an interval for taking Mind Snaps that should be between five to 15 minutes. Another important issue the paper addresses is how to separate actions into multiple micro-contexts in an environment where the analyst often concurrently engages in multiple tasks. The key to this issue is the ability to detect the change in contexts, i.e., context switch. We have developed an algorithm for separating context based on user modeling. Our experiment uses this algorithm to demonstrate the feasibility of capturing and disentangling the analytic micro-contexts. In particular, our results show that context switches can be successfully detected using as few as 10 analysis log event (ALE) windows. Better detection is achieved with larger windows. At a widow size of 30 ALE, we achieved a precision of 73 percent and a recall of 70 percent.
Nowadays, Web Applications (WAs) are complex software systems, used by multiple users with different roles and often developed to support and manage business processes. Due to the changing nature of the supported processes, WAs need to be easily and quickly modified, to adapt and align them to the processes they support. In recent years, Model Driven Engineering (MDE) approaches have been proposed and used to develop and evolve WAs. However, the definition of appropriate MDE approaches for the development of flexible process-centric WAs is still limited. In particular, (flexible) workflow models have never been integrated with the models (e.g., presentation, information models) used in MDE approaches to develop this type of applications. In this paper, we present M3D (Model Driven Development with Declare), a tool for developing WAs that integrates three MDE metamodels used to represent the main components of a WA with the metamodel of Declare, a declarative language to model business processes. The tool exploits and combines the declarative nature of Declare and the advantages of MDE to get an efficient roundtrip engineering support to develop and evolve flexible process-centric WAs.
Just a few years ago at the height of the Web 2.0 era, it was hard to believe that a significant amount of user generated data would be created and consumed without the use of the Web. In June 2011, researchers reported that time spent on native apps began to outpace time spent on the desktop or mobile Web [9]. People work, play, communicate and create content using narrowly focused apps, generating data the size of which is now comparable to the size of the Web. For example, Twitter users post 400 million tweets per day, most of them come from mobile clients [5], Foursquare boasts millions of check-ins everyday, more than 5 million photos per day are uploaded to Instagram [3], Spotify users listened to 13 billion songs during the first year the music streaming service was available in the US [1]. Surfing websites through hyperlinks to discover new information or using Web search engine for exploratory queries now seems so ineffective. Instead, we use social networks and specific apps as our daily news source, as well as sources of recommendations for places, music, films and other aspects of our life. This has become possible due to the smart instrumentation of modern social networks: users can connect not only to their friends, but also subscribe or follow interesting people, without being friends in real life, as well as brands, celebrities and news channels. Typically, data generated in apps are spread via friend connections in social networks. Many apps share their data (explicitly or implicitly) to Facebook. With Facebook Open
Social networking, and sophisticated wireless and positioning systems are fast developing and ever increasing technologies. Mobile social applications have the ability to increase the social connectivity by capturing automatically users' daily routines with Global Positioning System (GPS) receivers. These applications allow to record users' trajectories based on daily travel routes as well as to share experiences and interests among friends. However, there is always an increasing demand for providing an easy way to manipulate trajectory data, to generate and compare user profiles. Effective analysis of spatial trajectories has become an essential requirement to explore and understand the behavior of moving objects. In this paper, we highlight the importance of capturing users' daily routines in the form of trajectories in order to strengthen social connectivity. We also present the conceptual approach to multi-layer data representation in order to extract points of interest of correlated trajectories. Finally, we show how the data model could provide mobile social applications with direct support for trajectories at different abstraction levels.
In this paper, we describe the project SNOPS, a smart-city environment based on Future Internet technologies. We focused on the context-aware recommendation services provided in the platform, which accommodates location dependent multimedia information with user's needs in a mobile environment related to an outdoor scenario within the Cultural Heritage domain. In particular, we describe a recommendation strategy for planning browsing activities exploiting objects features, users' behaviors and context information gathered by apposite sensor networks. Preliminary experimental results, related to user's satisfaction, have been carried out and discussed.
Middleware is an important part of many search engine web crawling processes. We developed a middleware, the Crawl Document Importer (CDI), which selectively imports documents and the associated metadata to the digital library CiteSeerX crawl repository and database. This middleware is designed to be extensible as it provides a universal interface to the crawl database. It is designed to support input from multiple open source crawlers and archival formats, e.g., ARC, WARC. It can also import files downloaded via FTP. To use this middleware for another crawler, the user only needs to write a new log parser which returns a resource object with the standard metadata attributes and tells the middleware how to access downloaded files. When importing documents, users can specify document mime types and obtain text extracted from PDF/postscript documents. The middleware can adaptively identify academic research papers based on document context features. We developed a web user interface where the user can submit importing jobs. The middleware package can also work on supplemental jobs related to the crawl database and respository. Though designed for the CiteSeerX search engine, we feel this design would be appropriate for many search engine web crawling systems.
In this paper we introduce a novel approach for the thematic organization of bibliographic records that builds upon a semantic relatedness measure we have implemented for this task. In particular, we introduce the Omiotis measure, which captures the semantic relatedness between text segments and enables the thematic organization of the bibliographic data stored in online databases. Experimental evaluation demonstrates that Omiotis can significantly improve the performance of several data mining tasks, such as publications' classification and clustering, compared to existing approaches; even when considering a limited amount of information, i.e., the paper titles.
Generalization of web sessions is an effective approach used to overcome two major challenges in web usage mining, namely quality and scalability. Given a concept hierarchy, such as a website, generalization replaces actual page-clicks with their general concepts, i.e., nodes at higher levels. Presently known methods do this by choosing a level in the hierarchy, below which all the nodes are generalized to nodes at this level. The problem with this is that significant items may be coalesced, and insignificant ones may be retained. We present a usage driven generalization algorithm, which coalesces less significant pages into more general ones, independent of their level in the hierarchy. Based on actual usage set of sessions, item significance is estimated approximately but fast, using a small stratified sample of the large dataset. While providing scalability, the proposed generalization technique results in improved efficiency and quality of the discovered usage model, demonstrated through numerous experiments in our work.
Question-Answering (QA) service is a growing area of research study, and commercial QA systems have recently been developed. We are motivated to provide complementary QA service that answers questions in advertisements (ads). These days with almost all businesses online, potential buyers who search for merchandises to purchase through the Internet are also flourishing. When a Web user looks for products online, he may have many questions on his mind for which he would be eager to receive answers prior to finalizing his purchasing decision. Although some ads Web sites are complemented with FAQs, their QA services either are non-existent or do not provide answers to inquires in real time automatically. We address these problems by answering user's questions such as "Which is the cheapest car?", "Are there any entry-level, software developer positions?", etc., spontaneously in real time. Existing general-purpose QA systems, such as Ask.com, provide answers to a user's question Q in a list format. A more sophisticated approach is to order the answers to Q according to their degrees of relevance to Q . We propose a QA system which deals with the challenge of interpreting users' questions and retrieves correct, as well as partially-matched ranked, answers. Experimental results have verified that the proposed QA system is highly accurate in answering users' questions on car ads.
A large number of wrappers generate tables without column names for human consumption because the meaning of the columns are apparent from the context and easy for humans to understand, but in emerging applications, labels are needed for autonomous assignment and schema mapping where machine try to understand the tables. Autonomous label assignment is critical in volume data processing where ad hoc mediation, extraction and querying is involved. We propose an algorithm Lads for Labeling Anonymous Datasets, which can holistically label tabular web document. The algorithm has been tested on anonymous datasets from a number of sites, e.g music, movie, political, demographic, athletic obtained through different search engines such as Google, Yahoo and MSN. The comparative probabilities of attributes being candidate labels are presented which seem to be very promising, achieved as high as 93% probability of assigning good label to anonymous attribute. To the best of our knowledge, this is the first of its kind for label assignment based on multiple search engines' recommendation.
The first task any individual faces after joining an online social network (OSN) is locating friends that are present on that particular site. Most OSNs offer some variation of a tool that imports email contact lists to facilitate the task of finding one's friends. However, given that OSNs attempt to reconnect individuals with past acquaintances, one might not have access to the email address for a long lost friend. Furthermore, people tend to utilize a number of aliases online, meaning that an email address cannot always be used to reliably find a friend. Thus, new members must still manually search for friends based on a number of biographical attributes, such as gender, age, hometown, etc. It is not clear, however, what attributes are useful for conducting the search. Even after the search has been performed, the person performing the search might be left with a number of candidate profiles. In this paper, we develop a system for searching and matching individuals in OSNs. We evaluate the efficacy of our person matching techniques by measuring the overlap between social networks, and comparing our results to those published by compete.com. We then look at several interesting properties of overlapping profiles in both networks.
As the volume of multimedia data available on internet is tremendously increasing, the content-based similarity search becomes a popular approach to multimedia retrieval. The most popular retrieval concept is the k nearest neighbor (kNN) search. For a long time, the kNN queries provided an effective retrieval in multimedia databases. However, as today's multimedia databases available on the web grow to massive volumes, the classic kNN query quickly loses its descriptive power. In this paper, we introduce a new similarity query type, the k distinct nearest neighbors (kDNN), which aims to generalize the classic kNN query to be more robust with respect to the database size. In addition to retrieving just objects similar to the query example, the kDNN further ensures the objects within the result have to be distinct enough, i.e. excluding near duplicates.
Many bursty topics which are difficult to summarize and search exist in web forums. Most existing topic detection and tracking (TDT) methods deal with the news stories, but the language used in web forums are much casual, oral and informal compared with news data. In this paper, we present a noise-filtered model to extract bursty topics from web forums using terms and participations of users. Conducting experiments in ShuiMu community we demonstrate the efficiency of our model. Our model not only extracts bursty topics which are better organized for search and visualization, but also discoveries communities corresponding to these topics.
In this paper, we present how to use graph model for clustering Vietnamese document incrementally. Graph based model allows us to model completely the structure of not only each document but also the whole collection of documents. The graph structure is easily updated when there is a new document. When building the graph incrementally we can identify representative subgraph features, which are later used for calculating hybrid pair-wise document similarity. These subgraph features make clustering process less sensitive to the Vietnamese word segmentation step. Based on the hybrid similarity measure, the documents are groups into clusters on-the-fly without any assumptions on the number of clusters and without retrieving previous documents.