
This work focuses on understanding the user intent in the medical domain. The combination of Semantic Web and information retrieval technologies promises a better comprehension of user intents. Mapping queries to entities using Freebase is not novel, but so far only one entity per query could be identified. We overcome this limitation using annotations provided by Metamap. Also, different approaches to map queries to Freebase are explored and evaluated. We propose an indirect evaluation of the mappings, through user intent defined by classes such as Symptoms, Diseases or Treatments. Our experiments show that by using the concepts annotated by Metamap it is possible to improve the accuracy and F1 performances of mappings from queries to Freebase entities.
In this paper we describe the Searching as Learning Workshop (SAL 2014) taking place at IIiX 2014 in Regensburg, Germany.
This paper suggests a new interactive system based on visualization of the user's knowledge schema to aid the process of information search and knowledge discovery. The system, inspired from the human memory model, constructs an external representation of the user's conceptual knowledge structure from the user's existing information such as the folder structure and associated tags. In contrast to existing systems based on user models, the system allows users to specify their information needs by selecting a part of their knowledge schema, understand their search results in the context of their existing knowledge organization, and store new information based on their knowledge schema. Through a preliminary evaluation, we show that the visualization of the extracted knowledge structure and the presented relevance of new information was perceived to be useful, even with a basic model of text processing. The participants were able to search for relevant information effectively and expand their information collection with consistency.
We demonstrate X-Rec, a novel system for entity recommendation. In contrast to other systems, X-Rec can recommend entities from diverse categories including goods (e.g., books), other physical entities (e.g., actors), but also immaterial entities (e.g., ideologies). Further, it does so only based on publicly available data sources, including the revision history of Wikipedia, using an easily extensible approach for recommending entities. We describe X-Rec's architecture, showing how its components interact with each other. Moreover, we outline our demonstration, which foresees different modes for users to interact with the system.
Cultural heritage materials are increasingly being made available through standard search facilities. However, it is challenging to automatically organize these materials in a way that is well aligned with users' specific interests. We report on the development of a social bookmaking system to collect human annotations that are used to measure the performance of three different clustering algorithms. We find that there is a discrepancy between the latent structure present in the data and the clusters annotated by humans. However, it is difficult to detect such discrepancies explicitly.
The Internet-enabled smartphones are readily enabling ubiquitous and continuous access to information. Recent reports showed that Hispanics are more likely to own smartphones and use the mobile Internet than other racial groups in the U.S.A. However, little is known about the mobile access and use of smartphones in seeking health information for this group. This study conducted semi-structured interviews with 20 low SES (socioeconomic status) Hispanics in the U.S.A. Mobile context and situations prompting the adoption of smartphones for health information seeking were explored. The results shed light on how smartphones could help the underserved Hispanics search for health information, narrowing a gap in health disparity. Furthermore, this exploratory study contributes to a more in-depth understanding of mobile context and situations in mobile health information seeking behavior.
Searching for online health information has been well studied in web search, but social media, such as public microblogging services, are well known for different types of tacit information: personal experience and shared information. Finding useful information in public microblogging platforms is an on-going hard problem and so to begin to develop a better model of what health information can be found, Twitter posts using the word "depression" were examined as a case study of a search for a prevalent mental health issue. 13,279 public tweets were analysed using a mixed methods approach and compared to a general sample of tweets. First, a linguistic analysis suggested that tweets mentioning depression were typically anxious but not angry, and were less likely to be in the first person, indicating that most were not from individuals discussing their own depression. Second, to understand what types of tweets can be found, an inductive thematic analysis revealed three major themes: 1) disseminating information or link of information, 2) self-disclosing, and 3) the sharing of overall opinion; each had significantly different linguistic patterns. We conclude with a discussion of how different types of posts about mental health may be retrieved from public social media like Twitter.
We present a system that supports Interactive Information Retrieval user studies on the Web. Our system provides support for user and task management, for processing web-based task specific interfaces and for Web-event logging. It also offers functionality useful to IIR studies that capture eye-movement on Web page elements. The system complements logging functionality offered by a typical usability/eye-tracking software packages and is designed to act in concert with such software.
The ever expanding digital information universe makes us rely on search systems to sift through immense amounts of data to satisfy our information needs. Our searches using these systems range from simple lookups to complex and multifaceted explorations. A multitude of models of the information seeking process, for example Kuhlthau's ISP model, divide the information seeking process for complex search tasks into multiple stages. Current search systems, in contrast, still predominantly use a "one-size-fits-all" approach: one interface is used for all stages of a search, even for complex search endeavors. The main aim of this paper is to bridge the gap between multistage information seeking models, documenting the search process on a general level, and search systems and interfaces, serving as the concrete tools to perform searches. To find ways to reduce the gap, we look at existing models of the information seeking process, at search interfaces supporting complex search tasks, and at the use of interface features over time. Our main contribution is that we conceptually bring together macro level information seeking stages and micro level search system features. We highlight the impact of search stages on the flow of interaction with user interface features, providing new handles for the design of multistage search systems.
The wealth of digital information available in our time has become indispensable for a rich variety of tasks. We use data on the Web for work, leisure, and research, aided by various search systems, allowing us to find small needles in giant haystacks. Despite recent advances in personalization and contextualization, however, various types of tasks, ranging from simple lookup tasks to complex, exploratory and analytical ventures, are mainly supported in elementary, "one-size-fits-all" search interfaces. Web archives, keepers of our future cultural heritage, have gathered petabytes of valuable Web data, which characterize our times for future generations. Access to these archives, however, is surprisingly limited: online Web archives usually provide a URL-based Wayback Machine interface, sometimes extended with rudimentary search options. As a result of limited access, Web archives have not been widely used for research so far. For emerging research using Web archives, there is a need to move beyond URL-based and simple search access, towards providing support for complex (re)search tasks. In my thesis, I am exploring ways to move beyond the "one-size-fits-all" approach for search systems, and I work on systems which can support the flow of complex search, also in the context of archived Web data. Rich models of search and research can be incorporated into adaptive search systems, supporting search strategies in various stages of complex search tasks. Concretely, I look at the use case of the Humanities researcher, for which the large, Terabyte-scale Web archives can be a valuable addition to existing sources utilized to perform research.
Data visualization and exploration tools are crucial for data scientists, especially during a pilot study. In this paper, we present an extensible open-source workbench for aggregating, summarizing and filtering social network profiles derived from tweets. We briefly demonstrate its range of basic features for two use cases: geo-spatial profile summarization based on check-in histories and social media based complaint discovery in water management.
In this paper, we explore the relationship between time constraints and users' assessment of their search. A user experiment was conducted. Participants were asked to search under two conditions: with time constraint (TC) and with no time constraint (NTC). The results showed that time constraint did not significantly influence participants' assessment of task difficulty, but significantly influenced users' search confidence and their evaluation of search performance. Particularly, participants were less confident and considered their search performance worse in TC than in NTC conditions. We also found users acquired more new knowledge and had more positive affective states after searching in NTC than in TC conditions. Interestingly, we found time constraints also affect participants' anticipation of time needed to complete the task; participants thought they would need significantly less time to complete the search task when they were given time constraints than without time constraints. These preliminary results suggested that time constraints had remarkable influence upon users' perception of search tasks and their search experience.
In this paper I would like to outline the theoretical concept, methods and research questions of my dissertation project. Following the theory of Information Use Environment by Taylor (1991) and its interpretation in the context of Web search by Detlor (2003), this investigation looks into the information behavior of teachers when searching the Web for instruction-related information. Typical problem situations of the work context are identified as well as relevant information traits and approaches to information search and use. Furthermore, the use of several search environments with different degrees of interactive character are analyzed comparatively. Taking into consideration a variety of data sources to gain a multi-perspective view on the issue, usage data of the German Education Server (GES) are analyzed in addition to discussion fora on teacher-specific websites and complemented by qualitative interviews.
This paper reports on a user-centered analysis of video digital libraries. Video digital libraries enable "in the loop" retrieval and playback from centralized and organized collections. As a time-based and multi-channeled format, video digital library systems warrant different considerations for design and information delivery. The purpose of the present study is to collect and initially analyze users' criteria of video digital libraries as part of their interactive experiences. Fifty-two journalism and political science college majors were surveyed, resulting in a total of 242 individual collected responses. Content analysis was performed on the survey responses, and the emergent coding method produced 28 criteria (subcategories) under 5 major categories. Criteria corresponding to Retrieval functions of video digital libraries emerged as the highest priority of the participants, based on its frequency across the responses for the major categories. Criteria corresponding to the User Interface, Collection Quality, User Support, and Organization of Collection followed respectively, in terms of frequency. Cohen's Kappa was .87, indicative of high-level of inter-coder reliability. Findings of the present study provide an initial baseline for design and evaluation of video digital libraries and motivate further research.
This research addresses the need for faceted search systems that can support task-based searching. We report on a systems review carried out to identify the most prevalent facets in current use across three domains and an online questionnaire with 83 responses conducted to assess the perceived usefulness of search facets for different types of search tasks. Results include a ranked list of commonly used search facets. Facets are perceived to be more useful for search tasks motivated by learning goals than those with functional goals (doing tasks). Usefulness scores for specific facets were quite consistent across tasks, so findings do not support the concept of dynamic, task-based display of search facets.
While social networking and microblogging platforms have received considerable research attention, little work has been done to understand how users preserve, manage and reaccess content they acquire from these sources. In this paper we present initial analyses of a large-scale survey (n=606) to understand Personal Information Management (PIM) practices with social networking and microblogging systems, and Twitter in particular. Our results indicate that re-finding information in tweets is a common Twitter activity and can be frustrating. Using questionnaire responses, we investigate the influence of several factors, including how the user tends to preserve tweets of interest, and the level of frustration involved in re-finding such tweets when they are required later.
Prior text mining studies have documented a causal link between human emotions and stock market patterns, yet relatively little research exists into what triggers these emotions. This paper aims to bridge the gap by empirically testing a social psychology theory of human behavior. Underlying our approach lies Attribution Theory, which addresses how observers form causal inferences and moral judgments to explain human behavior, particularly those with negative outcomes. The system presented here works in three stages. The first phase computes a measure of media pessimism by counting negative terms from the General Inquirer dictionary to detect acts of corporate irresponsible behavior. The second phase extends the term-counting approach to capture contextual information. Emotion topic priors are incorporated in a Latent Dirichlet Allocation (LDA) model to infer the financial media's expression of negative affect. Finally, the system combines the two components in an ensemble tree to classify the impact of financial media allegations on a company's stock market patterns. The paper underlines the potential benefit of text mining technology for the support of investor strategies, and more generally demonstrates the power of combining multiple methods for applications in specific domains.
The usability of web search engines is an important factor that influences user experience and correlates with users' success in finding the relevant information. Currently, there are different search engines online available whose main target audience are children. In this paper, we investigate the differences between children and adults in terms of usability and perception of targeted search engines, i.e. search engines designed specifically for that audience. To this end, an eye-tracking study was conducted to compare children's and adults' search behavior and perception of search interface elements on search engine results pages (SERPs) during an informational and a navigational search with a standard search engine and a search engine for children. We identified differences in the information-seeking behavior and perception of search engines SERPs between children and adults. Based on these findings we propose criteria on how to design search user interfaces that are more appropriate for children.
Children's search behavior has mainly been studied in single search sessions where the search task is new to the children. In this proposal it is suggested that an investigation into children's search on a single topic across multiple search sessions may reveal different search behaviour to what is already known.