ABSTRACTThe purpose of this article is to capture pregnancy-related information seeking behavior of women and its correlation with socio-economic status and with Internet health information literacy profiles. A user survey based on Wilson's' macro-model of information seeking behavior was designed and carried out. The survey included 289 women at reproductive age who provided valid feedback on information needs, motives for seeking pregnancy-related information online, information resources employed and Internet information literacy. Responses were collected via an anonymous questionnaire and processed using SPSS (Statistical Package for Social Sciences) version 25.0. The results demonstrate that women have increased information needs during pregnancy, that they are primarily concerned about the health of their embryos and that they tend to turn mainly to their doctors, their family and to the Web for obtaining pregnancy-related information. Internet is most used by younger women and online information is generally deemed useful. The analysis of survey responses indicated that pregnancy-related information literacy profiles are influenced by women's educational status and age. Official online information sources should be made available in Greece, where women could find useful and trusted information about pregnancy issues.KEYWORDS: Barriers/obstacles when seeking informationGreeceinformation needsinformation sourcesInternet information literacypregnancysurvey Disclosure statementNo potential conflict of interest was reported by the author(s).
The purpose of this article is to investigate the information-seeking behavior of hospitalized patients in a peripheral hospital in Greece, based on Wilson’s theoretical model. After a thorough literature review, we report on the results of a cross-sectoral sur\vey that was conducted based on Wilson’s macro-model of information-seeking behavior. The survey exploits the feedback we obtained from 150 hospitalized patients in the General Hospital of Corfu island in Greece. For the feedback collection, we relied on a structured questionnaire that was distributed to patients and correlates information-seeking behavior of hospitalized patients’ clinical data and eHealth Literacy. According to our findings, the most commonly recorded factors that pertain to information needs of hospitalized patients relate to treatment information, procedural information as well as information about the patients’ rights and support. The main information sources to which patients turn for receiving information are digital scientific sources, health professionals, other patients, insurance funds, family members and friends and the media. Concerning the most frequent obstacles associated with the information-seeking process, these account primarily to individual, interpersonal and environmental factors. Finally, the research came to a conclusion with the adaptation of Wilson’s theoretical model to hospitalized patients by confirming the original-initial model. The findings of our survey reveal the necessity to equip hospital units with adequate staff and to develop programs that offer updated information services to the hospitalized population. Moreover, our survey suggests that the availability of advanced health-related information services can empower the general population toward establishing e-Health literacy skills as well as toward increasing their level of satisfaction from the quality of the health services offered. In this direction, a clinical librarian program would empower hospitals worldwide to make their patients literate about health-related information.
Web data is constantly increasing at a very high pace. So does the need to come up with methods and tools that are able to process, organize and store this data effectively. To meet this need, several approaches have been proposed in the literature over the last decades, a critical amount of which focus on methods for classifying Web content in order to be able to retrieve relevant information in a cost-effective yet effortless manner. Motivated by the observation that the Web is changing not only with respect to content but also with respect to structure, we designed a combined classification method that encounters both textual and structural elements in the Web pages under examination. Our classification approach, presented here, investigates a number of parameters before assigning a Web page to a suitable category(-ies). A preliminary experimental evaluation of our method indicates that it accurately classifies web content both thematically and structurally.
In this paper, we experimentally study the degree to which the length of a short text affects its comprehensiveness and readability, within quantitative linguistics. The quantitative linguistics focus mainly in analysis of large text collections and one of the major scientific theories in use is the Menzerath-Altmann law. In this paper we attempt to define the quantitative analysis framework for short texts consisting approximately of one or two sentences, due to the fact that they are considered very important in many scientific fields. To achieve the aim of this paper, a coherence statistical testing process of three variables was created for short texts. The implementation of that was possible through experimental and statistical evaluation. Upon completion of the above-mentioned evaluation, the statistical results showed that short text coherence, comprehensiveness and readability are fully achieved in short texts consisting of 14 words, when three predetermined variables are associated and vice versa. To prove the above hypothesis the theory of Vector Space Model and Kendall's Coefficient of Concordance were used. The assessment of statistical results concluded that the above hypothesis can be fully met for a number of cases with a probability p > 99%. Moreover, in the experiment were used short texts in English language but it was proven that language can be considered irrelevant. To corroborate this, a smaller scale experiment with short texts in the German language was conducted and hypothesis was confirmed that the proposed model of this paper can be applied in all short texts regardless of their linguistic origin.
It is well-known that the web contains many duplicate and near-duplicate documents. Despite the efforts that have been put towards equipping search engines with duplicate detection algorithms, still there are cases where the documents retrieved in response to web queries contain redundant information. In this paper, we are concerned with effectively identifying and reducing redundant information in search results. In particular, we describe how we automatically detect content that is lexically and/or semantically duplicated across search results and we introduce a novel algorithm that upon the detection of significant (i.e., above a given threshold) content duplication, it filters out redundant information. Information filtering takes place in two-steps depending on whether we are dealing with documents of (nearly) identical lexical content or with documents of lexically distinct but semantically equivalent content. In the first case, our algorithm retains in the result list the document that is the most relevant to the query intention and removes duplicates. In the second case, our algorithm merges into a single text, which we call SuperText, the documents of redundant information in a way that every document contributes diverse semantic content to the generated SuperText. Additionally, the algorithm re-ranks the remaining documents based on their contextual relevance to the query intention. The experimental evaluation of our approach demonstrates that it is very effective in identifying lexical and semantic information redundancy across search results. In addition, we have found that our algorithm manages to filter out successfully content duplication from the results list and the SuperTexts it generates for reducing information redundancy are syntactically and semantically coherent texts.
In this chapter we propose the exploration of text and data mining techniques for empowering e-government applications and services for the citizen’s benefit. In particular, we start by providing a field overview with respect to the current trends in e-government services and we demonstrate via proofs of concept the limited adaptation existing e-government applications entail. Stimulated by the need to transform e-government services to e-inclusion applications, we suggest the utilization of data mining techniques for processing the governmental data so as to extract and associate information fragments with real citizen needs and thus enable the encapsulation of the latter in future governmental decisions. To demonstrate the usability and added value of our proposed approach we have designed an interactive e-government infrastructure, the architecture of which we will present and discuss in our chapter. Moreover, we will elaborate on the system details, its adaptation capacity and we will discuss its usage benefits for both citizens and public sector bodies.
Wikipedia is one of the most successful worldwide collaborative efforts to put together user-generated content in a meaningfully organized and intuitive manner. Currently, Wikipedia hosts millions of articles on a variety of topics, supplied by thousands of contributors. A critical factor in Wikipedia's success is its open nature, which enables everyone to edit, revise, and/or question (via talk pages) the article contents. Considering the phenomenal growth of Wikipedia and the lack of a peer review process for its contents, it becomes evident that both editors and administrators have difficulty in validating its quality on a systematic and coordinated basis. This difficulty has motivated several research works on how to assess the quality of Wikipedia articles. In this chapter, the authors propose the exploitation of a novel indicator for the Wikipedia articles' quality, namely information credibility. In this respect, the authors describe a method that captures the polarized (i.e., biased) information across the article contents in an attempt to infer the amount of credible (i.e., objective) information every article communicates. This approach relies on the intuition that an article offering non-polarized information about its topic is more credible and of better quality compared to an article that discusses the editors' (subjective) opinions on that topic.
In many situations, searching the web is synonymous to information seeking. Currently, web search engines are the most popular vehicle via which people get access to the web. Their popularity is partially due to the intrinsic way that people interact with them, i.e., by typing some keywords to the corresponding input box. Despite their popularity, search engines often fail to satisfy certain information needs, especially when the latter are haze and poorly articulated. In this paper, we focus on the occasions when large-scale web search engines find it difficult to cope with specific information-seeking behaviors and we accordingly introduce a query construction service that is targeted towards the solution of this problem. The proposed service leverages information coming from various DBpedia datasets and provides an intuitive GUI via which searchers determine the semantic orientation of their queries before these are addressed to the underlying search engine. The evaluation of the query construction service justifies the motive of this paper and indicates that it can considerably improve the searchers’ querying ability when search engines fail to provide adequate help.
In this chapter, the authors propose a novel framework for the support of multi-faceted searches over distributed Web-accessible databases. Towards this goal, the authors introduce a method for analyzing and processing a sample of the database contents in order to deduce the topical, the geographic, and the temporal orientation of the entire database contents. To extract the database topics, the authors apply techniques leveraged from the NLP community. To identify the database geographic footprints, the authors first rely on geographic ontologies in order to extract toponyms from the database content samples and then employ geo-spatial similarity metrics to estimate the geographic coverage of the identified toponyms. Finally, to determine the time aspects associated with the database entities, the authors extract temporal expressions from the entities’ contextual elements and utilize a time ontology against which the temporal similarity between the identified entities is estimated.
In this article, we propose new word sense disambiguation strategies for resolving the senses of polysemous query terms issued to Web search engines, and we explore the application of those strategies when used in a query expansion framework. The novelty of our approach lies in the exploitation of the Web page PageRank values as indicators of the significance the different senses of a term carry when employed in search queries. We also aim at scalable query sense resolution techniques that can be applied without loss of efficiency to large data sets such as those on the Web. Our experimental findings validate that the proposed techniques perform more accurately than do the traditional disambiguation strategies and improve the quality of the search results, when involved in query expansion.
In this paper, we report on a preliminary study we carried out for identifying patterns that characterize the genre type of Greek texts. In the course of our study, we address four distinct genre types, we record their observable stylistic elements and we indicate their exploitation for automatic genre-based document classi-fication. The findings of our study demonstrate that texts contain lexical features with discriminative power as far as genre is concerned, however modeling those features so that they can be explored by computer-based applications is still in early stages.
In this paper, we report on a preliminary study we carried out for identifying patterns that characterize the genre type of Greek texts. In the course of our study, we address four distinct genre types, we record their observable stylistic elements and we indicate their exploitation for automatic genre-based document classification. The findings of our study demonstrate that texts contain lexical features with discriminative power as far as genre is concerned, however modeling those features so that they can be explored by computer-based applications is still in early stages.
To effectively manage the proliferating online content, it is imperative that we come up with efficient data structuring and organization methods. Based on the findings of previous research (6) (7) that the most flexible and useful way to organize the online content is via the use of taxonomies and/or ontologies, we carried out the present study, which aims at structuring the content of the Greek Wikipedia via the use of the Greek WordNet. In particular, our study objective is to design a model that can automatically organize the Greek Wikipedia categories into a thematic taxonomy and based on the derived organization, to implicitly assign hierarchical structure to the encyclopedia articles that have been classified to the respective categories. To this end, we relied on the data encoded in Greek WordNet out of which we harvested the hierarchical relations that hold between the terms used to verbalize the Wikipedia categories. The effectiveness of our model is verified by the findings of several experimental evaluations conducted, which demonstrate that semantic networks are powerful resources for hierarchically organizing large volumes of dynamic data.
We present a novel approach in machine learning by combining naı̈ve Bayes classifiers with tree kernels. Tree kernel methods produce promising results in machine learning tasks containing treestructured attribute values. These kernel methods are used to compare two tree-structured attribute values recursively. Up to now tree kernels are only used in kernel machines like Support Vector Machines or Perceptrons. In this paper, we show that tree kernels can be utilized in a naı̈ve Bayes classifier enabling the classifier to handle tree-structured values. We evaluate our approach on three datasets containing tree-structured values. We show that our approach using tree-structures delivers significantly better results in contrast to approaches using non-structured (flat) features extracted from the tree. Additionally, we show that our approach is significantly faster than comparable kernel machines in several settings which makes it more useful in resource-aware settings like mobile devices. Naı̈ve Bayes Classifier; Tree Kernel; Lazy Learning; Tree-structured Values
Web searches are driven by information needs and intend the accomplishment of specific tasks. Information needs are determined by the topical subject of queries, i.e. what we search, while tasks are determined by the user motives that induce the submission of queries, i.e. why we search. Though there exist numerous studies on how to assist searchers specify queries that are expressive of their underlying information needs, little has been done to help searchers specify queries that describe the tasks they pursue via their searches. In this paper we propose a query reformulation method to empower task-oriented web searches. Given a query, our method starts with the identification of terms that could serve as descriptors of the potential search tasks the query represents. Based on the identified terms, it generates query re-formulations that explicitly verbalize the possible search tasks. Query reformulations are presented to the user in order to select the one that best suits her search intention.
In this paper we propose a method that identifies and extracts keywords within URLs, focusing on the Greek Web and especially on URLs containing Greek terms. Although there are previous works on how to process Greek online content, none of them focuses on keyword identification within URLs of the Greek web domain. In addition, there are many known techniques for web page categorization based on URLs but, none addresses the case of URLs containing transliterated Greek terms. The proposed method integrates two components; a URL tokenizer that segments URL tokens into meaningful words and a Latin–to–Greek script transliteration engine that relies on a dictionary and a set of orthographic and syntactic rules for converting Latin verbalized word tokens into Greek terms. The experimental evaluation of our method against a sample of 1,000 Greek URLs reveals that it can be fruitfully exploited towards automatic keyword identification within Greek URLs
In the last decades the explosion of ICT has opened up new avenues regarding peoples' accessibility to new job opportunities. Current technological advances in conjunction with people's online presence provide a great opportunity to automate the recruitment process and make it more effective. In this paper, we propose a novel approach for improving the efficiency of e-recruitment systems. Our approach relies on the linguistic analysis of data available for job applicants, in order to infer the applicants' personality traits and rank them accordingly. To showcase the functionality of our method, we employed it in a web based e-recruitment system that we implemented.
: Wikipedia is a unique source of information that has been collectively supplied by thousands of people. Since its nascence in 2001, Wikipedia is continuously evolving and like most websites it is interconnected via hyperlinks to other web information sources. Wikipedia articles contain two types of links: internal and external. Internal links point to other Wikipedia articles, while external links point outside Wikipedia and normally they are not used in the body of the article. Although there exist specific guidelines about both the style and the purpose of the article external links, no approach has been recorded that tries to capture in a systematic manner the quality of Wikipedia external links. In this paper, we study the quality of Wikipedia external links by assessing the degree to which these conform to their intended purpose; that is to formulate a comprehensive list of accurate information sources about the article contents. For our study, we estimate the decay of Wikipedia external links and we investigate their distribution in the Wikipedia articles. Our measurements give perceptible evidence for the value of external links and may imply their corresponding articles' quality in a holistic Wikipedia evaluation.
eGovernment refers to the use of information and communications technologies (ICTs) to improve the quality of services and information offered to citizens, to make government more accountable to and ad- vance public sector transparency. As already pointed out by other researchers, one of the most important issues for making eGovernment effective is to enable to participate in the decision-making process. Nowadays, topics related to governmental decisions are among the most widely discussed ones within digital societies. This is not only because web 2.0 has empowered people with the ability to communicate remotely but also because governments all around the globe publish a great volume of their decisions and regulations online. In this paper, we propose the exploration of text and data mining techniques towards capturing the publics opinion communi- cated online and concerning governmental decisions. The objective of our study is twofold and focuses on under- standing the citizen opinions about eGovernment issues and on the exploitation of these opinions in subsequent governmental actions. We examine several features in the user-generated content discussing governmental de- cisions in an attempt to automatically extract the citizen opinions from online posts dealing with public sector regulations and thereafter be able to organize the extracted opinions into polarized clusters. Our goal is to be able to automatically identify the publics stance against governmental decisions and thus be able to infer how the citizens viewpoints may affect subsequent government actions. To demonstrate the usability and added value of our proposed approach we have designed an interactive eGovernment infrastructure, the architecture of which we will present and discuss in our paper. Moreover, we will elaborate on the system details, its adaptation capac- ity and we will discuss its usage benefits for both and public sector bodies. In this paper, we try to fill this void by proposing a novel eGovernment mechanism that captures the societal impact of public sector regulations in an attempt to decipher the publics stance towards gov- ernmental decisions. In particular, we propose the exploitation of data mining techniques towards firstly capturing the publics opinions (communicated online) about governmental decisions and sec- ondly analysing the polarity of the mined opinions so that they are considered in subsequent govern- mental decisions. Specifically, we introduce a method for decomposing citizens opinions and com- ments that are posted in online fora and blogs, in order to evaluate how governmental decisions are perceived by the public and thereafter how the publics implicit feedback should be interpreted by governmental bodies in their subsequent actions. What motivates our study is that up-to-date gov-
Iraklis Varlamis合作论文数Department of Informatics and Telematics, Harokopio University of Athens2
Andreas Nürnberger合作论文数Department for Technical & Operational Information Systems, Faculty of Computer Science, Otto-Von-Guericke-University Magdeburg1
Michalis Vazirgiannis合作论文数Computer Science Laboratory, Ecole Polytechnique;Mohamed bin Zayed University of Artificial Intelligence1
Lars Schmidt-Thieme合作论文数Institute of Computer Science, Department of Mathematics, Natural Science, Economics and Computer Science, University of Hildesheim1
Steffen Oeltze合作论文数Department of Simulation and Graphics
Faculty of Computer Science
University of Magdeburg1