With the shift from traditional to digital media, the online landscape now hosts not only reliable news articles but also a significant amount of unreliable content. Digital media has faster reachability by significantly influencing public opinion and advancing political agendas. While newspaper readers may be familiar with their preferred outlets' political leanings or credibility, determining unreliable news articles is much more challenging. The credibility of many online sources is often opaque, with AI-generated content being easily disseminated at minimal cost. Unreliable news articles, particularly those that followed the Russian invasion of Ukraine in 2022, closely mimic the topics and writing styles of credible sources, making them difficult to distinguish. To address this, we introduce SemCAFE (Semantically enriched Content Assessment for Fake news Exposure), a system designed to detect news reliability by incorporating entity-relatedness into its assessment. SemCAFE employs standard Natural Language Processing (NLP) techniques, such as boilerplate removal and tokenization, alongside entity-level semantic analysis using the YAGO knowledge base. By creating a "semantic fingerprint" for each news article, SemCAFE could assess the credibility of 46,020 reliable and 3,407 unreliable articles on the 2022 Russian invasion of Ukraine. Our approach improved the macro F1 score by 12% over state-of-the-art methods. The sample data and code are available on GitHub(1).
The TempWeb workshop (series) is an established co-located event at The Web Conference that aims at bringing together researchers and practitioners across various domains, taking the constantly evolving Web as a primary object of research. In this context, research questions aim at the investigation of infrastructures, scalable methods, and innovative software for aggregating, querying, and analyzing heterogeneous data at Web scale. To this end, submissions commonly address core fields of studies in computer science and aspects such as temporal IE/IR, Web mining, Web archiving and large scale data analysis to just name a few. However, TempWeb does not only address computer science related research, but also attracts not necessarily technical innovations from application domains such as the social sciences, marketing, economics, etc. As a consequence, TempWeb has developed as a forum for a community from science and industry across various disciplines. For its 2024 edition, topics of TempWeb covered again a wide spectrum of Web related research ranging from studies of control mechanisms, via terminological studies, up to community analysis and content recommendation. Date: 14 May 2024. Website: http://temporalweb.net/.
The Web Science conference series (WebSci) is an established venue for Web related research across disciplines. The co-located PhD symposium is an endeavor in fostering interdisciplinary thinking and collaboration at an early stage. Bringing PhD students from various disciplines together is, according to our understanding, an important step towards a more comprehensive investigation of the Web. As such, the PhD symposium is intended to establish a forum for the next generation of Web Scientists. Following the spirit of "unity in diversity" the many facets of Web related research ranging from computer science, the social sciences, via marketing and economics up to application domains such as, e.g., geography are addressed in this unique event. By doing so, a joint understanding beyond the limits of one's own discipline can emerge that might allow a more interdisciplinary and inclusive future of the Web. Date: 21 May 2024. Website: https://websci24.webscience.org/program/phd-symposium/.
The Ethical Web Science Workshop at WebSci’25 addresses the ethical challenges that arise from the rapidly evolving relationship between the Web and AI. With increasing reliance on Web-sourced data, issues of fairness, transparency, and consent have become central to technological innovation. To this end, we aim at bringing together researchers, practitioners, ethicists, and policymakers in order to create tools and guidelines for ethical compliance in Web-based research. By doing so, standards for ethical data sourcing and usage might be discussed and lead to guidelines for best practices.
The Web Science conference series (WebSci) has become a prime venue for Web related research across disciplines. The interdisciplinary PhD symposium is an important component of this endeavor. Bringing PhD students from various disciplines together is, according to our understanding, an important step towards a comprehensive investigation of the Web. As such, we are trying to establish a forum for the next generation of Web Scientists. Following the spirit of ’unity in diversity’ we hope to bring together the many facets of Web related research covering computer science, the social sciences, marketing, economics, and other fields. We believe that a joint understanding beyond the limits of one’s own discipline will allow the Web to develop into a more interdisciplinary and inclusive future.
The TempWeb workshop series has a long-standing tradition as a co-located event at The Web Conference. Its main objectives are to provide a venue for researchers of all domains (IE/IR, Web mining, etc.) where the temporal dimension opens an entirely new range of challenges and possibilities. For this, research published at TempWeb covers (and has covered in its previous editions) the full spectrum of longitudinal studies using web content. This ranges from “low-level” network analysis (e.g., Internet traffic) and graph analysis (e.g., in social media) up to “high-level” content and concept analysis ranging from micro-content (e.g., tweets) up to entire web archives. As such, TempWeb has become a potential object of longitudinal studies on its own. Aiming at the investigation of infrastructures, scalable methods, and innovative software for aggregating, querying, and analyzing heterogeneous data at Web scale, TempWeb has developed as a forum for a community from science and industry. Since longitudinal aspects in web content analysis is becoming more and more relevant for analysts from various domains, the studies, tools, and demonstrations are not only relevant for computer science, but also potentially interesting for sociology, marketing, environmental studies, politics, etc. Keeping up with its “tradition”, the current edition covers again the full spectrum of temporal Web analytics at various levels of granularity.
The TempWeb workshop series has a long-standing tradition as a co-located event at The Web Conference. Its main objectives are to provide a venue for researchers of all domains (IE/IR, Web mining, etc.) where the temporal dimension opens an entirely new range of challenges and possibilities. For this, research published at TempWeb covers (and has covered in its previous editions) the full spectrum of longitudinal studies using web content. This ranges from "low-level" network analysis (e.g., Internet traffic) and graph analysis (e.g., in social media) up to "high-level" content and concept analysis ranging from micro-content (e.g., tweets) up to entire web archives. As such, TempWeb has become a potential object of longitudinal studies on its own. Aiming at the investigation of infrastructures, scalable methods, and innovative software for aggregating, querying, and analyzing heterogeneous data at Web scale, TempWeb has developed as a forum for a community from science and industry. Since longitudinal aspects in web content analysis is becoming more and more relevant for analysts from various domains, the studies, tools, and demonstrations are not only relevant for computer science, but also potentially interesting for sociology, marketing, environmental studies, politics, etc. Keeping up with its "tradition", the current edition covers again the full spectrum of temporal Web analytics at various levels of granularity.
TempWeb is an established Workshop (series) with a long-standing tradition as a co-located event at The Web Conference. Considering the constantly evolving Web as a primary object of research, TempWeb provides a platform to present longitudinal studies on the Web (structure), its contents, and its communities, etc. by aiming at the investigation of infrastructures, scalable methods, and innovative software for aggregating, querying, and analyzing heterogeneous data at Web scale. Since longitudinal aspects in Web analytics is becoming more and more inter- and transdiciplinary relevant for analysts from various domains, the studies, tools, and demonstrations are not only limited to computer science, but also open to sociology, marketing, environmental studies, politics, etc., to name just a few. As such, TempWeb has developed as a forum for a community from academia and industry covering temporal Web analytics at all levels of granularity and across various disciplines. Date : 1 May 2023. Website : http://temporalweb.net/.
User interest tracing is a common practice in many Web use-cases including, but not limited to, search, recommendation or intelligent assistants. The overall aim is to provide the user a personalized “Web experience” by aggregating and exploiting a plenitude of user data derived from collected logs, accessed contents, and/or mined community context. As such, fairly basic features such as terms and graph structures can be utilized in order to model a user’s interest. While there are clearly positive aspects in the before mentioned application scenarios, the user’s privacy is highly at risk. In order to highlight inherent privacy risks, this paper studies Semantic User Interest Tracing (SUIT in short) by investigating a user’s publishing/editing behavior of Web contents. In contrast to existing approaches, SUIT solely exploits the (semantic) concepts [categories] inherent in documents derived via entity-level analytics. By doing so, we raise Web contents to the entity-level. Thus, we are able to abstract the user interest from plain text strings to “things”. In particular, we utilize the inherited structural relationships present among the concepts derived from a knowledge graph in order to identify the user associated with a specific Web content. Our extensive experiments on Wikipedia show that our approach outperforms state of the art approaches in tracing and predicting user behavior in a single language. In addition, we also demonstrate the viability of our semantic (language-agnostic) approach in multi-lingual experiments. As such, SUIT is capable of revealing the user’s identity, which demonstrates the fine line between personalization and surveillance, raising questions regarding ethical considerations at the same time.
TempWeb focuses on investigating infrastructures, scalable methods, and innovative software for aggregating, querying, and analyzing heterogeneous data at Web scale. Emphasis is given to data analysis along the time dimension for web data that has been collected over extended time periods. A major challenge in this regard is the sheer size of the data it exposes and the ability to make sense of it in a useful and meaningful manner for its users. It is worth noting that this trend of using big data to make inferences is not specific to Web content analytics, so work presented here might be useful in other areas, too. As such, longitudinal aspects in Web content analysis become relevant for analysts from from various domains, including, but not limited to sociology, marketing, environmental studies, politics, etc. Studies in this context range from "low-level" structural network log analysis over time, up to "high-level" entity-levelWeb content analytics and terminology evolution. While both before mentioned aspects represent the extremes of the spectrum, they have one thing in common: Web scale data analytics needs to develop infrastructures and extended analytical tools in order to make use of that data. TempWeb has been created for this purpose and is the ideal venue to discuss about all its facets.
TempWeb focuses on investigating infrastructures, scalable methods, and innovative software for aggregating, querying, and analyzing heterogeneous data at Web scale. Emphasis is given to data analysis along the time dimension for web data that has been collected over extended time periods. A major challenge in this regard is the sheer size of the data it exposes and the ability to make sense of it in a useful and meaningful manner for its users. It is worth noting that this trend of using big data to make inferences is not specific to Web content analytics, so work presented here might be useful in other areas, too. As such, longitudinal aspects in Web content analysis become relevant for analysts from from various domains, including, but not limited to sociology, marketing, environmental studies, politics, etc. Studies in this context range from “low-level” structural network log analysis over time, up to “high-level” entity-level Web content analytics and terminology evolution. While both before mentioned aspects represent the extremes of the spectrum, they have one thing in common: Web scale data analytics needs to develop infrastructures and extended analytical tools in order to make use of that data. TempWeb has been created for this purpose and is the ideal venue to discuss about all its facets.
TempWeb focuses on investigating infrastructures, scalable methods, and innovative software for aggregating, querying, and analyzing heterogeneous data at Web scale. Emphasis is given to data analysis along the time dimension for web data that has been collected over extended time periods. A major challenge in this regard is the sheer size of the data it exposes and the ability to make sense of it in a useful and meaningful manner for its users. It is worth noting that this trend of using big data to make inferences is not specific to Web content analytics, so work presented here might be useful in other areas, too. TempWeb has been created for this purpose and is the ideal venue to discuss all its facets. Date: 25 May, 2022. Website: http://temporalweb.net/.
Efforts by national libraries, institutions, and (inter-) national projects have led to an increased effort in preserving textual contents - including non-digitally born data - for future generations. These activities have resulted in novel initiatives in preserving the cultural heritage by digitization. However, a systematic approach toward Textual Data Denoising (TD 2 ) is still in its infancy and commonly limited to a primarily dominant language (mostly English). However, digital preservation requires a universal approach. To this end, we introduce a “Framework for Enabling Textual Data Denoising via robust contextual embeddings” (FETD 2 ). FETD 2 improves data quality by training language-specific data denoising models based on a small number of language-specific training data. Our approach employs a bi-directional language modeling in order to produce noise-resilient deep contextualized embeddings. In experiments we show the superiority compared with the state-of-the-art.
Gerhard Weikum合作论文数Department of Databases and Information Systems, Max-Planck Institute for Informatics30
Rynson W. H. Lau (劉永雄)合作论文数Department of Computer Science, College of Engineering, City University of Hong Kong;Swansea University2