
Ever since the Internet Archive started large-scale web archiving in 1996, historians, sociologists, and journalists have found web archives to be an important source of information for their work. Archive-It, a service focused on creating collections, allows curators to generate their own web archive collections. Many of these collections are vast, consisting of thousands of documents. This makes collection understanding a difficult, if not impossible, task. We seek to improve collection understanding by summarizing these collections and visualizing the summaries. Focusing on Archive-It, we seek to identify the different types of web archive collections, the algorithms that can be used summarize those collections, and the best visualizations of those summaries to support better collection understanding. ACM Reference Format: ShawnM. Jones. 2018. Improving Collection Understanding inWebArchives: Extended Abstract for the JCDL 2018 PhD Consortium. In Proceedings of ACM-IEEE Joint Conference on Digital Libraries Doctoral Consortium (JCDL’18). ACM, New York, NY, USA, 9 pages.
This paper provides an overview of my PhD thesis which in a first phase, focuses on developing an ontology which exposes the various text types, both literary and non-literary in Romanian language and in a second phase (the most complex one), on formalizing a temporal annotation scheme that can be used in order to reorder the events or more precisely, the sequences of events in novels and also, the switches that may occur in literary texts such as flashbacks, flashforwards, embedded fabulae, temporal ruptures, and transitions. The classification of texts in species and genres proposed by this ontology can be used in to create suggestions for readers, to identify a certain type of text, to extract relevant information, to analyse the format differences between two different types of texts (as those between the normative texts and the news) as a first step in automatic text writing techniques, etc., while the temporal model considers a natural order of events that the reader seems to perceive once she or he continues to read or, at the end of the book despite timelines (the representation of all the events chronologically exposed in a story) and storylines (the main story or plot of a literary text) in order to capture the actions of a character and his or her chronological evolution (time track) throughout a book.
The purpose of the proposed research is to explore how editors participate in Wikidata and how they organize their work. This research addresses the need to have greater knowledge of how Wikidata editors curate data for use. Three theories will be used as a methodological and conceptual framework for data collection and analysis: Activity Theory, Social Identity Theory, and Self-determination Theory. The study uses an exploratory case study methodology. An understanding of the activities in Wikidata, including the roles played, tools used, norms and rules followed, and solutions sought to address contradictions among different components of the activities will help inform communities wishing to contribute data to or reuse data from Wikidata. Furthermore, the findings of this study can go beyond understanding how Wikidata curates knowledge and potentially inform the design of other similar online production communities, scientific research institutional repositories, digital archives, and libraries.
The RAGE research project will provide access to a wide range of software assets for Applied Gaming (AG) enabling AG software developers to better collaborate, share their knowledge, and be able to react to new requirements and trends more efficiently. But, RAGE is facing a greater challenge which is Information Overload (IO) because of the permanent flow of new documents published daily on the Web including their different formats. To solve IO, existing contributors have used complicated mechanisms that index the Web and harvest large piles of documents to detect relevant materials. Others have used Semantic Web Technologies to enrich web resources with additional facts and meaning, but have let the system (and not the end-user) decide about the relevance of the search result failing to satisfy the users’ requirements. Thus, we propose a novel software architecture that combines NLP (by mean of Named Entity Recognition) and Domain Ontology to support searching and browsing (faceted-search), as well as reasoning about named entities on the Semantic Web. AG software developers will get only relevant information according to their specific need without much time spent on information harvesting.
Cultural Heritage Information (CHI) is vital for understanding heritage assets and it is available in various formats, standards, qualities and quantities on the web. Primarily CHI is created, organized and delivered by memory institutions through delivery portals, referred to as Digital Archives in this research. Apart from carefully crafted CHI there are many other third parties who provide information through the web. This research deals with aggregating this diverse CHI on the web to help enrich poorly made information related to South and Southeast Asian region. Many drawbacks related to CHI on the web were realized after investigating several portals and websites, resulting in the proposal of a comprehensive metadata model known as Cultural Heritage in Digital Environment (CHDE). CHDE explicitly identifies instances to curate tangible and intangible cultural heritage resources into digital archives and to help identify metadata structures for resource aggregation. Further, the author investigated the metadata description identification based on their objectives, with the help of One-to-One Principle of Metadata through Description Module mapping aligned with the Dublin Core Application Profiles (DCAP).
Images are dominant in the communication field, but the photograph still has a diverse treatment, in terms of description, interpretation and systematic use. The difficulty of searching in collections of images is known, but is diminishing with the potentialities of automatic analysis.This work focuses on the use of the photographic record as a support for research, whether as data, collected or processed, or as metadata, contributing to the description and interpretation of the data, giving them context. With a survey of the production and use of photographic in science, the informational behavior of researchers producing photography will be studied, as well as its relevance for research. Thus, models of image metadata will be developed, and tested with researchers. Through the combination of these models with automatic image processing, we expect to show that images are rich elements to promote the retrieval of research data, fostering semantic interoperability and data reuse.
The product of scientific research is the production of data. Nowadays researchers are faced with the growing challenge of how to manage, preserve and publish large sets of data, in way that allows for it to be reproduced and reused. This paper proposes to explore the concept of machine-actionable data management plan. In particular, the usage of semantic techniques to both express and exploit the features of machine-actionable data management plans will be analysed. As semantic techniques have been proven to be a reliable means to represent and analyse data in other contexts.
Research data management is required to enable data sharing and reuse. In this context, data description plays a central role, by providing researchers with sufficient information to interpret their data, yet it is a time consuming task. Therefore, it is important to provide them with appropriate tools that can facilitate the description process. Preliminary work with researchers shows that flexible metadata models help them to create detailed descriptions, making data easier to interpret. Given a metadata model, using controlled vocabularies can facilitate the introduction of descriptor values and improve metadata quality. However, few repositories offer support for flexible, domain-specific metadata and the corresponding controlled vocabularies. Thus, this proposal addresses the creation of metadata models, combining descriptors in multiple domains, and their articulation with controlled vocabularies. The expected results are methods for the development, implementation, and support of metadata models and the corresponding vocabularies, with application in research data repositories.
Author(s): Abrams, Stephen | Abstract: Digital information is indispensable to contemporary commerce, culture, science, and education. No future understanding of a prior time in the digital age is possible without proactive preservation of our digital heritage. But how can one know whether or not that preservation has been effective? There are two primary assessments of digital preservation efficacy: trustworthiness of managerial systems and programs, and successful use of preserved resources. The first has received extensive treatment in the literature, but the second has been little investigated. This stems from a too narrow conceptualization of the preservation domain as synonymous with data management. Given that the goal of that management is to facilitate future use, and that use is inherently contingent with respect to time, place, person, and purpose, digital preservation should be seen more broadly as facilitating human communication across time. My research asks what measures can meaningful evaluate the efficacy of such communicative acts. It proposes a communicological theory in which success is evaluated with regard to situational verisimilitude. Evaluation metrics are derived from a semiotic-phenomenological model of preservation-enabled communication and the affordances supported by preserved digital resources. This work contributes new conceptual clarity to the theory and practice of digital preservation, a more rigorous basis for demarcating the limits of preservation efficacy, and a more nuanced means of stating, measuring, and evaluating intentions, expectations, and outcomes.
Researchers must be aware of the current developments and novel findings in their field. To gain an overview of relevant prior work and ongoing research, academics must review the published literature. However, the search for related academic literature is tedious. Furthermore, researchers can easily oversee potentially valuable information in today’s increasing volume of academic literature. While academic search and recommendation engines have greatly simplified the information acquisition process, current recommendation approaches are not adequately considering specialized semantic similarity measures, such as citation-based similarity, mathematical formulae-based similarity, or image-based similarity to recommend and visualize academic literature. This paper proposes to take into account combinations of semantic features that have previously not been considered for the use case of literature recommendation. Additionally, supporting researchers in the sense-making of semantic similarities present in recommended literature remains a largely unsupported task. Researchers must review the content of each academic paper, manually identify or compare the sections that are of interest to them, and then arrive at a judgment regarding the relevance of the recommendation. We propose a supportive process that allows researchers to more quickly compare academic literature with regard to the user-selected semantic features that are of interest to them. Such user-selected features can be the citations to other literature, the contained mathematical formulae, graphs, or figures, i.e. a range of semantic features that go beyond pure text-based similarity. The proposed semantically-enriched literature recommendation and visualization concept could help researchers more quickly identify the specific content that is of interest to them within a large set of recommended literature.
This doctoral work aims to explore the possible avenues in academic peer review where Artificial Intelligence (AI) could have a positive impact. Present day peer review is suffering from many shortcomings, foremost being that it is a very time-delayed process. We identify three potential factors: Novelty, Scope and Quality which are central to the theme of scholarly communication process and seek to investigate various techniques encompassing Natural Language Processing, Information Extraction and Retrieval, Data Mining, Bibliographic Analysis, etc. to address those. The objective is to develop an AI assisted decision support system for the editors, reviewers as well as the authors to get preliminary meta information about a prospective manuscript which may aid in decision making and hence speed-up the overall process. The project is challenging, vast and incorporates several features which closely resembles human behavior to identify quality and appropriate manuscripts. Initial set of experiments on curated and scholarly data show promising results. We are sure that there more issues to address, many scope of improvements which would eventually lead us one step closer to this ambitious vision: to cut through the clutter of bad literature and accelerate scientific discovery and eventually bring AI more close to the academic peer review system.
Human-generated collections of archived web pages are expensive to create, but provide a critical source of information for researchers studying historical events. Hand-selected collections of web pages about events shared by users on social media offer the opportunity for bootstrapping archived collections. We investigated if collections generated automatically and semiautomatically from social media sources such as Storify, Reddit, Twitter, and Wikipedia are similar to Archive-It human-generated collections. This is a challenging task because it requires comparing collections that may cater to different needs. It is also challenging to compare collections since there are many possible measures to use as a baseline for collection comparison: how does one narrow down this list to metrics that reflect if two collections are similar or dissimilar? We identified social media sources that may provide similar collections to Archive-It human-generated collections in two main steps. First, we explored the state of the art in collection comparison and defined a suite of seven measures (Collection Characterizing Suite - CCS) to describe the individual collections. Second, we calculated the distances between the CCS vectors of Archive-It collections and the CCS vectors of collections generated automatically and semi-automatically from social media sources, to identify social media collections most similar to Archive-It collections. The CCS distance comparison was done for three topics: "Ebola Virus," "Hurricane Harvey," and "2016 Pulse Nightclub Shooting." Our results showed that social media sources such as Reddit, Storify, Twitter, and Wikipedia produce collections that are similar to Archive-It collections. Consequently, curators may consider extracting URIs from these sources in order to begin or augment collections about various news topics.
Digital libraries, especially academic digital libraries, are more important than ever. With the increase in the utilization of technology in educational institutions and online courses and degrees, academic digital libraries have become a necessity. In 2010, Saudi Arabia invested in the development of a digital library, called the Saudi Digital Library (SDL), to serve the needs of all university staff and students, offering resources in both Arabic and English. The purpose of this study is to investigate how successfully Saudi Digital Library users are able to find Arabic resources using the SDL database and whether or not they face challenges in their search. The research will utilize a qualitative method, Stimulated Recall Interviews, to evaluate whether or not the SDL meets its users’ needs. This study aims to address these challenges and find ways to overcome them. Information should be accessible to Arabic-speaking information-seekers and indeed to many other informationseekers in different languages and cultures around the world. General Terms Digital libraries, Information-seeking Behavior
In my dissertation, I will explore the following hypothesis: Considering mathematical formulae using the open knowledgebase Wikidata improves content-based literature exploration and recommendation for STEM documents.
e number of public and private web archives has increased, and we implicitly trust content delivered by these archives. Currently, users can access web archives without the ability to check xity of content. Checking xity is performed to ensure archived resources have remained unaltered since the time they were captured. We have noticed that most web archives do not allow users to access xity information. More importantly, even if xity information is available, it is provided by the same archive delivering the content. In this research, we propose a framework to establish and check xity of archived resources. is framework does not require changes in the current web archiving infrastructure, and it will be built based on well-known web archiving standards, such as the Memento protocol. In addition, the proposed framework will result in two main benets. First, any user can generate xity information, not only the archive serving the content. Also, it allows anyone to check xity of archived content regardless of which archive/user generated the xity information. Second, the proposed framework denes a process to store xity information independently from the archive of the associated content.
In mathematics, LaTeX is the de facto standard to prepare documents, e.g., scientific publications. While some formulae are still developed using pen and paper, more complicated mathematical expressions used more and more often with computer algebra systems. Mathematical expressions are often manually transcribed to computer algebra systems. The goal of my doctoral thesis is to improve the efficiency of this workflow. My envisioned method will automatically semantically enrich mathematical expressions so that they can be imported to computer algebra systems and other systems that can take advantage of the semantics, such as search engines or automatic plagiarism detection systems. These imports should preserve the essential semantic features of the expression.