Despite substantial investment in research data infrastructure, data discovery remains a fundamental challenge in the era of open science. The proliferation of repositories and the rapid growth of deposited data have not resulted in a corresponding improvement in data findability. Researchers continue to struggle to find data that are relevant to their work, revealing a persistent gap between data availability and data discoverability. Without rich, high-quality metadata, robust and user-centred data discovery systems, and a deeper understanding of how different researchers seek and evaluate data, much of the potential value of open data remains unrealised. This paper presents a set of practical, evidence-based recommendations for data repositories and discovery service providers aimed at improving data discoverability for both human and machine users. These recommendations emphasise the importance of 1) understanding the search needs and contexts of data users, 2) addressing the roles that data repositories play in enhancing metadata quality to meet users’ data search needs, and 3) designing discovery interfaces that support effective and diverse search behaviours. By bridging the gap between data curation practices, discovery system design, and user-centred approaches, this paper argues for a more integrated and strategic approach to data discovery.
Hugging Face models are on the rise currently nearly touching 3 million mark. This substantial increase has brought numerous challenges such as relevant model card discovery focusing tasks, architecture and license. For example, text classification as filtering criterion brings 112,077 models, and the user may select other filters to reduce that number ^1 . Moreover, current Hugging Face searching mechanism does not support certain metadata granularities i.e., base-model, dataset, and publications. This practically overlooks the relevant model discovery process considering these interconnections among models. In this demo paper, we exploit the Hugging Face model card metadata features and showcase i) valuable insights, ii) interconnections across heterogeneous metadata features as knowledge graph instance based on schema.org mappings, and iii) user queryable mechanism. Hence, these aspects support granular, metadata-driven model card discovery across research communities.
In today's data-driven research landscape, dataset visibility and accessibility play a crucial role in advancing scientific knowledge. At the same time, data citation is essential for maintaining academic integrity, acknowledging contributions, validating research outcomes, and fostering scientific reproducibility. As a critical link, it connects scholarly publications with the datasets that drive scientific progress. This study investigates whether repository visibility influences data citation rates. We hypothesize that repositories with higher visibility, as measured by search engine metrics, are associated with increased dataset citations. Using OpenAlex data and repository impact indicators (including the visibility index from Sistrix, the h-index of repositories, and citation metrics such as mean and median citations), we analyze datasets in Social Sciences and Economics to explore their relationship. Our findings suggest that datasets hosted on more visible web domains tend to receive more citations, with a positive correlation observed between web domain visibility and dataset citation counts, particularly for datasets with at least one citation. However, when analyzing domain-level citation metrics, such as the h-index, mean, and median citations, the correlations are inconsistent and weaker. While higher visibility domains tend to host datasets with greater citation impact, the distribution of citations across datasets varies significantly. These results suggest that while visibility plays a role in increasing citation counts, it is not the sole factor influencing dataset citation impact. Other elements, such as dataset quality, research trends, and disciplinary norms, can also contribute to citation patterns.
Sharing and reusing research artifacts, such as datasets, publications, or methods is a fundamental part of scientific activity, where heterogeneity of resources and metadata and the common practice of capturing information in unstructured publications pose crucial challenges. Reproducibility of research and finding state-of-the-art methods or data have become increasingly challenging. In this context, the concept of Research Knowledge Graphs (RKGs) has emerged, aiming at providing an easy to use and machine-actionable representation of research artifacts and their relations. That is facilitated through the use of established principles for data representation, the consistent adoption of globally unique persistent identifiers and the reuse and linking of vocabularies and data. This paper provides the first conceptualisation of the RKG vision, a categorisation of in-use RKGs together with a description of RKG building blocks and principles. We also survey real-world RKG implementations differing with respect to scale, schema, data, used vocabulary, and reliability of the contained data. We also characterise different RKG construction methodologies and provide a forward-looking perspective on the diverse applications, opportunities, and challenges associated with the RKG vision.
Search and harvesting use cases on harmonised metadata play an important role in several activities on National Research Data Infrastructures (NFDI). The working group Search and Harvesting of the NFDI section (meta)data, terminologies and provenance works on a common understanding of user needs (for search) and service requirements (for harvesting), analysis of the data sources landscape, and recommendations concerning common and specific needs, e.g., for spatial or sensitive data. Here, we present search and harvesting gaps and challenges across NFDI consortia and beyond, which were identified and structured in the Search and Harvesting Working Group, and the recommendations for the NFDI we derive from them. Our goal is to foster a common vision for search and harvesting in the NFDI.
The rapid development of social media in recent years has encouraged the sharing of vast amounts of data and the propagation of fake news. This has pushed the scientific community to focus on this phenomenon, particularly those working on natural language processing, by developing detection tools to combat fake news. At the same time, most studies have focused on languages with a high resource content (corpus). This paper aims to shed light on low-resource languages, particularly the Algerian dialect, through an experimental study with two objectives. The first one is to verify if the automatic translation from Modern Standard Arabic (MSA) to the Algerian dialect can be considered an approach to increase the resources in the Algerian dialect, especially with the rise of large language models (LLMs). The second is to verify the impact of the translation-based data augmentation method on fake news detection by using transformer-based Arabic pre-trained models in different data augmentation configurations. We have discovered that LLMs can generate translations that closely resemble human translations. In this study, we demonstrate that data augmentation can result in saturation and a decline in model performance due to the introduction of noise and variations in writing styles.
Purpose: Data discovery practices currently tend to be studied from the perspective of researchers or the perspective of support specialists. This separation is problematic, as it becomes easy for support specialists to build infrastructures and services based on perceptions of researchers' practices, rather than the practices themselves. This paper brings together and analyzes both perspectives to support the building of effective infrastructures and services for data discovery. Methods: This is a meta-synthesis of work the authors have conducted over the last six years investigating the data discovery practices of researchers from different disciplines, with a focus on the social sciences, and support specialists. We bring together and re-analyze data collected from in-depth interview studies with 6 support specialists in the field of social science in Germany, with 21 social scientists in Singapore, an interview with 10 researchers and 3 support specialists from multiple disciplines, a global survey with 1630 researchers and 47 support specialists from multiple disciplines, an observational study with 12 researchers from the field of social science and a use case analysis of 25 support specialists from multiple disciplines. Results: We found that there are many similarities in what researchers and support specialists want and think about data discovery, both in social sciences and in other disciplines. There are, however, some differences which we have identified, most notably the interconnection of data discovery with web search, literature search and social networks. Conclusion: We conclude by proposing recommendations for how different types of support work can address these points of difference to better support researchers' data discovery practices.
In the context of low-resource languages, the Algerian dialect (AD) faces challenges due to the absence of annotated corpora, hindering its effective processing, notably in Machine Learning (ML) applications reliant on corpora for training and assessment. This study outlines the development process of a specialized corpus for Fake News (FN) detection and sentiment analysis (SA) in AD called FASSILA. This corpus comprises 10,087 sentences, encompassing over 19,497 unique words in AD, addresses the language's significant lack of linguistic resources, and covers seven distinct domains. We propose an FN detection and SA annotation scheme detailing the data collection, cleaning, and labeling. The remarkable Inter-Annotator Agreement indicates that the annotation scheme produces high-quality and consistent annotations. Subsequent classification experiments using BERT-based and ML models are presented, demonstrating promising results and highlighting avenues for further research. The dataset is currently freely available to facilitate future advancements in the field.
The poster presents FAIR assessment experiences in the context of the two NFDI consortia KonsortSWD and BERD@NFDI, employing the established Research Data Alliance's FAIR Data Maturity Model (RDA-FDMM) and the F-UJI Tool, an automated solution. RDA-FDMM, a manual technique, is more comprehensive, while the automated F-UJI tool effectively detects areas of improvement in metadata presentation that automated means can address. Our experiences highlight the need to examine both machine-readable as well as non-machine-readable elements and acknowledge automated tools' limitations, while valuing their insights. As the research ecosystem advances, metadata representation should be made increasingly machine-readable. We recommend a "FAIR by design" approach from the beginning to ensure alignment with FAIR principles in project outcomes. Continuous assessments during a project’s lifetime promote ongoing research data infrastructure improvements within the NFDI consortia context, contributing to NFDI infrastructure innovation and optimization.
Data discovery is important to facilitate data re-use. In order to help frame the development and improvement of data discovery tools, we collected a list of requirements and users’ wishes. This paper presents the analysis of these 101 use cases to examine data discovery requirements; these cases were collected between 2019 and 2020. We categorized the information across 12 ‘topics’ and eight types of users. While the availability of metadata was an expected topic of importance, users were also keen on receiving more information on data citation and a better overview of their field. We conducted and analysed a survey among data infrastructure specialists in a first attempt at ranking the requirements. Between these data professionals, these rankings were very different, excepting the availability of metadata and data quality assessment.
The poster presents FAIR assessment experiences in the context of the two NFDI consortia KonsortSWD and BERD@NFDI, employing the established Research Data Alliance's FAIR Data Maturity Model (RDA-FDMM) and the F-UJI Tool, an automated solution. RDA-FDMM, a manual technique, is more comprehensive, while the automated F-UJI tool effectively detects areas of improvement in metadata presentation that automated means can address. Our experiences highlight the need to examine both machine-readable as well as non-machine-readable elements and acknowledge automated tools' limitations, while valuing their insights. As the research ecosystem advances, metadata representation should be made increasingly machine-readable. We recommend a "FAIR by design" approach from the beginning to ensure alignment with FAIR principles in project outcomes. Continuous assessments during a project’s lifetime promote ongoing research data infrastructure improvements within the NFDI consortia context, contributing to NFDI infrastructure innovation and optimization.
The dataset refers to the poster, which presents FAIR assessment experiences in the context of the two NFDI consortia KonsortSWD and BERD@NFDI, employing the established Research Data Alliance's FAIR Data Maturity Model (RDA-FDMM) and the F-UJI Tool, an automated solution. RDA-FDMM, a manual technique, is more comprehensive, while the automated F-UJI tool effectively detects areas of improvement in metadata presentation that automated means can address. Our experiences highlight the need to examine both machine-readable as well as non-machine-readable elements and acknowledge automated tools' limitations, while valuing their insights. As the research ecosystem advances, metadata representation should be made increasingly machine-readable. We recommend a "FAIR by design" approach from the beginning to ensure alignment with FAIR principles in project outcomes. Continuous assessments during a project’s lifetime promote ongoing research data infrastructure improvements within the NFDI consortia context, contributing to NFDI infrastructure innovation and optimization.
In this panel, the sections of the NFDI initiative will report on their organizational structures and their collaboration with the NFDI consortia to design cross-disciplinary services for a German national research data infrastructure. The panelists will share early experiences, insights and best practices from their interdisciplinary work. They will also highlight how their work links to European activities in the EOSC and other international initiatives.
NFDI is a German initiative to set up research data infrastructures across all disciplines. Within NFDI, Base4NFDI is a unique joint effort of all NFDI consortia to develop and deploy NFDI-wide basic services. Within this talk, we will give an overview of Base4NFDI, especially its structures and emerging work program, and inform about ways to participate and contribute ideas for potential basic services.
Andreas Kupfer合作论文数Technische Universität Braunschweig
Institut für Informationssysteme14