
The mastery of musical-instrument playing skills is significantly shaped by individual traits and instructional methodologies. Since teachers rely on personal subjectivity and empirical knowledge in their guidance, scientifically effective teaching methods remain partially understood. Our project sought to investigate this by conducting a series of activities to examine where teachers direct their attention during performances and how they formulate critiques. This paper presents the Critique Documents (CROCUS) database, which we created for music performance. CROCUS includes musical performances for three instruments, textual comments as critiques, assessments of the critique’s utility, and annotated semantic types for critique sentences. This comprehensive database offers a new substantial and valuable resource for researchers exploring music education and learning activities.
Abstract This article presents findings from a survey conducted in Lima, Peru, aimed at understanding the relationships between education, financial literacy, financial inclusion, and informal financial business practices among small female vendors. The study, which collected 118 valid responses, focused on the impact of these factors on vendors’ intentions toward formalization. Formality was assessed based on legal registration with tax authorities, emphasizing the informal practices viewed on a continuum. These practices were evaluated using a five-point gradation scale that depicted varying levels of formality. Financial literacy, financial inclusion, and formalization intentions were measured using a five-point Likert scale, while a dichotomous question captured the formality-informality of the businesses. The demographic variables included age, gender, business tenure, employee count, and business activity. Educational level, typically treated as demographic, was considered an antecedent to financial literacy. The dataset linked to this study included raw survey data. It serves as a valuable resource for researchers, industry representatives, public authorities, and stakeholders from developing countries to deal with informality and formalization. The survey methodology and data are adaptable for use in different national contexts, facilitating comparative analysis in developing countries.
Abstract In recent years, interest in three-dimensional visualizations as a complement to traditional text-based humanities scholarship has surged. These visualizations allow for new analysis methods and may serve as an interpretative tool, potentially leading to new insights for creators and viewers. However, contrary to more traditional publications, a myriad of preservation issues occurs from the moment a 3D visualization is created. This data paper illustrates the complexity of preserving 3D visualizations in a humanities context using six interactive visualizations produced in the Dynamic Drawings in Enhanced Publications project (2013). The visualizations serve as a concrete case study for the restoration, preservation and renewed publication of 3D content. This data paper sketches the general preservation challenges that led to non-functional interactive visualizations. The authors detail the process of restoring the visualizations from the original data and opening them up for potential reuse. To underline reuse possibilities, the authors provide tentative examples for vr and ar platforms. After presenting the dataset, they discuss the implications of the restoration process for future preservation practices of unstable interactive content. The authors introduce a ‘composite elements approach’, combining visual and non-visual levels of documentation necessary for the maintenance of interactive visualizations.
Abstract In genetic criticism, scholarly editing, authorial philology and, more generally, for the study of authorial manuscripts and writing processes, it is essential to order and classify the textual witnesses and their relationships. This article presents two datasets of so-called ‘genetic networks’, that is representation of the genetic entities (witnesses, publications, dossiers) and their relationships, modelled according to the geno 1.0 ontology. The datasets contain genetic networks of the works of two Swiss authors: the main publications of Gustave Roud (1897–1976) and the short story “En mer” by Bernard Comment (1960).
This paper presents a database on film programming in Moscow cinemas between 1946 and 1955. It outlines the place of this research at the intersection of new cinema history and academic debates on film distribution in the field of Soviet history. The paper describes the data collection, the coding to present the data, and the structure of the database, which consists of the three datasets on Moscow film programming (1946–1955), Moscow cinemas (1946–1955), and the 1952 film calendar. Concluding remarks summarize the knowledge obtained from the database and introduce the Soviet case into the international context of digital data collections for historical cinema studies.
This article introduces a newly constructed database: the Memories 1921 Dataset. The database is a representative sample drawn from the Tafel v-bis Dataset and consists of complete inheritance taxation records for 2,321 individuals who died in 1921. To increase the amount of information about these individuals, the authors connected the Memories 1921 Dataset to the original Tafel v-bis Dataset. This added information about the name, birthplace, place of residence, age, marital status, and profession of the deceased to his or her wealth portfolio. The database is open access and available via the Social Sciences and Digital Humanities Archive (sodha).
Abstract Although ‘the family’ is arguably the most fundamental of all social networks, surprisingly little data are available that enable researchers to study the full web of relationships between family members. Mapping family relationships from multiple – preferably ‘all’ – family members’ perspectives enables understanding relational dependencies, such as how parental divorce reverberates through the network. This article introduces a multi-actor family network survey method aimed at collecting ‘complete’ family network data. It discusses the design and implementation of the Lifelines Family Ties project. In this data collection project, a total of 160 children, parents, grandparents, aunts, uncles, and stepfamily members reported on their current and past well-being and their family relationships (contact, support, affection) with 524 family members, resulting in a dataset covering nearly 900 relationships. The article concludes by providing a preview of possible analysis techniques for future users of the Lifelines Family Ties dataset or other future multi-actor family network data.
Abstract This article presents a database focusing on settlement patterns, demographics, and economic transformations within one of the largest forested areas of early modern Poland, the White Forest, situated at the confluence of two significant rivers, the Narew and Bug. These waterways historically facilitated the export of grain and forest products from eastern Poland and certain regions of the Grand Duchy of Lithuania. This database was compiled according to the guidelines established by the “Historical Atlas of Poland” which gathers spatially-oriented data pertinent to the historical geography of Polish lands. The database pertains to a region that was somewhat peripheral in terms of the overall national development, rendering it suitable for comparative analysis with other European regions that exhibit similar trends—intensive settlement expansion, environmental pressure in the early modern era, and a degree of distance from contemporary political and economic hubs.
Abstract The indigo Impact Bond Dataset is an open-access dataset that describes a specific form of impact-focused cross-sector partnership adopted worldwide since 2010. These partnerships are data-rich in principle, yet historically, little data is shared and re-used. The dataset is the result of an engaged, collaborative process where different organisations involved in impact bond projects share data with the indigo initiative data stewards so that practitioners and researchers can analyse and learn from these partnerships. This article introduces the dataset in terms of scope, data collection methods, and data model. The authors provide descriptive summaries of the current landscape and demonstrate a practical application of the dataset. In closing, they discuss future avenues for research and dataset development, as well as the limitations of working with a collaborative approach.
Abstract This article introduces a pioneering dataset from a survey of civil society organizations (cso s) in the metropolitan region of Vienna, Austria. The survey was conducted between October 2019 and December 2020 and provides a comprehensive overview of the current state of the civil society sector in Vienna. It comprises a representative sample of 358 cso s and an additional targeted sample of 235 large cso s. The anonymized dataset is stored at the Austrian Social Science Data Archive (aussda). It can be freely accessed after the end of the embargo period in May 2025. The survey includes more than 60 questions covering a wide range of topics, including organizational goals and activities, beneficiary and staff demographics, different forms of organizing and related practices, performance metrics, budgeting, funding sources, and collaborative efforts. The dataset is a valuable resource for scholars interested in studying the inner workings, relationships, and societal contributions of civil society organizations, and it appeals to a variety of scholarly debates.
Abstract This data paper presents Arkas 2.0, a national research database and research infrastructure containing data on all archaeological sites and monuments in Slovenia. The new database is a hybrid cloud microservice built on low-code platforms (Caspio and ArcGIS Experience builder) and augmented by generative ai (ChatGPT-3.5). The data paper describes the Arkas 2.0 dataset and how it fits into the research context by discussing the challenges archaeologists face in setting up and curating datasets and the associated digital infrastructure. In response to these challenges, the data paper highlights the benefits of low-code platforms and ai-augmented code for archaeological research. It also describes the Arkas 2.0 development workflow, its new data structure, and its archiving process. The data paper concludes by suggesting that the use of low-code platforms combined with generative ai can democratise access to cutting-edge digital research infrastructure, bringing positive disruption to archaeology and the humanities.
Abstract In this article, the authors present a dataset of the text of the Samaritan Pentateuch with word-level linguistic annotations. The Samaritan Pentateuch is an important early witness of the Pentateuch or Torah. This dataset is based on a transcription generally taken from manuscript Dublin, Chester Beatty Library 751 (Genesis 1:1–Deuteronomy 32:36) and supplemented from manuscript Nablus (Kiryat Luza), Samaritan Synagogue, Garizim 1, where the former manuscript has not preserved the text (Deuteronomy 32:36b–34:10). The dataset is a Text-Fabric dataset. Text-Fabric is a Python package for processing annotated text corpora, which means that the dataset comes with an app, where the text can be inspected and queried using the annotations. It is also easy to perform textual and linguistic research using Python scripts and to make comparisons with other relevant textual datasets with the same annotation conventions.
Abstract The War dummies dataset offers structured data on Dutch-involved organised armed confrontations from 1566 to 1812. Comprising 1216 records detailing the participation of 95 entities in 548 encounters across 50 wars, it fills a crucial need for well-structured, accessible, and reusable pre-1815 historical warfare data. Based on the comprehensive Military History of the Netherlands book series, it aligns with post-1815 conflict datasets like the Inter-State War Database of the Correlates of War Project and the Georeferenced Event Dataset of the Uppsala Conflict Data Program. This article outlines the data collection, structure, and potential research applications, and discusses data quality and potential biases. The War dummies codebook offers comprehensive variable descriptions.
Abstract Studying Russian society is challenging, especially during the period of the Russian military invasion. However, it takes on special significance during a period of economic and social transformation. Studying the career and educational trajectories of Russians in the context of East Studies offers a multifaceted perspective on the state of the job market and education sector and gives an understanding of the current situation of the country’s economy and social structure. The lack of data with a high level of granularity is critical, especially for studying people with a focus on their career and educational trajectories. In this article, the authors respond to this request and present two datasets that can be useful for studying spatial and temporal patterns associated with people’s life trajectories in the context of work and education. The authors utilised open data on cv s created or updated by employment portal users over the period 2015–2023 from the Federal Service for Labor and Employment (Rostrud) and prepared two cleaned datasets covering 83 regions of Russia. Dataset 1 is on the educational and career trajectories (N = 6,221,439) and Dataset 2 is on the activity of unemployed and job-seeking candidates (N = 7,662,089).
Abstract Indigenous peoples are among the most vulnerable, ignored, and marginalized groups in society. Poverty is the oldest social problem and difficult to counter. The Indigenous people with which the authors live and work, the Agta Tabangnon, suffer from poverty and multidimensional socioeconomic deprivations. Indigenous peoples’ studies are qualitative, while poverty studies are typically generic, exposed to large sampling errors, and intended for nationwide decisions. Therefore, measuring poverty for specific tribes through complete enumeration with multifaceted disaggregation is critical for economic development. There is no comprehensive census specifically designed for Indigenous peoples to encompass the multidimensional aspects of their way of life. Nonetheless, the authors are resourceful in generating useful datasets from their partners. The locale is situated in the poorest district of the poorest province in the poorest region of Luzon, Philippines. The datasets contain multidimensional poverty indicators that are readily usable, along with complementary analytics to visualize the data. They may serve to measure poverty in Indigenous communities across different regions and countries. By utilizing this data, further empirical analysis, regressions, machine learning, and econometric modeling can be conducted. It can be freely utilized to target policies that address the multifaceted poverty and promote economic development within tribal communities.
Abstract The authors present a dataset containing transcriptions of manuscripts of the Middle Dutch strophic poem Martijn Trilogy by the Flemish poet Jacob van Maerlant. Of his very large oeuvre, Maerlant’s Strophic Poems had the longest tradition: originally written in the thirteenth century, copyists and printers continued to disseminate them until about 1500. These ten shorter poems on social, religious, and ethical issues stand out for their unusual and complex stanza form. The Martijn Trilogy was his most successful strophic poem: 17 text witnesses are extant, and the trilogy was imitated and even translated into French and Latin. This dataset contains hyperdiplomatic transcriptions of all witnesses, amounting to a total of 15,814 verses or 79,337 tokens. This open-access dataset abides by the fair principles, is licensed under a cc-by-sa license, and is made available in multiple, complementary file formats. Since these transcriptions are strictly diplomatic, this corpus offers valuable possibilities for research on scribal attributions (scribal profiling), abbreviations, stemmatology, textual stability and more.
Abstract The authors describe the Qualitative Election Study of Britain (qesb) Party Leader Evaluation Database, a database containing 4,119 words and phrases evaluating British political party leaders. The data were collected during pre-election focus groups and interviews with participants from England, Scotland, and Wales during the General Election campaigns of 2010, 2015, 2017, and 2019. A supplementary dataset of leaders’ evaluation data from Dundee residents after the Scottish Independence Referendum in 2014 is also provided. To collect the data, participants viewed headshot pictures of major and minor party leaders (depending on where in Britain they lived) taken from party websites. Participants wrote down words or phrases they associated with each leader and coded their assessment as positive, negative, or neutral. These data are suitable for content, sentiment, and discourse analysis or analytic generalization.
Abstract This article presents a novel, extensive, and thoroughly documented dataset describing Australian feature films and the personnel filling ten key production roles on those films. The dataset is curated from public information in multiple sources and draws on further supplemental resources to verify, validate and consolidate this information. In total, the data describes 22,720 roles filled by 9,397 distinct people across 1,877 films, covering an important 47-year period in the Australian film industry. The authors outline how the dataset solves several problems for scholars interested in data that provides a historical record of the collaborative filmmaking process. In particular, to address concerns about known coverage problems with popular sources such as the Internet Movie Database, this dataset has undergone extensive manual checking to ensure that it is reliable as a source of information on a national film industry. Moreover, the authors have carefully and manually linked each person appearing in the dataset, which allows the dataset to provide a rich source of information for exploring the relationality of filmmaking collaborations. The inclusion of ten key filmmaking roles further expands the utility of the dataset beyond existing datasets which tend to focus on actors and/or directors, writers and producers.
Abstract With the recognition of content discoverability and information analytics in the scientific publishing industry, more and more effort is dedicated to digitization and automated analysis of scholarly publications. As part of this effort, the authors designed a pipeline to extract structured information from bibliography and index lists of existing scholarly publications, as well as to disambiguate and export it as linked data. In this article, the authors present the Brill Knowledge Graph (kg), obtained by applying this pipeline to a corpus of books in the Arts, Humanities and Social Sciences (ahss) provided by the publisher Brill.
The Tagged Corpus of Early English Correspondence Extension Sampler (TCEECES) is the third public release from the 18th-century part of the Corpora of Early English Correspondence (CEEC-400). The TCEECES forms one part of the full CEECES. The other parts are the CEECES part 1 (released 1 April 2021), and the CEECES part 2 (released 11 April 2022). The TCEECES is an extract from the full Tagged Corpus of Early English Correspondence Extension (TCEECE), which remains unpublished. See the accompanying manual for more information on the TCEECES; see the manuals for the CEECES 1 and the CEECES 2 for more information on the CEECES. See https://varieng.helsinki.fi/CoRD/corpora/CEEC/ for more on the CEEC-400. Citation: TCEECES = Tagged Corpus of Early English Correspondence Extension Sampler. Compiled by Terttu Nevalainen, Helena Raumolin-Brunberg, Samuli Kaislaniemi, Mikko Laitinen, Minna Nevala, Arja Nurmi, Minna Palander-Collin, Tanja Säily and Anni Sairio at the Department of Languages, University of Helsinki. Spelling standardised by Mikko Hakala, Minna Palander-Collin, Minna Nevala, Emanuela Costea, Anne Kingma and Anna-Lina Wallraff. Annotated by Lassi Saario and Tanja Säily. XML conversion and encoding by Lassi Saario. Helsinki: VARIENG, 2022.