
This paper introduces the concept of "translation mining", a cross-lingual embedding approach that bridges close reading and large-scale quantitative analysis of historical translations. Focusing on eighteenth-century British (ECCO) and French (Gallica) corpora, we systematically identify semantically aligned passages across hundreds of thousands of documents. The approach provides a comprehensive AI-driven methodology for locating translations of various types. This enables us to build a multi-scalar, data-driven taxonomy of translation practices that offers finer granularity than prior dichotomies of "full" versus "partial" translation. By embedding all texts once using high performance computing, we can iteratively detect subtle textual connections across languages without re-running the entire process. This paper demonstrates how embedding models derived from large language models can capture cross-lingual translation pairs, challenge rigid classifications, and illuminate multi-level cultural transfer, offering an adaptable framework for historical research.
In 1901, the Canadian government published a summary of the Indigenous population resident within the boundaries of what would become present-day Canada. Until now, the census microdata underlying this summary has not been available to researchers. We use these newly available data from the enumerator manuscripts to provide a detailed, quantitative profile of Indigenous Peoples at this time, taking advantage of the fact that the 1901 census is the most comprehensive account of Indigenous Peoples published before WWI. We additionally document the limitations of these data and provide evidence of bias in the coverage of Indigenous economic life. Even with these concerns, we argue that the 1901 microdata provide an important benchmark that can be used to identify the effects of large-scale institutional and demographic change on Indigenous Peoples, particularly in Western Canada.
There has been a plethora of work produced in the fields of public participatory GIS (PPGIS), spatial humanities, and historical GIS (HGIS) that illustrate the benefits of engaging public stakeholders in digital and deep mapping projects. Scholars have promoted many different public outreach activities and programs that attempt to engage the public; however, little work has been done to evaluate the impact these activities and programs have had in creating and sustaining public engagement. Using the Keweenaw Time Traveler (KeTT) project as a case study, this paper analyses six community outreach activities and programs used to engage the public in the exploration of the KeTT deep map over a one-year period. This analysis uncovers the depth and breadth of engagement users have with the deep mapping environment and provides a set of best practices other projects can use to create and sustain public engagement.
Although a vast body of literature identifies an indissoluble link between Indigenous families and their organization in households, the residential arrangements of Indigenous Peoples in Canada have received little attention to date because of the lack of available data. The Canadian People (TCP) project's complete-count census data offers a unique opportunity to fill this gap, since one of its main goals is to support demographic analyses of smaller population groups, notably Indigenous Peoples. In this paper, we take advantage of the TCP 1901 census data for Manitoba to draw the first historical portrait of Indigenous households' structures at the beginning of the twentieth century, which is key to quantify the devastating effects of assimilation policies put in place by the newly established Canadian government, including children's removal to residential schools. Our portrait of Indigenous households' living arrangements provides a more precise picture of nuclear households and multihouseholds in Manitoba in 1901 than previous analyses of census sample datasets have allowed. Our comparison between the TCP data and the original manuscripts on microfilm, which was needed to validate household boundaries and relationships, emerge as a vital step of analyses not only about household structures but also about Indigenous enumeration in historical censuses.
This article reports on a large-scale historical criminology project to create longitudinal datasets of Australian prosecutorial data. We reflect on the decade-long process of digitization, transcription, classification, and aggregation that transformed hundreds of thousands of handwritten manuscript records into open-source, machine-readable data now readily available for analysis by researchers. Secondly, we explore the possibilities of developing historically relevant offense, verdict and sentence classification systems to analyze longitudinal data globally and offer comparisons to other available classification systems. We propose that this approach is a necessary next step to ensure we develop robust and replicable scholarship in historical criminology, but the development of such a framework needs to be highly attentive to issues of missing data and changing historical contexts.
In this paper we use newly digitized complete-count micro data from the 1901, 1911 and 1921 Canadian population censuses to document foreign born assimilation into Canadian labor markets during the early twentieth century. We find that the average immigrant faced a large earnings penalty relative to their native born peers when they first entered Canadian labor markets, and this penalty was very persistent-lasting for more than 30 years. However, the assimilation experiences of the foreign born in Canada were highly variable. We document significant differences in foreign born workers' age-earnings profiles depending on their country of origin, occupation skill level, age at arrival, arrival cohort, the region in which they live, the size of their local labor market, and the density of foreign born workers in their local labor market.
We introduce a generic machine learning-based pipeline for nominative linkage of records within and across large-scale Chinese historical datasets. The pipeline addresses key challenges, including character variations, incomplete data, and scalability issues specific to historical datasets in which names and other attributes are recorded with Chinese characters, not just for China, but potentially for Korea, Japan and Vietnam. Techniques developed for attributes recorded in phonetic alphabets are of limited usefulness for Chinese characters not only because homonyms are common, but characters that are similar enough in appearance to be frequently mistaken for each other may sound completely different. Our approach integrates stroke-based character embeddings for efficient blocking, supervised classification with active learning for record matching, and graph-based clustering for final linkage. We demonstrate the effectiveness of this pipeline using the career records of officials in the China Government Employee Database-Qing Jinshenlu (CGED-Q JSL) as a test case. We achieve improved linkage quality compared to standard probabilistic methods, with substantially longer linked sequences of career records and fewer aberrant transitions. To validate the generalizability, we also successfully apply the pipeline to another database and a cross-database linkage task. By minimizing the need for manual tuning, our pipeline offers a more accessible and effective solution for Chinese historical data linkage.
We describe a method to estimate child mortality using linked IPUMS full-count census data for the decades following the 1850, 1860, 1870, 1900, 1910, 1920, and 1930 censuses. We compare the estimates to several external sources and conclude that the method represents a viable way to estimate and model child mortality at individual and aggregate levels over the course of the mortality transition. We demonstrate an example use of the method by mapping county-level mortality estimates for the 1910-1920 period. Finally, with the help of model life tables, we construct complete life tables by race and decade at seven levels of geography-urban/rural, city, county, state economic area, state, census division and nation-and make these data available for public use.
We use the new complete count Canadian Census records spanning from 1871 to 1901 to explore measures of socioeconomic status including income, occupation, and literacy, which scholars commonly use to investigate economic mobility and inequality. We first explore the availability and comparability of measures across census years. Then, using individuals in the 1901 Census which contains wages, we document selection into the wage sample as well as the characteristics of individuals found in different segments of the wage distribution. Lastly, auxiliary status measures from social registers and probate records are explored for those in occupations where income is under-enumerated.
The digitization of county deed records creates new opportunities and challenges for historical research, rendering private deed or property restrictions-one of the root mechanisms of racial segregation in the United States-more discoverable, and more quantifiable. This essay surveys the baseline and varying organization of county deed records. It traces strategies for identified racial restrictions contained within those records, using a variety of research methods and tools. And it summarizes strategies and methods for analyzing and mapping the results.
The demographic decline of the rural world is usually seen as solid evidence of long-run economic and social development. Its correlate, urbanization, plays a key role in studies of comparative development. And yet, this is often founded on shaky empirical grounds: what exactly counted as rural in the past(s)? We compare the many thresholds used by historians to define the rural/urban boundary and, through the case of modern southern Spain, show that even apparently small changes to the criteria result in substantial variation in both levels and trends of rural population. This leads us to suggest that sensitivity analyses should be standard in empirical work that relies on historical urbanization rates or exploits differences between rural and urban places (in fertility, productivity, etc.). We warn against the rigidities implied in the criteria to distinguish between rural and urban, and ask that historians (and historically-minded social scientists) are mindful of the actual magnitude of the gap between the estimates produced by different methods.
This paper enhances the study of medieval kingdom governance through digital analytical methods, drawing on approaches developed by German historians in the 1980s and 1990s. We focused the research on P & rcaron;emysl Otakar II's reign (1251-1278), using a dataset of his activities, geocoded localities, and reconstructed travel routes with estimated times. We identified the primary centers of governance by applying twelve metrics from statistical, geostatistical, and network analysis. These metrics are further compared to provide insights from methodological and historical perspectives. In addition to providing a methodological extension to the research of historical itineraries, the outcomes of the analyses brought a new contribution to the historical study of the administration of the medieval kingdom.
Representative data on historical minority populations are often absent or costly to collect, while previously used sampling methods have led to small or biased samples. This paper presents a method to systematically identify Jews by leveraging the distinctiveness of Jewish and Gentile names in Dutch population data. Using two external sources for verification, I evaluate efficiency and accuracy of Jewish samples identified through different combinations of given names and surnames. The results show that incorporating additional family names, such as those of parents, substantially increases sample sizes while simultaneously minimizing errors. Using name distributions by ethno-religious background from mid-nineteenth-century Amsterdam population registers, between 81 and 86 percent of Amsterdam-born grooms and brides are accurately identified as either Jewish or Gentile on Dutch civil certificates between 1811 and 1932. I further demonstrate that the method is effective across time and across the Dutch provinces, producing Jewish samples that closely match their local population shares. Finally, I illustrate the potential of the approach by documenting the socioeconomic ascent of Amsterdam Jews from the second half of the nineteenth century onward, and discuss ways to extend the method through data linking and the incorporation of additional information.
IPUMS has finalized databases for each of the United States population censuses from 1850 to 1880. These data are the result of collaborations between FamilySearch and Ancestry.com, which provided the raw data, and IPUMS, which enhanced the data with editing, standardized coding, inter-census harmonization, and documentation. We discuss the data capture process conducted by the nineteenth-century United States Census Office, construction of the modern datasets, and variable availability. We conclude by briefly discussing the potential and limitations of these data for social science research. The public data are distributed by IPUMS and available for researchers to use free of charge.
This article presents a comprehensive database featuring the digitized, cleaned, geocoded, and linked data of the Swedish manufacturing censuses between 1863 and 1900. The data covers close to the universe of Swedish manufacturing activity and includes establishment-level information on workers, the sum of production value, and toll as output value. The article describes how the data was originally collected and the steps taken to go from raw data to the digital database. We discuss each variable's definition, how it changed over time, and provide an assessment of the reliability of the data pertaining to each variable. We also assess the quality of the data by comparing it to various other data sources from the same time period. The level of detail in the data makes the users able to both detect and address potential weaknesses of the data. The database offers a unique resource for scholars to study the manufacturing sector during a time of significant transformation in the Swedish industry. To the best of our knowledge, this is among the earliest sources of annual, establishment-level data worldwide. We discuss potential applications for researchers and potential extensions of the database.
The quality and representativeness of longitudinal datasets play a central role in historical migration research. In this study, we apply the child-ladder (CL) method to a population-scale family tree dataset to analyze U.S. interstate family migration from 1850 to 1920. The CL method infers moves from changes in birthplaces between successive children, allowing for more precise dating of migration events. However, it is limited to families with at least two children. To evaluate the representativeness and utility of family trees for migration research, we compare the CL data to the IPUMS Multigenerational Longitudinal Panel (MLP), which tracks household moves across census decades and serves as a proxy for broader population migration. The CL data reveal higher migration rates, suggesting a likely closer approximation to migration levels in the overall population. Also, by capturing intercensal and return migrations, the CL method provide a detailed view of migration patterns across space and time. Despite differences in migration rates, both datasets reveal similar regional migration structures, especially in the earlier periods. These findings show that population-scale family trees when analyzed using the CL method, offer a valuable complement to linked census data by enhancing our understanding of long-term U.S. migration patterns and regional divisions.
Artificial Intelligence (AI) is rapidly transforming all scientific disciplines. Among its many applications, AI can facilitate data retrieval from a wide range of sources. We evaluate the performance of large language models in extracting data from local heritage books-valuable sources for economic and demographic history. We compare the results of AI-driven, Python code-based, and manual data retrieval for random samples of observations from three heritage books. Our analysis shows that Python code-based retrieval consistently outperforms AI, particularly in minimizing issues such as omitted or hallucinated data. Furthermore, we show that, with minor modifications our Python code-based methods can be adapted to other local heritage books, highlighting the robustness and scalability of this traditional approach.
IPUMS recently released final versions of full count census data for the United States 1900-1930. The information contained in these files is the product of three broad work stages: historical census enumeration, digitization, and IPUMS processing. The data were produced within an evolving institutional context and subjected to subsequent processes that had important ramifications on the final product. This paper documents these histories and processes and their implications for research. Because of the datasets' sheer size and scale, the development of these files necessitated applying different methods and approaches to assess data quality and correct the data. We document cases where data quality was affected not only by choices made by the Census historically, but also by data transcription errors in the modern day. Finally, we describe our approaches to processing the data, and we note some of the implications for research these various decisions have. As with any dataset, researchers should use this resource critically for their particular research questions and consider the data creation process from respondent to digital dataset. Despite some limitations and liabilities, the IPUMS full count data provides a powerful and valuable resource to study demographic effects on a variety of health and socioeconomic questions.