In the last 65 years several major historical databases with reconstructed life courses of large populations have been launched. Around 1990, we could find two important types of databases with longitudinal micro-data. The first type were event databases aimed on family reconstructions and usually based on baptism, marriage and funeral registers or on civil certificates introduced after 1800. The second type were databases with life courses: persons are observed on a more permanent basis using church examination or population registers. After 1990 a third type, databases with census data, really took off. In first instance in the form of samples, in second instance by entering full count samples which makes it possible to link the several censuses into one system, creating semi-longitudinal databases. Another development was the growth of special purpose samples into semi-longitudinal ones by following sampled persons from one source during their life course through linking with all kind of other sources. The development of these databases is indicative of considerable investments that have greatly expanded the possibilities for new research within the fields of history, demography, sociology, as well as other disciplines. In this paper I will compare 84 of these databases on several key figures like included sources, year of foundation, period of observation, area of observation, sample fraction and number of included observations, families and unique persons. An overview of all databases with all key figures is presented in the Appendix.
The Historical Sample of the population of the Netherlands (HSN) is a database containing life histories from the 19th and 20th centuries. In total, around 85,600 individuals have been included, starting with the birth certificate. A distinctive feature of the civil registry is the role of declarants and witnesses in official records. In this article, I aim to provide greater insight into the nature and status of declarants in death certificates.
LINKS stands for 'LINKing System for historical family reconstruction' and is a software system to link nominal data from the Dutch archives and ultimately reconstruct historical individuals and families. We present the background and philosophy of this matching system and explain its data structure and functioning. Currently the core data of the LINKS system consists of indexed civil certificates. These certificates are available from 1812 — the start of the Dutch Vital Registration — until the year they are confidential based on privacy laws. For more than 20 years, thousands of volunteers have been working to build this index, which contains not only the names of newborn, married and deceased persons, but also the names of their parents, places of birth, ages and sometimes their occupational titles. The software system LINKS includes the standardization of all input before linking, nominal record linkage procedures and identification of all unique persons involved in the system. All processes are repeatable and a strict distinction is maintained between source data, standardized, linked and enriched data and released data. Moreover, LINKS also informs archives about all kinds of errors and inconsistencies found during the cleaning and matching process. We will discuss two matching systems, the first is the original querying system that runs within a MySQL database environment and the second is a newly developed system, called burgerLinker, which is based on knowledge graphs and which is designed as a system that can be used independently from LINKS and is made available as open source software. Finally, we present the most important releases of LINKS data so far: two national releases that link birth and parental marriage certificates, creating families and pedigrees and an integrated dataset of persons, families and family trees in four provinces.
In recent years the development of historical databases reconstructing the lives of large populations accelerated. These considerable investments of time and money have greatly expanded possibilities for new research in history, demography, sociology, economics, and other disciplines. This special issue describes the content and design of 23 important historical databases. Authors were given the freedom to discuss a range of practical and technical decisions from evaluating archival sources to crowdsourcing data entry. The most common issue is nominative record linkage, but we find different choices between semi-automatic and fully automatic linkage techniques and various approaches for connecting diverse sources. Some databases describe special problems, like linking Chinese names, handwritten text recognition or the construction of a release in IDS-format. Other databases offer detailed descriptions of sources or discuss prospects for including new datasets.
Over the last 60 years several major historical databases with reconstructed life courses of large populations have been launched. The development of these databases is indicative of considerable investments that have greatly expanded the possibilities for new research within the fields of history, demography, sociology, as well as other disciplines. In this volume spanning seven articles, eight databases are included that had a wide impact on research in various disciplines. Each database had its own unique genesis that is well described in the articles assembled in this volume. They inform readers about how these databases have changed the course of research in historical demography and related disciplines, how settled findings were challenged or confirmed, and how innovative investigations were launched and implemented. In the end we explore how research with this kind of databases will develop in future.
Loci associated with longevity are likely to harbor genes coding for key players of molecular pathways involved in a lifelong decreased mortality and decreased/compressed morbidity. However, identifying such loci is challenging. One of the most plausible reasons is the uncertainty in defining long-lived cases with the heritable longevity trait among long-living phenocopies. To avoid phenocopies, family selection scores have been constructed, but these have not yet been adopted as state of the art in longevity research. Here, we aim to identify individuals with the heritable longevity trait by using current insights and a novel family score based on these insights. We use a unique dataset connecting living study participants to their deceased ancestors covering 37,825 persons from 1,326 five-generational families, living between 1788 and 2019. Our main finding suggests that longevity is transmitted for at least two subsequent generations only when at least 20% of all relatives are long-lived. This proves the importance of family data to avoid phenocopies in genetic studies.
The HSN was initiated during the period 1987–1989 when an interdisciplinary and interuniversity group of Dutch scholars started discussing the foundation of one large database with data on individuals. Building one general prospective database with multiple research possibilities was considered as the only way to realize a cost-effective and properly documented tool for historical research from economic, social, demographic, epidemiological and geographic perspective. The birth registration was considered the most adequate sample framework. The new database should be 'open' in the sense that extension should be possible in all kinds of ways: more sources or variables, more persons and larger time periods. The HSN was deliberately created as a nationwide sample covering the whole 19th and 20th century. Since 1991 about 12 million Euro has been invested in the database and related projects. Besides the basic sample about 25 additional projects have been realized that created all kind of extensions to the database. A special project is LINKS by which the indices of names from the Dutch civil registration are used to reconstruct pedigrees (for the period 1780–1940) and complete families (1811–1900) for the whole of the Netherlands or parts of it. In this article we will present an overview of the research that was done with the original themes and the new fields that were introduced over the years. We will also go into methodological issues that were picked up by the 'HSN community' and we will point out the present and future challenges for the HSN.
Studies have shown that long-lived individuals seem to pass their survival advantage on to their offspring. Offspring of long-lived parents had a lifelong survival advantage over individuals without long-lived parents, making them more likely to become long-lived themselves. We test whether the survival advantage enjoyed by offspring of long-lived individuals is explained by environmental factors. 101,577 individuals from 16,905 families in the 1812-1886 Zeeland cohort were followed over time. To prevent that certain families were overrepresented in our data, disjoint family trees were selected. Offspring was included if the age at death of both parents was known. Our analyses show that multiple familial resources are associated with survival within the first 5 years of life, with stronger maternal than paternal effects. However, between ages 5 and 100 both parents contribute equally to offspring’s survival chances. After age 5, offspring of long-lived fathers and long-lived mothers had a 16-19% lower chance of dying at any given point in time than individuals without long-lived parents. This survival advantage is most likely genetic in nature, as it could not be explained by other, tested familial resources and is transmitted equally by fathers and mothers.
It remains unknown how different types of sources affect the reconstruction of life courses and families in large-scale databases increasingly common in demographic research. Here, we compare family and life-course reconstructions for 495 individuals simultaneously present in two well-known Dutch data sets: LINKS, based on the Zeeland province's full-population vital event registration data (passive registration), and the Historical Sample of the Netherlands (HSN), based on a national sample of birth certificates, with follow-up of individuals in population registers (active registration). We compare indicators of fertility, marriage, mortality, and occupational status, and conclude that reconstructions in the HSN and LINKS reflect each other well: LINKS provides more complete information on siblings and parents, whereas the HSN provides more complete life-course information. We conclude that life-course and family reconstructions based on linked passive registration of individuals constitute a reliable alternative to reconstructions based on active registration, if case selection is carefully considered.
During the 19th and early 20th century about 220,000 Dutch born persons migrated to the USA. The Historical Sample of the Netherlands (HSN) contains about 85,500 persons born in the Netherlands between 1812 and 1922. In this article we report the way we have matched persons from the HSN with the American censuses from the period 1850 till 1940. For this purpose, a linking process was designed, comprising of three stages: harmonization, matching and validation. The different nature of the two datasets (HSN and the USA Censuses) asked for some harmonization prior to the matching. Once the data had been properly prepared, two strategies were applied in order to link the data sets. The first one, called Similarity Approach, matched individuals from both datasets by comparing on the basis of resemblance of first and last names. The second approach, called Transformation Approach, made use of dictionaries with Anglicized versions of Dutch first and last names and their most common or most likely Dutch original(s). Because of the sample character of the HSN even exact matches showed ambiguity that needs to be resolved. For this reason, a validation process comparing the household context was run to provide a more trustworthy result. In the end we identified 484 individuals present in the HSN database with reliable links to the American censuses. We also evaluated the result in the light of what we know from emigration patterns to the USA over time and period and we concluded that our efforts have produced a reasonable result. Nevertheless, we are aware that we may have missed links. We also found that at least 45% of the emigrants returned to the Netherlands at some point during their life course.
For the Netherlands, a rich new data source has become available which contains indexed civil certificates for multiple generations of individuals: LINKS. The current version of the dataset contains information on 1.7 million demographic events for the province of Zeeland in the 19th and early 20th centuries and will be extended to other provinces in the Netherlands in the near future. To be able to study demographic behaviour, life courses and family relations need to be reconstructed from the civil certificates. This paper describes the steps that are taken to move from the LINKS database, which contains digitised birth, marriage, and death certificates and relational information between individuals on these certificates, to LINKS-gen, which contains over six hundred thousand life courses, family reconstructions for up to seven generations, and fertility, marital, mortality, and occupational status information, ready for analysis. We present procedures for variable construction and data cleaning. Furthermore, we give a short overview of the LINKS database, discuss quality checks, and give advice on selection of relevant cases necessary to move from LINKS to LINKS-gen. The paper is accompanied by R-scripts to convert and construct the datafiles.
Survival to extreme ages clusters within families. However, identifying genetic loci conferring longevity and low morbidity in such longevous families is challenging. There is debate concerning the survival percentile that best isolates the genetic component in longevity. Here, we use three-generational mortality data from two large datasets, UPDB (US) and LINKS (Netherlands). We study 20,360 unselected families containing index persons, their parents, siblings, spouses, and children, comprising 314,819 individuals. Our analyses provide strong evidence that longevity is transmitted as a quantitative genetic trait among survivors up to the top 10% of their birth cohort. We subsequently show a survival advantage, mounting to 31%, for individuals with top 10% surviving first and second-degree relatives in both databases and across generations, even in the presence of non-longevous parents. To guide future genetic studies, we suggest to base case selection on top 10% survivors of their birth cohort with equally long-lived family members.
In demographic research large-scale individual-level data have become increasingly available. At the same time, it remains unknown how varying sources affect the reconstruction of individual life courses and families in databases. In this paper, we conduct individual-level comparisons of family and life course reconstructions of 495 individuals simultaneously present in two well-known Dutch datasets: LINKS-Zeeland and the HSN. The first dataset is based on a province’s full population vital event registration data; the other is based on a national sample of birth certificates, after which individuals were followed in population registers. We compare indicators of fertility, marriage, mortality, and measurements of occupational status of individuals found in both databases and conclude that reconstructions in both the HSN and LINKS reflect each other well. LINKS provides more complete family information on siblings and parents, whereas the HSN provides more complete life course information, especially for individuals who migrate out of Zeeland.
Survival to extreme ages clusters within families. However, identifying genetic loci conferring longevity and low morbidity in such longevous families is challenging. There is debate concerning the survival percentile that best isolates the genetic component in longevity. Here, we use three-generational mortality data from two large datasets, UPDB (US) and LINKS (Netherlands). We studied 21,046 unselected families containing index persons, their parents, siblings, spouses, and children, comprising 321,687 individuals. Our analyses provide strong evidence that longevity is transmitted as a quantitative genetic trait among survivors up to the top 10% of their birth cohort. We subsequently showed a survival advantage, mounting to 31%, for individuals with top 10% surviving first and second-degree relatives in both databases and across generations, even in the presence of non-longevous parents. To guide future genetic studies, we suggest to base case selection on top 10% survivors of their birth cohort with equally long-lived family members.
The burden of infant mortality is not shared equally by all families, but clusters in high risk families. As yet, it remains unclear why some families experience more infant deaths than other families. Earlier research has shown that the risk of early death among infants may at least partially be transmitted from grandmothers to mothers. In this paper, we focus on the intergenerational transmission of mortality clustering in the Netherlands in the province of Zeeland between 1833 and 1912, using LINKS Zeeland, a dataset containing family reconstitutions based on civil certificates of birth, marriage and death. We assess whether intergenerational transmission of mortality clustering occurred in Zeeland, and if so, whether it can be explained on the basis of the demographic characteristics of the families in which the infants were born. In addition, we explore the opportunities for comparative research using the Intermediate Data Structure (IDS). We find that mortality clustering is indeed transmitted from grandmothers to mothers, and that the socioeconomic status of the family, the survival of mothers and fathers, and the demographic characteristics of the family affected infant survival. However, they explain the heterogeneity in infant mortality at the level of the mother only partially.
Historical censuses are the richest source of statistical data about a nation’s past. These censuses include valuable data about the population characteristics, housing information and socio-economic data. Not only are they taken consistently over time, in our case for almost two centuries, they also cover the entire nation geographically. The challenges of using these data for studies over time and space are well known, documented and shared by many projects focusing on the harmonization of historical census data. Napoleonic influences during the Batavian Republic were responsible for the first versions of the historical Dutch censuses. Starting with a general enumeration in 1795, over thirty years later the first official census was introduced in Netherlands and continued until 1971 in its traditional ‘door to door’ form. To make this valuable source of data better accessible and available for study, digitization effort where undertaken which resulted in the creation of thousands of scans representing the original census books. These images were later transcribed to Excel tables making them computer processable, but not solving the problem of harmonization. To solve this problem we apply Semantic Web technologies, more specifically the Resource Description Framework (RDF) and propose a specific (three tier) model to harmonize the Dutch historical censuses. We convert the data in RDF in such a way that we preserve all the peculiarities and hierarchies of the original tables. In other words, we provide a one to one representations of the Excel tables in RDF. Although now available and in one system (instead of having over two thousands heterogeneous Excel files) the data still has to be harmonized in order to allow comparisons across time and space. We acknowledge that even when using machine readable formats and Semantic Web technologies, the data still has to be formally defined by ‘expert users’, i.e. the historians working with the data. However, even for these users this is not a straightforward task. As we cannot claim to know upfront the ‘best’ harmonization model we recognize the need for a flexible approach which allow the users to build and test their harmonizations in a very efficient process, especially when dealing with only aggregated data. Current literature does not provide enough insights in the practice of data harmonization when dealing with aggregated census data. We have created a multilevel ‘flexible’ workflow, consisting of a set of specific harmonization practices. We do all of these transformation in RDF while not compromising the underlying sources. Being able to provide direct links from the harmonized results to the original Excel tables and images is a key requirement of our model.