As Artificial Intelligence (AI) systems and data-driven tools become integral to governmental decision-making, the ability to interpret and reason with visual information emerges as a critical competence for operating effectively in AI-mediated analytical environments. However, empirical evidence on the level of data visualization literacy within public administrations remains limited. To address this gap, the study provides a large-scale, diagnostic, and descriptive analysis of Data Visualization Literacy (DVL) performance in a real public organizational setting, using a standardized assessment instrument. A cross-sectional survey of 1,219 public employees was conducted using a bilingual Spanish–Valencian adaptation of the Mini-VLAT (12 items; 25 seconds per item), evaluating participants’ capacity to interpret, analyze, and reason with graphical representations of data. Mean performance reached 57.8% correct, with 27.1% omissions and 15.1% errors. Tasks involving proportional or relational reasoning—particularly stacked charts—produced the lowest accuracy and the highest nonresponse. Performance patterns were consistent: accuracy declined with age, improved with higher educational attainment, and varied across departments. Omissions under time pressure, rather than misinterpretation, were the predominant source of error. The findings underscore the importance of treating DVL as part of the institutional infrastructure, through periodic diagnostics, shared graphic-interoperability standards, targeted domain training under time constraints, and longitudinal monitoring to preserve epistemic control while harnessing AI’s speed and scale.
The Gini index is the most widely-used measure of inequality. Unfortunately, its computation is subject to error. Researchers and practitioners often fall into common methodological pitfalls, leading to inaccurate estimates and inferences, and ultimately hindering efforts to reduce inequality and improve societal quality of life. This paper clarifies the challenges of non-parametric estimation of the Gini index more comprehensively than previous contributions, and offers robust methodological recommendations to ensure accurate estimates. Additionally, we reference a free, easy-to-use R package which, together with the clear methodological insights, enhances the real-world applicability of our findings. First, we investigate the impact of common methodological pitfalls on point estimates, providing a complete review for both infinite and finite populations. We then examine variance estimation and the performance of confidence intervals. Among other issues, the findings reveal that, when a popular regression-based variance estimator is used, the variance of the Gini index is seriously underestimated in distributions with high skewness and inequality, as often observed in real-world applications. Jackknife variance estimates and jackknife intervals, based on studentized quantiles, prove to be the most accurate approaches. The analysis employs variables with varying degrees of skewness and inequality (as both characteristics influence the potential for bias), thereby encompassing most of the situations found in empirical research.
The article is about methods for estimating a table of vote transition counts between two elections, ensuring consistency with the observed marginals, given an estimated table of transition probabilities obtained by some method of ecological inference. We argue that count data are essential for conducting in-depth investigations into voting behavior. Several new methods are compared with Iterative Proportional Fitting (IPF), known for its speed and reliability. To evaluate their performance, we use both simulated data and real electoral results from the 2011 New Zealand general election. An algorithm for simulating artificial electoral data according to the modified Brown and Payne model is presented and models that account for strategic voting are applied to the New Zealand data to reduce ecological bias. Among the methods initially considered, only two valid competitors of IPF emerged, constrained maximum likelihood and minimum Chi-square, with the latter performing significantly better than IPF. We illustrate the potential insights which may be gained from tables of transition counts by an application to transitions from a parliamentary and a regional elections in Umbria, Italy.
Longevity and death rates are closely linked to socioeconomic conditions. However, the relationship between income and specific causes of death (CoD) remains insufficiently explored. This study examines socioeconomic inequalities in cause-specific mortality in Spain, simultaneously accounting for CoD, age, sex, and income. To the best of our knowledge, no previous study has jointly combined all these variables. Using data for the entire population residing in Spain from 2010 to 2019, we compute death rates by CoD, age, sex, and income decile by linking individual demographic records with census tract-level disposable income. The analysis covers 4 million deaths and over 466 million person-years at risk. Robust income-based relative risks and gradients by CoD are calculated for five-year age groups, stratified by sex, while Relative Index of Inequality (RII) values are derived for single years of age. Results show strong and consistent income-related inequalities in mortality across age, CoD, and sex, especially among younger and middle-aged groups. Particularly steep and significant gradients are found for circulatory, respiratory, digestive, infectious, endocrine, and genitourinary diseases, as well as for symptoms and abnormal findings, and other conditions, with lower-income individuals facing markedly higher risks. For neoplasms, nervous, mental and behavioural disorders, and external causes, patterns are more heterogeneous and gender-specific. Notably, neoplasms exhibit RIIs < 1 among women aged 55 and older. Overall, the results highlight widespread income-related inequalities in mortality in Spain and underscore the need to prioritise lower-income groups in public health efforts, particularly regarding chronic and behavioural health conditions.
This paper examines the implications of incorporating education as a risk factor into life insurance, focusing on its effects on insurers' income statements. While the relationships between education and mortality have been widely studied in demography, epidemiology, and public health, their relevance for the insurance sector remains largely unexplored. To address this gap, we introduce the Education-based Risk Indices (ERI), which quantify differential mortality risk by educational level across ages and sexes relative to the population average, and analyse how the ERI can inform premium setting and shape competition in a market of rational insureds. The complete approach is illustrated with data from Spain. Individual-level population and mortality microdata from 2021 to 2022 are employed to construct life tables by education and derive the ERI by modelling the relationship between stratified and overall death rates. An application is presented through a simulation of a competitive insurance market in which one company adopts ERI-based pricing while another relies solely on age and sex. The findings indicate that integrating education into pricing strategies yields a measurable competitive advantage, enhancing insurers' financial performance relative to traditional approaches. Robustness analyses and a discussion about issues to be considered in real applications are also included.
We review some early contributions by Corrado Gini to modified binomial models and predictive probability and highlight their role, once extended to the multivariate context, for modelling voting behaviour in Ecological Inference. A collection of overdispersed multinomial models are described, their properties investigated and a connection to Gini's results for the corresponding binomial model, when available, is provided. After a concise introduction to Ecological Inference, we discuss recent developments aiming at more realistic models of voting behaviour and some connections to Gini's work.
Geographical access to medical services is a key determinant of population health, yet evidence on how hospital travel time affects mortality remains limited, particularly in countries with universal health coverage. Using Spain as a case study, this research explores a generalizable framework to quantify the association between hospital travel time and mortality. In particular, it examines the 2010–2019 period, assessing differences by age, sex, and cause of death (CoD), and discussing implications for healthy ageing and equitable healthcare access. We combine population and mortality microdata with a detailed measure of residential hospital accessibility, based on car travel time from census-section centroids to the average of the five nearest hospitals. Analyses are conducted for all-cause mortality and major cause-of-death groups, including circulatory, neoplasm, respiratory, and external causes. Longer travel times are associated with higher mortality, with the strongest effects observed among older adults—particularly those aged 65 years and over, for whom residential location is expected to provide a more informative geographical reference—and for time-sensitive conditions. Circulatory diseases show measurable associations even at relatively short travel times, whereas external causes exhibit steeper gradients at longer distances. Results are robust to sensitivity analyses using the nearest hospital alone. Overall, territorial disparities in residential proximity to hospital services persist in Spain and are associated with inequalities in mortality.
Reconciling multidimensional count data across multiple sources is a common challenge in social and economic research. Iterative proportional fitting is widely used for this purpose, but aligning indices under weighted sum-convex constraints calls for a more fexible approach. We introduce the weighted iterative proportional fitting algorithm, which incorporates sum-weighted constraints to adjust indicators-such as death-risk indices by wealth, habitat, and climate-while preserving marginal consistency. Weighted iterative proportional fitting has been implemented in an R package of the same name, enabling scholars, statisticians, and policymakers' advisors, among others, to apply it easily to multidimensional data.
Background Despite the well-recognized relationship between income and life expectancy, much of the existing evidence is indirect, relying on cross-country comparisons, statistical modeling, analyses beginning at older ages, or data limited to a small number of specific socioeconomic groups. Objective To examine life expectancy at birth across the entire Spanish population using up to one hundred income groups and to analyze its variation by income level during the austerity period following the Great Recession. Methods The study adopts a whole-population microdata ecological approach and employs five-year rolling death rate estimators to construct life tables by income group and sex in Spain from 2012 to 2017. The computations, which involve over 14 billion person-year observations, are based on a dataset comprising 520 million unique demographic events recorded over a ten-year period. Results Life expectancy at birth increases progressively with income percentile, with differences between the poorest and richest groups exceeding 8 years among men and 5 years among women. The gap between the highest- and lowest-income groups widened during the austerity period under study, indicating a growing disparity. Conclusions These findings support the view that economic conditions are closely linked to mortality risk and underscore the importance of social policies. They also suggest that addressing income-related health disparities through targeted interventions and public healthcare spending may reduce inequalities across social groups and yield substantial gains in life expectancy.
Rural depopulation in Spain, reflecting broader European trends, presents a complex challenge with far-reaching economic, social, and ecological consequences. Population decline and ageing, compounded by imbalanced sex ratios characterised by masculinisation, contribute to a shrinking labour force, increased pressure on healthcare systems, and diminished generational renewal. Despite numerous policy attempts over past decades, these trends have persisted. Recently, Spain launched a more ambitious initiative, the Demographic Challenge, aimed at reversing rural depopulation. This study uses the cohort-component model to construct counterfactual demographic projections-unaffected by the new policy-as a benchmark for future evaluation of the age-sex distribution of the population residing in Spanish rural municipalities. Estimates are obtained by instrumentally projecting demographic rates under two alternative frameworks: constant rates and ARIMA-based models. Results from both approaches anticipate continued population decline and ageing in rural areas by 2050, with a progressively narrower base in the age pyramid if current trends are not reversed. These results underscore the urgency of effective policy interventions to address the ongoing demographic crisis in rural Spain.
The ecological inference canonical problem in the social sciences consists of estimating the unobserved internal counts of a global RxC table from the known margins of a set of units. This paper proposes a new, computation-based strategy designed to better exploit the information contained in the unit margins. This approach can be integrated into any ecological inference method that explicitly estimates unit tables and accounts for differences in unit size. We evaluate its performance using as a baseline the fastest ecological inference linear programming method, relying on real electoral data from over 550 datasets where true contingency tables are known. In this extensive assessment, the proposed strategy reduces average global errors by more than 21% relative to the baseline, outperforming it in 95% of cases. It also improves upon nslphom-identified in the literature as the most accurate algorithm for this dataset-reducing average global errors by over 5% and outperforming it in 60% of cases. The versatility of the approach is further illustrated by also integrating it into three more computationally intensive methods, including the two main statistical ecological inference models-ei.MD.bayes and BPF-and nslphom, yielding consistent improvements over their respective baselines in a small set of examples.
Adversities during foetal development and early life are important determinants of health across the life course, as exposures during sensitive developmental stages may lead to negative outcomes later in life. Mortality represents the most extreme consequence of such adversities. In this study, we focus on perinatal mortality, defined as registered deaths occurring from the 22nd week of gestation up to the first 24 hours of life. Understanding the factors that increase the likelihood of extreme events during this critical developmental stage is essential for prevention. We assess the impact of socioeconomic factors-maternal origin, educational attainment, and municipality size of residence-on perinatal mortality, controlling for maternal age. Joint conditional statistical analyses are based on a mixed-effects Poisson regression model for 2007-2022, with random effects accounting for unobserved heterogeneity across calendar years. Among the factors analysed, origin and education emerge as the most relevant. Results indicate that foreign-born mothers exhibit substantially higher perinatal mortality risks, with estimated IRRs ranging from 2.2 (CI: 2.1-2.4) to 6.5 (CI: 5.7-7.2), depending on the mother's continent of origin. Higher maternal education is associated with a reduced risk of perinatal mortality, with a mean relative risk (IRR) of 0.64 (CI: 0.60-0.68) compared to mothers with primary education or less. Maternal age also plays a key role: teenage mothers (≤ 20 years) and mothers of advanced age (≥ 37 years) experience markedly higher perinatal mortality. Overall, this study identifies maternal profiles most at risk and quantifies disparities in perinatal health across socioeconomic characteristics in a country with universal, free healthcare.
Cross-section and longitudinal spatial statistics and econometric models rely on spatially and temporally referenced data. Administrative units like cities, counties, and provinces provide stable data sources, enabling models to combine statistics collected at different times. In Spain, census sections serve as the smallest territorial units in which official statistics are delivered. These areas offer valuable statistics, such as population and housing censuses. Providing these statistics at the postcode level is also pertinent for conducting local analyses and surveys. The issue is that boundaries of census sections undergo regular updates, sometimes involving significant reorganization. The R-package sc2sc automates the transfer of variables between different census sections and postal codes. This paper introduces the package and outlines its methodology, which employs areal weighting to transfer counts and rates.
Purpose Data visualisation literacy (DVL) is increasingly recognised as a key component of data and information literacy in data-intensive learning environments. Existing assessment instruments such as the Visualisation Literacy Assessment Test (VLAT) and Mini-VLAT provide reliable measures of interpretative performance but offer limited insight into the reasoning processes underlying learners’ interpretations of visual data. This study aims to propose a formative assessment framework designed to capture how students interpret, evaluate and justify inferences from visual representations of data. Design/methodology/approach The study develops a reasoning-centred assessment framework that integrates diagnostic testing, structured critique of visualisations and reproducible redesign tasks, organised as a progressive two-phase protocol (Core I and Core II). The diagnostic phase and Core I were implemented in a mixed-methods pilot study in an undergraduate data analytics course (n = 48), while Core II is proposed as a subsequent assessment phase. Students completed a Spanish-adapted Mini-VLAT in pre- and post-test anonymous cohort cross-sections and participated in guided critical-analysis activities based on materials from the critical thinking assessment for literacy in visualisations (CALVI). Findings Mini-VLAT scores remained largely stable across administrations. However, qualitative evidence revealed increasingly explicit criterion-based reasoning, more critical assessment of inferential validity and better justified redesign proposals, suggesting that the piloted component of the framework captures reasoning processes not visible through conventional outcome-based measures. Research limitations/implications The pilot study was conducted in a single course context with a limited sample, and the anonymous cohort design precluded individual-level trajectory analysis. Core II constitutes a designed extension of the framework and awaits empirical validation. Future research should implement Core II and examine the framework across different disciplines and learning environments. Practical implications The framework provides educators with a structured approach to assessing students’ reasoning with visual data representations. Originality/value The study introduces a preliminary, empirically informed assessment framework that conceptualises DVL as a reasoning-centred form of data and information literacy.
Mortality and life expectancy statistics have profound implications across diverse facets of society, including the financial, economic, and actuarial fields. Mortality is shaped by various factors, with income becoming a particularly critical determinant once age and sex have been accounted for. This paper proposes a new methodology for easily integrating income-related (contextual wealth) differential risks into death probabilities and (general and insured) life tables, showcasing its deep impact on financial and insurance products and risk management. The approach is based on estimating general population Income-Indices, whose estimation is illustrated via an extensive dataset composed of 253 million demographic events of the Spanish population. Employing the Income-Indices-based methodology in a competitive market emerges as a strategic advantage, positioning companies favorably compared with those adhering to conventional pricing methodologies.
In the demographic, actuarial, and economic fields, projections and forecasts provide valuable insights into the future evolution of death rates, population growth, and age distributions. Population projections and forecasts result from a comprehensive integration of migration flows, death rates, and birth rates over a given population stock. These figures furnish information on the size, age, and gender compositions of the entire population, as well as specific subgroups defined through conditioning on other variables of interest, such as income. Understanding mortality and population dynamics equips economic and social agents with tools for planning and decision-making across social, insurance, and health dimensions, to name a few. This work employs an extensive dataset comprising 533 million microdata entries from the Spanish population between 2010 and 2019, covering information on population stock, deaths, migrants, and births, with spatial markers at the highest level of territorial disaggregation-census sections. The aim of this study is to project the mortality dynamics of the Spanish population across four income levels, determined after splitting people according to the average income of their census section of residence. We achieve this by applying stochastic mortality projection models to the smoothed realized series, selecting the three established death forecasting models that best align with the statistical properties of our data.
Studying ideological preferences is essential for understanding the social and political changes a society undergoes over time. Generational ideology tables provide a structured framework to analyze ideological shifts within and across generations, offering valuable insights into societal evolution and political behavior. This paper introduces the Spanish Cohort Ideology Database (SCID), detailing the data sources and methodology used for its construction. By leveraging more than five million individual observations (which expand to over 100 million when employing double 5 × 5 moving windows) from over 1,800 surveys conducted since 1977 by the Spanish official center for sociological research (CIS), we construct 1,554 period and cohort ideology tables, including breakdowns by gender, education level, gender and education level, and region. SCID comprises age-year and age-generation tables with mean values, sample sizes, and variances, enabling the analysis of the dispersion/polarization of ideological self-placement. This work facilitates the analysis of social and political change processes from a cross-section and longitudinal perspective, creating a unique database that could also be developed in other countries, thereby enabling international comparative studies.
Purpose Public pension systems in advanced countries are characterised as being generous, as they present high replacement rates and real rates of return (pension-to-contribution ratios adjusted for differences in purchasing power over time) at values greater than one. They are also considered to be progressive, being slightly more in favour of the lower incomes. In this paper, we evaluate the (in)appropriateness of this last statement in the context of Spain, focusing exclusively on contributory benefits. Design/methodology/approach We use a microdata set of the Spanish population composed of 48.5 million entries, disaggregated at the census section level, and calculate real rates of return on contributions based on salary for four income levels. Findings This study shows the inappropriateness of assuming the progressiveness of Spain's public pension system, which arises from the erroneous assumption of independence between income levels and (residual) life expectancy. The results reveal that contributors with higher incomes receive, on average, relatively higher returns. Research limitations/implications The conclusions are true on average. Deviations within and between groups are expected at the individual level. Practical implications These findings can help to better understand the so-called solidarity quota introduced in the latest legislative reform of the pension system in Spain. Social implications The results could contribute to formulating more equitable public policies that consider sociodemographic disparities. Originality/value This is a previously unrevealed result. While the relationship between income levels and longevity is well established, our research examines how these factors specifically influence the redistributive character of contributory pension benefits in Spain, offering a deeper understanding of how these inequalities manifest within a public pension system. Further research could be conducted in other countries to explore the (in)appropriateness of assuming the progressiveness of their public pension systems.
Funeral insurance is an insurance product designed to cover the expenses and administrative procedures associated with a person's death. Its main purpose is to relieve the insured person's family of the financial and bureaucratic burdens arising from the funeral. Despite its widespread presence in the Spanish market-currently, more than 22 million people in Spain have coverage under this type of insurance-academic and scientific literature on funeral insurance remains surprisingly scarce. The aim of this study is to help fill this gap by providing a comprehensive overview of funeral insurance in Spain, based on the analysis of the sociodemographic profiles of policyholders, including their geographic distribution, income levels, and municipality sizes. The study examines a representative sample of 2.1 million policies, using the insured individuals' postal codes of residence. The results reveal an uneven penetration across provinces, income groups, and habitat sizes, with lower rates in urban areas and among higher-income individuals. However, education level, more than income, is the variable with the greatest impact, showing a statistically significant interaction effect with income. The combination of a high level of education and high income reduces the likelihood of holding a funeral insurance policy.