
Access to high-resolution long-term earth system projections is essential for advancing research in hydrology, agriculture, disaster management, and adaptation planning. This Data Descriptor presents an Artificial Intelligence (AI)-refined ensemble of Earth system model (ESM) projections for the conterminous United States at 1/24° (~4 km) spatial resolution. The dataset was generated by downscaling ten Coupled Models Intercomparison Project phase 6 (CMIP6) Earth system models (ESMs) under two emissions scenarios (SSP245 and SSP585) for an 80-year period (1980–2059). Two AI-driven methods, Super-Resolution Convolutional Neural Networks (SRCNN) and Super-Resolution Generative Adversarial Networks (SRGAN), were applied to produce high-resolution daily precipitation and minimum/maximum temperature. This descriptor documents the input datasets, processing workflow, bias-correction steps, file structure, variables, spatial and temporal coverage, and technical validation of the released data. Validation analyses compare the generated products with Daymet observations and existing downscaled datasets to characterize spatial patterns, biases, temporal consistency, and differences among products. The dataset provides daily high-resolution historical and future earth system projections that can support regional, impact assessment, and related applications, and details the AI downscaling framework, training protocols, and evaluation strategies to ensure reproducibility and usability.
The whiteleg shrimp (Litopenaeus vannamei) is one of the most economically significant species in global aquaculture. We present a high-quality, chromosome-level genome assembly for L. vannamei, generated by integrating PacBio HiFi, Nanopore ultra-long, and Hi-C sequencing data. The final assembly spans 1.95 Gb (contig N50: 1.97 Mb), with approximately 97.25% of sequences anchored to 44 pseudochromosomes. Repetitive sequences constitute approximately 1.195 Gb, representing 71.95% of the genome. We annotated a total of 26,830 protein-coding genes, 23,952 of which were functionally characterized. The assembly exhibits exceptional completeness, achieving BUSCO scores of 70.98% (genome mode) and 69.80% (protein mode). This reference genome will serve as a fundamental resource for future genetic research on key economic traits and molecular breeding applications in L. vannamei.
China’s rapid urbanization has led to a substantial increase in material stocks, particularly in road construction with cement being a major component. Accurate knowledge of road material stocks and their distribution are essential for constructing material stocks and flows accounts, and promoting sustainable resource management. However, existing estimates mainly focus on the use of multiple materials in particular case study areas or the use of a limited number of materials in certain road types, failing to form a comprehensive understanding of different road construction materials across China. This study presents China’s Roads Network and Material Stocks Database (CRANMS), a high-precision road network database integrating road structure and material stocks data. With the help of multiple data sources and GIS techniques, CRANMS provides detailed spatial data on China's road network, including material use, length, and administrative classification, and location. The dataset was validated through comparisons with national statistical data and existing literature. The database covers approximately 97% of China’s total road length relative to official statistical records and provides road-segment-level estimates of construction input material quantities, enabling users to aggregate material stocks by road class, province, or material category. The dataset is intended for use in material stocks accounting, infrastructure-related material flow analysis, and spatial comparison of road material stocks.
We present a harmonised dataset of soil morphological, physicochemical, and nutrient biogeochemistry for riparian zones across northern Europe. The dataset comprises 262 soil samples from 54 positions across 12 riparian observatories spanning boreal and temperate climates in Scotland, England, Germany and Sweden. The data collected at each site includes soil profile descriptions, bulk density, pH, total and water-extractable and microbial biomass Carbon (C), Nitrogen (N), Phosphorus (P), and selective Fe/Al oxide extractions. To address the limited availability of coordinated, cross-climate data on riparian soil properties, we present one of the first interdisciplinary efforts to characterise riparian soils both in situ and through comprehensive laboratory analyses. Riparian soils are often spatially heterogeneous and poorly mapped due to their mosaic structure. In this study they were systematically described to capture variability in hydromorphic and biogeochemical properties across landscape positions. The dataset enables cross-climate comparison of riparian function and sensitivity to environmental change, providing a valuable foundation for modelling nutrient retention, release, and transformations at the land-water interface. The data are openly available via the Environmental Information Data Centre.
Tau pathology is a defining feature of several neurodegenerative disorders, but reusable proteomic resources from non-human primate tauopathy models remain limited. Here we describe a four-dimensional label-free quantitative proteomic dataset from cortex and spinal cord tissues of wild-type and Tau-P301L transgenic cynomolgus macaques. The dataset comprises 16 neural tissue samples from four wild-type and four Tau-P301L animals, with paired cortical and spinal cord specimens collected from each monkey. Samples were analysed by LC-MS/MS on a Bruker timsTOF Pro platform, and label-free quantification was performed with MaxQuant against a Macaca fascicularis UniProt protein database. In total, 6,632 protein-group entries were identified, of which 6,520 protein groups were retained in the analysis-ready matrix for quality control and reuse. The accompanying records include raw mass spectrometry files, processed expression matrices, sample metadata and analysis scripts. This Data Descriptor provides a reusable non-human primate proteomic resource for tissue-resolved analysis and comparison of tau-associated molecular features.
Comprehensive, high-resolution, and accurate datasets are critical for advancing COVID-19 research. However, most early pandemic datasets were limited to demographic data and sparse clinical information, often lacking diagnostic imaging tests. We present COVID Data for Shared Learning (CDSL), a multimodal publicly accessible database of patients hospitalized with COVID-19 in Spain. This dataset includes de-identified medical data from 4,479 hospitalization episodes, represented by distinct patient_id values, covering electronic health records with demographics, admissions, diagnoses, clinical and treatment information, and radiological images, specifically 4,608 X-ray images from 1,616 episodes and 1,440,156 computed tomography slices from 784 episodes. Comprehensive anonymization techniques were implemented by removing protected health information to ensure patient privacy, and extensive technical validation was conducted to confirm data completeness, internal consistency, and image conversion quality. Open-source code for data preprocessing, metadata extraction, and database deployment is publicly available to support reproducibility and community-driven improvement. Early in the pandemic, CDSL was shared to support timely COVID-19 research and became one of the earliest publicly accessible multimodal COVID-19 datasets available to the international research community. CDSL aims to promote global collaboration by sharing clinical data to advance healthcare research and drive innovation in personalized medicine and public health.
Air pollution poses significant risks to public health and the environment, highlighting the need for long-term, spatially comprehensive air quality datasets. Here, we present a large-scale ground-based air pollutant dataset comprising 4,453,372 daily records from 2,040 monitoring stations across mainland China during 2015-2020. The dataset includes concentrations of six major air pollutants: ozone (O3), carbon monoxide (CO), nitrogen dioxide (NO2), sulfur dioxide (SO2), and particulate matter (PM2.5 and PM10). To address missing observations in long-term monitoring records, we applied a LightGBM-based imputation framework that integrates model selection, feature enrichment, and iterative imputation to leverage spatiotemporal dependencies and inter-pollutant relationships. Each record is enriched with 128 features, including meteorological variables, remote sensing products, geographic attributes, anthropogenic indicators, and temporal descriptors.Comprehensive technical validation demonstrates that the reconstructed dataset preserves realistic temporal dynamics, spatial patterns, statistical distributions, and inter-pollutant correlation structures without introducing systematic biases. Benchmark experiments further illustrate the dataset’s usability for air pollution prediction tasks. This dataset is designed with machine-learning readiness, providing a comprehensive, high-quality, and standardized resource with spatiotemporally explicit data to advance air quality research, develop and benchmark machine learning models, conduct exposure assessment, and enable data-driven environmental management across China.
Structural remodeling of the vascular wall is a universal response to chronic damage, pressure overload or flow overload. Therefore, this process occurs in a wide range of diseases affecting both the systemic and pulmonary circulations. Vascular wall fibrosis is one of the most important factors in the pathogenesis of various types of pulmonary hypertension. In this regard quantitative assessment of collagen deposition in the vascular wall is crucial for both preclinical studies and directly affects the life prognosis for patients with pulmonary hypertension. Also, histological analysis is the gold standard for assessing vascular structural remodeling. This dataset includes 705 micrographs of vascular wall region with varying degrees of vascular wall fibrosis that developed in various types of pulmonary hypertension. All included studies were devoted to the investigation of pulmonary arterial hypertension and chronic thromboembolic pulmonary hypertension in male Wistar rats. Lung sections were stained with Picro-Mallory and digitized using a whole slide scanner. Pulmonary artery branches were subdivided into vessel segments, and exported as fixed-size micrographs. For each micrograph, the dataset provides a region-of-interest mask and two independent expert annotations, including binary masks of the vascular wall and intramural fibrosis zones. The presented database will be useful for the development of new software solutions for the analysis of histological micrographs.
People at different ages and with different educational background have different migration patterns. However, international migration flows by age and educational attainment are scarce and not always comparable. We use a super learning algorithm with multiple initial models to estimate age and education composition of male and female immigrants for 199 countries over six five-year periods from 1990-1995 to 2015-2020. We access the performance of our approach through cross-validation and comparing initial learners with the super learner algorithm using multiple evaluation metrics. We also compare our age composition estimates against equivalent measures from Eurostat. The estimates indicate that while the proportion of immigration flows with higher education increase in all regions, Europe has the highest immigration flow with secondary or higher education.
This paper introduces DevEmo+, a new dataset of facial expression recordings from students engaged in programming tasks. Collected “in-the-wild” from participants’ personal computing environments, this dataset addresses the critical need for ecologically valid data to research student affective states during complex computer-based learning. The data was produced using a three-phase approach that integrates automatic emotion recognition, crowdsourcing for human based filtering, and a final selection stage by expert annotators to ensure high-quality labels. The final dataset contains 304 video clips from 51 participants, balanced between 152 emotional and 152 neutral expressions. A consensus protocol was applied, requiring agreement from at least two of three experts on both the emotion category and timing. The resulting annotations are dominated by cognitive states such as confusion (51.9%), happiness (22.4%), and surprise (16.4%), making the dataset particularly well-suited for investigating the cognitive-affective dynamics of problem-solving. The publicly available dataset includes detailed metadata to facilitate its use.
High-quality multimodal datasets are essential for developing vision-language models, yet publicly available figure-text resources in specialized scientific domains remain limited. To address this gap, we present a large-scale figure-text pair dataset constructed from map-related scientific literature (FTPD-ML). The dataset was derived from 96,859 scientific publications in cartography, geography, remote sensing, and related disciplines, and provides 75,702 publicly redistributable figure-text pairs under Creative Commons Attribution (CC BY) licenses with multiple levels of textual representations, including original figure captions, standardized caption variants, and contextual paragraph descriptions, together with bibliographic metadata. To ensure compliance with copyright and licensing requirements, records from non-redistributable publications are represented by metadata only. The dataset supports a range of multimodal research tasks, including image-text retrieval, image captioning, scientific document understanding, and cartography-oriented vision-language studies.
High quality data on natural hazard damages are crucial for effective disaster risk management. Yet, existing impact datasets remain limited and often biased toward Northern countries and monetary losses. To help address these gaps, we present ROUGE (Redcross Operations Unified Global Emergency database); a new socio-economic impact database obtained using textual operational reports from the International Federation of Red Cross and Red Crescent Societies. These reports are systematically collected and provide broad coverage of regions that are commonly underrepresented in existing impact datasets. Using large language models, we extract qualitative and quantitative information on a wide range of non-monetary impacts at national and sub-national scales. The resulting dataset documents socio-economic impacts of natural hazards on the population, infrastructure and economy with a spatial detail reaching the subregional level. This resource is designed to support research and applications that require geographically explicit information on socio-economic impacts of disasters, enabling more precise and inclusive analyses of socio-economic consequences of natural hazards worldwide.
Land surface temperature (LST) is a fundamental variable in hydrological and climate research. However, cloud contamination and sparse temporal sampling in satellite thermal infrared observations cause extensive data gaps, hindering generation of continuous long-term high-resolution daily mean LSTs. Although global gap-filled and reanalysis-based LST products already exist, they often face one of the challenges: coarse spatial resolutions, monthly mean/daily instantaneous scales, or struggling to balance accuracy and spatial texture. To address these challenges, we integrate empirical, semi-physical, and machine learning approaches to generate gapless 1 km daily mean LSTs for 70°N–60°S and 2003–2023. We implement a pixelwise, daily-mean–then–annual-cycle, preliminary-to-optimized workflow that harmonizes MODIS observations with all-weather in situ measurements. Validation demonstrates progressive accuracy improvements across steps, with final RMSEs of 2.073 K (daily) and 1.511 K (monthly) against independent ground observations. Comparisons with existing gap-filled products and reanalysis datasets demonstrate improved statistical accuracy—particularly over cropland, grassland, and cold regions—as well as enhanced seasonal robustness and spatial texture, particularly over rugged terrain and urban areas.
Fish escapes from aquaculture facilities pose significant risks, including genetic introgression or competition with wild populations. Accurately identifying the origin of fish, whether wild, farmed, or escaped, is therefore essential for impact assessment, traceability, and sustainable aquaculture management. In this work, we present GLORiA, a curated image dataset specifically designed for fish origin identification. The main dataset contains 9,511 high-quality images of three economically and ecologically important Mediterranean species: Sparus aurata, Dicentrarchus labrax, and Argyrosomus regius. These images correspond to repeated acquisitions of individual specimens under standardized laboratory conditions, rather than to 9,511 independent biological samples. In total, the laboratory dataset includes 423 unique S. aurata specimens, 309 unique D. labrax specimens, and 108 unique A. regius specimens. All images were acquired following a consistent imaging protocol and were annotated by domain experts according to their origin. In addition to the laboratory dataset, GLORiA includes an initial test subset acquired under more realistic commercial conditions, comprising fish-market tray images and cropped individual specimens, which is intended to support evaluation beyond standardized laboratory settings. Unlike existing fish image datasets, which primarily focus on species recognition, GLORiA explicitly targets origin classification, enabling the study of morphological and visual traits associated with aquaculture practices. The dataset is systematically organized by species and origin category and includes preprocessing scripts to support reproducibility. GLORiA addresses a critical gap in publicly available resources and provides a foundation for the development and evaluation of computer vision methods aimed at supporting sustainable aquaculture and mitigating the ecological impact of fish escapes.
The striped flea beetle (SFB), Phyllotreta striolata, is a globally devastating cruciferous pest. A key challenge to its control lies in its complex life cycle—with egg, larval, and pupal stages confined to soil, while adults infest aerial plant parts—coupled with distinct antennal differences between males and females that underpin sex-specific ecological traits. Transcriptomic profiling across developmental stages offers a powerful approach to decode the genetic basis of these stage- and sex-related traits and inform targeted pest management. Here, we present the first comprehensive transcriptomic atlas of P. striolata, encompassing four key life stages (egg, larva, pupa, adult) with separate male and female. RNA sequencing of 15 samples (biological triplicates per group) generated 334.5 million raw reads. Dataset reliability was robust, as validated by correlation analysis, principal component analysis, and qRT-PCR. The deposited threshold-based gene lists may support future investigations of developmental-stage- and adult-sex-associated gene expression in P. striolata. These data provide valuable insights into the biology and developmental dynamics of P. striolata and establish a robust foundation for functional genetic studies. Moreover, the P. striolata transcriptome can be uploaded to the dsRIP platform, enabling systematic identification of effective RNAi targets and optimization of species-specific dsRNA designs with minimized off-target risks. Together, these resources may accelerate the translation of molecular insights into effective and sustainable strategies for controlling this pest.
This work presents a Laser-induced Breakdown Spectroscopy (LIBS) dataset of regolith simulant samples and reference XRF composition data from the manufacturer. 4 Mars and 6 Moon regolith simulants were systematically mixed to have a set of 21 samples for each simulant category. The samples primarily contain Al2O3, SiO2, MgO, CaO, Fe2O3 or FeO, K2O, TiO2, Na2O and P2O5. All the samples were measured under three different atmospheric conditions: (1) Earth’s ambient gas; (2) Mars ambient gas, 7 mbar pressure and CO2 atmosphere; (3) low vacuum simulating Moon conditions, 0.1 mbar. The dataset’s potential use lies in two key areas. The first is quantitative analysis, addressing challenges in analyzing a broad composition range of primary oxides. Enhancing concentration prediction accuracy for Al2O3 and SiO2 is particularly important, as their high concentrations and resonant spectral lines complicate analysis, according to current literature. The second area is sample classification. Both tasks are limited in the size of the dataset, simulating the data-throughput of the rovers for in-situ analysis. The dataset aims to support LIBS performance improvements under such restrictions.
Tobacco (Nicotiana tabacum) is both a major industrial crop and a foundational model organism for plant biology. While alternative splicing significantly diversifies transcriptome complexity, tobacco isoform annotation remains incomplete, hampered by the limitations of previous short-read sequencing efforts. To address this gap, we generated a long-read transcriptome dataset across five tobacco tissues using the PacBio Sequel IIe platform, producing 198.4 million subreads. From these data, we successfully reconstructed 64,260 isoforms, including 35,013 novel transcripts. For the novel isoforms, we also performed systematic examinations of their structural features, biotypes and expression profiles. This dataset expands the tobacco isoform atlas and provides a valuable resource for genome annotation and regulatory studies.
This dataset contains longitudinal clinical data from 68 patients hospitalized with a clinical diagnosis of snakebite at a hospital in Gansu Province, China, over a 3-year period. The dataset contains 10,361 time-stamped laboratory test results, including coagulation tests, clinical chemistry tests, complete blood counts, and other laboratory measurements, together with detailed time-stamped records of antivenom administration and supportive treatments. To protect patient privacy while preserving relative temporal relationships, all absolute dates were randomly shifted to future calendar years. The dataset also includes derived Blaylock local effect scores. Laboratory test items were mapped to the corresponding LOINC codes using the Logical Observation Identifiers Names and Codes (LOINC) v2.73 reference table. The dataset is accompanied by a comprehensive data dictionary and data-validation code and can be used for trajectory clustering, exploratory analyses of antivenom use patterns, and research on early warning of organ dysfunction. We plan to update the dataset annually, with the goal of covering all patients hospitalized with a clinical diagnosis of snakebite at the study hospital over a 10-year period; version 1.0 includes cases from 2023 to 2025.
Causal machine learning represents a promising paradigm for estimating intervention effects in operations management, yet its adoption is constrained by the scarcity of structured datasets suitable for method development and validation. This paper presents a synthetic production dataset comprising 10,000 production order records from a simulated metalworking factory with ten welding lines and ten product types. The dataset incorporates a binary treatment variable representing the implementation of drum-buffer-rope synchronization from the Theory of Constraints on two production lines. Data generation follows established statistical distributions with explicit causal structure, including line-product interaction effects, Little’s Law relationships, and treatment-dependent outcome variations. The dataset includes 15 variables capturing order characteristics, operational parameters, performance outcomes, and a line-level carry-over state that induces temporal serial correlation between successive orders on the same line. Primary reuse value lies in benchmarking causal machine learning methods, particularly meta-learners, tree-based algorithms, and neural network–based approaches for heterogeneous treatment effect estimation in operations management contexts. The dataset supports methodological research on ex-ante effect estimation without requiring explicit simulation models of production systems.
Marine microbes drive global biogeochemical cycles, yet most remain recalcitrant to cultivation and poorly characterized. Microbial community structure varies markedly along environmental gradients, particularly between oligotrophic open-ocean and nutrient-rich coastal regions. However, metagenomic insights across such natural gradients remain limited, particularly in the western Pacific Ocean. Here, we applied genome-resolved metagenomics to seven surface seawater samples collected along an open-ocean–to–coastal transect in the western Pacific Ocean. Sequencing generated 80.16 Gbp of data, enabling the reconstruction of 194 species-level bacterial and archaeal metagenome-assembled genomes (MAGs). Of these, eight met the criteria for high-quality draft genomes and 186 were classified as medium-quality draft genomes according to MIMAG standards ( ≥50% completeness, ≤10% contamination). Both read-based taxonomic profiling and genome-resolved binning revealed differences in bacterial and archaeal community structure among the samples. Pseudomonadota, Bacteroidota, and Marinisomatota were relatively more abundant in the transitional and coastal samples, whereas Cyanobacteriota were more abundant in the open ocean samples. Archaeal phyla were relatively more abundant in the transitional and coastal samples. Collectively, this dataset provides a region-specific, genome-resolved resource for the western Pacific Ocean and enables investigation of microbial community structure across a continuous open-ocean-to-coastal gradient.