
People at different ages and with different educational background have different migration patterns. However, international migration flows by age and educational attainment are scarce and not always comparable. We use a super learning algorithm with multiple initial models to estimate age and education composition of male and female immigrants for 199 countries over six five-year periods from 1990–1995 to 2015–2020. We access the performance of our approach through cross-validation and comparing initial learners with the super learner algorithm using multiple evaluation metrics. We also compare our age composition estimates against equivalent measures from Eurostat. The estimates indicate that while the proportion of immigration flows with higher education increase in all regions, Europe has the highest immigration flow with secondary or higher education.
This paper introduces DevEmo+, a new dataset of facial expression recordings from students engaged in programming tasks. Collected “in-the-wild” from participants’ personal computing environments, this dataset addresses the critical need for ecologically valid data to research student affective states during complex computer-based learning. The data was produced using a three-phase approach that integrates automatic emotion recognition, crowdsourcing for human based filtering, and a final selection stage by expert annotators to ensure high-quality labels. The final dataset contains 304 video clips from 51 participants, balanced between 152 emotional and 152 neutral expressions. A consensus protocol was applied, requiring agreement from at least two of three experts on both the emotion category and timing. The resulting annotations are dominated by cognitive states such as confusion (51.9%), happiness (22.4%), and surprise (16.4%), making the dataset particularly well-suited for investigating the cognitive-affective dynamics of problem-solving. The publicly available dataset includes detailed metadata to facilitate its use.
High-quality multimodal datasets are essential for developing vision-language models, yet publicly available figure-text resources in specialized scientific domains remain limited. To address this gap, we present a large-scale figure-text pair dataset constructed from map-related scientific literature (FTPD-ML). The dataset was derived from 96,859 scientific publications in cartography, geography, remote sensing, and related disciplines, and provides 75,702 publicly redistributable figure-text pairs under Creative Commons Attribution (CC BY) licenses with multiple levels of textual representations, including original figure captions, standardized caption variants, and contextual paragraph descriptions, together with bibliographic metadata. To ensure compliance with copyright and licensing requirements, records from non-redistributable publications are represented by metadata only. The dataset supports a range of multimodal research tasks, including image-text retrieval, image captioning, scientific document understanding, and cartography-oriented vision-language studies.
High quality data on natural hazard damages are crucial for effective disaster risk management. Yet, existing impact datasets remain limited and often biased toward Northern countries and monetary losses. To help address these gaps, we present ROUGE (Redcross Operations Unified Global Emergency database); a new socio-economic impact database obtained using textual operational reports from the International Federation of Red Cross and Red Crescent Societies. These reports are systematically collected and provide broad coverage of regions that are commonly underrepresented in existing impact datasets. Using large language models, we extract qualitative and quantitative information on a wide range of non-monetary impacts at national and sub-national scales. The resulting dataset documents socio-economic impacts of natural hazards on the population, infrastructure and economy with a spatial detail reaching the subregional level. This resource is designed to support research and applications that require geographically explicit information on socio-economic impacts of disasters, enabling more precise and inclusive analyses of socio-economic consequences of natural hazards worldwide.
Land surface temperature (LST) is a fundamental variable in hydrological and climate research. However, cloud contamination and sparse temporal sampling in satellite thermal infrared observations cause extensive data gaps, hindering generation of continuous long-term high-resolution daily mean LSTs. Although global gap-filled and reanalysis-based LST products already exist, they often face one of the challenges: coarse spatial resolutions, monthly mean/daily instantaneous scales, or struggling to balance accuracy and spatial texture. To address these challenges, we integrate empirical, semi-physical, and machine learning approaches to generate gapless 1 km daily mean LSTs for 70°N–60°S and 2003–2023. We implement a pixelwise, daily-mean–then–annual-cycle, preliminary-to-optimized workflow that harmonizes MODIS observations with all-weather in situ measurements. Validation demonstrates progressive accuracy improvements across steps, with final RMSEs of 2.073 K (daily) and 1.511 K (monthly) against independent ground observations. Comparisons with existing gap-filled products and reanalysis datasets demonstrate improved statistical accuracy—particularly over cropland, grassland, and cold regions—as well as enhanced seasonal robustness and spatial texture, particularly over rugged terrain and urban areas.
Fish escapes from aquaculture facilities pose significant risks, including genetic introgression or competition with wild populations. Accurately identifying the origin of fish, whether wild, farmed, or escaped, is therefore essential for impact assessment, traceability, and sustainable aquaculture management. In this work, we present GLORiA, a curated image dataset specifically designed for fish origin identification. The main dataset contains 9,511 high-quality images of three economically and ecologically important Mediterranean species: Sparus aurata, Dicentrarchus labrax, and Argyrosomus regius. These images correspond to repeated acquisitions of individual specimens under standardized laboratory conditions, rather than to 9,511 independent biological samples. In total, the laboratory dataset includes 423 unique S. aurata specimens, 309 unique D. labrax specimens, and 108 unique A. regius specimens. All images were acquired following a consistent imaging protocol and were annotated by domain experts according to their origin. In addition to the laboratory dataset, GLORiA includes an initial test subset acquired under more realistic commercial conditions, comprising fish-market tray images and cropped individual specimens, which is intended to support evaluation beyond standardized laboratory settings. Unlike existing fish image datasets, which primarily focus on species recognition, GLORiA explicitly targets origin classification, enabling the study of morphological and visual traits associated with aquaculture practices. The dataset is systematically organized by species and origin category and includes preprocessing scripts to support reproducibility. GLORiA addresses a critical gap in publicly available resources and provides a foundation for the development and evaluation of computer vision methods aimed at supporting sustainable aquaculture and mitigating the ecological impact of fish escapes.
The striped flea beetle (SFB), Phyllotreta striolata, is a globally devastating cruciferous pest. A key challenge to its control lies in its complex life cycle—with egg, larval, and pupal stages confined to soil, while adults infest aerial plant parts—coupled with distinct antennal differences between males and females that underpin sex-specific ecological traits. Transcriptomic profiling across developmental stages offers a powerful approach to decode the genetic basis of these stage- and sex-related traits and inform targeted pest management. Here, we present the first comprehensive transcriptomic atlas of P. striolata, encompassing four key life stages (egg, larva, pupa, adult) with separate male and female. RNA sequencing of 15 samples (biological triplicates per group) generated 334.5 million raw reads. Dataset reliability was robust, as validated by correlation analysis, principal component analysis, and qRT-PCR. The deposited threshold-based gene lists may support future investigations of developmental-stage- and adult-sex-associated gene expression in P. striolata. These data provide valuable insights into the biology and developmental dynamics of P. striolata and establish a robust foundation for functional genetic studies. Moreover, the P. striolata transcriptome can be uploaded to the dsRIP platform, enabling systematic identification of effective RNAi targets and optimization of species-specific dsRNA designs with minimized off-target risks. Together, these resources may accelerate the translation of molecular insights into effective and sustainable strategies for controlling this pest.
This work presents a Laser-induced Breakdown Spectroscopy (LIBS) dataset of regolith simulant samples and reference XRF composition data from the manufacturer. 4 Mars and 6 Moon regolith simulants were systematically mixed to have a set of 21 samples for each simulant category. The samples primarily contain Al2O3, SiO2, MgO, CaO, Fe2O3 or FeO, K2O, TiO2, Na2O and P2O5. All the samples were measured under three different atmospheric conditions: (1) Earth’s ambient gas; (2) Mars ambient gas, 7 mbar pressure and CO2 atmosphere; (3) low vacuum simulating Moon conditions, 0.1 mbar. The dataset’s potential use lies in two key areas. The first is quantitative analysis, addressing challenges in analyzing a broad composition range of primary oxides. Enhancing concentration prediction accuracy for Al2O3 and SiO2 is particularly important, as their high concentrations and resonant spectral lines complicate analysis, according to current literature. The second area is sample classification. Both tasks are limited in the size of the dataset, simulating the data-throughput of the rovers for in-situ analysis. The dataset aims to support LIBS performance improvements under such restrictions.
Tobacco (Nicotiana tabacum) is both a major industrial crop and a foundational model organism for plant biology. While alternative splicing significantly diversifies transcriptome complexity, tobacco isoform annotation remains incomplete, hampered by the limitations of previous short-read sequencing efforts. To address this gap, we generated a long-read transcriptome dataset across five tobacco tissues using the PacBio Sequel IIe platform, producing 198.4 million subreads. From these data, we successfully reconstructed 64,260 isoforms, including 35,013 novel transcripts. For the novel isoforms, we also performed systematic examinations of their structural features, biotypes and expression profiles. This dataset expands the tobacco isoform atlas and provides a valuable resource for genome annotation and regulatory studies.
This dataset contains longitudinal clinical data from 68 patients hospitalized with a clinical diagnosis of snakebite at a hospital in Gansu Province, China, over a 3-year period. The dataset contains 10,361 time-stamped laboratory test results, including coagulation tests, clinical chemistry tests, complete blood counts, and other laboratory measurements, together with detailed time-stamped records of antivenom administration and supportive treatments. To protect patient privacy while preserving relative temporal relationships, all absolute dates were randomly shifted to future calendar years. The dataset also includes derived Blaylock local effect scores. Laboratory test items were mapped to the corresponding LOINC codes using the Logical Observation Identifiers Names and Codes (LOINC) v2.73 reference table. The dataset is accompanied by a comprehensive data dictionary and data-validation code and can be used for trajectory clustering, exploratory analyses of antivenom use patterns, and research on early warning of organ dysfunction. We plan to update the dataset annually, with the goal of covering all patients hospitalized with a clinical diagnosis of snakebite at the study hospital over a 10-year period; version 1.0 includes cases from 2023 to 2025.
Causal machine learning represents a promising paradigm for estimating intervention effects in operations management, yet its adoption is constrained by the scarcity of structured datasets suitable for method development and validation. This paper presents a synthetic production dataset comprising 10,000 production order records from a simulated metalworking factory with ten welding lines and ten product types. The dataset incorporates a binary treatment variable representing the implementation of drum-buffer-rope synchronization from the Theory of Constraints on two production lines. Data generation follows established statistical distributions with explicit causal structure, including line-product interaction effects, Little’s Law relationships, and treatment-dependent outcome variations. The dataset includes 15 variables capturing order characteristics, operational parameters, performance outcomes, and a line-level carry-over state that induces temporal serial correlation between successive orders on the same line. Primary reuse value lies in benchmarking causal machine learning methods, particularly meta-learners, tree-based algorithms, and neural network–based approaches for heterogeneous treatment effect estimation in operations management contexts. The dataset supports methodological research on ex-ante effect estimation without requiring explicit simulation models of production systems.
Marine microbes drive global biogeochemical cycles, yet most remain recalcitrant to cultivation and poorly characterized. Microbial community structure varies markedly along environmental gradients, particularly between oligotrophic open-ocean and nutrient-rich coastal regions. However, metagenomic insights across such natural gradients remain limited, particularly in the western Pacific Ocean. Here, we applied genome-resolved metagenomics to seven surface seawater samples collected along an open-ocean–to–coastal transect in the western Pacific Ocean. Sequencing generated 80.16 Gbp of data, enabling the reconstruction of 194 species-level bacterial and archaeal metagenome-assembled genomes (MAGs). Of these, eight met the criteria for high-quality draft genomes and 186 were classified as medium-quality draft genomes according to MIMAG standards ( ≥50% completeness, ≤10% contamination). Both read-based taxonomic profiling and genome-resolved binning revealed differences in bacterial and archaeal community structure among the samples. Pseudomonadota, Bacteroidota, and Marinisomatota were relatively more abundant in the transitional and coastal samples, whereas Cyanobacteriota were more abundant in the open ocean samples. Archaeal phyla were relatively more abundant in the transitional and coastal samples. Collectively, this dataset provides a region-specific, genome-resolved resource for the western Pacific Ocean and enables investigation of microbial community structure across a continuous open-ocean-to-coastal gradient.
This study presents a synthetic dataset designed to support the validation of joint-level trajectory extraction for pedestrian dynamics measurements. Existing pedestrian datasets commonly provide pedestrian-level trajectories, but rarely include frame-wise 3D body-joint coordinates that can be used to evaluate joint-level trajectory extraction and microscopic step measurements. To address this gap, this study developed a dataset containing 147 rendered video sequences, comprising 18,669 video frames in total, of a single animated pedestrian walking along a straight path at three walking speeds. The same walking motions were rendered from 49 virtual cameras placed at different distances, heights, and viewing angles, enabling investigation of camera configuration effects. For each video, the dataset provides frame-wise ground-truth 3D body-joint coordinates in both the world and camera coordinate systems, together with camera intrinsic and extrinsic parameters. By combining rendered videos, known camera parameters, and frame-wise body-joint coordinates, this dataset provides a controlled benchmark for developing and validating vision-based methods for extracting 3D joint-level pedestrian trajectories and deriving step measurements from video. The controlled variation in camera configurations can provide guidance for camera placement in real-world pedestrian experiments.
Climate change has increased the frequency and intensity of extreme rainfall, driving widespread clusters of landslides. These landslides pose growing risks to lives and infrastructure, highlighting the need for datasets to support hazard assessment. Here, we present a point-based dataset of shallow landslide initiation locations triggered by recent rainfall events in China. Using high-resolution satellite imagery captured before and after rainfall events, we assembled 18,496 landslides triggered by the July 2023 rainfall event in the Northeastern Taihang Mountains (Hebei and Beijing). This inventory combines published data for the western mountainous regions of Beijing with new mapping for the remaining area. We also mapped 15,016 landslides triggered by the July 2024 rainfall event in Zixing, and 3,386 landslides triggered by the June 2024 rainfall event in Longyan. Comparison between field photographs and satellite imagery confirms the spatial reliability of the mapped landslides, while validation against a published landslide inventory further supports the positional consistency of dataset. The dataset is openly available in a public Zenodo repository to promote related landslide research.
The United States National Marine Fisheries Service determines “essential fish habitat (EFH)” for federally managed species in coordination with regional fishery management councils, considers adverse effects to those habitats, and provides information to further habitat conservation and enhancement. Identifying discrete subsets of EFH as “habitat areas of particular concern (HAPC)” can help focus conservation, management, and research efforts. In 2006, the Pacific Fishery Management Council designated rocky reefs along the United States (U.S.) West Coast as HAPCs for groundfishes because of their ecological significance, sensitivity to human impacts, and relative rarity. To better understand where rocky reefs occur, we (1) located rocky reef areas that were not included in the 2006 rocky reef dataset, and (2) incorporated best available data into a refined geographic dataset that enables visualization. Our update shows that rocky reefs are distributed throughout the U.S. West Coast continental margin, are patchier than previously known, and comprise 8% of the extent of all data inputs. This updated dataset will inform resource management decisions in coastal and marine environments.
The MEDA A and B elastic beacons, located in the Gulf of Naples (southern Tyrrhenian Sea) and managed by Stazione Zoologica Anton Dohrn (RIMAR Department), has been operating since 2015 with continuous retrieved of meteo-oceanographic parameters. The dataset of 10 years of acquisitions is presented: meteorological data, wave data, ADCP data, and CTD data. This timeseries can be used in different application studies in the study area and represent an enormous resource for scientists and stakeholders.
Eelpouts (family Zoarcidae) serve as an ideal model for investigating adaptive radiation in polar fishes and the parallel evolution of functional genes. However, collection of Antarctic specimens remains highly constrained by logistical difficulties. Here, we present a high-quality, chromosome-level genome assembly for Lycenchelys nigripalatum, reconstructed by combining Illumina short reads, PacBio high-fidelity (HiFi) long reads, and Hi-C chromatin interaction data. The final assembly spans 700.14 Mb with a scaffold N50 of 28.04 Mb, successfully anchored onto 24 pseudo-chromosomes. Technical validation demonstrated exceptional genome integrity and accuracy, yielding a Benchmarking Universal Single-Copy Orthologs (BUSCO) completeness score of 99.31% and a base-level quality value (QV) of 62.44. Genome annotation identified 18,956 protein-coding genes and revealed that transposable elements comprise 22.79% of the genome. This high-resolution genome provides a foundation for elucidating the molecular mechanisms underlying polar adaptation and environmental transitions in Antarctic eelpouts.
Abstract Far-right activism materializes in music events, marches, and violence. Yet, the empirical links between these types of activism remain unclear. Quantifying these manifestations is difficult because of fragmented reporting and inconsistent geocoding at the local level. We introduce the Far-Right Activism (FARAC) dataset, a panel (2013–2024) that aggregates German federal parliamentary inquiries and civil society records at the county level. The dataset addresses the problem of regional assignment by standardizing all event records in accordance with 2024 administrative boundaries. This harmonized structure permits the first systematic, local-level comparison of far-right activism. It uses standard administrative identifiers, enabling direct integration with electoral results, sociodemographic indicators, and other county-level datasets.
Triosteum pinnatifidum is a perennial herbaceous plant endemic to East Asia, whose roots, leaves, and fruits have long been utilized in traditional herbal medicine, thereby demonstrating substantial medicinal value. The root of T. pinnatifidum (commonly referred to as “Tianwangqi”) exhibits multiple pharmacological effects: enhancing Qi-blood circulation, dispelling wind-dampness, tonifying the spleen and stomach, and relieving inflammation and pain. However, the absence of a reference genome has impeded genetic research on this species. In the present study, we report a high-quality chromosome-level genome assembly of T. pinnatifidum generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. The final assembled genome spans 698.94 Mbp with a contig N50 of 46.55 Mbp, which was successfully anchored onto 9 chromosomes. Genome annotation revealed that repetitive sequences account for 65.84% of the genome, and a total of 26,027 protein-coding genes were predicted, among which 98.19% were functionally annotated. The final gene set of T. pinnatifidum exhibited a BUSCO completeness of 98.5%, indicating a high level of assembly accuracy and completeness. This genome assembly not only contributes to the genetic conservation of Triosteum L. species but also provides a valuable genomic resource for evolutionary studies within the Caprifoliaceae family.
Cutaneous spindle cell (CSC) neoplasms are a notoriously challenging diagnostic group within the spectrum of skin malignancies. While specialized CSC neoplasm datasets like AI4SKIN have enabled the development of AI whole-slide image (WSI) classification methods, these methods are largely constrained by closed-set designs. However, in real-world clinical practice, biopsies often contain secondary tumors or rare entities unknown to the developed AI classifiers. When faced with such Out-of-Distribution (OoD) cases, automated systems can produce highly confident but incorrect predictions, posing a significant risk to patient safety and hindering the reliable deployment of AI in routine pathology. To address this limitation, we present ASSIST, a multicenter dataset of 410 WSIs comprising both spindle-cell tumors and a diverse set of metastatic and rare lesions explicitly included as OoD samples. ASSIST expands AI4SKIN by increasing sample diversity, reinforcing underrepresented categories, and introducing clinically meaningful OoD cases, thereby enabling the development of more robust and safety-oriented computational pathology models. We describe the dataset acquisition pipeline, annotation structure, and technical validation using multiple instance learning (MIL) models combined with a suite of OoD detection methods. ASSIST provides an essential benchmark for advancing open-world pathology and supports future research in dependable AI systems for dermatopathology.