
This dataset contains estimates of embodied carbon in building structures for a total of 53 residential building cohorts in Finland from the 1940s to the 2010s, which have been reconstructed using present-day GWP coefficients. The building types covered are one dwelling houses and blocks of flats. The dataset was created by combining and extending two existing datasets: (1) a building part specific material inventory dataset for residential buildings in Finland, and (2) the generic Finnish national emissions database CO2data. Materials presented in the former dataset were matched with the most appropriate corresponding material in the latter. The volumes and densities of materials were multiplied by the GWP coefficients of present-day Finnish construction materials to approximate the embodied carbon of materials present in structures of Finnish housing archetypes. The resulting embodied carbon estimates are presented at the scale of a building, a horizontal building level, and at smaller scales of components. The embodied carbon estimates are also presented as intensities that divide by the floor area of building archetypes; this enables comparing structural embodied carbon between archetypes. The data can be used for life cycle analysis (LCA) on the neighbourhood level or higher, or combined with the original material inventory dataset, for a combined material stock and flow analysis (MSFA) and LCA.
Agricultural tractors operate under highly variable load conditions that depend on the field operation, the attached implement, soil characteristics, and operator behavior. Because no standardized, publicly available load cycles exist that reflect this variability, the thermal and durability validation of tractors is commonly performed with boundary-condition scenarios (e.g., continuous maximum torque) that do not represent real field usage. This article presents a dataset of raw, high-resolution Controller Area Network (CAN) bus recordings acquired from an agricultural tractor (Tümosan 81.110 Powershuttle) during real-world operation on fields in the Konya Plain, Türkiye. The tractor was instrumented with a custom-developed CAN-based telemetry system with integrated Global Navigation Satellite System (GNSS) positioning, which continuously streamed the data in real time over a GSM/WebSocket link to a remote server, where the raw frames were decoded into physical values using J1939 signal databases and stored in a time-series database. Approximately one year of operation was recorded at a nominal sampling rate of 10 Hz, covering all activities of the tractor (field work, road transport and idling). From this campaign, the authors filtered the recordings down to the in-field soil tillage segments, which impose the highest engine torque demand, and only these high-load segments are published. The resulting dataset comprises 10 session files, one per worked field, totaling approximately 24.7 hours (889,255 samples), so that researchers can focus directly on the most demanding operating condition of the tractor. Each tab-separated value file contains the decoded raw CAN signals. The cyclic alternation between soil-engagement passes under sustained high torque and headland turning events is directly observable in the signals, so that the average duration of working passes under load and of headland turns can be determined. The data can be reused for block load cycle construction, thermal validation on PTO dynamometers, tractor powertrain electrification studies, fuel consumption and productivity analyses and the development of machine state detection and data mining methods for agricultural machinery.
The increasing availability of Web data creates new opportunities for studying the history and organization of sports. However, it also raises major challenges regarding data integration, completeness, and consistency. In this work, we introduce RugbyScope, a comprehensive database of rugby union players, teams, and stints. It was built by aggregating data from two open knowledge bases (Wikidata, DBpedia) and five language editions of Wikipedia, carefully selected to cover all the main rugby union nations. Building such a resource requires addressing several challenges, including heterogeneous data structures across sources, incomplete coverage, and inconsistencies between records describing the same entities. We propose a processing pipeline that extracts data from different sources, formats, and languages, that implements a merging strategy based on data consistency constraints, source agreement, and manually validated heuristics. The resulting database provides a structured and unified representation of rugby union careers, including player information, team histories, and career stints. RugbyScope currently focuses on men's rugby union and achieves a higher coverage of professional clubs worldwide. Although some limitations remain regarding source completeness and historical depth, we provide a reproducible framework and an openly documented methodology for constructing large-scale sports databases from heterogeneous Web sources. RugbyScope is intended as a research resource for studying the evolution, structure, and global dynamics of rugby union.
Cylindrical workpieces made of AISI 4140 steel and A97075 aluminum alloy were machined in an external longitudinal turning process using uncoated cemented carbide cutting tools without metal working fluid. During machining, the process forces were measured by means of piezoelectric sensors within a dynamometer. In addition, temperature on the workpiece surface was measured using infrared thermography. The dataset provides the process forces as raw data over the process time, along with the calculated mean values and standard deviations for each test point within the investigated design of experiments, in which the depth of cut, spindle rotation speed, and feed rate are varied. The process temperatures measured along a focused line in feed direction are also provided as raw data and mean values over time and corresponding uncertainty for all experiments. In metal cutting, an understanding of the process forces, which act as external loads on the workpiece, and the workpiece temperature, as an internal material load, is closely linked to the subsequent functional performance of the final component. Reuse potential thus lies primarily in explaining modifications of material properties in the surface layer of the workpiece and its geometric deviations due to cutting and also in using the data for validation of modelling approaches. For analyzing tool wear, process forces and temperatures are also important factors. This dataset enables such analyses without having to repeat the time-consuming calibration of temperature measurement systems and preparation of machining tests.
High resolution solar irradiance data are essential for the reliable modelling and validation of solar photovoltaic (PV) systems. This article presents a dataset of one-minute resolution Global Horizontal Irradiance (GHI) measurements collected at the UCSI University, Kuala Lumpur Campus, Malaysia throughout 2025. The standalone GHI measurement system was powered by a solar PV panel and battery. It utilised a Kipp & Zonen CMP3 pyranometer, an ISO 9060 spectrally flat class C thermopile pyranometer. To ensure precise data logging, the pyranometer’s analogue signals were digitalised using an ADS1115 16-bit analogue-to-digital converter (ADC) with a programmable gain amplifier. An ESP32-WROOM-32U microcontroller was used to process these signals, compute the irradiance measurements, and transmit the data to a ThingSpeak Internet of Things (IoT) server at a one-minute intervals from 07:00 to 18:59 local time. A total of 259,858 observations were retrieved out of the 262,800 expected observations. The 2,942 missing observations (1.12%) were linear interpolated. Additionally, a quality control screening was conducted using Baseline Surface Radiation Network tests based on Extremely Rare Limits and Physically Possible Limits, with no observations flagged. The dataset has a combined standard uncertainty of ±4.42% and an expanded standard uncertainty of ±8.84% (k=2). This one-minute resolution GHI dataset captures rapid irradiance variations in Kuala Lumpur, Malaysia. The dataset contained daytime GHI measurements only and does not include other meteorological variables. It may support PV system modelling, simulation, and validation by enabling representation of irradiance-driven power fluctuations. The data can also be used for short-term GHI and PV power forecasting, including machine learning-based approaches. In addition, the dataset may support studies of PV power variability, PV battery energy storage system scheduling, PV output smoothing, and grid integration under high PV penetration.
Cavenderia pseudoaureostipes, a dictyostelid species from Thailand described in 2018, is distinguished by the size and complexity of its fruiting bodies. Here, we report its complete mitochondrial genome, assembled from Illumina short-read data. The AT-rich circular genome is 68,470 bp in length, making it larger than previously reported dictyostelid mitogenomes. It contains 41 protein-coding genes, 17 tRNA genes, and three rRNA genes, all encoded on a single strand. Introns were detected in cox1/2 and cytb, and the genome exhibits the conserved A–B–C segmental arrangement characteristic of Cavenderia. Phylogenetic analysis based on 30 conserved protein-coding genes supported the monophyly of Cavenderia and other genera, consistent with order-level classifications. Within Cavenderia, C. pseudoaureostipes was most closely related to C. subdiscoidea, in agreement with the SSU-based phylogeny. This mitogenome represents a valuable resource for comparative which will help resolving key questions regarding evolution of dictyostelids mitochondrial genome.
Atomic charge descriptors are widely used to interpret chemical bonding, quantify charge transfer, and construct data-driven models for inorganic materials. This data article describes a first-principles dataset containing Hirshfeld and iterative-Hirshfeld outputs for 5,080 inorganic materials, corresponding to 31,288 atomic sites and 77 chemical elements. For each material, the database stores the atomic structure, Materials Project identifier, calculation metadata, standard Hirshfeld charges, iterative Hirshfeld charges, Hirshfeld relative volumes, and iterative-Hirshfeld relative volumes. The data were generated using density-functional theory calculations in VASP with a consistent local-density-approximation setup, a 500 eV plane-wave cutoff, 0.25 inverse angstrom k-point spacing, and an electronic convergence threshold of 1e-6 eV. The dataset is distributed as an ASE database with accompanying JSON statistics, analysis scripts, figures, and a Flask-based web interface. It can support charge-transfer analysis, benchmarking of charge-partitioning methods, and charge-informed machine-learning workflows in which atomic populations, charge-transfer magnitudes, or volume descriptors are used as features for materials-property prediction.
This dataset contains field measurements of soil-sweep tool interactions in clay loam soil acquired at the research field of the Institute of Technology of the Hungarian University of Agriculture and Life Sciences in Gödöllő, Hungary. Measurements were conducted during three field campaigns representing distinct soil moisture conditions over the course of one year. The objective of the measurement campaign was to document soil conditions, tillage parameters, and soil response using a consistent experimental methodology at the same agricultural field. Cone Penetration Tests (CPT) were performed at multiple locations within each campaign using an Eijkelkamp 06.15.SA penetrologger, while soil moisture was measured at the same locations. Tillage experiments were carried out using three sweep tools with different working widths operated at three target working speeds. During each tillage run, draft force and speed were continuously recorded using a custom instrumented tractor. Surface profiles were measured before and after tillage using a laser profilograph for the dry and medium-moisture soil conditions. The repository is organized according to soil moisture conditions and includes raw measurement files and accompanying JSON metadata. In addition, scanned 3D models of the investigated sweep tools are provided in STL format. The dataset can be reused for studies of soil moisture effects on tillage performance, development and validation of discrete element method (DEM) and finite element method (FEM) models, and for benchmarking soil-tool interaction simulations.
This data article presents a synthetic stereo vision dataset composed of synchronized RGB stereo images, ground-truth depth maps, and precise 6-DoF ground-truth poses generated using Unreal Engine 4 integrated with the AirSim simulation framework. Four indoor virtual scenes resembling industrial and enterprise-like environments were created, including structural elements such as pipes, pillars, walls, furniture, and occlusions. Multiple acquisition configurations were executed as closed-loop drone trajectories, resulting in a total of 36,838 stereo image pairs. The dataset includes multiple trajectory smoothness conditions, stereo baselines, and camera convergence configurations. All RGB images have a fixed resolution of 640 × 480 pixels and are provided alongside pixel-aligned depth maps and time-stamped ground-truth poses. The dataset also includes association files linking stereo images, depth maps, and poses, following formats commonly used in visual SLAM benchmarks. The acquisition is scripted and free of stochastic components, enabling benchmarking and development of stereo depth estimation, visual odometry, and SLAM algorithms under controlled and repeatable indoor conditions.
This article presents a pilot Long Term Evolution (LTE) drive-test dataset collected from a real urban cellular network in Dhaka, Bangladesh. The dataset contains field measurements of radio signal conditions, Layer-3 measurement reports, and handover statistics obtained using professional drive-testing tools including XCAL-M software and Samsung Exynos-based user equipment. Measurements were recorded during multiple drive-test sessions along a 13 km urban route, producing timestamped observations of key radio parameters such as Reference Signal Received Power (RSRP), Reference Signal Received Quality (RSRQ), Carrier-to-Interference-plus-Noise Ratio (CINR), and serving/neighbor cell identifiers. The dataset is organized into raw and processed directories containing dynamic radio management (DRM) log files, Layer-3 measurement reports, handover event statistics, and tabular datasets to support reproducibility and flexible analysis. The processed dataset demonstrates an example preprocessing workflow including filtering, timestamp alignment, interpolation of missing signal values, and handover labeling with Time-to-Trigger (TTT) context. Due to the limited size of the dataset, it is primarily intended for exploratory analysis, mobility characterization, small-scale machine learning experiments, reinforcement learning environments, and educational research in cellular network engineering. By providing both raw drive-test logs and structured datasets, this resource supports reproducible studies of LTE handover behavior and urban radio environments [1].
Medieval and Renaissance Latin geographical works constitute a major source for understanding how space, places, and territories were described and conceptualised in pre-modern Europe. However, information about these works, their manuscript transmission, and the places they mention remains dispersed across catalogues, archives, and specialised scholarship. Here we present the IMAGO knowledge graph, a semantically structured dataset representing 343 Latin geographical works written between the 6th and the 15th centuries. The dataset integrates curated information provided by domain experts, including authors, works, manuscripts, printed editions, libraries, literary genres, and mentioned places. Data were initially collected in tabular form and subsequently enriched through semi-automatic reconciliation with external authority sources such as Wikidata and the MIRABILE digital archive. Domain experts further expanded the dataset using a dedicated annotation tool. The curated data were transformed into an OWL 2 DL knowledge graph aligned with the IMAGO ontology and published following FAIR and Linked Open Data principles. The knowledge graph was validated through automated reasoning, expert review, and query-based evaluation. The resulting dataset enables systematic exploration of textual, bibliographic, and spatial relationships within medieval and Renaissance geographical literature and supports reuse in historical, philological, and digital humanities research.
In this article, we present a morphologically annotated lexical dataset designed to support the translation of texts between the Khorezm dialect of Uzbek and standard Uzbek. The database consists of 1445 Khorezm dialect–standard Uzbek word pairs. For each entry, the dialect form, its corresponding standard form, and part of speech are provided, along with a morphological structure segmented into affixes according to such grammatical categories as number (singular/plural), possessive, and person/number properties. During the tagging process, the grammatical system of Uzbek and the specific inflectional properties of the Khorezm dialect were taken into account, resulting in a clear and machine-processable layer of correspondences between the dialect and the standard language. The resulting lexical resource is intended to serve as additional input data for AI-based translation and normalization models, as well as in applications for morphological analysis, spelling and grammar checking, and educational tools for teaching the Khorezm dialect. This dataset constitutes the first systematic lexical-morphological resource for corpus-based research on the Khorezm dialect and lays the groundwork for studies on machine translation and automatic alignment between dialects and the standard language in low-resource Turkic varieties.
Most of the cyber attacks are initiated through phishing URLs, which are shared with the victims through multiple media. In spite of the research community proposing varied solutions, the volume and nature of such attacks have evolved unpredictably. To develop effective solutions, researchers require comprehensive datasets that encompass a wide range of attack types rather than focusing on a narrow subset. We present CompPhish, a processed dataset, comprising of 15,358 URLs paired with their respective HTML sources. 7,204 URLs are phishing, and 8,154 are legitimate. A set of 70 features, specifically curated to capture the properties exhibited by phishing URLs, is extracted . These features represent a diverse range of phishing attacks, including generic URL-based phishing attacks, phishing through brand-jacking, phishing sites hosted on compromised domains, and auto-downloadable malicious files embedded in webpages. Phishing URLs are gathered from PhishTank and OpenPhish, while legitimate URLs are compiled from multiple independent sources, like DataforSEO and GitHub.By publishing the raw URLs and HTML codes along with the feature vectors, CompPhish attempts to aid the researchers in devising generalized, robust, and reliable machine learning based solutions for real-world phishing attempts.
Environmental sound classification (ESC) has been used to implement various applications such as smart city, environmental monitoring systems, context-aware systems, and assistive technologies. The existing benchmark datasets, however, mainly focus on the acoustic environment of the west and limited acoustic scenes are available in Bangladesh. In this paper, we introduce a real-world environmental audio dataset named AcousticSceneBD, with a total of 5535 recordings across 7 classes of acoustic scenes: Bus, Metro, Metro Station, Park, Restaurant, Shopping Mall, and University. All recordings were conducted using the standard operating conditions of the smartphone's microphone and with all files standardized following the same procedure:10-second, mono, 16-bit WAV audio recorded at 44.1/48 kHz, with a total duration of 15.38 hours and a dataset size of 4.87 GB.To guarantee consistency and reproducibility, the dataset has been created using a structured approach for data collection and processing consisting of eight stages. Four baseline models, namely Random Forest, YAMNet, Wav2Vec2, and PANNs CNN14 were tested to validate the data, with respective test accuracy of 58.81%, 82.99%, 47.11%, and 78.15%, where Random forest is trained on extracted MFCC features from the raw audio The results indicate that AcousticSceneBD is well annotated, acoustically discriminative and appropriate for the development and evaluation of traditional, deep learning, transfer learning, and self-supervised environmental sound classification methods.
Acidisoma sibiricum DSM 21000T, the type strain of an obligately aerobic, acidophilic, and psychrotolerant bacterium, was isolated from a Sphagnum-dominated peat bog in the Bakchar region of western Siberia, Russia. Here we report the complete genome sequence data of this type strain. Genomic DNA was sequenced using both Illumina NovaSeq X (paired-end, 2×150 bp) and PacBio Revio platforms. Hybrid assembly yielded a complete genome consisting of one chromosome (3,925,831 bp) and two plasmids, pAS1 (577,300 bp) and pAS2 (19,537 bp), with a total genome size of 4,522,668 bp and a sequencing coverage of 312×. The genome was found to be 98.97% complete with an estimated contamination of 2.25%. This is the first report of a complete genome sequence for the type strain of A. sibiricum. Digital DNA-DNA hybridisation (dDDH) analysis confirmed that DSM 21000 is genomically distinct from the closest described Acidisoma species tested, with the highest dDDH value of 25.0% with A. cellulosilyticum HW T5.17, well below the 70% species delineation threshold. Genome annotation by NCBI PGAP revealed 4,225 genes, including 4,127 protein-coding sequences. A complete set of polyhydroxyalkanoate (PHA) biosynthesis genes was also identified, consistent with the presence of intracellular poly-β-hydroxybutyrate granules previously reported for this species. Furthermore, a putative aerobic anoxygenic photosynthesis (AAP) gene cluster was identified on the chromosome, comprising genes encoding a type II photosynthetic reaction center, a light-harvesting antenna complex LH1, and a complete bacteriochlorophyll a biosynthesis pathway. All sequence data were deposited in the NCBI GenBank database under BioProject PRJNA1302119.
This article presents the draft genome sequence of Actinoalloteichus caeruleus strain 372, isolated from saline soil in Kazakhstan. Whole-genome sequencing was performed using the Illumina platform, followed by de novo assembly with SPAdes. Genome quality assessment using CheckM indicated 99.25% completeness with no detectable contamination. The draft genome is 6.18 Mb in size and contains 5,024 predicted protein-coding genes, 47 tRNA genes, and 22 CRISPR arrays. Average nucleotide identity (ANI) analysis showed 96.27% identity to the closest publicly available A. caeruleus genome. Biosynthetic gene cluster (BGCs) analysis using antiSMASH v8.0.4 identified 31 predicted clusters, including NRPS, PKS, NRPS–PKS hybrids, terpenes, and RiPPs. These data expand the genomic resources available for the genus Actinoalloteichus.
The High-Speed Road Network Vector Data with Construction Year Attribution for Poland (HiRoN-PL, version 1.1) provides spatially and temporally explicit data documenting the development of the high-speed road network in Poland between 1936 and 2023. The dataset is distributed in GeoPackage format containing vector line features representing the time-stamped high-speed road network in Poland.HiRoN-PL v1.1 road geometries are derived from OpenStreetMap (OSM) ‘way’ objects corresponding to high-speed road classes in Poland, namely motorways and expressways (dual carriageways). The selected OSM road sections were aligned with authoritative data provided by the General Directorate for National Roads and Motorways (GDDKiA), the central authority responsible for national roads in Poland. First, the authoritative GDDKiA road geometries were attributed with commissioning dates based on expert interpretation of information from GDDKiA records, public registers, and verified crowdsourced data sources. Second, the geometry of OSM road sections was adjusted to reflect distinct construction phases defined by the authoritative data. The resulting OSM-derived road network was then assigned with temporal attributes, including the year, month, and day of commissioning, as well as the road segment name, class, and the names of the road junctions that designate the starting and ending nodes of each segment. The dataset includes additional attributes, such as the length of each section and identifiers linking to the source OSM and authoritative data.HiRoN-PL v1.1 is a valuable source of data for spatial-historical analyses of transport infrastructure development at a national level, enabling multi-temporal accessibility modelling and assessment of transportation efficiency over time. The dataset can be used to support analyses relating to the monitoring of land use changes, the impact of transport and planning policies, or the effects and impacts of European Union investment policies. Data in the GeoPackage format are publicly available via the Zenodo platform: https://doi.org/10.5281/zenodo.19570041.
This article presents a multi-label Sentinel-2 image dataset addressing the limited availability of remote sensing benchmarks that reflect the real-world co-occurrence of energy, transport, and storage infrastructure. The collection contains 10,000 true-color RGB PNG image chips from across Germany. Each image is independently annotated for electrical substations, power plants, storage tanks, railway infrastructure, and major road infrastructure, allowing none, one, or multiple infrastructure types to occur within the same scene. Candidate locations and labels were derived from OpenStreetMap geometries. Samples were selected from retained linear and polygonal features, spatially balanced across Germany, and accepted only when their complete image footprints satisfied the defined intersection and non-overlap criteria. Contextually similar no-label samples were generated near infrastructure locations to provide challenging negative examples. Predefined spatial folds keep neighboring samples, associated no-label samples, and samples from the same local OpenStreetMap feature together, supporting reproducible evaluation while avoiding spatial leakage. The image chips were generated from Copernicus Sentinel-2 Level-2A bands B04, B03, and B02 using a maximum cloud-cover threshold of 10 percent and least-cloud-cover mosaicking. Each 224 × 224 pixel image is a resampled representation of a 512 m × 512 m footprint. The released files include the image collection and CSV metadata containing filenames, binary labels, and spatial fold assignments. Spatial five-fold baseline experiments demonstrate the dataset’s suitability for supervised multi-label classification, reproducible model comparison, and infrastructure mapping.
The dataset provides a complete record of microbial, functional, physicochemical, and contaminant dynamics from controlled co-composting microcosms designed to remediate petroleum refinery sludge using targeted animal manure amendments. Five treatments—cow, pig, horse, and poultry manures, as well as an unamended control—were monitored over 300 days. The study incorporated amplicon-based 16S rRNA gene sequencing, functional gene inference, culture-based validation, bulk chemistry, and chromatographic analyses. Illumina MiSeq profiling of the V1–V3 regions identified 359 bacterial genera (raw OTU-level assignments prior to quality and abundance filtering) across 17 phyla, with taxonomic inventories structured from phylum to genus. Alpha and beta diversity measures demonstrated treatment-dependent community assembly, with the highest richness and diversity observed in cow-manure microcosms—pig and poultry amendments selectively enriched hydrocarbon-degrading taxa, including Pseudomonas. Functional predictions generated using PICRUSt2 (NSTI = 0.02–0.16) indicated enrichment of pathways involved in xenobiotic degradation, aromatic compound metabolism, and benzoate catabolism. These predictions were supported by culture-based evidence, including redox indicator screening and detection of the cbzE gene, which encodes catechol 2,3-dioxygenase, a key enzyme in chlorobenzoate/chlorocatechol degradation pathways and functionally analogous to the widely recognised xylE gene in aromatic hydrocarbon-degrading microorganisms. Collectively, these determinations provide complementary genomic and phenotypic evidence for the biodegradation potential of the microbial community and its capacity to transform aromatic and other environmentally relevant xenobiotic compounds. Additional datasets document total organic carbon, nitrogen, and phosphorus profiles of feedstocks and sludge. At the same time, Soxhlet extraction GC–MS measurements quantify polycyclic aromatic hydrocarbon (PAH) attenuation, with up to 99.9% removal achieved for multiple compounds in pig, horse, and poultry-amended systems. Through the combination of taxonomic profiling, PICRUSt2-based functional inference, chemical transformation analyses, and degradation kinetic measurements, this dataset affords a comprehensive characterisation of microbial community structure, predicted metabolic potential, and biodegradation performance associated with crude oil sludge co-composting at the sampled time point.