
Pineapple (Ananas comosus) is a tropical fruit of great economic value. It is grown in several areas of India and makes a good contribution to horticultural production and farmer’s income. Unfortunately, pineapples are very vulnerable to a variety of leaf diseases that reduce crop yield. Due to this, there is a pressing need for innovative ways to monitor plant health, such as deep learning-based disease detection using well-prepared datasets. The main aim of the dataset is to support research in the areas of condition analysis and early disease detection including machine learning models that reduce manual inspections, enable early detection, facilitate timely treatments. The dataset consists of 4,476 original images collected from 200 unique pineapple leaf specimens, representing four classes. These specimens were collected in the field from Vazhakulam, Kerala, India, during June to July 2026. Fifty unique specimens were included for each class, with at least 20 distinct original images acquired per specimen. They were then pre-processed and augmented with image modifications, while the independent test images were retained without augmentation. The images were used to evaluate five transfer-learning-based deep learning architectures- Xception, InceptionV3, MobileNetV2, ResNet50, and EfficientNetB0, using a specimen-level 80:20 training–testing strategy. Researchers and experts can make great use of this dataset because of its enormity and quality. It can be used to develop and test techniques for leaf classification, disease detection, and severity estimation that will aid to the development of deep learning-enabled plant health assessment and precision farming systems. Such systems allow for more precise treatment of the affected area, reduce chemical use, and overall crop management.
This article presents labeled data collected from video footage of additive manufacturing processes. This computer vision dataset contains 90 labeled additive manufacturing video segments across four distinct additive manufacturing processes, amounting to 900 labeled, object instance video frames. The technologies covered by the dataset include laser hot-wire additive manufacturing (LHW-DED), plasma arc welding (PAW), tungsten inert gas wire arc additive manufacturing (TIG-WAAM), and polymer-based additive manufacturing (visPolymer and irPolymer). For each of the manufacturing processes, build deposition video was captured and then randomly sampled into 10 frame-long video segments. Depending on the manufacturing process, two of the following four object instance classes are labeled in each frame: Melt Pool, Feed Wire, Nozzle, and Material. Since the dataset’s directory organization follows that of common video object segmentation (VOS) datasets, such as coMplex video Object SEgmentation (MOSE), MOSEv2, Densely Annotated Video Segmentation (DAVIS), and YouTube-VOS, it can be easily integrated into other VOS model training pipelines for foundation model fine-tuning, training, or inference capability testing. This dataset provides researchers access to labeled additive manufacturing object instances for VOS tasks.
Predictive Maintenance has become an essential objective in industry as it has the potential to decrease unplanned downtime, reduce inventories of replacement parts and increase operational safety, by identifying and predicting the time of failures. To develop accurate predictive maintenance systems, sufficient historical data to failure must be acquired. While such datasets are publicly available for components (i.e. bearing and motors), their public availability for industrial manufacturing systems is limited. This article presents a time-series dataset for the run-to-failure of single-screw polymer extruder filters. The dataset includes the temperatures along the barrel of the extruder, screw motor voltage and current at 1Hz. The current intake of the extruder is also provided and acquired at 1 kHz, to support feature engineering works. Data is collected for 60 run-to-failures during the extrusion of low-density polyethylene. Two failure modes were recorded: full filter clogging, preventing the extrusion of the material, and filter failure, where the material accumulates until the filter bursts. The established dataset can be used to develop and verify state-of-the-art predictive maintenance models under industrial conditions, physics-based models, feature engineering and multi-rate data fusion methodologies.
This article describes a nationwide, georeferenced entomological dataset of disease vectors compiled from the historical archives of the Servicio Nacional de Erradicación del Paludismo (SENEPA) in Paraguay from 1998 to 2023. The dataset consolidates 20,619 standardized occurrence records derived from routine programmatic surveillance, targeted ecological-zone monitoring, and reactive outbreak responses. Each record is mapped to the Darwin Core standard and includes taxonomic identification, administrative location, geographic coordinates (WGS84), georeferencing provenance, and spatial uncertainty. Data cleaning, harmonization, and taxonomic curation were executed in R using reproducible workflows, incorporating GBIF Backbone Taxonomy alignment and independent expert validation. Designed to support FAIR data reuse, this repository provides a high-density spatial baseline. Because collections originate from institutional surveillance rather than probabilistic sampling designs, users must account for spatial and temporal sampling biases in downstream ecological analyses.
This article presents a dataset integrating Headspace Gas Chromatography–Mass Spectrometry (HS-GC-MS) analyses and descriptive sensory evaluation data for three commercially available maize-based snack products from the Colombian market. The dataset was developed to support research in food science, sensory analysis, chemometrics, and predictive modeling of food perception. HS-GC-MS analyses were performed in triplicate for each product, generating nine chromatographic files in (.cdf) format and a processed dataset containing retention times, retention indices, peak areas, and tentative volatile compound identifications. In parallel, a semi-trained sensory panel composed of 32 participants evaluated 17 sensory descriptors associated with odor, taste, and texture using a structured intensity scale. To assess the statistical value of chromatographic data, a Partial Least Squares Discriminant Analysis (PLS-DA) model was developed using volatile compound profiles. The model showed a clear discrimination among the three products, with the first two components explaining 97.3 % of the total variance. Variable Importance in Projection (VIP) analysis identified terpene-related compounds, including β-pinene, limonene, γ-terpinene, α-pinene, and p-cymene, as the main contributors to sample differentiation. Cross-validation demonstrated excellent model performance with R² = 0.99 and Q² = 0.98. The dataset provides openly accessible chromatographic and sensory information that may support comparative studies, food quality assessment, sensory prediction models, and the development of novel maize-based food products.
MeteoCoast-EC is an open-access, multi-variable meteorological dataset comprising 62,815 hourly records collected between January 2022 and May 2026 from three automatic weather stations operated by the Universidad Técnica de Manabí (UTM) in the coastal region of Ecuador: Santa Ana, Portoviejo, and Calderón, located at 50, 64, and 150 m a.s.l. respectively. The dataset addresses a critical observational gap: the coastal region of Ecuador holds the lowest density of operational meteorological stations of any region in the country, and INAMHI (Instituto Nacional de Meteorología e Hidrología) does not currently maintain active stations at the Portoviejo and Calderón sites. This gap is particularly consequential given that the hydrographic systems encompassing these stations experienced extreme precipitation anomalies affecting approximately 93% of basin areas during the 1997-1998 El Niño event, and extreme El Niño events are projected to increase in frequency under ongoing climate change. The monitored area also constitutes the primary agricultural and artisanal fishing zone of Ecuador, where locally observed, sub-daily meteorological records are essential for agroclimatic risk modeling, hydrological forecasting, and early warning system development. Two files are released. The raw export preserves the 75 columns delivered by the WeatherLink API together with the SQL export layer, 68 API fields plus 7 export fields, with no variable withheld. The processed release retains 23 analytically distinct variables, 4 identifiers and 19 measurements, and materializes them as 33 physical columns by adding 9 SI unit conversions and 1 auxiliary grouping key. Data are acquired using Davis Vantage Pro 2 sensors via a JavaScript-based ETL pipeline and stored in a PostgreSQL database at hourly resolution. Neither file is quality-controlled beyond deduplication, so the native sensor values are preserved. MeteoCoast-EC is released under CC BY 4.0 and deposited in Mendeley Data together with the complete processing and figure-generation code.
This dataset contains estimates of embodied carbon in building structures for a total of 53 residential building cohorts in Finland from the 1940s to the 2010s, which have been reconstructed using present-day GWP coefficients. The building types covered are one dwelling houses and blocks of flats. The dataset was created by combining and extending two existing datasets: (1) a building part specific material inventory dataset for residential buildings in Finland, and (2) the generic Finnish national emissions database CO2data. Materials presented in the former dataset were matched with the most appropriate corresponding material in the latter. The volumes and densities of materials were multiplied by the GWP coefficients of present-day Finnish construction materials to approximate the embodied carbon of materials present in structures of Finnish housing archetypes. The resulting embodied carbon estimates are presented at the scale of a building, a horizontal building level, and at smaller scales of components. The embodied carbon estimates are also presented as intensities that divide by the floor area of building archetypes; this enables comparing structural embodied carbon between archetypes. The data can be used for life cycle analysis (LCA) on the neighbourhood level or higher, or combined with the original material inventory dataset, for a combined material stock and flow analysis (MSFA) and LCA.
Agricultural tractors operate under highly variable load conditions that depend on the field operation, the attached implement, soil characteristics, and operator behavior. Because no standardized, publicly available load cycles exist that reflect this variability, the thermal and durability validation of tractors is commonly performed with boundary-condition scenarios (e.g., continuous maximum torque) that do not represent real field usage. This article presents a dataset of raw, high-resolution Controller Area Network (CAN) bus recordings acquired from an agricultural tractor (Tümosan 81.110 Powershuttle) during real-world operation on fields in the Konya Plain, Türkiye. The tractor was instrumented with a custom-developed CAN-based telemetry system with integrated Global Navigation Satellite System (GNSS) positioning, which continuously streamed the data in real time over a GSM/WebSocket link to a remote server, where the raw frames were decoded into physical values using J1939 signal databases and stored in a time-series database. Approximately one year of operation was recorded at a nominal sampling rate of 10 Hz, covering all activities of the tractor (field work, road transport and idling). From this campaign, the authors filtered the recordings down to the in-field soil tillage segments, which impose the highest engine torque demand, and only these high-load segments are published. The resulting dataset comprises 10 session files, one per worked field, totaling approximately 24.7 hours (889,255 samples), so that researchers can focus directly on the most demanding operating condition of the tractor. Each tab-separated value file contains the decoded raw CAN signals. The cyclic alternation between soil-engagement passes under sustained high torque and headland turning events is directly observable in the signals, so that the average duration of working passes under load and of headland turns can be determined. The data can be reused for block load cycle construction, thermal validation on PTO dynamometers, tractor powertrain electrification studies, fuel consumption and productivity analyses and the development of machine state detection and data mining methods for agricultural machinery.
The increasing availability of Web data creates new opportunities for studying the history and organization of sports. However, it also raises major challenges regarding data integration, completeness, and consistency. In this work, we introduce RugbyScope, a comprehensive database of rugby union players, teams, and stints. It was built by aggregating data from two open knowledge bases (Wikidata, DBpedia) and five language editions of Wikipedia, carefully selected to cover all the main rugby union nations. Building such a resource requires addressing several challenges, including heterogeneous data structures across sources, incomplete coverage, and inconsistencies between records describing the same entities. We propose a processing pipeline that extracts data from different sources, formats, and languages, that implements a merging strategy based on data consistency constraints, source agreement, and manually validated heuristics. The resulting database provides a structured and unified representation of rugby union careers, including player information, team histories, and career stints. RugbyScope currently focuses on men's rugby union and achieves a higher coverage of professional clubs worldwide. Although some limitations remain regarding source completeness and historical depth, we provide a reproducible framework and an openly documented methodology for constructing large-scale sports databases from heterogeneous Web sources. RugbyScope is intended as a research resource for studying the evolution, structure, and global dynamics of rugby union.
Cylindrical workpieces made of AISI 4140 steel and A97075 aluminum alloy were machined in an external longitudinal turning process using uncoated cemented carbide cutting tools without metal working fluid. During machining, the process forces were measured by means of piezoelectric sensors within a dynamometer. In addition, temperature on the workpiece surface was measured using infrared thermography. The dataset provides the process forces as raw data over the process time, along with the calculated mean values and standard deviations for each test point within the investigated design of experiments, in which the depth of cut, spindle rotation speed, and feed rate are varied. The process temperatures measured along a focused line in feed direction are also provided as raw data and mean values over time and corresponding uncertainty for all experiments. In metal cutting, an understanding of the process forces, which act as external loads on the workpiece, and the workpiece temperature, as an internal material load, is closely linked to the subsequent functional performance of the final component. Reuse potential thus lies primarily in explaining modifications of material properties in the surface layer of the workpiece and its geometric deviations due to cutting and also in using the data for validation of modelling approaches. For analyzing tool wear, process forces and temperatures are also important factors. This dataset enables such analyses without having to repeat the time-consuming calibration of temperature measurement systems and preparation of machining tests.
High resolution solar irradiance data are essential for the reliable modelling and validation of solar photovoltaic (PV) systems. This article presents a dataset of one-minute resolution Global Horizontal Irradiance (GHI) measurements collected at the UCSI University, Kuala Lumpur Campus, Malaysia throughout 2025. The standalone GHI measurement system was powered by a solar PV panel and battery. It utilised a Kipp & Zonen CMP3 pyranometer, an ISO 9060 spectrally flat class C thermopile pyranometer. To ensure precise data logging, the pyranometer’s analogue signals were digitalised using an ADS1115 16-bit analogue-to-digital converter (ADC) with a programmable gain amplifier. An ESP32-WROOM-32U microcontroller was used to process these signals, compute the irradiance measurements, and transmit the data to a ThingSpeak Internet of Things (IoT) server at a one-minute intervals from 07:00 to 18:59 local time. A total of 259,858 observations were retrieved out of the 262,800 expected observations. The 2,942 missing observations (1.12%) were linear interpolated. Additionally, a quality control screening was conducted using Baseline Surface Radiation Network tests based on Extremely Rare Limits and Physically Possible Limits, with no observations flagged. The dataset has a combined standard uncertainty of ±4.42% and an expanded standard uncertainty of ±8.84% (k=2). This one-minute resolution GHI dataset captures rapid irradiance variations in Kuala Lumpur, Malaysia. The dataset contained daytime GHI measurements only and does not include other meteorological variables. It may support PV system modelling, simulation, and validation by enabling representation of irradiance-driven power fluctuations. The data can also be used for short-term GHI and PV power forecasting, including machine learning-based approaches. In addition, the dataset may support studies of PV power variability, PV battery energy storage system scheduling, PV output smoothing, and grid integration under high PV penetration.
Cavenderia pseudoaureostipes, a dictyostelid species from Thailand described in 2018, is distinguished by the size and complexity of its fruiting bodies. Here, we report its complete mitochondrial genome, assembled from Illumina short-read data. The AT-rich circular genome is 68,470 bp in length, making it larger than previously reported dictyostelid mitogenomes. It contains 41 protein-coding genes, 17 tRNA genes, and three rRNA genes, all encoded on a single strand. Introns were detected in cox1/2 and cytb, and the genome exhibits the conserved A–B–C segmental arrangement characteristic of Cavenderia. Phylogenetic analysis based on 30 conserved protein-coding genes supported the monophyly of Cavenderia and other genera, consistent with order-level classifications. Within Cavenderia, C. pseudoaureostipes was most closely related to C. subdiscoidea, in agreement with the SSU-based phylogeny. This mitogenome represents a valuable resource for comparative which will help resolving key questions regarding evolution of dictyostelids mitochondrial genome.
Atomic charge descriptors are widely used to interpret chemical bonding, quantify charge transfer, and construct data-driven models for inorganic materials. This data article describes a first-principles dataset containing Hirshfeld and iterative-Hirshfeld outputs for 5,080 inorganic materials, corresponding to 31,288 atomic sites and 77 chemical elements. For each material, the database stores the atomic structure, Materials Project identifier, calculation metadata, standard Hirshfeld charges, iterative Hirshfeld charges, Hirshfeld relative volumes, and iterative-Hirshfeld relative volumes. The data were generated using density-functional theory calculations in VASP with a consistent local-density-approximation setup, a 500 eV plane-wave cutoff, 0.25 inverse angstrom k-point spacing, and an electronic convergence threshold of 1e-6 eV. The dataset is distributed as an ASE database with accompanying JSON statistics, analysis scripts, figures, and a Flask-based web interface. It can support charge-transfer analysis, benchmarking of charge-partitioning methods, and charge-informed machine-learning workflows in which atomic populations, charge-transfer magnitudes, or volume descriptors are used as features for materials-property prediction.
This dataset contains field measurements of soil-sweep tool interactions in clay loam soil acquired at the research field of the Institute of Technology of the Hungarian University of Agriculture and Life Sciences in Gödöllő, Hungary. Measurements were conducted during three field campaigns representing distinct soil moisture conditions over the course of one year. The objective of the measurement campaign was to document soil conditions, tillage parameters, and soil response using a consistent experimental methodology at the same agricultural field. Cone Penetration Tests (CPT) were performed at multiple locations within each campaign using an Eijkelkamp 06.15.SA penetrologger, while soil moisture was measured at the same locations. Tillage experiments were carried out using three sweep tools with different working widths operated at three target working speeds. During each tillage run, draft force and speed were continuously recorded using a custom instrumented tractor. Surface profiles were measured before and after tillage using a laser profilograph for the dry and medium-moisture soil conditions. The repository is organized according to soil moisture conditions and includes raw measurement files and accompanying JSON metadata. In addition, scanned 3D models of the investigated sweep tools are provided in STL format. The dataset can be reused for studies of soil moisture effects on tillage performance, development and validation of discrete element method (DEM) and finite element method (FEM) models, and for benchmarking soil-tool interaction simulations.
This data article presents a synthetic stereo vision dataset composed of synchronized RGB stereo images, ground-truth depth maps, and precise 6-DoF ground-truth poses generated using Unreal Engine 4 integrated with the AirSim simulation framework. Four indoor virtual scenes resembling industrial and enterprise-like environments were created, including structural elements such as pipes, pillars, walls, furniture, and occlusions. Multiple acquisition configurations were executed as closed-loop drone trajectories, resulting in a total of 36,838 stereo image pairs. The dataset includes multiple trajectory smoothness conditions, stereo baselines, and camera convergence configurations. All RGB images have a fixed resolution of 640 × 480 pixels and are provided alongside pixel-aligned depth maps and time-stamped ground-truth poses. The dataset also includes association files linking stereo images, depth maps, and poses, following formats commonly used in visual SLAM benchmarks. The acquisition is scripted and free of stochastic components, enabling benchmarking and development of stereo depth estimation, visual odometry, and SLAM algorithms under controlled and repeatable indoor conditions.
This article presents a pilot Long Term Evolution (LTE) drive-test dataset collected from a real urban cellular network in Dhaka, Bangladesh. The dataset contains field measurements of radio signal conditions, Layer-3 measurement reports, and handover statistics obtained using professional drive-testing tools including XCAL-M software and Samsung Exynos-based user equipment. Measurements were recorded during multiple drive-test sessions along a 13 km urban route, producing timestamped observations of key radio parameters such as Reference Signal Received Power (RSRP), Reference Signal Received Quality (RSRQ), Carrier-to-Interference-plus-Noise Ratio (CINR), and serving/neighbor cell identifiers. The dataset is organized into raw and processed directories containing dynamic radio management (DRM) log files, Layer-3 measurement reports, handover event statistics, and tabular datasets to support reproducibility and flexible analysis. The processed dataset demonstrates an example preprocessing workflow including filtering, timestamp alignment, interpolation of missing signal values, and handover labeling with Time-to-Trigger (TTT) context. Due to the limited size of the dataset, it is primarily intended for exploratory analysis, mobility characterization, small-scale machine learning experiments, reinforcement learning environments, and educational research in cellular network engineering. By providing both raw drive-test logs and structured datasets, this resource supports reproducible studies of LTE handover behavior and urban radio environments [1].
Medieval and Renaissance Latin geographical works constitute a major source for understanding how space, places, and territories were described and conceptualised in pre-modern Europe. However, information about these works, their manuscript transmission, and the places they mention remains dispersed across catalogues, archives, and specialised scholarship. Here we present the IMAGO knowledge graph, a semantically structured dataset representing 343 Latin geographical works written between the 6th and the 15th centuries. The dataset integrates curated information provided by domain experts, including authors, works, manuscripts, printed editions, libraries, literary genres, and mentioned places. Data were initially collected in tabular form and subsequently enriched through semi-automatic reconciliation with external authority sources such as Wikidata and the MIRABILE digital archive. Domain experts further expanded the dataset using a dedicated annotation tool. The curated data were transformed into an OWL 2 DL knowledge graph aligned with the IMAGO ontology and published following FAIR and Linked Open Data principles. The knowledge graph was validated through automated reasoning, expert review, and query-based evaluation. The resulting dataset enables systematic exploration of textual, bibliographic, and spatial relationships within medieval and Renaissance geographical literature and supports reuse in historical, philological, and digital humanities research.
In this article, we present a morphologically annotated lexical dataset designed to support the translation of texts between the Khorezm dialect of Uzbek and standard Uzbek. The database consists of 1445 Khorezm dialect–standard Uzbek word pairs. For each entry, the dialect form, its corresponding standard form, and part of speech are provided, along with a morphological structure segmented into affixes according to such grammatical categories as number (singular/plural), possessive, and person/number properties. During the tagging process, the grammatical system of Uzbek and the specific inflectional properties of the Khorezm dialect were taken into account, resulting in a clear and machine-processable layer of correspondences between the dialect and the standard language. The resulting lexical resource is intended to serve as additional input data for AI-based translation and normalization models, as well as in applications for morphological analysis, spelling and grammar checking, and educational tools for teaching the Khorezm dialect. This dataset constitutes the first systematic lexical-morphological resource for corpus-based research on the Khorezm dialect and lays the groundwork for studies on machine translation and automatic alignment between dialects and the standard language in low-resource Turkic varieties.
Most of the cyber attacks are initiated through phishing URLs, which are shared with the victims through multiple media. In spite of the research community proposing varied solutions, the volume and nature of such attacks have evolved unpredictably. To develop effective solutions, researchers require comprehensive datasets that encompass a wide range of attack types rather than focusing on a narrow subset. We present CompPhish, a processed dataset, comprising of 15,358 URLs paired with their respective HTML sources. 7,204 URLs are phishing, and 8,154 are legitimate. A set of 70 features, specifically curated to capture the properties exhibited by phishing URLs, is extracted . These features represent a diverse range of phishing attacks, including generic URL-based phishing attacks, phishing through brand-jacking, phishing sites hosted on compromised domains, and auto-downloadable malicious files embedded in webpages. Phishing URLs are gathered from PhishTank and OpenPhish, while legitimate URLs are compiled from multiple independent sources, like DataforSEO and GitHub.By publishing the raw URLs and HTML codes along with the feature vectors, CompPhish attempts to aid the researchers in devising generalized, robust, and reliable machine learning based solutions for real-world phishing attempts.