
Conceptual metaphors structure abstract domains through concrete, perceptual, affective, or bodily experience, but they are accessed through metaphorical expressions. Controlled experimental work requires normed materials characterizing expressions, not only the conceptual mappings they instantiate. This is particularly relevant for European Portuguese, where research on metaphors lacks multidimensional normative resources for everyday metaphorical expressions. We introduce ME6D-PT, a database of 213 European Portuguese metaphorical expressions instantiating 20 conceptual metaphors and accompanied by English counterparts. Portuguese native speakers rated the expressions on six dimensions relevant to language processing: familiarity, concreteness, valence, arousal, embodiment, and transparency. Items were presented individually without context, organized in seven lists in a within-subject design. Aggregated item-level reliability ranged from moderate to excellent across dimensions. Descriptive results showed that the expressions were generally familiar and transparent while varying substantially in concreteness, valence, arousal, and embodiment. Further analyses showed strong positive associations among familiarity, transparency, and concreteness, whereas embodiment was weakly related to experiential and semantic dimensions. Exploratory quadratic models revealed nonlinear patterns involving valence, including the expected U-shaped association with arousal, and a decelerating positive association between transparency and familiarity. ME6D-PT provides a controlled resource for stimulus selection and research on metaphors, abstract language, embodiment, and future cross-linguistic research.
Mobile self-reports provide subjective and contextual information that passive sensors cannot recover, but frequent prompts compete with everyday activities and are often answered late. Here, we show that situational context and recent response history predict whether an ecological momentary assessment response begins within 15 min. We analyzed 70,375 records from 170 participants in the Italian arm of DiversityOne; 43.19% met the operational timeliness criterion. Participant-clustered generalized estimating equations identified associations with activity, social setting, location, mood, weekday, hour and survey day, with corrected Cramér’s V values of 0.044–0.175. In leakage-resistant evaluation, gc-Forest achieved an accuracy of 0.719 for personalized forward prediction, whereas history-augmented LightGBM achieved an accuracy of 0.704 and area under the receiver operating characteristic curve of 0.779 when participants were held out. Retrospectively ranking ten candidate moments increased the participant-balanced timely-response rate from 44.60% to 51.43%. These results establish prompt timeliness as a learnable scheduling outcome, while distinguishing it from content accuracy or cognitive effort.
Alpine meadows in the eastern Qinghai–Tibet Plateau serve as core water retention zones within China’s Three-River Source Region, while simultaneously functioning as vital grassland pastures. Livestock breeding activities impose severe disturbances on aquatic environments across this plateau area, raising concerns regarding microbial community shifts and associated environmental risks. Previous studies have explored how pastoral farming reshapes microbial assemblages within alpine meadow water bodies, yet few investigations have addressed such effects on aquatic microorganisms across an entire watershed scale. In the present study, we constructed a metagenomics dataset based on next-generation sequencing to survey spatial shifts in aquatic microbial composition and community structure along a livestock pollution gradient spanning the entire course of the Baihe River, a tributary of the upper Yellow River. Taxonomic classification revealed that domain bacteria dominated all microbial communities, with their relative abundances positively correlated with the intensity of livestock contamination. Additionally, microbial richness was markedly higher in stagnant headwater habitats than in lotic downstream reaches. The generated metagenomic dataset advances our mechanistic understanding of how livestock pollution drives spatial variations in aquatic microbial assemblages in alpine meadow waters and provides a critical baseline for environmental risk management and watershed microbial monitoring on the eastern Qinghai–Tibet Plateau.
Traditional Earned Media Value (EMV) calculations suffer from methodological limitations by depending on oversimplified and outdated proxies like Advertising Value Equivalency (AVE). In service marketing, where offerings are intangible and consumers cannot inspect an offering before purchase, trust relies heavily on authentic influencer communication and peer advocacy; these measurement flaws are therefore especially problematic. This paper introduces two innovations tailored for the modern media ecosystem: (1) Mention Quality and Impact (MQI), a transparent scoring system (1–10 scale) that evaluates sentiment, engagement, credibility, and AI visibility, with a bounded adjustment for conversion intent; and (2) Mention Earned Media Value (mEMV), a CPM-based valuation model that integrates MQI with platform-specific benchmarks. By accounting for participation inequality, influencer dynamics, and the rise of virtual search environments, this framework advances media measurement for service brands toward greater accuracy, transparency, and scalability.
Quantitative flowering phenotypes are needed to support breeding and harvest management in Hypericum perforatum L. (St. John’s wort), but manual flower assessment is slow and difficult to standardize under field conditions. UAV-Hyp is a multi-temporal UAV RGB dataset containing 12,653 high-resolution images acquired at 26 measurement dates across the complete flowering period of 15 H. perforatum accessions. The images represent variable illumination, soil moisture, weed pressure, and developmental stages. The dataset provides 59,163 plant bounding boxes and 107,054 flower bounding boxes. As an application example, cascaded YOLOv8 plant and flower detectors achieved mAP@0.50:0.95 values of 0.977 and 0.950, respectively. UAV-Hyp supports scalable flower quantification and the development of time-series phenotyping methods for genotype comparison and quality-oriented medicinal-plant breeding.
The development of intrusion detection and network security solutions for securing Internet of Things (IoT) networks is constrained by the limited availability of representative network security datasets. Many existing datasets rely on centralised traffic collection and do not capture the non-Independent and Identically Distributed (non-IID) characteristics inherent to edge environments. To address this limitation, this work presents a device-level IoT network dataset generated using the open-source Gotham testbed, a virtualised smart city environment. Network traffic is collected in a distributed manner at the interfaces of 78 heterogeneous IoT devices operating across multiple protocols, including MQTT, CoAP, and RTSP. The dataset comprises over 31.8 million packet-level records, each described by 22 features. It includes both benign traffic and multiple attack classes, namely Network Scanning, Brute Force, Infection, Denial of Service (DoS), and Command and Control (C&C) Communication. Ground-truth labels are assigned using a deterministic process based on orchestration logs. The dataset preserves device-level traffic distributions and captures non-IID characteristics without artificial partitioning. It is publicly available and can be used to support reproducible evaluation of intrusion detection approaches and network analysis tasks in both centralised and distributed learning settings.
Frost prediction in tropical high-mountain agricultural regions is difficult because sparse meteorological networks must represent strong terrain-driven microclimatic variability. This article presents a topoclimatic graph dataset for frost prediction in the Altiplano Cundiboyacense, Colombia. The released core dataset contains 23 agricultural weather stations and is seasonally focused on recurrent November–February frost periods rather than year-round continuous monitoring. It includes four consistently available meteorological variables at 30 min resolution: air temperature, relative humidity, dew point temperature, and solar radiation. The data were consolidated from multiple operational sources, harmonized to a common temporal grid, subjected to physical and consistency-based quality control, and completed through temporal and spatial reconstruction with traceability labels. The final release also provides binary frost_event and frost_warning_6h labels, point-based topographic descriptors, 1 km buffer-based raster summaries, land-cover proportions, station-level static feature vectors, and graph products including edge lists and adjacency matrices. These data products support graph-based deep learning, multimodal spatiotemporal analysis, and frost early warning experiments in a tropical mountain agroecosystem. The dataset offers a reproducible framework for integrating heterogeneous environmental observations into graph-ready representations while preserving sufficient environmental context for benchmarking frost prediction methods in data-sparse regions.
This work presents conditioned and normalized vibration signal datasets acquired from the spanwise axis of the three blades of a rotating bladed system operating at a constant rotational speed of 240 rpm. The conditioned dataset was obtained using piezoelectric accelerometers mounted at the blade roots. The accelerometer output signals were conditioned and recorded by a dedicated data acquisition system. The signals were acquired under both healthy and damaged operating conditions. Baseline vibration signals were first recorded with all three blades in a healthy condition. Subsequently, cracks were deliberately introduced at three different locations along the blade span—the root, middle, and tip zones. Each crack location was independently evaluated on each of the three blades, resulting in a comprehensive dataset that includes healthy operation and all combinations of blade–damage locations. The datasets enable analysis of the system’s vibratory response and of dynamic information propagation toward the blade root, depending on the crack zone. Their main contribution is to provide reliable experimental data for the development, validation, and benchmarking of vibration-based diagnostic and structural health monitoring techniques. Furthermore, the datasets serve as valuable resources for advancing early crack detection strategies and enhancing the reliability of rotating industrial equipment with blades, such as fans, compressors, turbines, and aerogenerators.
The growing complexity of urban mobility requires datasets that integrate dynamic traffic observations with meteorological, geometric, and urban-context information. This study presents MUTra-CDMX, a multisource urban traffic dataset covering a 14.72 km section of the Insurgentes Sur corridor in Mexico City. Traffic data were obtained from TomTom at five-minute intervals for 20 consecutive road segments from 1 November 2024 to 28 February 2025. Hourly meteorological data were retrieved from Meteosource, while segment-level geometry, topology, signalized locations, and nearby points of interest were derived from TomTom metadata and OpenStreetMap. The primary analytical file contains 691,200 segment–timestamp records and 12 variables describing traffic and free-flow conditions, meteorological information, derived operational indicators, and reconstruction status. Of these records, 682,264 are original observations and 8936 are reconstructed segment–timestamp combinations, identified by the Boolean variable is_imputed. Technical validation confirmed complete temporal coverage, preservation of original traffic observations, consistent weather alignment, and reconstruction performance through artificial masking. Predictive utility was evaluated through chronological travel-time forecasting under a leakage-controlled protocol. At the 30 min horizon, XGBoost achieved a mean absolute error of 12.84 s, a root mean squared error of 37.91 s, and a coefficient of determination (R2) of 0.771, outperforming a persistence baseline. MUTra-CDMX supports congestion analysis, imputation studies, spatiotemporal modeling, and travel-time forecasting.
Background: Borderline dengue immunoglobulin M (IgM) enzyme-linked immunosorbent assay (ELISA) results create diagnostic and epidemiological uncertainty. Collapsing the manufacturer’s three qualitative categories into a binary outcome changes the effective decision rule and may materially alter the apparent positivity proportion. This Data article describes an anonymized laboratory dataset from Trinidad and Tobago and evaluates both category-handling uncertainty and test-misclassification uncertainty. Methods: The dataset comprises 161 consecutive serum specimens submitted for routine dengue IgM testing between 1 September 2025 and 28 February 2026 and processed in four analytical batches. The results were analyzed under three prespecified scenarios: borderline classified as negative, borderline excluded, and borderline classified as positive. Exact Clopper–Pearson 95% confidence intervals (CIs) were calculated. Rogan–Gladen adjustment was applied only to the borderline-excluded scenario because the ELISA test validation estimates of sensitivity and specificity were calculated after excluding borderline results. Results: Twenty specimens (12.4%) were positive, 29 (18.0%) borderline, and 112 (69.6%) negative. The apparent IgM positivity was 12.4% (20/161; 95% CI 7.8–18.5%) when borderline results were classified as negative, 15.2% (20/132; 95% CI 9.5–22.4%) when they were excluded, and 30.4% (49/161; 95% CI 23.4–38.2%) when they were classified as positive. For determinate results, the conditional Rogan–Gladen estimates were 13.2% using a sensitivity of 100.0% and specificity of 97.7%, and 15.1% using a sensitivity of 82.2% and specificity of 96.8%. Conclusions: Misclassification adjustment is informative but remains conditional on the transportability of external assay-performance estimates. The dataset supports transparent category-level reanalysis, while the absence of continuous index values, confirmatory testing, and detailed clinical metadata limits numerical cut-off recalibration and patient-level inference.
Software-Defined Networking (SDN) separates the control and data planes, introducing a logically centralized controller that is itself a high-value attack target. Despite growing interest in SDN intrusion detection, publicly available datasets either restrict evaluation to binary normal-vs-DDoS classification or lack control-plane telemetry, leaving multi-class detection of SDN-architectural attacks without a dedicated benchmark. This work presents LAN-SDN-NIDS, a publicly available, multi-class flow-level dataset of 1,125,059 records generated in a fully containerized Containernet/OpenDaylight testbed across five standard network topologies. Each flow record combines 29 traffic-level features with 11 control-plane-aware metrics—including Packet-In and Flow-Mod counts and first-seen delay. The dataset covers five attack classes in two categories: three that exploit SDN control-plane mechanisms (link fabrication, host injection, and port hijack) alongside DDoS and port scan, plus normal traffic. An XGBoost classifier trained on the full feature set achieved a macro F1 of 0.94; an ablation study showed that removing OpenFlow features causes link fabrication F1 to collapse from 0.97 to 0.19, indicating that control-plane telemetry is decisive for detecting SDN-architectural attacks under the conditions evaluated. A UMAP embedding is consistent with class separability, except for a structural overlap between host injection and normal traffic attributable to their shared ARP protocol.
Automatic extraction of genealogical information from historical archival-genealogical documents in Uzbek is an understudied problem for low-resource languages. Multi-layer NLP benchmarks are not sufficient to automatically identify individuals, family relationships, dates, place names, and archival identifiers in such texts. Also, the same people are mentioned in various forms: full name, pronoun (18.8%), initial, surname-name order, indirect expression (9.4%), and title. Existing NER and relation extraction corpora are mainly focused on high-resource languages or general domain texts and do not sufficiently cover the FAMILY_ROLE signals, historical spelling variants, and fond–opis–delos identifiers specific to Uzbek archival-genealogical texts. Proposed resource: We present the ArchiveGene Corpus, a controlled, fully synthetic, and reproducible five-layer resource consisting of 1000 Uzbek archival-genealogical-style documents, divided into 700 training, 150 validation, and 150 test documents. The corpus contains 8366 named entities, 10,625 person mentions, 2000 coreference chains, and 1000 genealogical relation triples. The dataset was generated using a deterministic template-based pipeline and a lexicon of Uzbek names, and is fully reproducible. Inter-annotator agreement values were 0.847 for NER, 0.793 for coreference, and 0.821 for RE, according to Cohen’s κ. Comparative results are presented with four baseline models (rule-based, BiLSTM-CRF, mBERT, and XLM-RoBERTa). The dataset is openly hosted on the Zenodo platform under the CC BY 4.0 license; concept DOI: 10.5281/zenodo.20670360, v1.1.1 version DOI: 10.5281/zenodo. 21429998. Scientific significance: To the best of our knowledge, ArchiveGene is among the first openly released, controlled synthetic resources for Uzbek that integrates named-entity recognition, person-mention detection, coreference resolution, genealogical relation extraction, and final tuple generation within a single annotation framework. The baseline analysis provides three main conclusions: (1) on the clean synthetic test set, the transformer models already reach 100.00 Micro-F1 for NER and 100.00 Macro-F1 for coreference-aware relation extraction, so coreference aggregation adds little on synthetic data (+2.25 for mBERT and +0.04 for XLM-RoBERTa) but its contribution is expected to grow on real archival text; (2) the rule-based and heuristic baselines lag far behind (Macro-F1 50.28 and 70.73) and fail entirely on spouse_of, showing the limits of lexical rules; and (3) a zero-shot evaluation on a real-document pilot reduces NER Micro-F1 from 100.00 to 22.17, indicating that the synthetic corpus is trivially learnable and that real-archival validation is essential.
Background: Dental insurance claims data are vital for research in oral health, epidemiology, and policy. However, issues like data quality, coding standards, validity, interoperability, and analytical approaches hinder their use. This review outlines these challenges. Methods: Following PRISMA-ScR and Arksey-O’Malley, we searched Web of Science, Scopus, and PubMed through April 2026 for peer-reviewed studies on dental insurance data issues. Two reviewers screened and extracted data, identifying key challenges and implications. Results: Out of 563 records, 389 remained after deduplication; 45 studies met criteria. Data sources included Medicaid, Medicare, insurers, and national systems from various countries. Six main challenges emerged: (1) coding errors and lack of standardization; (2) data validity and quality concerns; (3) interoperability and linkage barriers; (4) fraud detection issues; (5) analytical limitations; (6) policy insights on disparities. Validation showed variable accuracy, with diagnosis codes more reliable than procedure codes. Conclusions: Challenges limit data use in research and policy. Standardized coding, validation, interoperability, transparency, and causal inference are essential for leveraging these data to improve oral health research and policies.
The Eddy Covariance station provides observations of agrometeorological variables and surface energy fluxes, collected from 2019 to 2022, in a citrus orchard located in the Souss-Massa plain, Morocco. The present dataset comprises measurements recorded via a set of aboveground and subsurface sensors. The aboveground setup consistently measures air temperature, relative humidity, wind speed, net radiation, and precipitation. Additionally, the subsurface setup continuously tracks soil temperature, moisture, and electrical conductivity at depths from 5 to 80 cm, along with soil heat flux. Moreover, these setups enable the measurement of turbulent fluxes (sensible and latent heat). Given the limited availability of long-term agrometeorological data in semi-arid regions of the Mediterranean, this paper addresses a critical data gap by providing a reliable agrometeorological dataset. The latter consists of two types of data: 30 min interval files and high-frequency files (20 Hz, i.e., one measurement every 50 ms). The processing of this data involved Card Convert, MATLAB EC-Pack, and Excel, with data quality control performed by removing outliers and excluding nighttime fluxes. The dataset is organized in a table and provided in a .csv format with standard metadata. It is designed for a wide range of applications, including evapotranspiration modeling, satellite product validation, agroclimatic monitoring, determining crop irrigation requirements, precision irrigation planning, and water management. Additionally, the dataset can be reused for crop and hydrological model calibration, as well as soil moisture and crop stress prediction using machine learning algorithms.
The majority of currently available hand kinematic databases have been gathered using expensive marker-based systems or are restricted to a particular gesture-recognition task, failing to capture the dynamic nature of joints when the hand is engaged with an object. To address this gap, we introduce the RGB-based Hand Joint Kinematics (RGB-HJK) dataset, a publicly available collection of continuous, frame-level 3D joint angle trajectories, recorded while ten healthy adults (six male, four female; age 25.8 +/- 3.2 years; BMI 22.8 +/- 2.0 kg/m(2)) performed five standardized object interaction grasps: Power Grasp (cylindrical bottle), Tripod Grasp (pen), Static Power Hold (smartphone), Precision Pinch (thin paper), and Lateral Pinch (book). Data were collected using a standard RGB camera and the MediaPipe Hands markerless pipeline at 26.95 +/- 0.29 Hz, a rate that was stable across all subjects. Each participant completed five trials for each grasp type. After filtering using active hold, 28,111 validated frames remained, with a 100% detection rate for all 250 trials. Intra-subject repeatability was good (mean SD <= 7.9 degrees across all joint grasp combinations) and inter-subject variability was within the range expected based on normal anatomical diversity. Importantly, kinematic validation of the Index Proximal Interphalangeal (PIP) joint (61.8 degrees +/- 18.4 degrees) showed values consistent with ranges reported in previous studies using instrumented gloves and depth sensors. Principal Component Analysis (PCA) confirmed clear linear separability among the five grasp configurations. Unlike existing datasets, the RGB-HJK method does not compromise the natural sense of touch and is free of hardware occlusions, thereby providing an easily accessible ecological baseline.
LeafScans-Orchard is a curated, multi-year RGB image dataset of orchard plant leaves designed to support research in computer vision, machine learning, and plant phenotyping. The dataset comprises 9708 high-quality leaf scans acquired during collection campaigns conducted between 2015 and 2025, covering seven orchard crop species: apple, pear, sweet cherry, sour cherry, plum, peach, and apricot. In total, the dataset includes 67 cultivar labels. All samples were acquired using flatbed scanning under controlled conditions on a uniform background, ensuring high visual consistency and minimal background variability. The original scans were captured at 1200 dpi and subsequently converted into a public release format at 300 dpi, stored as lossless TIFF images to preserve morphological and textural details. Each image corresponds to a single leaf and is organized in a hierarchical directory structure by species, cultivar, and acquisition year, accompanied by image-level metadata and aggregated species–cultivar–year counts. LeafScans-Orchard is suitable for plant species classification, cultivar recognition, leaf morphology analysis, texture analysis, and general visual feature extraction. In addition to the main release, a representative subset of 300 original 1200 dpi scans is provided to support high-resolution analyses. The dataset is particularly suited for fine-grained classification, morphology-driven analysis, and methodological studies under controlled imaging conditions.
We present the use of iOrganoAssay (images of Organoid Assay) to connect microscopy images with organoid assessment assays such as live-dead, immunocytochemistry, and drug treatment assays. The iOrganoAssay consists of an R script-based application (App) interface and datasets encompassing (1) microscopy images, (2) segmentation results, (3) morphometric data, (4) a metadata file, and (5) a validation dataset. The microscopy image collection includes 234 large-area images of intestinal organoids cultured in Matrigel dome region (similar to 3 mm), acquired using an automated stage-equipped microscopy system. Upon treatment with dextran sulfate sodium (DSS), microscopy images of morphological changes in intestinal organoids were captured and quantified. Image segmentation was performed to extract organoid morphological data, including area, perimeter, and circularity. These metrics were plotted to visualize daily variations, enabling systematic tracking of drug-induced morphological changes over time. Statistical comparisons were also provided using violin plots. To evaluate segmentation quality, we established a validation dataset of 28 manually annotated organoids (14 control, 14 DSS-treated) and calculated Dice scores, accuracy (Acc), segmentation error (SegErr), and centroid error (CenErr). This integrated dataset-covering organoid images, segmentation outputs, morphometric data, and validation metrics-provides a resource for organoid image-based studies in morphological monitoring and segmentation validation.
Early prediction of academic outcomes is vital to enabling timely intervention, supporting at-risk students, and improving educational planning and institutional performance. However, this task becomes particularly challenging when data availability is limited, such as in small or graduate-level programs. This study explores the potential of data augmentation techniques, specifically the Synthetic Minority Oversampling Technique, to enhance the performance of machine learning models applied to such constrained educational datasets. We conduct a comparative analysis using four datasets derived from prior research, each representing a distinct educational use case: one focused on predicting academic success in graduate programs, another on student dropout in virtual learning environments, a third on dissertation performance prediction, and a fourth addressing multi-class performance prediction in undergraduate coding courses. By applying consistent machine learning methods in the original and augmented datasets, we systematically evaluate the impact of data augmentation on classification performance using accuracy, precision, recall, and the F1 score. The results demonstrate marked improvements, with accuracy increases up to 21% and precision gains exceeding 25% in some models, notably with KNN and MLP. While not all algorithms benefit equally, our findings highlight data augmentation as a practical and impactful strategy for improving early prediction capabilities in Educational Data Mining (EDM). By leveraging multiple datasets and diverse educational contexts, this contribution provides robust evidence supporting the broader goal of enhancing decision-making and personalized support in digital learning environments.
Significant progress in legal natural language processing (NLP) has enabled advancements in tasks such as legal judgment prediction, case retrieval, and question answering. However, the development of analogous technologies for Arabic legal texts remains severely constrained by the scarcity of large-scale, publicly available benchmarks for summarisation and classification. This paper addresses this gap by introducing a novel, comprehensive dataset of 9699 Arabic legal cases sourced from the Saudi Board of Grievances. This corpus is unique in pairing full-length court decisions with expertly human-crafted abstractive summaries and multi-class category labels (Administrative, Commercial, and Criminal), establishing a dedicated benchmark for Arabic legal NLP. The dataset was constructed via a robust, reproducible pipeline that ensures high textual fidelity, incorporating specialised optical character recognition (OCR) via Google Document AI and precise structural segmentation into facts, reasons, and summaries. To establish robust baselines, we conduct an extensive empirical evaluation of seven summarisation models—encompassing four extractive algorithms (TextRank, LexRank, Latent Semantic Analysis, and Luhn) and three transformer-based abstractive architectures (AraT5v2, AraBART, and mBART)—each evaluated in both base and fine-tuned configurations. Results across ROUGE, BERTScore, BLEU metrics and human evaluation demonstrate substantial performance gains achieved through domain-specific fine-tuning, with the fine-tuned AraBART model achieving the strongest performance among all evaluated models. Furthermore, we present a novel analysis of the downstream utility of generated summaries by evaluating their performance on legal category classification using five machine learning models. This investigation reveals a strong positive correlation between summarisation quality and classification accuracy, empirically demonstrating that domain-adapted abstractive summarisation not only enhances intrinsic evaluation scores but also significantly boosts extrinsic task performance. By providing this essential dataset and comprehensive benchmarking, our work contributes a much-needed resource to the field, facilitating future research and innovations in Arabic legal text analysis.
The Marmara region of T & uuml;rkiye, situated along the North Anatolian Fault Zone (NAFZ), constitutes one of the most seismically active and densely monitored zones globally. Given the region's high vulnerability and the catastrophic impacts of historical events-notably the 1999 & Idot;zmit and 2023 Kahramanmara & cedil;s sequences-there is a critical need for advanced seismic hazard risk assessment (SHRA) methods that move beyond static models. This review examines the paradigm shift from traditional geophysics to big data seismology, characterized by the "Five Vs": volume, velocity, variety, veracity, and value. Critically, we distinguish between two fundamentally different problems: Earthquake Early Warning (EEW), which operates on sub-second timescales after rupture initiation, and probabilistic earthquake forecasting, which operates on timescales of years to decades. The study discusses how cloud-native platforms such as Azure Databricks, combined with data pipelines using Apache Kafka (version 3.5.1) and Apache Spark (version 4.1.2), enable the real-time processing of petabyte-scale seismic sensor streams. Key technological tools, including Physics-Informed Neural Networks (PINNs) and deep learning models such as PhaseNet, are analyzed for their demonstrated ability to enhance EEW systems through sub-second phase picking and automated event detection. Seismic tomography is also undergoing AI-enabled transformation, yielding higher-resolution subsurface imaging. We present statistical validation metrics and uncertainty quantification methods essential for credible hazard assessment. By addressing computational bottlenecks through hybrid computing architectures and edge computing, this framework aims to improve the warning lead time for Istanbul's critical infrastructure. This work provides a structured roadmap for bridging the gap between traditional seismic data analysis and operational predictive analytics in the Marmara region.