The Arabic language is characterized by a rich tapestry of regional dialects that differ substantially in phonetics and lexicon, reflecting the geographic and cultural diversity of its speakers. Despite the availability of many multi-dialect datasets, mapping speech to fine-grained dialect sources, such as cities, remains underexplored. We present ARCADE (Arabic Radio Corpus for Audio Dialect Evaluation), the first Arabic speech dataset designed explicitly with city-level dialect granularity. The corpus comprises Arabic radio speech collected from streaming services across the Arab world. Our data pipeline captures 30-second segments from verified radio streams, encompassing both Modern Standard Arabic (MSA) and diverse dialectal speech. To ensure reliability, each clip was annotated by one to three native Arabic reviewers who assigned rich metadata, including emotion, speech type, dialect category, and a validity flag for dialect identification tasks. The resulting corpus comprises 6,907 annotations and 3,790 unique audio segments spanning 58 cities across 19 countries. These fine-grained annotations enable robust multi-task learning, serving as a benchmark for city-level dialect tagging. We detail the data collection methodology, assess audio quality, and provide a comprehensive analysis of label distributions. The dataset is available on: https://huggingface.co/datasets/riotu-lab/ARCADE-full
The exponential growth of the Internet of Things (IoT) ecosystem demands accurate and explainable intrusion detection systems (IDS). This study introduces NEXUS-IDS, a framework for trustworthy IoT security analytics that integrates graph-based deep learning, Explainable AI (XAI), and Large Language Models (LLMs). Central to the framework is a Flow-Node Graph Neural Network (IDS-FN-GNN), which models network flows as nodes and communication relationships as edges to capture relational attack patterns. To enhance transparency, a dual-layer explainability mechanism is proposed that combines GraphLIME-based structural feature attributions with LLM-driven semantic explanations. This approach translates the model outputs into analyst-readable threat narratives supported by entropy-aware uncertainty calibration. Experiments on the Edge-IIoTSet and IoTID20 datasets yielded detection accuracies of 94.30% and 97.42%, respectively. At the default threshold, τ=0.3, entropy-based filtering forwards 19.33% of total incoming flows on the Edge-IIoTSet and 32.08% on IoTID20 directly to the semantic layer. Within this forwarded subset, the semantic cache resolves the majority of requests, yielding cache hit rates of 14.33% and 27.29% of the total incoming flows, respectively, thereby successfully reducing fresh LLM invocations to a minimal 5.00% and 4.78%. An automated LLM-as-a-judge assessment and a structured human expert evaluation across four architectures (82M–70B parameters) demonstrated high alignment. Human evaluation of a random sample verified that the foundation models achieved mean expert quality scores of up to 4.2/5.0. However, expert-flagged hallucination rates varied widely (0%–77%), indicating that perceived quality and factual reliability require independent assessment. Consequently, the LLM-driven semantic layer is positioned as an assistive research prototype. While the underlying detection and structural attributions remain robust, operationalizing these semantic explanations currently necessitates Human-in-the-Loop (HITL) oversight to ensure appropriate trust calibration.
Medical cities serve large populations and experience high parking demand due to heterogeneous visitor profiles and time constraints. This paper addresses the parking allocation problem in such environments by proposing a grouping-based optimization framework that allocates spaces according to the total number of passengers per vehicle rather than vehicle count alone. The proposed system integrates passenger-count information and real-time occupancy data through a coordinated architecture involving booking interfaces, IoT-based sensing, and a centralized allocation engine. Six novel algorithms are developed to achieve equitable load distribution across parking zones. Their performance is evaluated using nine instance classes and 2,790 test cases. Experimental results demonstrate that the $GVF$ algorithm achieves the highest best-solution percentage (68.2%) and the lowest average gap among the proposed algorithms, confirming its scalability and robustness in large-scale medical city environments
Link failures are a critical issue that disrupts consistent communications in the fast-evolving Internet-of-Vehicles (IoV) environment. In response to the aforementioned limitation, we introduce the rapid discovery algorithm for routes (RADAR), a novel scheme designed for rapid route discovery in the event of abrupt link failures or topological changes. Built on a heuristic Theta⁎ algorithm, RADAR operates within a three-layered Software-Defined Networking (SDN) architecture to ensure seamless communications. The first layer collects vital vehicle information, including location, speed, direction, and vehicle ID. The second layer uses an SDN edge controller to carry out local and intra-zone path discovery with the Theta⁎ heuristic algorithm. The third layer introduces a global controller that oversees inter-zone route discovery, thereby expanding the network's reach and capabilities. To validate RADAR's effectiveness, we compare its performance to those of Dijkstra-based and A⁎ path planning algorithms across key metrics—Route Discovery Messages (RDM), Path Length (PL), and Route Discovery Time (RDT). Our analysis shows significant improvements with RADAR over Dijkstra, achieving up to an 80% increase in RDM, 20% in PL, and 70% in RDT. It also surpasses A⁎ by up to 70% in RDM and 60% in RDT. These results indicate that RADAR enhances routing reliability in IoV systems, thereby improving overall network performance.
The integration of Internet of Things (IoT)-enabled consumer medical devices, platform and professional expertise has significantly improved clinical diagnosis accuracy. However, privacy concerns and variations in consumer device-specific or hospital-specific data analyses hinder centralized processing of massive multi-label medical image data. This paper proposes a Cluster-Fair Federated Learning (CFFL) framework to address data heterogeneity in distributed IoT-based consumer healthcare systems. Using an end-to-end Dual Decoder Network (DDN), medical consumers are grouped into clusters by data similarity, with cluster-specific models optimized locally and aggregated through intra-cluster volume-weighted and inter-cluster contribution-based strategies. Experiments on the Chest X-ray14 and CheXpert datasets show that the proposed CFFL improves diagnostic accuracy by 6.2% on average compared to baselines and achieves performance close to centralized learning while ensuring privacy. These results underscore the framework potential for secure, efficient, and scalable distributed consumer medical image analysis.
Remote sensing (RS) has witnessed significant progress over the past decade, driven by the increasing spatial, spectral, and temporal resolution of satellite imagery. Yet, scene classification models trained on a specific domain often exhibit severe performance degradation when deployed in new environments due to distribution shifts arising from variations in geography, acquisition conditions, sensor specifications, and atmospheric noise. While Unsupervised Domain Adaptation (UDA) has been widely explored to alleviate this challenge, most UDA methods require concurrent access to source data and target samples during adaptation. This assumption remains unrealistic in many operational RS scenarios where source data may be confidential, storage-restricted, or unavailable after deployment. To address this limitation, we introduce CAB-SFDA, a Class-Aware Balanced Source-Free Domain Adaptation framework that adapts a pretrained classifier to an unlabeled target domain without access to source data. The proposed method preserves source decision boundaries by freezing the classifier while refining the feature extractor using a momentum-based prototype memory. Stable pseudo-label learning is achieved through inverse-frequency class weighting and information maximization, whereas weak-strong consistency regularization enhances robustness against target distribution shifts. We evaluated our framework across 12 cross-scene adaptations and 2 cross-sensor RS benchmarks. Our method achieves superior performance over existing models, with accuracy gains of up to 2.47% and 3.32% across the two primary evaluation settings, respectively.
Self-supervised learning (SSL) is a standard approach for representation learning in aerial imagery. Existing methods enforce invariance between augmented views, which works well when augmentations preserve semantic content. However, aerial images are frequently degraded by haze, motion blur, rain, and occlusion that remove critical evidence. Enforcing alignment between a clean and a severely degraded view can introduce spurious structure into the latent space. This study proposes a training strategy and architectural modification to enhance SSL robustness to such corruptions. It introduces a per-sample, per-factor trust weight into the alignment objective, combined with the base contrastive loss as an additive residual. A stop-gradient is applied to the trust weight instead of a multiplicative gate. While a multiplicative gate is a natural choice, experiments show it impairs the backbone, whereas our additive-residual approach improves it. Using a 200-epoch protocol on a 210,000-image corpus, the method achieves the highest mean linear-probe accuracy among six backbones on EuroSAT, AID, and NWPU-RESISC45 (90.20
Vehicle-to-Everything (V2X) networks face significant challenges in achieving optimal, reliable routing due to their large scale, high mobility, and rapidly changing topologies. Traditional routing approaches have been applied to address this highly dynamic behavior, but they often fail to deliver steady performance in dynamic environments. In this study, we introduce Quantum-Heuristic Routing for V2X (QHR-V2X), a novel paradigm that integrates the heuristic efficiency of the A* algorithm with quantum-inspired amplitude amplification. The proposed scheme is designed to enable rapid and optimal route discovery while reducing redundant exploration. This makes it particularly suitable for latency-sensitive vehicular environments. To validate its effectiveness, QHR-V2X was evaluated against two classical baselines, Dijkstra and A*, within grid-based simulation environments of increasing size and obstacle density. Experimental results show that QHR-V2X achieves significantly lower route-discovery time and reduced message overhead compared to classical algorithms, while maintaining path lengths nearly identical to those of the optimal routes produced by Dijkstra. These findings demonstrate the potential of QHR-V2X to improve scalability and responsiveness in next-generation V2X communications.
Arabic is a linguistically and culturally rich language with a vast vocabulary that spans scientific, religious, and literary domains. Yet, large-scale lexical datasets linking Arabic words to precise definitions remain limited. We present MURAD (Multi-domain Unified Reverse Arabic Dictionary), an open lexical dataset with 96,243 word-definition pairs. The data come from trusted reference works and educational sources. Extraction used a hybrid pipeline integrating direct text parsing, optical character recognition, and automated reconstruction. This ensures accuracy and clarity. Each record aligns a target word with its standardized Arabic definition and metadata that identifies the source domain. The dataset covers terms from linguistics, Islamic studies, mathematics, physics, psychology, and engineering. It supports computational linguistics and lexicographic research. Applications include reverse dictionary modeling, semantic retrieval, and educational tools. By releasing this resource, we aim to advance Arabic natural language processing and promote reproducible research on Arabic lexical semantics.
We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a continuous geographic space rather than discrete categories. Speaker origin is predicted as continuous latitude-longitude coordinates using a hierarchical neural architecture that fuses frame-level XLS-R-300M and Whisper-large-v3 encoder representations with phonotactic descriptors through a Transformer encoder and a learnable attention-pooled query. A spherical geodesic loss directly optimizes great-circle distance on Earth's surface, avoiding distortions inherent to planar coordinate regression. Under a leakage-free 5-fold GroupKFold protocol grouped by source recording, our model attains a pooled median localization error of 481.2 km. Auxiliary country and city heads reach 64.5
The evolution of Large Language Models (LLMs) from passive text generators to autonomous, goal-driven systems represents a fundamental shift in artificial intelligence. This chapter examines the emergence of agentic AI systems that integrate planning, memory, tool use, and iterative reasoning to operate autonomously in complex environments. We trace the architectural progression from statistical models to transformer-based systems, identifying capabilities that enable agentic behavior: long-range reasoning, contextual awareness, and adaptive decision-making. The chapter provides three contributions: (1) a synthesis of how LLM capabilities extend toward agency through reasoning-action-reflection loops; (2) an integrative framework describing core components perception, memory, planning, and tool execution that bridge LLMs with autonomous behavior; (3) a critical assessment of applications and persistent challenges in safety, alignment, reliability, and sustainability. Unlike existing surveys, we focus on the architectural transition from language understanding to autonomous action, emphasizing the technical gaps that must be resolved before deployment. We identify critical research priorities, including verifiable planning, scalable multi-agent coordination, persistent memory architectures, and governance frameworks. Responsible advancement requires simultaneous progress in technical robustness, interpretability, and ethical safeguards to realize potential while mitigating risks of misalignment and unintended consequences.
This study addresses the critical gap in Arabic natural language processing by developing an effective Arabic Reverse Dictionary (RD) system that enables users to find words based on their descriptions or meanings. We present a novel transformer-based approach with a semi-encoder neural network architecture featuring geometrically decreasing layers that achieves state-of-the-art results for Arabic RD tasks. Our methodology incorporates a comprehensive dataset construction process and establishes formal quality standards for Arabic lexicographic definitions. Experiments with various pre-trained models demonstrate that Arabic-specific models significantly outperform general multilingual embeddings, with ARBERTv2 achieving the best ranking score (0.0644). Additionally, we provide a formal abstraction of the reverse dictionary task that enhances theoretical understanding and develop a modular, extensible Python library (RDTL) with configurable training pipelines. Our analysis of dataset quality reveals important insights for improving Arabic definition construction, leading to eight specific standards for building high-quality reverse dictionary resources. This work contributes significantly to Arabic computational linguistics and provides valuable tools for language learning, academic writing, and professional communication in Arabic.
Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print media into digital formats. However, training modern OCR models, especially powerful vision-language models, is hampered by the lack of large, diverse, and well-structured datasets that mimic real-world book layouts. Existing Arabic OCR datasets often focus on isolated words or lines or are limited in scale, typographic variety, or structural complexity found in books. To address this significant gap, we introduce SARD (Large-Scale Synthetic Arabic OCR Dataset). SARD is a massive, synthetically generated dataset specifically designed to simulate book-style documents. It comprises 843,622 document images containing 690 million words, rendered across ten distinct Arabic fonts to ensure broad typographic coverage. Unlike datasets derived from scanned documents, SARD is free from real-world noise and distortions, offering a clean and controlled environment for model training. Its synthetic nature provides unparalleled scalability and allows for precise control over layout and content variation. We detail the dataset's composition and generation process and provide benchmark results for several OCR models, including traditional and deep learning approaches, highlighting the challenges and opportunities presented by this dataset. SARD serves as a valuable resource for developing and evaluating robust OCR and vision-language models capable of processing diverse Arabic book-style texts.
This paper addresses critical gaps in Arabic language model evaluation by establishing comprehensive theoretical guidelines and introducing a novel evaluation framework. We first analyze existing Arabic evaluation datasets, identifying significant issues in linguistic accuracy, cultural alignment, and methodological rigor. To address these limitations in LLMs, we present the Arabic Depth Mini Dataset (ADMD), a carefully curated collection of 490 challenging questions spanning ten major domains (42 sub-domains, see Figure 1. Using ADMD, we evaluate five leading language models: GPT-4, Claude 3.5 Sonnet, Gemini Flash 1.5, CommandR 100B, and Qwen-Max. Our results reveal significant variations in model performance across different domains, with particular challenges in areas requiring deep cultural understanding and specialized knowledge. Claude 3.5 Sonnet demonstrated the highest overall accuracy at 30%, showing relative strength in mathematical theory in Arabic, Arabic language, and islamic domains. This work provides both theoretical foundations and practical insights for improving Arabic language model evaluation, emphasizing the importance of cultural competence alongside technical capabilities.
The inherent complexities of Arabic script; its cursive nature, diacritical marks (tashkeel), and varied typography, pose persistent challenges for Optical Character Recognition (OCR). We present Qari-OCR, a series of vision-language models derived from Qwen2-VL-2B-Instruct, progressively optimized for Arabic through iterative fine-tuning on specialized synthetic datasets. Our leading model, QARI v0.2, establishes a new open-source state-of-the-art with a Word Error Rate (WER) of 0.160, Character Error Rate (CER) of 0.061, and BLEU score of 0.737 on diacritically-rich texts. Qari-OCR demonstrates superior handling of tashkeel, diverse fonts, and document layouts, alongside impressive performance on low-resolution images. Further explorations (QARI v0.3) showcase strong potential for structural document understanding and handwritten text. This work delivers a marked improvement in Arabic OCR accuracy and efficiency, with all models and datasets released to foster further research.
The rich linguistic landscape of the Arab world is characterized by a significant gap between Modern Standard Arabic (MSA), the language of formal communication, and the diverse regional dialects used in everyday life. This diglossia presents a formidable challenge for natural language processing, particularly machine translation. This paper introduces \textbf{SHAMI-MT}, a bidirectional machine translation system specifically engineered to bridge the communication gap between MSA and the Syrian dialect. We present two specialized models, one for MSA-to-Shami and another for Shami-to-MSA translation, both built upon the state-of-the-art AraT5v2-base-1024 architecture. The models were fine-tuned on the comprehensive Nabra dataset and rigorously evaluated on unseen data from the MADAR corpus. Our MSA-to-Shami model achieved an outstanding average quality score of \textbf{4.01 out of 5.0} when judged by OPENAI model GPT-4.1, demonstrating its ability to produce translations that are not only accurate but also dialectally authentic. This work provides a crucial, high-fidelity tool for a previously underserved language pair, advancing the field of dialectal Arabic translation and offering significant applications in content localization, cultural heritage, and intercultural communication.