
Artificial intelligence-based image analysis is transforming medical diagnosis; however, its clinical usefulness is limited by inherent noise and the quality of available medical datasets. Recent advances in generative AI promise to improve image quality through denoising, enhanced resolution, and more. This study applies the Segment Anything Model 2 (SAM2) and the Enhanced Super-Resolution Generative Adversarial Networks (ESRGAN) to dental radiographic datasets. A dataset of 6,413 radiographs across 35 dental implant systems was used to explore the pre-processing paradox, the relationship between human visual clarity and machine-learning accuracy. First, YOLOv11 detected and cropped implant areas, establishing the baseline for two classification models: ConvNeXt and ResNet50. These were compared with three experimental variants: SAM2 segmentation, ESRGAN super-resolution, and a fully integrated YOLO+SAM2 + ESRGAN pipeline. Performance was assessed using accuracy, F1-score, PSNR, and SSIM. In the baseline YOLO+CNN approach, the accuracy of ConvNeXt (83%) outperformed that of ResNet50 (79%). Adding SAM2 segmentation lowered performance (ConvNeXt: 81%, ResNet50: 76%), suggesting that removing the bone-implant interface eliminates important contextual signals. Both ConvNeXt and ResNet maintained their average accuracies (83 and 80%) when ESRGAN was applied to raw and cropped images. The largest difference appeared in the fully integrated pipeline, where ConvNeXt achieved higher accuracy (82%) compared to a significant decline in ResNet performance (64%). The study confirms the pre-processing paradox. While SAM2 and ESRGAN enhance visual clarity, the artifacts they introduce can reduce ResNet’s accuracy. The YOLOv11-CNN approach is optimal, and ConvNeXt is recommended for its robustness to image preprocessing. However, the proposed preprocessing pipeline did not significantly improve accuracy.
IntroductionChatGPT's conversational interface invites anthropomorphic interpretations that conflict with its underlying probabilistic architecture. Prior research has identified folk theories and conceptions of generative AI among technical laypersons, but has largely examined these mental concepts in isolation or at the group level. Less is known about how conceptualizations of ChatGPT as a deterministic, probabilistic, anthropomorphic or dynamic AI system coexist, interact, and contradict one another within technical laypersons. This study examined (RQ1) recurring qualitative mental model configurations of ChatGPT among participants and (RQ2) whether technical AI terminology was supported by mechanism-specific understanding or coexisted with misconceptions concerning the same or another functional mechanism.MethodsWe conducted semi-structured interviews with technical laypersons (N = 21) to ask them about their understanding of how ChatGPT works. Statements were deductively coded into four mental concepts (deterministic, probabilistic, anthropomorphic, and dynamic). Mental models were constructed for each participant. Statements containing technical terminology were additionally classified according to whether the term was contradicted by the participant's explanation of the same mechanism (“Chauffeur Knowledge”) or co-occurred with a misconception in another mechanism.ResultsThree recurring mental model configurations emerged: (i) Predominantly Probabilistic Thinking Patterns (n = 8); (ii) Anthropomorphic and Dynamic Thinking Patterns (n = 4), characterized by limited probabilistic reasoning; and (iii) Predominantly Deterministic Retrieval-based Thinking Patterns (n = 9), characterized mainly by deterministic conceptions alongside limited probabilistic awareness. All participants showed at least one misconception about ChatGPT and among 10 participants, who used technical terminology in their explanations, seven showed chauffeur knowledge, seven demonstrated correct or partially correct understanding of one mechanism alongside a misconception about another, and four participants met both criteria.DiscussionParticipants' mental models of ChatGPT are fragmented and often internally contradictory. Technical terminology may mask misunderstanding of the mechanism it describes or coexist with correct understanding in one domain and misconceptions in another. These findings challenge AI literacy assessments that rely predominantly on knowledge of technical terminology. Identifying recurring qualitative configurations provides a basis for developing measures that capture the complexity of mental models and may improve predictions of over-reliance on ChatGPT output.
IntroductionHybrid Variational Quantum Algorithms (VQAs) present a highly viable pathway to near-term quantum utility; however, their performance is fundamentally bottlenecked by classical-quantum communication latency in Quantumas-a-Service (QaaS) environments.MethodsThis paper proposes an optimized classical-quantum orchestration architecture designed to minimize cloud-induced latency and maximize Quantum Processing Unit (QPU) active compute time. By implementing edge-colocated classical optimizers alongside batched parameter-shift gradient evaluations, the system circumvents stateless cloud API barriers.ResultsBenchmarking across parameterized quantum circuits ranging from 15 to 50 qubits demonstrates an 84% reduction in network-induced QPU idle time. The framework yields a 3.2 × speedup in overall convergence time for the Quantum Approximate Optimization Algorithm (QAOA) and up to a 98% reduction in classical API call overhead compared to standard RESTful QaaS execution models.DiscussionThese quantitative findings demonstrate that tightly coupled hybrid co-processing, physically adjacent to the control electronics, is critical for extending the computational bound of Noisy Intermediate-Scale Quantum (NISQ) devices.
Although gesture-based explanations for robotic behavior seem promising for enhancing the transparency of robotic systems, careful design is crucial to their success. Based on a pilot laboratory study (N = 42) that highlights the need for a systematic framework for gesture-based design of transparent robotic systems, we propose an initial 12-dimensional design space taxonomy, clustered in three levels (communicational, spatial, and motor), to develop gesture-based robotic explanations. To explore the design space taxonomy, we used it to create five gesture-based explanations for robotic behavior deviating from the users' goal (i.e., a robot arm in a smart kitchen environment serving a different type of soda can than the one ordered and indicating that the ordered can was empty by showing it to users, pointing at the lid, shaking it, or (hinting at) disposing of it). We investigated the effect of these gestures on users' understanding, perception of technology transparency, and further interaction-related variables in a video-based between-subjects user study (N = 235), including a control condition without a gesture-based explanation. Results showed that the show gesture tended to be most effective at creating understanding and transparency perception, while other gestures were ineffective, and the shake gesture even reduced facets of perceived technology transparency compared to the control condition. The findings point to gestures as a potential means of making robot actions understandable, while at the same time highlighting the design challenges due to still insufficient user understanding of the robot's behavior across conditions and a lack of significant differences compared to the control condition. Moreover, as improvements in understanding were not correspondingly reflected in increased transparency perceptions, our studies indicate that these two outcomes may yield different results and that both should be considered separately in explanatory robot design. Nevertheless, the results suggest that direct gestures in users' proximity with explicit evidence presentation might be most promising. This highlights the relevance of our design space taxonomy for systematically exploring robotic gesture designs and investigating their effectiveness in specific use contexts.
To optimize the early detection of breast cancer, a strong integration of heterogeneous biomedical signals is essential to improve diagnostic reliability and reduce false negatives. This study proposes a Tri-Fusion deep learning model for timely breast cancer screening, which integrates White Blood Cell (WBC) morphological features, mammogram image embeddings, and gene mutation signatures within a unified multimodal framework. WBC features are extracted from peripheral blood smear images using a CNN-based morphologic encoder, mammographic representation is ensured by a pretrained deep convolutional foundation based on breast imaging information, and genomic details are modelled using a completely connected mutation-signature encoder deduced from a breast cancer–related gene panel. To capture cross-domain correlation, a dense categorization head is used for feature-level fusion. The proposed model uses three publicly available datasets: a WBC image dataset containing 12,500 samples, a CBIS-DDSM mammogram dataset containing 3,102 annotated instances, and a curated Genomic Dataset containing a mutant profile of 1200 patients from TCGA-BRCA. The experimental results show high accuracy compared with the unimodal and bimodal baselines, corresponding to 96.2% accuracy, 95.4% correctness, 94.8% recall, a F1-score of 95.1%, and an AUC-ROC of 98.6%. Furthermore, the model exhibits strong generalization under cross-validation and robustness to the missing-modality scenario.
IntroductionBy 2030, an estimated 40% of current cloud infrastructures may be rendered vulnerable by cryptanalytically relevant quantum computers (CRQCs).MethodsThis paper introduces a 4-tier security framework tailored for Quantumas-a-Service (QaaS) deployments, focusing on securing data-in-transit. Integrating 3 NIST-standardized post-quantum algorithms (ML-KEM, ML-DSA, and SLHDSA), our architecture mitigates interception threats based on Shor's algorithm.ResultsEmulation across an enterprise cloud topology demonstrates a maximum latency overhead of 17.7 milliseconds per TLS handshake under high-latency WAN conditions, ensuring high-availability operations without catastrophic fragmentation failure.DiscussionThe proposed model demonstrates the viability of modern latticebased cryptography for live environments and provides a structured 3-phase transition roadmap for cloud service providers (CSPs) to achieve quantum-resilience seamlessly.
IntroductionUnderwater visual sensing is severely hindered by wavelength-dependent absorption and scattering, which induce non-linear color shifts and structural haze. Existing enhancement methods often face a fundamental trade-off between chromatic restoration and geometric integrity, where color correction frequently occurs at the expense of blurring fine-grained details or amplifying noise.MethodsA lightweight physics-inspired dual-stream network (LPD-Net) is proposed for structure-preserving enhancement, establishing an explicit mapping from physical degradation phenomena to targeted neural operators. A dual-stream architecture is designed to decouple chromatic restoration and geometric reconstruction by utilizing a raw sensor stream as a fidelity anchor and a physics-inspired prior stream for guidance. A key innovation is the spatial-difference gated fusion mechanism, which dynamically optimizes the pixel-wise weights between raw sensor observations and empirical priors through data-driven perception. By employing end-to-end joint optimization, the network adaptively suppresses prior-induced artifacts in clear water while leveraging physics-inspired guidance in turbid scenarios.ResultsThe performance of the proposed model was evaluated using the UIEB and LSUI benchmarks. Experimental results indicate that LPD-Net achieves superior restoration quality, outperforming the fully aligned, competitive baseline by up to 0.67 dB in PSNR and securing the lowest perceptual distortion (0.1227 LPIPS). With a minimalist footprint of only 1.84 M parameters, the framework attains a core neural inference throughput of 27.70 FPS at 256 × 256 resolution on a mid-range GPU.ConclusionThe proposed LPD-Net provides a robust, efficient, and high-fidelity solution for structure-preserving underwater image enhancement. Its resolution-agnostic design and competitive throughput demonstrate significant potential for real-time deployment on autonomous underwater platforms and structural survey systems.
Accurate prediction of the remaining useful life (RUL) of rolling element bearings under target-bearing data scarcity remains a critical challenge in prognostics and health management (PHM). This paper proposes a multi-representation domain generalization framework for unseen-bearing RUL prediction. The framework uses a Bidirectional Multi-scale ConvLSTM architecture to capture temporal degradation dependencies and multi-granular spatial patterns from two aligned representations of the same vibration signal: the raw time-domain signal and its continuous wavelet transform (CWT) time-frequency representation. In addition, a lightweight linearly constrained temporal prior is integrated into the prediction layer to encourage a monotonically decreasing degradation trend, following the common damage-irreversibility assumption in bearing prognostics. The target bearing is excluded from training, validation, normalization-parameter estimation, and model selection, and is used only for final testing. Comprehensive experiments on two public benchmark datasets show that the proposed framework improves RUL prediction accuracy and provides competitive cross-bearing generalization.
Lightweight security mechanisms are essential for Internet of Things (IoT) edge environments, where devices operate under strict constraints in computation, memory, and energy. The emergence of TinyML-enabled edge intelligence introduces new communication security requirements, particularly for protecting compact inference outputs (typically 1-32 bytes, including class labels, confidence scores, or anomaly flags) transmitted over potentially insecure networks. This paper presents a lightweight confidentiality-focused encryption framework based on an enhanced variant of the Tiny Encryption Algorithm (TEA), tailored for TinyML-driven IoT communication. The proposed Enhanced TEA incorporates a plaintext-dependent dynamic key diversification mechanism using SHA-256, improving empirical diffusion and ciphertext randomness while preserving the computational efficiency of ARX-based cipher structures. Beyond algorithmic design, the study develops a system-level secure TinyML-enabled IoT communication architecture, integrating encryption directly into edge inference workflows. The framework is implemented and evaluated on an edge computing platform to assess performance in terms of execution time, memory usage, CPU utilization, communication latency, and energy behavior. Experimental results demonstrate an avalanche effect of 54.69% and ciphertext entropy of 7.66 bits/byte, while maintaining less than 5% throughput degradation and minimal latency overhead in MQTT-based communication. The proposed approach provides a practical lightweight confidentiality mechanism for TinyML-enabled edge computing environments operating under moderate resource constraints. However, as the design focuses on confidentiality, it should be combined with lightweight authentication mechanisms to ensure comprehensive security in real-world deployments.
As global urbanisation concentrates the majority of people in cities, encounters between humans and non-human lifeworlds increasingly take place within the built environment. Plants are key living organisms integrated into urban systems at scale, and some architectural spaces incorporate plants to benefit from their presence through sensing, actuation, and biosensing technologies. Buildings are increasingly understood as mediated sensory environments, in which plants are not static, aesthetic additions but active elements of these systems—often coupled with technology. Media artists and designers have long used interactive installations and prototypes to explore and establish relationships among humans, plants, and the built environment; exploring configurations ranging from unidirectional human benefit to mutual coexistence and ecological extension. These creative practices offer insights into how plants could become a key element of human-building interaction (HBI). This paper examines these emerging practices of technology-mediated human-plant interaction within the built environment through a dual contribution: a scoping review and the development of a relational taxonomy. Drawing on a corpus of interactive works sourced from academic and creative practice, the review analyses how plants are engaged as sensory media and how interactions are structured through technological systems. Addressing the question of how benefits are distributed among humans, plants, and the wider environment, the paper identifies five recurring orientations of benefit flow, from which it derives a Beneficiary Taxonomy: Restoring, Nurturing, Autonomising, Entangling, and Extending. Rather than prescribing optimal approaches, the taxonomy provides a conceptual framework for understanding variation in how human-plant relationships are designed and experienced. The paper contributes an analytical reference by linking creative practice with design-oriented analysis. The Beneficiary Taxonomy offers an interdisciplinary tool for identifying variations in benefit flow and reflecting on the role assigned to plants within technologically mediated sensory environments. This reflection fosters more informed dialogue and encourages a shift from viewing plants as ornamental elements to recognising them as ecological co-occupants in the design of future built environments.
This systematic review assessed 64 fashion datasets identified through a PRISMA 2020-guided two-stream search, evaluating each on FAIR compliance, a composite AI-Readiness Score, and a five-level ontology maturity model. The AI-Readiness assessment yielded a grade distribution concentrated in the middle tiers (B: 60.3%, C: 36.5%), while 79.7% of datasets provided only flat attribute annotations (Level 2 or below) on the ontology maturity scale. Beyond diagnosis, the review distills its findings into evidence-derived recommendations for future dataset construction — most consequentially, that new fashion datasets should adopt machine-readable open licenses with persistent identifiers, and should structure annotations at least at the Taxonomy level (Level 3) of the proposed maturity model. Together, the three assessment perspectives — FAIR compliance, AI-readiness, and ontology maturity — and these recommendations provide actionable tools for evaluating future data investments and guiding the development of next-generation fashion AI datasets.
Vehicular Ad Hoc Networks (VANETs) and Flying Ad Hoc Networks (FANETs) are essential for safety-critical and data-intensive applications in intelligent transportation and aerial systems. However, high node mobility, dynamic topologies, heterogeneous communication technologies, and limited energy resources in FANETs cause frequent handovers, resulting in increased latency, packet loss, and reduced network performance. Conventional threshold-based and heuristic mechanisms struggle to adapt to these highly dynamic environments. This paper proposes a unified Q-learning based intelligent handover framework for both VANETs and FANETs operating in heterogeneous communication settings. The framework models the handover decision as a Markov Decision Process (MDP) and employs a lightweight, model-free reinforcement learning approach. It dynamically learns optimal policies by jointly considering multiple real-time parameters: received signal strength, node mobility, traffic load, network occupancy, and residual energy. In VANETs, the framework enables intelligent vertical handovers between high-bandwidth Li-Fi and reliable RF (IEEE 802.11p) links. In FANETs, it incorporates energy-aware decision-making to extend network lifetime under 3D mobility. Extensive simulations using OMNeT++ integrated with simulation of urban mobility (SUMO) demonstrate significant performance gains over RSS-based, heuristic, SDN-assisted, and existing RL methods. The proposed framework achieves a handover success rate of up to 91.6% under high mobility, improves throughput to 15.8 Mbps in hybrid Li-Fi/RF VANET scenarios, and maintains a packet delivery ratio of 92.8% under heavy traffic. It further reduces end-to-end delay by up to 38%, lowers handover latency to 41 ms (VANET) and 54 ms (FANET), decreases energy consumption per packet to 0.48 J, and extends FANET network lifetime to 168 min. The proposed Q-learning-based approach provides a scalable, adaptive, and energy-efficient solution for seamless handover management in next-generation heterogeneous vehicular and aerial ad hoc networks.
The Conceive–Design–Implement–Operate (CDIO) Syllabus expresses graduate competencies in natural language developed through project-based learning, yet the link between declared competencies and delivered content is maintained manually. This study proposes the CDIO Competency–Knowledge Ontology (CDIO–CKO): a decidable OWL 2 Description Logic (DL) formalization of the Syllabus connected to programs, projects, courses, learning outcomes, and knowledge-content units, paired with a parameterized transformation that converts educational projects into competency-linked units via decomposition, extraction, semantic annotation, and aggregation. Applied to one 240-ECTS mechatronics and robotics program (42 courses, 18 projects, 156 competencies), the methodology produced 1,287 individuals and 9,540 triples classifying in 2.8 s. On a 240-unit held-out test set against a three-expert gold standard, automated mapping reached precision 0.91, recall 0.86, and F1 0.88 (95% CI 0.84–0.91; system–standard κ = 0.79, expert pre-adjudication κ = 0.74). Traceable coverage rose from 0.59 to 0.84, and 23 prerequisite inconsistencies missed by routine manual review were surfaced (23 of 26 issues in a dedicated expert audit: precision 1.00, recall 0.88). Against five comparator families, the method achieved the highest F1 among the five evaluated baselines while uniquely supporting consistency checking, proficiency propagation, and an auditable evidence trail. Results are a single-program proof-of-concept; broader validation remains future work.
The rapid integration of large language model-based conversational systems has intensified longstanding questions in mediated communication, human–computer interaction, and social presence research by placing language-generating systems within ongoing conversational exchanges. Rather than treating these systems only as information-retrieval tools, users often orient to them through interactional cues associated with responsiveness, continuity, and interlocutor recognition. However, existing analytical approaches do not adequately capture the multidimensional nature of these interactions. This study addresses this gap by proposing a computational framework grounded in communication theory, operationalized through three interaction indices: the Social Presence Index (SPI), the Social Bonding Index (SBI), and the Companion Communication Index (CCI). The framework is evaluated on more than 50,000 conversations from heterogeneous datasets (WildChat, LMSYS, DailyDialog, and MultiWOZ), enabling a comparative analysis between human–AI and human–human communication. Results show that human–human interactions exhibit higher social presence (SPI = 0.31 vs. 0.13), while human–AI interactions demonstrate greater structural continuity (CCI = 0.41 vs. 0.34), with statistically significant differences (p < 0.001). Segmentation analysis reveals non-linear interaction patterns, where conversational continuity peaks at intermediate levels of social presence in human–AI exchanges. These findings demonstrate that social presence, affective bonding, and conversational continuity are interdependent and context-sensitive, supporting a multidimensional framework for understanding and designing human–AI communication processes.
Intrusion detection in Industrial Internet of Things (IIoT) environments is a risk-asymmetric problem: false alarms increase analyst workload, but false negatives may allow malicious activity to persist in safety- and operation-sensitive systems. Although recent deep learning-based intrusion detection systems report high aggregate accuracy, near-ceiling performance can obscure the residual errors that remain under class imbalance and fine-grained label settings. This study investigates missed-attack-risk-oriented IIoT intrusion detection through CKAN-AFG, a compact non-recurrent Kolmogorov–Arnold Network (KAN)-centered architecture that combines KAN-based feature transformation, residual multi-head feature attention, feature gating, adaptive pooling, and a compact KANLinear classifier. The model is evaluated on the CIC IIoT Dataset 2025 (DataSense) as the main benchmark, with TON-IoT used as a secondary benchmark for cross-dataset comparison and continuity with prior IIoT evaluation. The evaluation uses a leakage-controlled repeated-seed protocol with fold-confined RF-RFE, a final 55-common-feature DataSense protocol, operational error decomposition, KAN-isolation and component-level ablation, repeated-seed stability analysis, and CPU-based deployment-oriented profiling. Under the final DataSense protocol, CKAN-AFG achieves a weighted F1-score of 99.8146% and an FNR of 0.2636%, corresponding to an approximately 5.6% relative FNR reduction compared with the CKAN–BiLSTM continuity baseline. Operational error decomposition shows that CKAN-AFG produces the lowest attack-to-benign count, with 33.8 mean true missed attacks, and the lowest total off-diagonal error count among the retained models. KAN-isolation results show that replacing the KAN-specific feature-transformation and classifier components with conventional CNN or LSTM alternatives increases FNR, while removing attention, feature gating, or both also degrades the false-negative-aware profile. CPU profiling shows that CKAN-AFG remains compact, with 247,005 trainable parameters, a 0.956 MB FP32 footprint, and throughput of approximately 1,691 samples/s under CPU-only profiling. These findings suggest that compact KAN-centered feature refinement can provide a practical operating point for missed-attack-risk-oriented IIoT intrusion detection while avoiding overclaims of universal dominance or direct zero-shot transfer.
This paper systematically examines monitoring tools for containerized applications, microservices, and DevOps environments, and provides an experimental evaluation across two deployment scenarios. The study was conducted across two infrastructures: on-premises (Docker Swarm, Kubernetes) and cloud-based (Google Kubernetes Engine). A test application was deployed, and load and stress tests were performed using Apache JMeter and k6 with virtual user counts ranging from 100 to 1,000. Metrics including CPU, memory, I/O, and network utilization were collected using Prometheus, Grafana, cAdvisor, and Docker stats. Based on exploration and evaluation, 69 unique monitoring-related tools were identified and grouped into four non-mutually exclusive categories: container-based, cloud-based, microservices-oriented, and DevOps-focused (18, 25, 13, and 38 assignments, respectively; 94 assignments in total because some tools belong to multiple categories). A classification is presented according to visualization capabilities, supported metrics, and quality attributes, including performance, security, interoperability, and usability. Experimental results demonstrate that both monitoring stacks captured variations in processor utilization under load at the resolution configured by their default sampling intervals (15–30 s for Prometheus + cAdvisor and 60 s for Google Cloud Monitoring), with the 1,000-virtual-user stress bursts visible as sharp CPU peaks in both stacks. The contribution of this work is a practical guideline for selecting monitoring tools, developed on the basis of a reference experimental environment.
IntroductionThe reliability of phishing Uniform Resource Locator (URL) detectors under adversarial URL rewriting, domain shift, and explanation instability remains insufficiently understood. This study proposes an integrated robustness evaluation protocol for URL-based phishing detection, which integrates structured adversarial perturbation, unseen attack-family generalization, compositional attack effects, explanation stability, and external vulnerability transfer.MethodsThe protocol tests four representative model families: Logistic Regression, XGBoost, CharCNN, and BERT-base, using 235,370 validated URLs from PHIUSIIL, consisting of 100,520 phishing and 134,850 benign URLs, along with 49,121 PhishTank-validated phishing URLs for external validation.ResultsAll models performed well on the clean test sets, ranging from 0.9962 to 0.9984, but their robustness decreased substantially under realistic URL mutations. Subdomain injection degraded the strong performance of Logistic Regression, XGBoost, and CharCNN to around 0.432, indicating collapse to the phishing-prevalence floor. BERT was highly susceptible to homoglyph, padding, and path-based perturbations. Leave-one-family-out evaluation also showed poor transfer to unseen subdomain attacks for both Logistic Regression and XGBoost, with Robustness Degradation Index values of 0.535 and 0.565, respectively. Explanation stability also suffered, with SHAP top-K Jaccard similarity dropping to 0.526-0.535 under subdomain perturbation.DiscussionThese results provide a solid benchmark for evaluating robustness-aware phishing URL detection for achieving deployable reliability under realistic adversarial and non-IID settings.
IntroductionWhile spatial Vision Transformers (ViTs) achieve high precision in urban scene parsing, their frame-by-frame application in autonomous driving suffers from severe temporal flickering and prohibitive retraining costs across decentralized vehicle fleets.MethodsTo overcome these dual bottlenecks, this paper introduces a unified Spatiotemporal Hierarchical Mask-Refinement (ST-HMR) framework integrated with a Byzantine-Robust Federated Learning (BR-FL) protocol. The ST-HMR module caches fine-grained prompt tokens via an asymmetrical Temporal Cross- Attention buffer to enforce inter-frame geometric continuity. Concurrently, the BR-FL pipeline employs Multi-Krum geometric distance filtration to aggregate 3 localized prompt gradients from decentralized fleets securely, updating only a 1.4% active parameter subset.ResultsEvaluated on the Cityscapes Video dataset, the ST-HMR framework improves the video segmentation mean Intersection over Union (mIoU) to 83.5%, elevates the Temporal Consistency (TC) score to 88.5, and reduces depth Absolute Relative Error (Abs Rel) to 0.085, all while maintaining real-time edge processing at 38 FPS. Under severe adversarial network conditions (up to 30% Byzantine/malicious sensor nodes), the BR-FL protocol achieves a 98.4% Byzantine detection rate and maintains a global mIoU of 81.9%, while reducing Over-The-Air (OTA) transmission payloads by over 99% (3.8 MB vs. 1.2 GB per round).DiscussionThese findings demonstrate that parameter-efficient prompt caching eliminates temporal boundary jitter without heavy 3D transformer overhead, while geometric gradient filtering provides robust defense against decentralized poisoning, establishing a scalable, secure, and temporally coherent perception paradigm for next-generation edge robotics.