Deep Research (DR) requires LLM agents to autonomously perform multi-step information seeking, processing, and reasoning to generate comprehensive reports. In contrast to existing studies that mainly focus on unstructured web content, a more challenging DR task should additionally utilize structured knowledge to provide a solid data foundation, facilitate quantitative computation, and lead to in-depth analyses. In this paper, we refer to this novel task as Knowledgeable Deep Research (KDR), which requires DR agents to generate reports with both structured and unstructured knowledge. Furthermore, we propose the Hybrid Knowledge Analysis framework (HKA), a multi-agent architecture that reasons over both kinds of knowledge and integrates the texts, figures, and tables into coherent multimodal reports. The key design is the Structured Knowledge Analyzer, which utilizes both coding and vision-language models to produce figures, tables, and corresponding insights. To support systematic evaluation, we construct KDR-Bench, which covers 9 domains, includes 41 expert-level questions, and incorporates a large number of structured knowledge resources (e.g., 1,252 tables). We further annotate the main conclusions and key points for each question and propose three categories of evaluation metrics including general-purpose, knowledge-centric, and vision-enhanced ones. Experimental results demonstrate that HKA consistently outperforms most existing DR agents on general-purpose and knowledge-centric metrics, and even surpasses the Gemini DR agent on vision-enhanced metrics, highlighting its effectiveness in deep, structure-aware knowledge analysis. Finally, we hope this work can serve as a new foundation for structured knowledge analysis in DR agents and facilitate future multimodal DR studies.
Modern deep reinforcement learning methods are commonly trained in the physics simulators for robot control tasks instead of on real hardware due to inefficiency and safety issues. However, the simulators hardly emulate physical work environment and complex dynamics changes. The incurred domain discrepancy negatively affects the policy’s performance in the real world. In this article, we propose a continuous dynamics adaptive meta-reinforcement learning (CDAMetaRL) framework to realize dynamic adaptation for flexible and intelligent control. CDAMetaRL introduces, first, an adversarial dynamic alignment method that involves actor and critic encoders and a continuous domain discriminator for learning domain-invariant features, and second, a meta-optimization method to further improve adversarial domain adaptation in meta-reinforcement learning. Extensive experiments show CDAMetaRL outperforms several state-of-the-art domain adaptive reinforcement learning methods on five classic sim-to-sim and one typical sim-to-real robot control benchmarks.
The Internet of Vehicles (IoVs) integrates vehicles to the enormous realm of cyberspace which introduces some intelligence and convenience in the transportation sector. However, real-time traffic related information is exchanged over the open public internet among the vehicles and with other infrastructures. This exposes these networks to a myriad of security threats that can lead to accidents and congestions. Although many solutions have been developed over the recent past, most of them are inefficient while others are still susceptible to attacks. In this paper, we leverage on the k-valued modified bilinear inverse Diffie-Hellman problem and one-way hashing function to develop an efficient authentication protocol for IoVs. To demonstrate the robustness of its semantic security, we deploy the Real or Random (ROR) model. In addition, we execute extensive informal security anaysis to show that our scheme resists typical IoVs attacks such as forgery, privileged insider, and replay. Moreover, its performance evaluation shows that it incurs the lowest computation and communication overheads among its peers. Specifically, the proposed protocol reduces the transmission overheads by 8.5%, while increasing the supported security functionalities by 88.9%. It is therefor suitable for deployment in the IoV environment to mitigate the numerous security threats at relatively lower computation and energy costs.
Stance detection is a crucial research direction in natural language processing, aiming to determine the author's supportive, opposing, or neutral attitude towards a specific target. To address challenges such as complex linguistic expressions, insufficient domain-specific knowledge, and weak model generalization in social media contexts, this paper proposed a stance detection method named KABERT, based on knowledge augmentation and adversarial training. The method integrated the advantages of generative and discriminative models. A generative model was first used to extract deep semantic relationships between text and the target from the perspectives of keywords, implicit sentiment, and rhetorical devices, generating implicit knowledge. Then, a discriminative model was used as the classification backbone network for supervised fine-tuning, and a fast gradient method was introduced for adversarial training. Experiments were conducted on two standard stance detection datasets, SEM16 and P-Stance. Results show that KABERT achieves Macro-F1 scores of 68.85% and 79.24% respectively, outperforming current mainstream approaches by varying margins, demonstrating the effectiveness of the proposed approach.
Knowledge Graphs (KGs) are widely used to mitigate the limitations of Large Language Models (LLMs), such as outdated knowledge and hallucinations. Existing LLM-KG integration frameworks typically rely on predefined operators to retrieve factual knowledge from KGs and inject it into prompts for answer generation. This paradigm faces two critical bottlenecks: 1) Inflexibility: The predefined operators are limited in scope and thus lack sufficient compositional expressiveness to fully capture the complex semantics required by KG questions. 2) Unscalability: Direct injection of factual knowledge into prompts limits scalability in handling large-scale factual knowledge. To address these two bottlenecks, we propose Code-on-Graph (CoG), a programmatic reasoning framework for LLM-KG integration. Specifically, given the factual knowledge retrieved at each reasoning step, CoG first identifies the corresponding KG schemas and represents these schemas as Python classes, which serve as abstract interfaces to the retrieved facts. It then generates executable code grounded in these classes, with the retrieved facts instantiated as objects of the corresponding classes during execution. This design enables flexible code-based reasoning while avoiding the direct injection of large-scale factual knowledge into prompts. Experiments on WebQSP, CWQ, and GrailQA demonstrate that CoG outperforms prior state-of-the-art models by up to 10.5
Vision–Language–Action (VLA) models often use intermediate representations to connect multimodal inputs with continuous control, yet spatial guidance is often injected implicitly through latent features. We propose CorridorVLA, which predicts sparse spatial anchors as incremental physical changes (e.g., Δ-positions) and uses them to impose an explicit tolerance region in the training objective for action generation. The anchors define a corridor that guides a flow-matching action head: trajectories whose implied spatial evolution falls outside it receive corrective gradients, while minor deviations from contacts and execution noise are permitted. On the more challenging LIBERO-Plus benchmark, CorridorVLA yields consistent gains across both SmolVLA and GR00T, improving success rate by 3.4%–12.4% over the corresponding baselines; notably, our GR00T-Corr variant reaches a success rate of 83.21%. These results indicate that action-aligned physical cues can provide direct and interpretable constraints for generative action policies, complementing spatial guidance encoded in visual or latent forms. Code is available at https://github.com/corridorVLA.
Stance detection aims to identify the attitude expressed in text towards a given target, with applications in public opinion analysis and misinformation mitigation. Despite recent advances in large language models (LLMs), two key challenges remain: (1) spurious correlations between superficial features and stance labels, and (2) lack of cognitive modeling that simulates the transition from intuitive perception to deliberate reasoning. To address these issues, we propose Cognitive-Driven Stance Detection (CDSD), inspired by Kahneman’s Dual-Process Theory. CDSD integrates fast intuitive judgment (System 1) and analytical reasoning (System 2), enhanced by three key modules: attention-based cognitive alignment to compare system focus, uncertainty-aware belief update using Bayesian inference, and self-doubt-triggered counterfactual reasoning for re-evaluation under low consistency or high uncertainty. Experimental results on SEM16, P-Stance, and VAST show that CDSD outperforms state-of-the-art methods across multiple LLMs. Notably, CDSD exhibits strong robustness against textual perturbations such as emotional word removal and rhetorical restructuring. By integrating cognitive theory with NLP, our work provides a promising path toward more reliable and interpretable stance detection systems.
High-resolution ground penetrating radar (GPR) is a critical tool for subsurface remote sensing, but the high cost and strict size, weight, and power constraints of traditional systems limit their widespread deployment, particularly on unmanned aerial vehicles (UAVs). While software-defined radios (SDRs) offer a flexible alternative, implementing high-resolution stepped-frequency continuous-wave (SFCW) radar on low-cost SDRs introduces fundamental challenges: synthesizer phase incoherence caused by independent phase-locked loops and severe data transport bottlenecks. This article introduces the stepped-frequency coherent radar architecture (SCORA), an open-source framework designed to transform a low-cost, commercial SDR into a fully coherent SFCW radar for near-surface geophysics. The framework features built-in support for both continuous-wave (CW) and linear frequency modulated (LFM) subpulses, utilizing an extensible architecture that allows users to easily integrate custom, advanced waveforms. SCORA utilizes reference-channel compensation and incorporates system latency characterizations to optimize PRF and prevent spatial aliasing. The framework is experimentally validated through free-space ranging, sandbox landmine detection, and an airborne UAV survey that successfully mapped the buried infrastructure. This open-source release significantly reduces the barrier to entry for agile GPR instrumentation.
Multimodal Stance Detection (MSD) aims to determine a user’s stance - support, oppose, or neutral - toward a target by analyzing multimodal content such as texts and images from social media. Existing MSD methods struggle with generalizing to unseen targets and handling modality inconsistencies. To address these challenges, we propose the Target-driven Multi-modal Alignment and Dynamic Weighting Model (T-MAD), which combines target-driven multi-modal alignment and dynamic weighting mechanisms to capture target-specific relationships and balance modality contributions. The model incorporates iterative reasoning to iteratively refine predictions, achieving robust performance in both in-target and zero-shot settings. Experiments on the MMSD and MultiClimate datasets show that T-MAD outperforms state-of-the-art models, with optimal results achieved using RoBERTa, ViT, and an iterative depth of 5. Ablation studies further confirm the importance of multi-modal alignment and dynamic weighting in enhancing model effectiveness.
Stance detection is a pivotal task in Natural Language Processing (NLP), identifying textual attitudes toward various targets. Despite advances in using Large Language Models (LLMs), challenges persist due to hallucination-models generating plausible yet inaccurate content. Addressing these challenges, we introduce MPVStance, a framework that incorporates Multi-Perspective Verification (MPV) with Retrieval-Augmented Generation (RAG) across a structured five-step verification process. Our method enhances stance detection by rigorously validating each response from factual accuracy, logical consistency, contextual relevance, and other perspectives. Extensive testing on the SemEval-2016 and VAST datasets, including scenarios that challenge existing methods and comprehensive ablation studies, demonstrates that MPVStance significantly outperforms current models. It effectively mitigates hallucination issues and sets new benchmarks for reliability and accuracy in stance detection, particularly in zero-shot, few-shot, and challenging scenarios.
Material sensing holds significant potential in areas such as environmental awareness and security monitoring. While technologies like RFID, WIFI, and UWB offer potential solutions for portable, non-contact material identification, the need to place targets in fixed positions for identification has limited the flexibility of material sensing. In this article, we first innovatively apply the Range-Angle heatmap (RAheatmap) to effectively represent the distance, placement angle, and inherent material attributes to pave the way for precise material identification. Then propose an innovative system called Material-ID to utilize Commercial-Off-The-Shelf (COTS) millimeter wave (mmWave) radar for material sensing. Additionally, we endow the system with cross-domain adaptability to make it tailored to identify material reflection attributes and minimize the effects of variables such as distance and placement angle. The experiments prove the effectiveness of the proposed system.
Accurate and timely 90-day survival prediction for critically ill COVID-19 patients is vital to optimize scarce ICU resources, yet single-source models often miss important pathophysiological cues. Emerging studies show that combining complementary modalities can reveal richer prognostic signatures than any modality in isolation. Motivated by this, we present DQU-CLIP, an advanced multimodal deep learning framework designed to overcome this limitation. Utilizing the CoCross dataset (comprising 171 ICU patients), our model integrates chest Xrays (CXRs) via a pre-trained Contrastive Language-Image Pretraining (CLIP) encoder with key clinical features, including Age, Charlson Comorbidity Index (CCI), APACHE II, and SOFA scores, processed by a neural network. DQU-CLIP achieves a robust ROC-AUC of 0.85, significantly outperforming unimodal baselines (CXR-only: 0.78, Clinical-only: 0.72) and competing multimodal approaches. Extensive validation and ablation studies confirm the synergistic benefit of this fusion. Furthermore, interpretability analysis using Grad-CAM identified relevant lung regions in CXRs, while feature importance pinpointed SOFA and APACHE II scores as critical indicators of disease severity. By effectively unifying radiological and clinical evidence, DQU-CLIP provides a more reliable prognostic assessment.
Stance detection, a critical task in Natural Language Processing (NLP), aims to identify the attitude expressed in text toward specific targets. Despite advancements in Large Language Models (LLMs), challenges such as limited interpretability and handling nuanced content persist. To address these issues, we propose the Multi-Path Reasoning Framework (MPRF), a novel framework that generates, evaluates, and integrates multiple reasoning paths to improve accuracy, robustness, and transparency in stance detection. Unlike prior work that relies on single-path reasoning or static explanations, MPRF introduces a structured end-to-end pipeline: it first generates diverse reasoning paths through predefined perspectives, then dynamically evaluates and optimizes each path using LLM-based scoring, and finally fuses the results via weighted aggregation to produce interpretable and reliable predictions. Extensive experiments on the SEM16, VAST, and PStance datasets demonstrate that MPRF outperforms existing models. Ablation studies further validate the critical role of MPRF’s components, highlighting its effectiveness in enhancing interpretability and handling complex stance detection tasks.
Document retrieval plays an essential role in many real-world applications especially when the data storage is outsourced. Due to the great advantages offered by cloud computing, clients tend to outsource their personal data to remote servers maintained by external service providers. This raises serious privacy concerns about outsourced data because such providers are usually considered untrusted entities. The majority of previous schemes of document similarity search share the same limitation: they focus mainly on static collections. Dynamic searchable schemes (DSE) allow adding or removing documents at the expense of more leakage than static schemes. To thwart certain attacks, DSE schemes should support forward privacy property, which ensures that newly added documents cannot be related to previously issued search queries. We design and implement dynamic secure similarity search schemes with forward privacy for textual documents utilizing simhash method for hamming similarity. Our scheme provides an efficient search time and a sufficient level of privacy. To show the practicality of our proposed scheme, we performed excremental results with large document collections.
This study presents a systematic literature review on blockchain-based authentication in smart environments that include smart city, smart home, smart grid, smart healthcare, smart farming and smart transportation. The review incorporated 39 articles presenting blockchain solutions for security and privacy issues through authentication mechanisms in these smart environments. Guided by three research questions to determine the main issues in smart environment, the availability of blockchain-based authentication solutions and identified research gaps and future research endeavours, this review used PRISMA method to provide insights on the use of blockchain-based authentication schemes. The research gap is that blockchain solutions are mostly at the proposal, and sometimes conceptual stage is in the reviewed articles. In addition, solutions presented in smart environments require exploration into blockchain. The findings show similar situational issues across different smart environments and the flexibility and adaptability of blockchain to provide solutions to the identified issues pertaining to security and privacy. More clearly, the authentication problem posed across different smart environments can be adapted to blockchain technology provided that it is combined with other technologies to increase efficiency, despite the existence of large-scale authentication mechanisms. Now, Blockchain still has the unique, distributed feature of converting the current database into Blockchain databases. This review guided future research directions which could further contribute to the sustainable management of smart environments.
Solar cells have been widely used for offering energy for Internet of Things (IoT) devices. Recently, solar cells have also been used as sensors for context awareness sensing due to their sensitivity to varying lighting conditions. In this article, we are the first to use solar cells for symmetric key generation. To generate symmetric keys, we take advantage of photovoltage measurements generated from solar cells equipped with a pair of IoT devices. Symmetric keys are essential for pairing IoT devices and further securing wireless communication. Despite the sensitivity to varying lighting conditions, challenges still remain for the use of solar cells for key generation, such as time unsynchronisation and noisy measurements. To solve these challenges, we design a novel key generation framework, SolarKey, which includes the starting point detection and a compressed sensing-based two-tier key reconciliation method. Extensive experiments have been conducted to evaluate the performance of our proposed key generation method in various environments, which shows the proposed method can improve the key matching rate by up to 25%. We also conduct security analysis and the randomness test, which shows that SolarKey is resilient to common attacks such as the eavesdropping attack and the imitating attack and sufficiently random.
Cracks are usually curve-like structures that are the focus of many computer-vision applications (e.g., road safety inspection and surface inspection of the industrial facilities). The existing pixel-based crack segmentation methods rely on time-consuming and costly pixel-level annotations. And the object-based crack detection methods exploit the horizontal box to detect the crack without considering crack orientation, resulting in scale variation and intra-class variation. Considering this, we provide a new perspective for crack detection that models the cracks as a series of sub-cracks with the corresponding orientation. However, the vanilla adaptation of the existing oriented object detection methods to the crack detection tasks will result in limited performance, due to the boundary discontinuity issue and the ambiguities in sub-crack orientation. In this paper, we propose a first-of-its-kind oriented sub-crack detector, dubbed as CrackDet, which is derived from a novel piecewise angle definition, to ease the boundary discontinuity problem. And then, we propose a multi-branch angle regression loss for learning sub-crack orientation and variance together. Since there are no related benchmarks, we construct three fully annotated datasets, namely, ORC, ONPP, and OCCSD, which involve various cracks in road pavement and industrial facilities. Experiments show that our approach outperforms state-of-the-art crack detectors.
Human has an unique gait and prior works show increasing potentials in using WiFi signals to capture the unique signature of individuals' gait. However, existing WiFi-based human identification (HI) systems have not been ready for real-world deployment due to various strong assumptions including identification of known users and sufficient training data captured in predefined domains such as fixed walking trajectory/orientation, WiFi layout (receivers locations) and multipath environment (deployment time and site). In this paper, we propose a WiFi-based HI system, MetaGanFi, which is able to accurately identify unseen individuals in uncontrolled domain with only one or few samples. To achieve this, the MetaGanFi proposes a domain unification model, CCG-GAN that utilizes a conditional cycle generative adversarial networks to filter out irrelevant perturbations incurred by interfering domains. Moreover, the MetaGanFi proposes a domain-agnostic meta learning model, DA-Meta that could quickly adapt from one/few data samples to accurately recognize unseen individuals. The comprehensive evaluation applied on a real-world dataset show that the MetaGanFi can identify unseen individuals with average accuracies of 87.25% and 93.50% for 1 and 5 available data samples (shot) cases, captured in varying trajectory and multipath environment, 86.84% and 91.25% for 1 and 5-shot cases in varying WiFi layout scenarios, while the overall inference process of domain unification and identification takes about 0.1 second per sample.
In recent years, as the rise of edge intelligence hand tracking applications have emerged in various IoT systems and applications, such as human-computer interaction, sign language translation, and motion rehabilitation, etc. Due to the hardware and energy limitation of embedded IoT devices, hand tracking is difficult to achieve real-time continuous operation. To achieve this goal, in this paper we propose a set of optimization methods such as space-based error loss and Gaussian process constraints to design a lightweight hand tracking model with low storage overhead and high recognition efficiency. We design and implement the system, and the effectiveness of our proposed methods is verified in the evaluation comparing with the state-of-the-art.