Abstract Background Rehabilitation medicine faces a significant challenge due to the rising demand for services coupled with a shortage of specialized professionals. Large Language Models (LLMs) show promise for enhancing clinical efficiency, but their evaluation has been largely limited to simulated scenarios, lacking direct performance comparisons with human experts in complex, real-world clinical tasks. Objective To systematically benchmark five state-of-the-art LLMs against senior physiatrists in formulating comprehensive rehabilitation plans for authentic clinical cases, evaluating their utility as clinical decision support tools. Methods We conducted a rigorous, blinded evaluation using 48 authentic cases across six subspecialties. Plans generated by five LLMs (Grok-4, Gemini−2.5-pro, ChatGPT-5-2025-08-07, Deepseek-r1-0528, and Claude-opus-4-20250514) were compared with expert-authored plans. A panel of 6 senior physiatrists evaluated the plans using a multi-dimensional framework covering four key domains: Clinical Applicability and Safety (primary safety endpoint), Scientific Rigor, Individualization, and Clarity. To address the data’s hierarchical structure, we employed Linear Mixed-Effects Models (LMM) with random intercepts for cases and raters, and fixed effects for models and language. Pairwise comparisons were adjusted using the Holm-Bonferroni correction. Results Quantitative analysis revealed that Grok-4 (mean 4.31) and Gemini−2.5-pro (mean 4.14) significantly outperformed the human benchmark (derived from standardized expert solutions) (mean 3.56; $$P<0.001$$ ). Notably, the open-source Deepseek-r1 (mean 3.69) also achieved a statistically significant advantage over experts ( $$P<0.001$$ ). Conversely, human experts scored numerically higher than Claude-opus-4 (mean 3.50), though this difference was not statistically significant ( $$P=0.099$$ ). Qualitative analysis further highlighted human experts’ distinct strengths in strategic pathway design and humanistic care. Conclusions Top-tier LLMs demonstrate capability in generating high-quality, evidence-based plans, positioning them as effective “executors” for drafting preliminary regimens. We propose a human-AI collaboration paradigm where experts function as “strategists,” focusing on optimization and humanistic care to elevate rehabilitation service quality.
JPEG is the most widely used image compression format on the Internet, and reversible data hiding (RDH) in JPEG images has attracted increasing attention for privacy protection. Existing two-dimensional (2D) histogram-based RDH schemes suffer from two main limitations: (1) fixed 2D mappings designed heuristically lack universality, and (2) adaptive 2D mappings often incur prohibitive computational costs. To address these issues, this paper proposes a low-complexity adaptive 2D-model-based RDH scheme for JPEG images. Zero AC coefficients are selectively used for embedding to increase capacity while controlling file size growth. A new block-sorting strategy is further designed, and the 2D model is optimized using three components: block usage length, 2D mapping, and frequency selection. By constraining the candidate space and estimating distortion through simulated embedding, the proposed scheme reduces computational complexity. Experimental results show that, on the USC-SIPI image database, the proposed scheme achieves an average PSNR of 40.39 dB and an average file size increment of 15,215 bits under QF = 60 and a payload of 12,000 bits. The proposed scheme achieves a better balance between embedding capacity, visual quality, and file size increment than state-of-the-art methods.
Electronic Health Records (EHRs) continuously monitor patients’ health status in Intensive Care Units (ICUs), capturing irregular numerical time-series data and unstructured clinical text. While existing studies primarily focus on handling modality irregularities, they often overlook the complex intra- and inter-sequence interactions as well as the dependencies between short-term and long-term features. Moreover, clinical notes are typically semantically sparse and structurally noisy, making them difficult to interpret. To address these challenges, we propose a novel multimodal predictive model. For irregular numerical time-series data, we design a cross-view multi-scale framework that integrates cross-attention mechanisms with multi-scale convolutions. This enables dynamic modeling of diverse temporal embeddings while precisely capturing intrinsic inter-variable interactions and cross-temporal dependencies, all with reduced computational complexity. For clinical text, we adopt a retrieval-augmented technique that leverages external medical knowledge graphs (KGs) and large language models (LLMs) to enrich text representations related to medical codes. These enhanced embeddings are then fused with clinical notes via a gated mechanism, effectively alleviating semantic sparsity. We validate the effectiveness of the proposed approach on two critical clinical prediction tasks. Experimental results show maximum relative F1 score improvements of 3.3%, 6.0%, and 3.4% for MISTS, clinical notes, and multimodal fusion tasks, respectively, demonstrating our method’s excellent medical predictive capability.
Cross-dataset Human Activity Recognition (HAR) suffers from limited model generalization, hindering its practical deployment. While prior research has predominantly focused on designing novel model architectures, the crucial role of training data composition is often overlooked. To address this gap, inspired by the success of data mixture optimization in Large Language Models (LLMs), we introduce this strategy to optimize the composition of multi-source training data for HAR models. This approach facilitates the creation of balanced and effective training datasets, thereby enhancing the models’ utility across diverse conditions. To tailor this strategy to continuous, multi-channel Inertial Measurement Unit (IMU) data, we propose HAR-DoReMi, a framework consisting of two complementary parts: a domain reweighting scheme and a pre-processing step for sensor alignment. The domain reweighting scheme includes a Smooth-Stable Strategy using a pre-computed global baseline to improve training stability, a Conditional Batch Score (CBS) for more accurate domain difficulty assessment, and a composite Mean Squared Error (MSE) and Soft-Dynamic Time Warping (Soft-DTW) loss to capture both signal fidelity and temporal dynamics. In addition, the pre-processing step employs the Mahony fusion algorithm to reduce sensor orientation heterogeneity, thereby complementing the domain reweighting scheme in cross-dataset HAR. Extensive experiments across multiple cross-dataset transfer tasks show that the full HAR-DoReMi framework improves average accuracy by 10.5% over baseline approaches while using about 30% to 50% of the training data, improving both robustness under distribution shift and data efficiency. Code will be available at https://github.com/ABan147/har-doremi.
With the rapid growth of data, redundancy among different users in cloud environments has become increasingly prominent. Detecting and removing these redundant parts can effectively improve storage efficiency. But these processes may dramatically degrade the system performance, especially when dealing with similar data. Although deduplication and delta compression are common data reduction techniques, their high overhead can outweigh the benefits. As a result, users often cannot determine in advance whether compression is worthwhile for their datasets. Some approaches have attempted to solve this, but each has important limitations. Danny Harnik et al. proposed a sampling-based deduplication estimation method using linear programming, which efficiently estimates redundancy from exact duplicates. However, it fails to capture redundancy arising from similar data, thus underestimating the full compression potential. To address this limitation, we propose Smart-to-Compress, a predictive compression decision framework. We introduce the Super Feature Frequency Histogram (SFH) to capture redundancy among similar data. Combined with the Duplication Frequency Histogram (DFH), our method estimates the overall Data Reduction Ratio (DRR) without scanning the entire dataset. Furthermore, we design a game-theoretic decision model to weigh compression benefits against predicted costs, providing users with guidance on whether compression should be applied. Experiments on real-world datasets show that our method accurately predicts compression value, reduces unnecessary overhead, and offers reliable decision-making support for users.
The transportation of nuclear radioactive materials is critical to the nuclear industry, yet any leakage during transit is unacceptable due to its severe consequences. To enable rapid localization and identification of a leaking container during transportation, we propose a deep learning (DL)-based method utilizing a sensor array system. A major challenge in deploying DL for real-world nuclear radiation monitoring during transportation is the distribution shift between training and deployment environments, which often leads to overfitting and undermines model effectiveness. To simulate this discrepancy, we intentionally introduce distribution inconsistency between the training and testing datasets. To improve generalization and reduce overfitting under such mismatches, we adopt a multitask (MT) learning framework in which several localization tasks are trained jointly. The localization results of these tasks are fused using a probability-informed neural network (PrINN), where the task-specific predictions are integrated based on their probabilistic relationships. Simulation results demonstrate that the proposed approach maintains strong generalization performance, even in the presence of significant discrepancies between the training phase and real-world deployment.
Large language models (LLMs) excel in various natural language processing tasks and are increasingly applied in specialized fields like medicine. However, their deployment in the medical domain is challenged by limited domain-specific data and the tendency to generate inaccurate information, known as “hallucinations.” While domainspecific fine-tuning has improved open-source LLMs, they still underperform compared to proprietary models like ChatGPT and PaLM. To address this gap, retrieval-augmented generation (RAG) techniques have been explored to enhance LLMs by integrating external knowledge bases. Nevertheless, the success of RAG depends on the quality of retrieved documents, and its application within the medical field remains in the early stages. In this paper, we introduce the “Bailicai” framework as an exploratory approach to integrating RAG with LLMs in the medical field. The framework employs fine-tuning to improve the RAG process, where “falsely relevant” and “completely irrelevant” interference documents are intentionally included in the training data. This enables Bailicai to develop the ability to assess the quality of retrieved documents and selectively incorporate them. The framework is organized into four modules: (1) medical knowledge injection, (2) self-knowledge boundary identification, (3) directed acyclic graph task decomposition, and (4) retrieval-augmented generation. Through the synergy of these modules, Bailicai achieves superior performance on multiple medical benchmarks, outperforming existing large models in the medical domain, RAG-based methods, and proprietary models such as GPT-3.5. Furthermore, Bailicai effectively mitigates the hallucination problem common in LLMs applied to medical tasks and enhances the robustness of RAG when dealing with irrelevant or misleading documents, enabling more accurate information retrieval and integration.
This paper tackles a critical bottleneck in Super-Structure-based divide-and-conquer causal discovery: the high computational cost of constructing accurate Super-Structures–particularly when conditional independence (CI) tests are expensive and domain knowledge is unavailable. We propose a novel, lightweight framework that relaxes the strict requirements on Super-Structure construction while preserving the algorithmic benefits of divide-and-conquer. By integrating weakly constrained Super-Structures with efficient graph partitioning and merging strategies, our approach substantially lowers CI test overhead without sacrificing accuracy. We instantiate the framework in a concrete causal discovery algorithm and rigorously evaluate its components on synthetic data. Comprehensive experiments on Gaussian Bayesian networks, including magic-NIAB, ECOLI70, and magic-IRRI, demonstrate that our method matches or closely approximates the structural accuracy of PC and FCI while drastically reducing the number of CI tests. Further validation on the real-world China Health and Retirement Longitudinal Study (CHARLS) dataset confirms its practical applicability. Our results establish that accurate, scalable causal discovery is achievable even under minimal assumptions about the initial Super-Structure, opening new avenues for applying divide-and-conquer methods to large-scale, knowledge-scarce domains such as biomedical and social science research.
The widespread adoption of multi-agent systems (MAS) in industrial automation has revealed three critical challenges: instruction safety assurance, operational feasibility validation, and cross-agent logical consistency maintenance. While conventional rule-based methods demonstrate two inherent deficiencieslimited semantic comprehension of contextual dependencies and fragmented verification mechanisms for emergent coordination issuesthis paper proposes MA-Shield, an innovative collaborative verification framework that architecturally integrates large language models (LLMs) with formal logic reasoning. The framework comprises three synergistic components: (1) An LLM-Based Intelligent Safety Inspector utilizing semantic parsing and Predefined Most Hazardous Operations to detect latent risks and generate corrective strategies; (2) A Counterfactual Feasibility Evaluator implementing hypothesis generation-verification cycles for empirical validation of production plans under resource constraints; (3) A Multi-Dimensional Consistency Verifier enforcing temporal logic constraints and goal alignment checks across agents. Comprehensive experiments in wine brewing automation demonstrate MA-Shield's superior performance: 98.3% safety inspection accuracy (23.3 percentagepoint improvement over rule-based baselines), 95.0% recall in feasibility evaluation, while maintaining 3.3% false negative and 5.0% false positive rates in consistency verification. These results substantiate the framework's effectiveness in enhancing the safety, reliability, and coordination precision of industrial MAS, establishing a new paradigm for secure multi-agent instruction management.
Gesture recognition plays a crucial role in Human-Machine Interaction (HMI) by enabling interaction with systems without physical contact. Nevertheless, current gesture recognition methods encounter various challenges, including suboptimal lighting conditions, low detection rates, slow processing speeds, and occlusion from protective gloves, which can impede sensor capture of hand movements and consequently degrade recognition accuracy. To overcome the high computational cost and limited robustness observed in existing gesture recognition algorithms, this paper introduces LGRDet, a novel gesture recognition model. LGRDet enhances NanoDet-Plus by integrating Coordinate Attention (CA) and Squeeze-and-Excitation (SE) attention mechanisms into its backbone network, thereby strengthening its capacity to capture long-range spatial dependencies and effectively detect small target gestures. This enhancement is crucial for capturing the fine-grained features of gestures, such as finger bends and palm shapes. Furthermore, the Filtration-Fusion (FF) attention mechanism is incorporated into the original Path Aggregation Network (PAN) to optimize feature fusion across diverse scales. Our proposed LGRDet algorithm represents a notable improvement in gesture recognition accuracy, achieved while upholding a remarkably small model footprint. This characteristic makes LGRDet ideally suited for practical, real-time gesture detection and recognition. Specifically, LGRDet achieved an accuracy of 92.4 for recognizing 9 gestures on our custom IHGD dataset (involving protective gloves), and a robust 92.9 accuracy on the publicly available HAGRID dataset. Crucially, these high-accuracy results are coupled with an outstandingly low inference latency of merely 8.32 ms. These compelling experimental findings underscore the efficacy and real-time capability of the LGRDet algorithm. With a compact model size of just 1.23 MB and its inherently streamlined nature, LGRDet demonstrates immense potential for integration into resource-constrained real-world environments.
To address the dual requirements of efficiency and professionalism in diabetes-related intelligent question-answering, this study presents DiaRAG, an innovative system that synergistically integrates knowledge graphs with retrieval-augmented generation (RAG) techniques. The proposed system is specifically tailored to the diabetes domain, in which both medical expertise and updated knowledge are critical. DiaRAG introduces an autoprompt generation (APG) method that automatically synthesizes diabetes-specific prompt templates. These templates are used to extract structured information from diabetes literature and clinical data, thus facilitating the construction of a comprehensive diabetes knowledge graph and a dedicated retrieval knowledge base. By applying APG, the system effectively generates candidate prompts that enhanced the extraction of relevant knowledge triples, addressing the challenges posed by ambiguous or complex medical queries and ensuring that the subsequent retrieval process is grounded in an accurate, domain-specific context.Furthermore, DiaRAG integrates a specialized text correction module based on PL-BART (Prompt learning and bidirectional auto-regressive transformers). This module is designed to correct semantic and syntactic errors in patient queries. By leveraging prompt-guided correction, PL-BART improves the clarity of input questions, thus enabling the retrieval module to perform more precise matching with the underlying diabetes knowledge graph.In the retrieval phase, a fine-tuned re-ranker model is introduced to further optimize the ordering of the candidate community summaries. This re-ranker, built on a cross-encoder architecture that employs BERT, evaluates the relevance of the retrieved documents to the patient’s query. The secondary filtering provided by this module not only enhances the alignment between the query intent and the retrieved content but also mitigates the common issue of hallucinations in large language models (LLMs) by ensuring that only high-quality, domain-relevant information is passed to the generation stage.Experimental evaluations were conducted on the DaCorp diabetes question-answering dataset, and the results showed that DiaRAG achieved superior performance compared to state-of-the-art models, such as GPT-3.5, HuatuoGPT, and other retrieval-augmented frameworks, such as NaiveRAG and SelfRAG. Key evaluation metrics, including ROUGE-1, ROUGE-2, and ROUGE-L, indicated that DiaRAG consistently outperformed baseline methods in terms of answer accuracy and community summary relevance.Ablation studies further demonstrated that each component—the APG module, PL-BART-based text correction, and fine-tuned re-ranker—contributed significantly to the overall system performance. Notably, iterative prompt optimization via APG and a specialized re-ranking process have been shown to be critical for handling the intricate and specialized language inherent in diabetes-related queries. In a detailed case study involving patient inquiries about the suitability of a traditional Chinese medicine for diabetic conditions, DiaRAG provided a comprehensive answer that not only considered the general pharmacological properties of the medicine but also incorporated detailed clinical insights. This nuanced explanation, which directly addressed the complexities of diabetic complications and the specific indications of the medicine, resulted in expert evaluations rating DiaRAG’s response significantly higher than those provided by competing models such as GPT-3.5 and HuatuoGPT. The experts praised DiaRAG for its precise and contextually appropriate advice, which ultimately highlighted the system’s potential for delivering personalized and reliable medical guidance.Overall, DiaRAG represents an important advancement in the design of domain-specific intelligent question-answering systems. Seamlessly integrating structured knowledge extraction, robust text correction, and refined retrieval strategies, it offers an innovative solution for personalized medical knowledge services in diabetes care.
In recent years, pseudo point clouds generated from depth completion of RGB images and LiDAR data have provided a robust foundation for multimodal 3D object detection. However, the generation process often introduces noise, reducing data quality and detection accuracy. Moreover, existing methods fail to effectively capture channel correlations and global contextual information during the 2D feature extraction stage after the 3D backbone network, limiting detection performance. To address these challenges, this paper proposes NRAP-RCNN, a pseudo point cloud-based 3D object detection method with two key innovations: (1) A noise-reduction sparse convolution network (NRConvNet), comprising NRConv (noise-resistant submanifold sparse convolution), SRB (sparse convolution residual block), and MHSA (multi-head self-attention). NRConv suppresses pseudo point cloud noise by jointly encoding 2D and 3D features, SRB enhances feature extraction depth and robustness, and MHSA optimizes global feature representation. (2) An attention fusion module (ECA_GCA) is introduced to enhance the feature representation of the 2D backbone network by combining channel and global contextual information. The experimental results demonstrate that NRAP-RCNN achieves 88.4% car AP (R40) on the KITTI validation set and 85.1% on the test set, significantly outperforming advanced 3D detection methods, showcasing its effectiveness in improving detection performance.
Recent advances in U-Net and its variants have significantly improved medical image processing; however, they still face issues such as excessive model complexity, limited extraction of fine details, and susceptibility to background noise in real-world applications. To overcome challenges related to the large number of parameters and the insufficient integration of low- and high-level features in skip connections, this paper introduces a novel U-Net architecture enhanced by a dual-attention mechanism. First, the feature propagation pathway is reengineered by incorporating an Attention Gate Module (AG) that applies an adaptive feature selection strategy. This module refines multi-scale features within the skip connections by effectively filtering out noise from low-level semantics while strengthening the expression of critical target features. Second, a Convolutional Block Attention Module (CBAM) is seamlessly integrated after each downsampling stage of the encoder. This addition allows the network to dynamically reassign the importance of feature maps using a combined channel and spatial attention strategy, greatly improving its capability to delineate target boundaries and capture fine structural details. Experimental results on several medical image datasets demonstrate that the proposed model outperforms existing approaches, thereby affirming its feasibility and enhanced performance in medical image segmentation tasks.
Facial action unit (AU) recognition involves predicting the activation states of AUs, which describe facial movements. In complex scenarios, AU relationships and facial features are challenging to capture effectively. This study proposes a novel AU recognition approach that supplements topological features with AU relationship learning. By integrating a channel-topology convolution feature generation structure (CCFG) with a multi-scale attention feature generation structure (MAFG) within a graph neural network, our method models dynamic AU associations and enriches feature representations. Experimental results demonstrate that our approach significantly outperforms state-of-the-art methods on benchmark datasets, achieving average F1 scores of 66.6 https://github.com/lkq52110/au-recognition .
Most existing reversible data hiding (RDH) algorithms for color images adopt a fixed-ratio capacity allocation approach, often resulting in allocation errors. This paper presents a reversible data hiding scheme for color images, integrating adaptive capacity allocation (ACA) and pixel-value similarity distance (PVSD) sorting. In this paper, the adaptive capacity allocation method avoids pre-allocating capacity for the three RGB channels, instead combining them for unified data embedding. When the data embedding is completed, the capacity allocation can be realized. ACA effectively enhances pixel value correlation, aiding data embedding. Existing global pixel value ordering (PVO) methods utilize pixel complexity for secondary sorting. Pixel complexity, which describes local texture smoothness, does not accurately represent pixel value magnitude, leading to inaccurate secondary sorting. This paper employs pixel value similarity distance for secondary sorting. PVSD accurately assesses pixel value differences, enabling precise arrangement of similar-valued pixels in adjacent positions. It enhances pixel sequence smoothness, thereby improving the accuracy of the PVO prediction method. Experimental results demonstrate that the stego-images generated by the proposed scheme exhibit significantly superior visual quality compared to those from other state-of-the-art schemes. Specifically, an average PSNR of 61.38 dB was achieved on the Kodak image dataset at a capacity of 20000 bits.
In the era of big data, deep learning models play a crucial role in identifying underlying patterns within data. However, the need for large volumes of training data, often scattered across various organizations with privacy constraints, poses a significant challenge. Federated Learning (FL) addresses this by enabling the collaborative training of models without sharing the underlying data. Despite its promise, FL encounters challenges with model privacy leakage and computational overhead, particularly when dealing with non-identically distributed (Non-IID) data. To overcome these challenges, we introduce Sym-CS-HFL, a novel Privacy-Preserving Federated Learning (PPFL) framework that combines Symmetric Homomorphic Encryption with a Local Adaptive Aggregation (LAA) scheme. Our approach minimizes the reliance on asymmetric keys, simplifying the encryption process and reducing computational overhead. We implement a DCT-Neural Network Compressive Sensing Scheme to decrease communication costs substantially. Furthermore, the LAA scheme addresses the heterogeneity in Non-IID data, enhancing model convergence and accuracy. Our experiments on diverse datasets, including MNIST, FashionMNIST, CIFAR-10/100, and AG News, demonstrate that Sym-CS-HFL achieves a Top-3 test accuracy while significantly reducing communication overhead by 15.2x to 74x compared to existing HE schemes. The computational overhead is also reduced, with training times only 1.1x to 1.8x that of plaintext training. These results underscore Sym-CS-HFL's effectiveness in maintaining high performance and privacy in PPFL.
Wearable sensor-based human activity recognition (HAR) is a critical research domain in activity perception. However, achieving high efficiency and long sequence recognition remains a challenge. Despite the extensive investigation of temporal deep learning models, such as CNNs, RNNs, and transformers, their extensive parameters often pose significant computational and memory constraints, rendering them less suitable for resource-constrained mobile health applications. This study introduces HARMamba, an innovative light-weight and versatile HAR architecture that combines selective bidirectional State Spaces Model and hardware-aware design. To optimize real-time resource consumption in practical scenarios, HARMamba employs linear recursive mechanisms and parameter discretization, allowing it to selectively focus on relevant input sequences while efficiently fusing scan and recompute operations. The model employs independent channels to process sensor data streams, dividing each channel into patches and appending classification tokens to the end of the sequence. It utilizes position embedding to represent the sequence order. The patch sequence is subsequently processed by HARMamba Block, and the classification head finally outputs the activity category. The HARMamba Block serves as the fundamental component of the HARMamba architecture, enabling the effective capture of more discriminative activity sequence features. HARMamba outperforms contemporary state-of-the-art frameworks, delivering comparable or better accuracy with significantly reducing computational and memory demands. It’s effectiveness has been extensively validated on 4 publically available datasets namely PAMAP2, WISDM, UNIMIB SHAR and UCI. The F1 scores of HARMamba on the four datasets are 99.74%, 99.20%, 88.23% and 97.01%, respectively.
Nuclear power, as a quintessentially complex system, is characterized by prolonged operational cycles and high operational costs. The application of intelligent health management can effectively reduce both operational and maintenance costs. This paper combines digital twin technology with generative adversarial networks, sparse denoising autoencoders, and long short-term memory networks to mitigate issues such as high sparsity in sensor data, limited adaptability of purely data-driven algorithms, and challenges faced by physics-driven models in simulating intrinsic system characteristics. Specifically, after establishing the digital twin model, twin data are generated using generative adversarial networks. The health management model was constructed by combining sparse denoising autoencoders and long short-term memory networks. This model encompasses health monitoring, fault diagnosis, and degradation prediction. In this way, the health management model enables full lifecycle health management of nuclear power systems and transforms passive periodic preventive maintenance into proactive predictive maintenance, thus saving maintenance costs.
Large-scale multi-label Text Classification (LMTC) is an advanced facet of NLP that entails assigning multiple labels to text documents from an extensive label space, often comprising thousands to millions of possible categories. This classification task is pivotal across various domains, including e-commerce product tagging, news categorization, medical code assignment, and legal document analysis, where accurate multi-label predictions drive search efficiency, recommendation systems, and regulatory compliance. However, LMTC poses significant challenges, the dynamic nature of label sets, which traditional supervised learning approaches find difficult to address due to their reliance on annotated data. In light of this challenge, this work introduces a novel approach leveraging Large Language Models (LLMs) for dynamic label alignment in LMTC tasks, based on counterfactual analysis, called DyLas (Dynamic Label Alignment Strategy). Through a multi-step strategy, we aim to mitigate the issues arising from dynamic label sets. We evaluate the performance of LMTC on the 8 LLMs by 4 datasets and apply DyLas to 3 closed-source and 3 open-weight LLMs. Compared to the single-step approach, our method, DyLas, achieves improvements in almost all metrics across the datasets. Our method can also work well in dynamic label set environments. This work not only demonstrates the potential of LLMs to address complex classification challenges, but is also, to the best of our knowledge, the first to address dynamic label set challenges in LMTC tasks with LLMs without requiring additional model training.
As the prevalence of cloud storage increases, many individuals and companies prefer outsourcing their data for backup and management. However, this has led to a significant increase in redundancy, decreasing storage utilization and wasting network bandwidth. While conventional resemblance detection methods remove redundancy among similar data by comparing the features extracted from each chunk's content. However, we observed that small changes between similar data chunks may cause false dissimilarity detection by conventional resemblance detection techniques. This is because features derived solely from the chunk content are highly susceptible to various modification patterns. Fortunately, we have discovered that two chunks are likely to be similar if their surrounding chunks are also similar, a concept we refer to as "chunk- context". Therefore, we propose a novel chunk-context aware resemblance detection method, called CARD, which includes a network-based chunk-context aware model and an N-sub-chunk shingles-based initial feature extraction strategy. By leveraging the Neural network, it can discover the complex patterns between the chunk- context and chunk content itself. A high-level understanding of the contextual information with chunk content can be synthesized into the representation of a chunk. The primary difference compared with others is that our design can significantly improves the accuracy or efficiency of resemblance detection by considering the chunk- context with chunk content itself. Furthermore, we implemented a CARD prototype and conducted extensive experiments using real workload, demonstrating that CARD can detect up to 75.03% more redundant data and accelerate the resemblance detection operations by 5.6 x to 86.7 x faster than state-of-the-art work.