
To address the challenges of exponential data growth in the Earth observation field, weak semantic relationships among multi-source heterogeneous data, difficulties in natural language search for users, and high demands for retrieval accuracy, we propose a semantic retrieval method that integrates knowledge graphs with multi-stage retrieval and re-ranking. First, we construct the Earth Observation Knowledge Graph (EORec-KG), which covers core entities such as satellites, data products, and research reports, to achieve a unified structured representation of multi-source heterogeneous knowledge; Second, we design a multi-stage retrieval workflow that combines parallel initial recall (graph querying and vector retrieval) with semantic re-ranking via BGE-Reranker; finally, we develop a Earth Observation scientific data semantic retrieval system based on the Dify platform and establish a multi-dimensional evaluation framework encompassing precision, recall, F1 score, and hit rate. Comparative experiments were conducted on a test set comprising 150 typical queries, the results show that the proposed method achieves the best performance across all four metrics—P@5, R@5, F1@5, and HR@5—and outperforms baseline methods such as pure large language models, pure vector retrieval, RAG without reranking, and external reranking models. Specifically, compared to RAG without re-ranking, P@5 improved by 29.4%, R@5 by 52.4%, and F1@5 by 44.4%, validating the effectiveness of the proposed method in enhancing retrieval accuracy and overall performance. This method provides a deployable technical framework for intelligent retrieval of Earth observation scientific data.
In geo-distributed big data computing based on Apache Spark, data skew may cause excessive workloads on certain partitions, thereby delaying the overall execution progress of computing tasks. Existing data skew handling approaches are mainly designed for single data center environments and are insufficient to achieve global partition load balancing in geo-distributed scenarios. To address this issue, this paper proposes a data skew handling method for geo-distributed big data computing, termed Skew Detection and Repartition (SDR), which consists of two key operations: skewed partition detection and balanced repartitioning. Specifically, SDR first collects statistical information, including the frequency and corresponding data size of each key generated during the Map phase across multiple data centers. A key-based pre-aggregation mechanism is introduced to reduce cross-domain communication overhead during statistic collection. Then, a median-driven skew detection strategy is employed to identify globally skewed partitions. Furthermore, a greedy bin-packing-based repartitioning algorithm is designed, where the data size associated with each key is considered as its weight. The algorithm sorts keys in descending order of their weights and iteratively assigns them to the currently maximum load partition that does not exceed the target partition size. While ensuring that records with the same key are not split across different partitions, the proposed algorithm dynamically balances partition workloads and improves subsequent task scheduling efficiency. Experimental results demonstrate that SDR effectively mitigates partition skewness and reduces task execution time across datasets with different skew levels. In highly skewed WordCount workloads, SDR reduces execution time by up to 32.20% compared with the original hash partitioning method. Moreover, across five datasets with varying skew levels, SDR achieves an average reduction of 10.74% in overall execution time.
Scientific evaluation of data quality is a core prerequisite for effective data governance, as it directly affects the utility of high-quality AI-Ready dataset for downstream tasks. However, current research mainly focuses on specific modalities or dimensions, leading to incomparable evaluation results across different modalities and dimensions, and making it difficult to establish a systematic data quality profile. To address this issue, this work proposes an Evaluatology-based multimodal data quality evaluation framework for AI-Ready materials dataset, by defining the evaluation subject, evaluation conditions, and evaluation method to construct a unified evaluation system. Then, the six rules and five-step construction process of "Inheritance–Development" are employed to extend the unified framework into modality-specific evaluation systems. Furthermore, a case study on the completeness of AI-Ready materials time-series data was conducted, validating the effectiveness of the evaluation system designed based on the proposed framework. This work provides a theoretical and scientific foundation for constructing high-quality AI-ready dataset and developing highly reliable data-driven models.
In recent years, the application of graph contrastive learning in recommendation systems has made significant progress. However, most current graph enhancement strategies still rely on manual experience design, lack flexibility and generalization ability, and are difficult to adapt to diverse recommendation tasks and complex graph structures. To this end, this paper proposes an adaptive multi-view contrastive recommendation algorithm (MAMGCL) based on a hybrid expert model. This method constructs an expert pool containing multiple enhancement strategies, introduces a two-way gating mechanism to achieve dynamic fusion of experts on the user side and the item side, further generates differentiated enhanced views, uses a multi-branch graph convolutional network to achieve node embedding expression, and finally combines sub-view-level contrastive learning for optimization. Experiments are carried out on three real-world datasets, verifying that the proposed method is superior to the existing mainstream contrastive recommendation models in terms of recommendation performance and robustness, demonstrating strong generalization ability and application potential.
Large model technologies are reshaping the fundamental paradigm of digital government construction. This paper proposes a dual-architecture model empowered by large models, namely the “Provincial Intensive Middle Platform + Industry Base Platform”. The Provincial Intensive Middle Platform is centrally planned and built by provincial digital government administration departments, hosting the province‑wide unified data middle platform and large‑model capability base (MaaS, Model as a Service), and exporting common capabilities to industry platforms via standardized APIs. The Industry Base Platform is constructed and maintained by individual provincial bureaus, which accumulate industry‑specific common business components and rapidly configure and assemble various business application scenarios based on the capability components provided by the provincial middle platform. The core argument of this paper is that, with the deep application of large models in the digital government domain, the mainstream model of government informationization construction will gradually shift from “building systems” to “configuring scenarios” – business personnel in various bureaus, without programming, can independently build “one after another” business scenarios using AI tools, agents, and low‑code configuration capabilities provided by the MaaS platform. This paper systematically elaborates the implementation path of this paradigm shift, and demonstrates it with cases such as Guangdong’s “Digital Housing and Urban-Rural Development” and “Yuefuyong” (reusable application marketplace), as well as the “application model” of Feishu Multidimensional Tables. It also proposes supporting reform plans from dimensions such as project approval and management mechanisms, providing a theoretical framework and operational reference for the transformation of the digital government construction model.
With the widespread application of speech technology, voice privacy protection faces severe challenges. A speech anonymization method (kLCAnoy++) based on latent space splicing and contrastive learning was proposed, aiming to effectively remove speaker identity information while preserving speech content intelligibility and naturalness. The method first employed the pretrained WavLM model to extract self-supervised speech features. Subsequently, it matched target features in a diverse reference audio feature library using the k-nearest neighbors algorithm, achieving latent space feature splicing to generate preliminary anonymized speech. Finally, a contrastive learning mechanism was introduced, treating natural splicing as positive samples and k-nearest neighbor splicing as negative samples, to optimize the vocoder decoding process and enhance the coherence and naturalness of anonymized speech. Experiments on the VCTK dataset demonstrated that the proposed method significantly outperformed baseline approaches, with visualization analysis further confirming its effectiveness in disentangling speaker identity information.
With the development of quantum computing, traditional public-key cryptographic systems based on integer factorization and discrete logarithm problems face security challenges. To systematically review the research progress of post-quantum cryptography, its theoretical foundations, algorithmic systems, engineering implementation, and deployment issues were surveyed. First, the vulnerabilities of traditional cryptographic systems under quantum attack models and the constraints of protocol migration were analyzed. Then, post-quantum cryptographic schemes, including lattice-based, hash-based, code-based, multivariate, and isogeny-based schemes, were classified and compared, and hardware adaptation, implementation optimization, and implementation security were discussed in combination with CPU, FPGA, and embedded platforms. Under a unified testing environment, the performance and communication overhead of traditional, post-quantum, hybrid, and selected signature and KEM schemes proposed and publicly implemented by domestic research teams were evaluated. The results show that different technical routes involve obvious trade-offs among security, computational efficiency, and communication cost. Finally, challenges in long-term security validation, implementation security, system migration, and engineering deployment were summarized, providing a reference for post-quantum cryptography research and application.
Addressing the shift in the development and utilization of public data from the question of "whether to open it up" to "how to govern it to realize its value", the dilemma that institutional rules, organizational coordination, resource management and security oversight have not yet formed a comprehensive governance system was investigated. Based on institutional texts, literature review and local cases, a five-layer governance framework was constructed, comprising the institutional rule layer, organizational coordination layer, resource management layer, circulation and application layer, and security assurance layer, and the mechanisms of rule transmission, conflict adjustment and feedback iteration among the layers were elaborated. Public data from eight provinces and municipalities were selected, and the entropy weight method and Pearson correlation analysis were employed to calculate the layer weights and correlation degrees. Shanghai, Henan and Guizhou were chosen to represent the eastern, central and western regions for comparative case analysis. The results showed that the circulation and application layer and the resource management layer contributed the most, and the strongest correlation was found between the institutional rule layer and the organizational coordination layer, which verified the general applicability of the framework. It is concluded that the key to unlocking the value of public data lies in forming a closed-loop mechanism of rule transmission, collaborative governance and controllable security through the five-layer framework.
High-quality government datasets can effectively empower intelligent capabilities in public services, law enforcement regulation, and urban governance. Focusing on the scenario of enterprise credit risk regulation, this paper proposes a high-quality dataset construction method based on pluggable domain knowledge and multi-agent collaboration. By encapsulating domain ontologies and business rules into pluggable knowledge plugins, this approach decouples the system architecture from domain knowledge, thereby facilitating low-code migration across domains. Furthermore, it constructs a multi-agent collaboration closed-loop of "detection, planning, action, and evaluation" to achieve fully autonomous and iterative workflows for data quality management. An auditable chained logging mechanism is also established to ensure end-to-end data lineage traceability and operational accountability. This method provides an interpretable, reusable, and cross-domain adaptive technical pathway for building high-quality datasets in the government sector.
Data governance is a fundamental basis for ensuring data quality and security, promoting efficient data circulation, and realizing data value. However, in practical machine learning (ML) model training scenarios, data preparation is still constrained by low resource utilization efficiency, uncontrollable costs, and insufficient optimization under budget constraints. To address these problems, a data preparation method oriented toward cost-benefit collaborative optimization was proposed. Mathematical programming was adopted to incorporate data preprocessing, cleaning, transformation, and other operations into a unified cost modeling and decision-making framework, and a cost-benefit analysis and planning method for the overall pipeline was designed. Experimental results showed that the proposed method achieved a better trade-off between cost and performance under limited budget conditions, thereby improving the overall performance and cost-effectiveness of the data preparation process. The results indicate that incorporating cost-benefit analysis and optimization into data preparation strategies facilitates efficient resource allocation and provides useful support for cost control and performance optimization in complex data processing scenarios.
To meet the growing data demands of AI systems, AI-Ready data has become a key focus area in current AI development and has received high priority from governments worldwide. Among these, the U.S. National Institutes of Health (NIH) has taken the lead in building AI-Ready scientific datasets. A systematic study of its approach can provide valuable reference and guidance for the construction of high-quality scientific datasets in China. Employing web-based investigation, content analysis, and case study methods, this paper systematically introduces NIH's data policies and funding policies, the flagship datasets of the Bridge2AI (bridge to artificial intelligence) research program, and the development of AI readiness frameworks. Furthermore, drawing on the practical realities of high-quality dataset construction in China, it distills beneficial experiences and insights, including improving the policy system, establishing standard systems, and building a talent support system.
With the intensifying trend of population aging, residential elderly care has become the predominant modality in China. Embodied AI robots are increasingly recognized as a pivotal technology for mitigating the scarcity of caregiving resources and the latency in emergency responses. The training of embodied models necessitates large-scale, high-fidelity interactive data; however, empirical data acquisition is severely constrained by high safety risks, prohibitive costs, and insufficient sample diversity. Furthermore, general-purpose simulation datasets frequently fail to encapsulate the idiosyncratic environmental features, human-robot safety imperatives, and core service tasks essential to elderly care, thereby hindering industrial deployment. To address these gaps, this paper proposes a systematic methodology for constructing embodied Intelligence simulation datasets specifically for residential care scenarios. Leveraging a synergistic dual-engine framework that integrates the Marble World Model with NVIDIA Isaac Sim, we establish a four-tier architecture comprising scene generation, physical modeling, control execution, and data acquisition. We further design specialized operational pipelines for quadrupedal robots in emergency response and bipedal robots in daily assistance, incorporating rigorous physical safety constraints. This approach facilitates the generation of multimodal, standardized, and highly generalizable domain-specific simulation data, effectively filling the data void in the field of embodied intelligence for elderly care and providing robust technical support for sim-to-real transfer and large-scale industrialization.
The management of electronic data in power material supply chains is complex. Traditional methods for grading trust intensity are limited by static rules and single evaluation dimensions, and lack guidance from domain knowledge, making them insufficient for fine-grained and differentiated applications. To address this, we propose Adap-LLM, an adaptive grading method that injects expert-constructed multi-dimensional scoring systems and grading rules into large language models through structured prompts, and further integrates domain knowledge via parameter-efficient fine-tuning. Experiments on real-world power datasets and public benchmarks show that the proposed method outperforms existing approaches across multiple metrics, demonstrating its effectiveness and potential application value in hierarchical management of power supply chain data.
International Chinese language education relies on heterogeneous resources such as proficiency standards, characters, words, grammar points, assessment items, learner errors, cultural knowledge and multimedia materials. However, these resources are often fragmented, weakly aligned and difficult to reuse in teaching applications. CFL-KG, a knowledge graph dataset for international Chinese language education, was constructed. Through data cleaning, field normalization, entity alignment, relation annotation and quality checking, multi-source resources were organized into computable and traceable graph data. The current snapshot contains 389,111 nodes and 1,017,421 relations. The dataset supports cross-standard lesson preparation, assessment-resource tracing, learner-error diagnosis, cultural explanation and intelligent question answering, providing a reusable data foundation for intelligent Chinese language education.
To address the lack of open and standardized fine-grained pronunciation error datasets for international Chinese language education and the limited adaptation of general models to non-native pronunciation diagnosis, a Chinese pronunciation error dataset for international students was constructed. A ten-layer structured annotation scheme was developed by integrating basic acoustic alignment with expert diagnostic labels, and a human-machine collaborative workflow of machine pre-annotation, manual refinement, and expert review was established. Data cleaning, quality review, and structured storage were also completed to ensure that the samples were traceable and trainable. As a result, 340 authentic speech samples, more than 120,000 time-aligned slices at the initial-final level, and tens of thousands of fine-grained diagnostic labels were obtained. To evaluate the dataset, two representative tasks were conducted: automatic speech recognition for non-native accented speech and end-to-end pronunciation error diagnosis. The results showed that, after fine-tuning on the dataset, Paraformer-large and Qwen2-Audio-7B achieved better task adaptation. The dataset provides supervised data for model fine-tuning and intelligent pronunciation feedback in international Chinese language education, and it also offers a reference for high-quality dataset governance in vertical domains.
High-quality dataset is a critical foundation for advancing artificial intelligence(AI). However, significant practical challenges were identified in existing development models, including static and preset quality evaluations, difficulties in forming closed-loop business models, and a disconnect from the real economy. In response, a new paradigm for building high-quality datasets was proposed and justified: the development of high-quality dataset was deeply integrated into the data factor market, and data quality was dynamically defined through market-based circulation and multi-dimensional user feedback. This paradigm was centered on“empowerment of the real economy as the main focus and AI training services as a supplement.” Trustworthiness, controllability, and measurability in data circulation were ensured under this paradigm using a “trusted data space based on real-virtual integration”, supported by full-flow circulation mapping and large models for data circulation. On this basis, a full-process technical solution was designed—from data resource market entry to precise distribution—and sustainable business models and incentive mechanisms were established. These efforts provided innovative ideas for shifting the development of high-quality dataset from a “static, supply-driven” approach to a “dynamic, market-driven, and ecosystem-based operational” model.
With the breakthrough progress of large language models (LLMs) in natural language understanding and generation, leveraging artificial intelligence to assist legal document drafting and judicial reasoning has emerged as a pivotal direction in legal AI research. However, existing methods applied to courtroom simulation tasks still face critical scientific challenges, including incomplete legal information acquisition, strong structural constraints of legal texts, uninterpretable judicial reasoning processes, and difficulties in modeling complex multi-stage court procedures. To address the aforementioned challenges, this paper proposes a role-driven, multi-stage court simulation framework (AgentCourts), which aims to effectively alleviate information gaps, improve the standardization of document structure, enhance the interpretability of reasoning, and optimize trial process modeling. Specifically, our method first employs a multi-turn dialogue-based information collection mechanism to incrementally extract key case elements, formulating complaint generation as an information extraction and integration problem; second, it adopts an adversarial legal text modeling strategy to generate defense statements that semantically respond to the plaintiff's claims; third, it explicitly models the civil trial workflow, including court session initiation, evidence investigation, debate, mediation, and judgment deliberation by introducing evidence-driven reasoning mechanisms; finally, it transforms legal document generation from "free-form generation" to "structured template filling" under template constraints, while invoking a legal article retrieval interface to enhance the accuracy of legal applicability. Experimental results on the CAIL 2025 Courtroom Simulation dataset demonstrate that the proposed method consistently generates legally compliant, structurally standardized, and logically coherent documents, achieving superior performance across three subtasks (complaint, defense statement, and judgment generation), thereby validating the effectiveness and practicality of multi-stage modeling and template-constrained generation in judicial scenarios.
The heterogeneity of method-level interactions, the high dynamics of container instances, and the scarcity of anomaly labels make anomaly detection in microservice systems difficult. Therefore, a dynamic graph and code execution context-aware microservice anomaly detection method (DGCEC-MAD) was proposed. Firstly, code information corresponding to service calls was collected through instrumentation techniques, including execution semantics such as class names, method names, parameters, and return values, which were encoded using a pre-trained sentence embedding model. Secondly, service calls were modeled as a continuous-time dynamic graph with timestamps and edge attributes; updatable node memories and a graph attention mechanism that fuses temporal and edge features were employed to capture the evolution of inter-service relations. Then, an improved link prediction task with dual negative sampling was designed to learn normal calling patterns purely from normal data and obtain high quality node representations. Finally, the learned representations were fused with KPIs, and anomalies were identified in an unsupervised manner via autoencoder reconstruction errors. Experiments on two microservice systems, including container-migration scenarios, demonstrate superior performance and robustness, with consistent gains over strong dynamic-graph baselines. Ablation studies further show that incorporating code execution context finely characterizes method-level behavior and significantly improves detection effectiveness.
As an emerging production factor in the digital economy era, data elements face numerous constraints caused by underdeveloped basic institutional frameworks during the advancement of market-oriented allocation reform. Taking the Yangtze River Delta Region and the Guangdong-Hong Kong-Macao Greater Bay Area—the two pioneering pilot zones for data element marketization as the research subjects, this paper systematically compared the development status and policy landscape of data element market-oriented allocation in these two regions. Through comparative study, it identifies the common institutional challenges and distinctive features of their respective basic data element institutions. It is found that leveraging their comparative advantages, both regions have created a series of practical mechanisms to address institutional bottlenecks and unlock the value of data elements, two typical paradigms for regional synergy and cross-border flow in the reform of the market-oriented allocation of data factors have been formed respectively. Furthermore, this paper puts forward a national applicable development approach for the reform of the market-oriented allocation of data elements, which adapts the unified national core basic systems with regionallyflexible implementation rules, specific mechanism recommendations are also proposed for each basic system.
To address the challenges of diverse airborne data types, difficulties in precise probe alignment, and high data leakage risks during data splitting in weather modification aircraft operations, a multi-source microphysical benchmark dataset oriented toward data governance and leakage-controlled evaluation is constructed. Through methods including dirty data rejection, adaptive parsing of multi-character encodings, 1 Hz fusion anchored by the DXAS probe, and non-negative physical ice thickness generation, airborne data from 43 flights were transformed into a high-quality fused table containing 336,570 samples. Furthermore, a strict nested Leave-One-Session-Out (LOSO) cross-validation protocol was established to block event-level leakage. Benchmark testing using the XGBoost algorithm verified the validity of icing onset signals in the airborne data and established a baseline for regression tasks under leakage control. This dataset provides a standardized data foundation for aircraft icing mechanism analysis and early-warning algorithm development, offering a rigorous evaluation reference for enhancing risk monitoring and ensuring flight safety.