Farm field units are the fundamental management units of modern precision agriculture, and accurate identification of their geographic boundaries is essential for the data management of farmland soil and nutrients. However, traditional manual vectorization methods are inefficient and costly. To address this issue, this study proposes an automated farm field unit identification framework based on remote sensing image segmentation and parameter optimization. Starting from the framework of Geographic Object-Based Image Analysis (GEOBIA), the framework employs the multiresolution segmentation (MRS) algorithm combined with intra-segment homogeneity and inter-segment heterogeneity indicators, and iteratively determines the optimal scale parameters through three unsupervised segmentation parameter optimization (USPO) methods, namely the methods proposed by Espindola et al. (EUSPO), Zhang et al. (ZUSPO), and Johnson et al. (JUSPO), to achieve automatic identification of farm field information. The framework was evaluated on eleven production teams of the Tenihe Farm, a large-scale intensive farm in Hulunbuir, China, using two scenes of Landsat 8 OLI imagery acquired on 18 September 2019 (multispectral bands pan-sharpened to 15 m, with the red, green, blue, and near-infrared bands used for segmentation), taking field-investigation-verified farm field vector maps as the reference. The results show that the identification accuracy of farm field units generally reaches over 80%. The main contribution of this study lies in extending USPO methods from segmentation quality evaluation to automated farm field vector extraction in precision agriculture, and in providing a systematic horizontal comparison of three USPO combination strategies under large-scale heterogeneous agricultural landscapes, which reveals that the JUSPO method, with its F-measure fusion strategy, achieves the most robust performance in most teams. These findings verify the effectiveness and application potential of the proposed framework.
Sampling design plays a crucial role in managing soil variability at the farm level by ensuring the collection of reliable soil samples. Current methods focus on optimizing sample size and locations within geographic or feature space, taking a global perspective. However, treating geographic and feature space separately diminishes the robustness of the sampling design against deviations from the assumptions of the prediction model. Furthermore, these approaches often overlook local spatial structures of soil, potentially leading to decreased representativeness of the soil samples. In this paper, a novel sampling design method that considers local soil heterogeneity in both geo-graphical and feature spaces was developed. First, the level of local soil heterogeneity was inferred. Second, the study area was divided into spatially continuous subregions, each characterized by similar levels of soil heterogeneity, to facilitate more objective sample assignment. Finally, key sample locations within each subregion were deter-mined by achieving a balance between coverage in both geographical and feature spaces. To validate the proposed method, it was compared to stratified random sampling, k-means sampling and conditional Latin hypercube sampling. The ordinary kriging method was employed to map five soil properties, including soil organic matter, pH, total nitrogen, available phosphorus and available potassium. The comparative experiment demonstrates that the proposed method is more effective in generating a sampling plan that accurately represents multiple soil properties by analyzing the local soil spatial structure.
The rapid rise of malware challenges traditional detection methods due to code obfuscation and polymorphism. While machine learning classifiers offer quick detection and can identify complex malicious features, they are susceptible to backdoor attacks. We introduce GAT, a genetic algorithm-based approach for generating effective and stealthy Android backdoors. Using the SHAP interpretability tool, we first select efficient features as primary backdoors. A fitness function then enables iterative optimization through a genetic algorithm. Additionally, we propose a method to integrate backdoor features into the source code, maintaining functionality while facilitating attacks on Android classifiers in real data outsourcing scenarios. Our evaluation of the Drebin and Mamadroid malware detectors in data outsourcing scenarios indicates that an attack success rate exceeding 70
Soil datasets, including soil sample data and soil map products, often contain outliers that can lead to inaccurate modeling and analysis of various soil-related issues. Existing methods for identifying potential outliers in soil datasets rely on simple statistical approaches and tend to overlook the geographical characteristics of the soil. Local indicators of spatial association (LISA) can address this limitation by examining the local spatial structures inherent in soil data. However, distinguishing some outliers remains challenging because of the varying levels of heterogeneity across different soil regions. In this paper, we present a novel method for recognizing potential outliers through local heterogeneity enhancement, which is aimed at improving the quality of soil datasets. In this method, stratified soil variations are first balanced to mitigate the effects of spatial discrepancies in different soil regions. Second, local heterogeneity enhancement is conducted to modify the outlier scores associated with abnormal soils exhibiting low heterogeneity. Third, a frequency histogram of outlier scores is applied to determine a suitable threshold at which to recognize potential abnormal values in soil datasets. To validate the proposed method, it was compared with the LISA and box-plot methods. Simulation data and soil data were adopted in the experiment, incorporating two types of irregular points and spatially continuous surfaces. The comparative experiments demonstrated that the proposed method more effectively identifies potential outliers by analyzing and balancing the local spatial structure of the soil than traditional methods do. It can be concluded that local heterogeneity enhancement is beneficial for recognizing potential outliers in soil datasets.
Large language models (LLMs), such as Codex and GPT-4, have recently showcased their remarkable code generation abilities, facilitating a significant boost in coding efficiency. This paper will delve into utilizing LLMs for code generation in private libraries, as they are widely employed in everyday programming. Despite their remarkable capabilities, generating such private APIs poses a formidable conundrum for LLMs, as they inherently lack exposure to these private libraries during pre-training. To address this challenge, we propose a novel framework that emulates the process of programmers writing private code. This framework comprises two modules: APIFinder first retrieves potentially useful APIs from API documentation; and APICoder then leverages these retrieved APIs to generate private code. Specifically, APIFinder employs vector retrieval techniques and allows user involvement in the retrieval process. For APICoder, it can directly utilize off-the-shelf code generation models. To further cultivate explicit proficiency in invoking APIs from prompts, we continuously pre-train a reinforced version of APICoder, named CodeGenAPI. Our goal is to train the above two modules on vast public libraries, enabling generalization to private ones. Meanwhile, we create four private library benchmarks, including TorchDataEval, TorchDataComplexEval, MonkeyEval, and BeatNumEval, and meticulously handcraft test cases for each benchmark to support comprehensive evaluations. Numerous experiments on the four benchmarks consistently affirm the effectiveness of our approach. Furthermore, deeper analysis is also conducted to glean additional insights.
Code large language models (Code LLMs) have demonstrated remarkable performance in code generation. Nonetheless, most existing works focus on boosting code LLMs from the perspective of programming capabilities, while their natural language capabilities receive less attention. To fill this gap, we thus propose a novel framework, comprising two modules: AttentionExtractor, which is responsible for extracting key phrases from the user's natural language requirements, and AttentionCoder, which leverages these extracted phrases to generate target code to solve the requirement. This framework pioneers an innovative idea by seamlessly integrating code LLMs with traditional natural language processing tools. To validate the effectiveness of the framework, we craft a new code generation benchmark, called MultiNL-H, covering five natural languages. Extensive experimental results demonstrate the effectiveness of our proposed framework.
In vertical federated learning (VFL), commercial entities collaboratively train a model while preserving data privacy. However, a malicious participant's poisoning attack may degrade the performance of this collaborative model. The main challenge in achieving the poisoning attack is the absence of access to the server-side top model, leaving the malicious participant without a clear target model. To address this challenge, we introduce an innovative end-to-end poisoning framework P-GAN. Specifically, the malicious participant initially employs semi-supervised learning to train a surrogate target model. Subsequently, this participant employs a GAN-based method to produce adversarial perturbations to degrade the surrogate target model's performance. Finally, the generator is obtained and tailored for VFL poisoning. Besides, we develop an anomaly detection algorithm based on a deep auto-encoder (DAE), offering a robust defense mechanism to VFL scenarios. Through extensive experiments, we evaluate the efficacy of P-GAN and DAE, and further analyze the factors that influence their performance.
The ecological quality of large-scale farms is a critical determinant of crop growth. In this paper, an ecological assessment procedure suitable for agricultural regions should be developed based on an improved remote sensing ecological index (IRSEI), which introduces an integrated salinity index (ISI) tailored to the salinized soil characteristics in farming areas and incorporates ecological indices such as the greenness index (NDVI), the humidity index (WET), the dryness index (NDBSI), and the heat index (LST). The results indicate that between 2013 and 2022, the mean IRSEI increasing from 0.500 in 2013 to 0.826 in 2020 before decreasing to 0.646 in 2022. From 2013 to 2022, the area of the farm that experienced slight to significant improvements in ecological quality reached 1419.91 km2, accounting for 71.94% of the total farm area. An analysis of different land cover types revealed that the IRSEI performed more reliably than did the original RSEI method. Correlation analysis based on crop yields showed that the IRSEI method was more strongly correlated with yield than was the RSEI method. Therefore, the proposed IRSEI method offers a rapid and effective new means of monitoring ecological quality for agricultural planting areas characterized by soil salinization, and it is more effective than the traditional RSEI method.
In the era of code large language models (code LLMs), data engineering plays a pivotal role during the instruction fine-tuning phase. To train a versatile model, previous efforts devote tremendous efforts to crafting instruction data that covers all the downstream scenarios. Nonetheless, this will incur significant expenses in data construction and model training. Therefore, this paper introduces CODEM, a novel data construction strategy, which can efficiently train a versatile model using less data via our newly proposed ability matrix. CODEM uses ability matrix to decouple code LLMs' abilities into two dimensions, constructing a lightweight training corpus that only covers a subset of target scenarios. Extensive experiments on HumanEvalPack and MultiPL-E reveal that code LLMs can combine the single-dimensional abilities to master composed abilities, validating the effectiveness of CODEM.
The task of code generation aims to generate code solutions based on given programming problems. Recently, code large language models (code LLMs) have shed new light on this task, owing to their formidable code generation capabilities. While these models are powerful, they seldom focus on further improving the accuracy of library-oriented API invocation. Nonetheless, programmers frequently invoke APIs in routine coding tasks. In this paper, we aim to enhance the proficiency of existing code LLMs regarding API invocation by mimicking analogical learning, which is a critical learning strategy for humans to learn through differences among multiple instances. Motivated by this, we propose a simple yet effective approach, namely DiffCoder, which excels in API invocation by effectively training on the differences (diffs) between analogical code exercises. To assess the API invocation capabilities of code LLMs, we conduct experiments on seven existing benchmarks that focus on mono-library API invocation. Additionally, we construct a new benchmark, namely PanNumEval, to evaluate the performance of multi-library API invocation. Extensive experiments on eight benchmarks demonstrate the impressive performance of DiffCoder. Furthermore, we develop a VSCode plugin for DiffCoder, and the results from twelve invited participants further verify the practicality of DiffCoder.
The impressive performance of large language models (LLMs) on code-related tasks has shown the potential of fully automated software development. In light of this, we introduce a new software engineering task, namely Natural Language to code Repository (NL2Repo). This task aims to generate an entire code repository from its natural language requirements. To address this task, we propose a simple yet effective framework CodeS, which decomposes NL2Repo into multiple sub-tasks by a multi-layer sketch. Specifically, CodeS includes three modules: RepoSketcher, FileSketcher, and SketchFiller. RepoSketcher first generates a repository's directory structure for given requirements; FileSketcher then generates a file sketch for each file in the generated structure; SketchFiller finally fills in the details for each function in the generated file sketch. To rigorously assess CodeS on the NL2Repo task, we carry out evaluations through both automated benchmarking and manual feedback analysis. For benchmark-based evaluation, we craft a repository-oriented benchmark, SketchEval, and design an evaluation metric, SketchBLEU. For feedback-based evaluation, we develop a VSCode plugin for CodeS and engage 30 participants in conducting empirical studies. Extensive experiments prove the effectiveness and practicality of CodeS on the NL2Repo task.
Vertical federated learning (VFL) enables multiple parties to collaboratively train a model while preserving privacy. However, recent studies have raised concerns about the susceptibility of VFL models, including those using logistic regression and neural networks, to feature inference attacks. Meanwhile, the non-differentiable characteristics of decision tree ensembles make conducting such attacks impractical. To address this challenge, we introduce a feature inference attack framework FIA-TE tailored for decision tree ensembles, including gradient boosted decision trees (GBDT) and random forest. Specifically, we distill the knowledge from trees into neural networks by leaf embedding and structure distillation to create a targeted model for the inference attack. We then employ a generative model based on the deconvolutional network for capturing correlation features and reconstructing the target features. Through extensive experiments on table and image data, we evaluate the effectiveness of our framework and provide an analysis of potential influencing factors.
GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs), but has so far only focused on Python version. However, supporting more programming languages is also important, as there is a strong demand in industry. As a first step toward multilingual support, we have developed a Java version of SWE-bench, called SWE-bench-java. We have publicly released the dataset, along with the corresponding Docker-based evaluation environment and leaderboard, which will be continuously maintained and updated in the coming months. To verify the reliability of SWE-bench-java, we implement a classic method SWE-agent and test several powerful LLMs on it. As is well known, developing a high-quality multi-lingual benchmark is time-consuming and labor-intensive, so we welcome contributions through pull requests or collaboration to accelerate its iteration and refinement, paving the way for fully automated programming.
Soil datasets with outliers lead to inaccurate farm-level digital soil mapping (DSM) results. Existing methods identify potential outliers in soil datasets based on expert experience or simple statistics that neglect the geographical characteristics of soil. In this paper, a novel potential outlier recognition method was developed from the perspective of geographical context. First, spatial search distance was automatically determined by the spatial distance among soil samples. Second, similarities of adjacent soil samples and the local spatial variation level were comprehensively considered to calculate outlier scores. Finally, a frequency histogram of outlier scores was generated to determine a suitable threshold for recognizing potential abnormal samples. To validate the proposed method, it was compared to Lambda and Box-Plot methods, and the ordinary kriging method was used to map five soil properties, including pH, soil organic matter, total nitrogen, available phosphorus and available potassium, in an agricultural region. Then, a synthetic study using artificially contaminated DEM data was also conducted. The comparative experiment shows that the proposed method is better able to recognize potential outliers by mining the local spatial structure, as indicated by lower mean absolute error (MAE) and root mean square error (RMSE) values. It can be concluded that consideration of local spatial autocorrelation and heterogeneity is helpful in recognizing potential outliers.
When human programmers have mastered a programming language, it would be easier when they learn a new programming language. In this report, we focus on exploring whether programming languages can boost each other during the instruction fine-tuning phase of code large language models. We conduct extensive experiments of 8 popular programming languages (Python, JavaScript, TypeScript, C, C++, Java, Go, HTML) on StarCoder. Results demonstrate that programming languages can significantly improve each other. For example, CodeM-Python 15B trained on Python is able to increase Java by an absolute 17.95% pass@1 on HumanEval-X. More surprisingly, we found that CodeM-HTML 7B trained on the HTML corpus can improve Java by an absolute 15.24% pass@1. Our training data is released at https://github.com/NL2Code/CodeM.
The task of generating code from a natural language description, or NL2Code, is considered a pressing and significant challenge in code intelligence. Thanks to the rapid development of pre-training techniques, surging large language models are being proposed for code, sparking the advances in NL2Code. To facilitate further research and applications in this field, in this paper, we present a comprehensive survey of 27 existing large language models for NL2Code, and also review benchmarks and metrics. We provide an intuitive comparison of all existing models on the HumanEval benchmark. Through in-depth observation and analysis, we provide some insights and conclude that the key factors contributing to the success of large language models for NL2Code are "Large Size, Premium Data, Expert Tuning". In addition, we discuss challenges and opportunities regarding the gap between models and humans. We also create a website https://nl2code.github.io to track the latest progress through crowd-sourcing. To the best of our knowledge, this is the first survey of large language models for NL2Code, and we believe it will contribute to the ongoing development of the field.
Integrating social relations into recommendation is an effective way to mitigate data sparsity. Most social recommendation methods encode user representations from a unified graph that includes user-user and user-item relations. Due to the enriched relations on this graph, a large fraction of users are aware of each other within only a few hops, and the user representations generated by existing methods may encode the information received from a large number of neighbors. Thus, many user representations are enforced to be too similar, which hinders modeling fine-grained user interest. Here, we name this phenomenon as user representation collapse. To address this problem, in this paper we propose a robust user representation learning method named RobustSR with social regularization and multi-view contrastive learning, which aim to enhance the model’s awareness of relation informativeness and the discriminativeness of user representations, respectively. Concretely, the social regularization mechanism encourages the model to learn from the relation importance weights derived from graph topologies, which helps recognize important observed relations meanwhile mining potential useful relations. To enhance the discriminativeness of user representations, we further perform multi-view contrastive learning between collaborative and social-enhanced user representations. Extensive experiments on four benchmark datasets show that RobustSR effectively alleviates user representation collapse and improves recommendation performance. Our code is deposited at https://github.com/paulpig/RobustSR.
Sampling design plays a critical role in farm-level digital soil mapping (DSM). In many cases, a soil mapping model may not have been decided upon at the sample design stage. Design-based sampling may be more appropriate than model-based sampling because it is independent of subsequent soil mapping models. However, existing sampling methods optimize the sample size and locations in geographical space or feature space without considering the impacts of environmental similarity in local geographical space. In this paper, a novel sampling design method based on local environmental similarity was developed. Image segmentation was introduced into the sampling design by partitioning agricultural soil into subregions with good spatial continuity, within-region homogeneity, and between-region heterogeneity to determine the optimal sample size and locations. First, the environmental similarity between adjacent soils was calculated. Second, the merging process was iteratively conducted, and a series of segmentations was generated. Finally, the optimal sample size and locations were determined based on the optimal segmentation results. To validate the proposed method, it was compared with stratified random sampling, k -means sampling, and spatially balanced sampling methods. Two mapping models, ordinary kriging and sandwich estimation, were employed to map five soil properties, including pH, soil organic matter, total nitrogen, available phosphorus, and available potassium. These comparative experiments showed that the proposed method had better potential to generate farm-level muti-soil property mapping results with good accuracy than the competing sampling methods. In conclusion, consideration of local environmental similarity and the use of image segmentation for soil sampling were helpful in determining the optimal sample size and key sample locations.
Based on the digital footprint data, exploring the differences in tourist market structure and driving factors before and after COVID-19 is important for identifying tourist market demand and optimizing tourism product supply in the post-pandemic era. Most of the existing studies have explored the impact of the pandemic on the tourist market in well-known or large cities and have provided suggestions for tourism recovery. However, these suggestions are not entirely applicable to smaller cities. Small cities have a single level of tourism product, high homogeneity of tourism resources, small tourist market scale, and high volatility of the tourism industry. Therefore, it is necessary to study the differences in the tourist market structure of small cities and its driving factors before and after the pandemic and to propose targeted measures for the tourism recovery in the post-pandemic period. This paper, taking small cities as the study area and using online travel diaries as the data source, analyzed the differences in the spatial and temporal structures of tourist markets and their driving factors in Dengfeng and Kaifeng, China, before and after the pandemic. Then, countermeasures for tourism industry recovery in the post-pandemic era were proposed. The results were as follows: the difference in the tourism off-peak season increased after the pandemic, and the concentration of tourist market spatial distribution in Dengfeng showed a decreasing trend while that in Kaifeng showed an increasing trend. In addition to region traffic, the driving effects of leisure time, climate comfort and residents' income level weakened after the outbreak. Dengfeng and Kaifeng can enhance the tourist market tendency and attractiveness by creating special indoor tourism projects, strengthening tourism product promotion and marketing and enhancing the facilities related to self-driving tours.
To adapt to the scenario characteristics of vertical federated learning (VFL) applications regarding high communication cost, fast model iteration, and decentralized data storage, a generalized adversarial sample generation algorithm named VFL-GASG was proposed.Specifically, an adversarial sample generation framework was constructed for the VFL architecture.A white-box adversarial attack in the VFL was implemented by extending the centralized machine learning adversarial sample generation algorithm with different policies such as L-BFGS, FGSM, and C&W.By introducing deep convolutional generative adversarial network (DCGAN), an adversarial sample generation algorithm named VFL-GASG was designed to address the problem of universality in the generation of adversarial perturbations.Hidden layer vectors were utilized as local prior knowledge to train the adversarial perturbation generation model, and through a series of convolution-deconvolution network layers, finely crafted adversarial perturbations were produced.Experiments show that VFL-GASG can maintain a high attack success while achieving a higher generation efficiency, robustness, and generalization ability than the baseline algorithm, and further verify the impact of relevant settings for adversarial attacks.