Natural and semi-natural soils are essential benchmarks for carbon accounting and restoration planning, yet they remain critically under-represented in regional and European soil datasets, particularly in intensively used agricultural settings of Southeast Europe. In the Pannonian Basin, existing soil monitoring systems focus predominantly on arable land, creating a major data gap for natural and semi-natural reference states required for climate mitigation reporting and digital soil mapping.This study addresses this gap by establishing the first natural-soil reference system for remnant natural and semi-natural ecosystems of Vojvodina (Serbia), where less than 10% of original steppe-forest-wetland habitats persist. Using a spatially explicit, machine-learning-assisted sampling framework, we identified 62 representative forest and grassland locations and collected 186 soil samples following harmonized international protocols. Soil Organic Carbon (SOC) stocks were quantified and analyzed across spatial clusters. Average SOC stocks were higher in grasslands than in forests, with variation driven primarily by soil texture, aridity gradients, and land-cover type. Comparison with historical datasets (1950-1960) revealed a consistent SOC decline across all clusters. Machine-learning models achieved moderate predictive performance for continuous SOC estimation, while classification into SOC categories showed stronger performance. Results showed significant mismatches between global SOC datasets and locally derived measurements, underscoring the necessity of regionally calibrated reference systems. The proposed spatially stratified sampling methodology ensures proportional representation of environmental strata and provides a scalable framework for soil monitoring in heterogeneous landscapes, strengthening carbon accounting under land use frameworks, and supports EU Soil Mission 2030 and nature-based climate solutions.
Remote sensing (RS) has evolved from occasional mapping to continuous, indicator-based monitoring of terrestrial ecosystems. This review synthesizes four decades of global progress in RS to characterize natural and semi-natural ecosystems, examining how study purposes, sensor types and analytical methods have diversified from 1985 to 2025. A systematic literature review of 6856 publications (1567 selected) documents the transition from expert-based visual interpretation using aerial photography and early Landsat missions, to harmonized, AI-driven workflows that enable scalable and replicable ecosystem assessments. Advances in cloud computing, data cubes and open-access archives now allow wall-to-wall time series of analyses across regions and biomes. Yet, important challenges persist, including the underrepresentation of biodiversity-rich areas, limited in-situ calibration data and uncertainties related to phenological variability, image correction, or temporal mosaicking pipelines. Building on case studies from a global perspective, we outline design principles for policy-ready ecosystem indicators traceable to raw observations, comparable through time and space, and aligned with biodiversity policy frameworks. Integrating multi-sensor data (optical, radar, LiDAR, thermal), standardized in-situ observations and artificial intelligence/machine learning algorithms, RS provides a robust pathway towards operational ecosystem accounting and large-scale functional mapping and monitoring, strengthening conservation planning and ecosystem management worldwide.
The paper presents design and prototype implementation of an edge based object detection system within the new paradigm of AI agents orchestration. It goes beyond traditional design approaches by leveraging on LLM based natural language interface for system control and communication and practically demonstrates integration of all system components into a single resource constrained hardware platform. The method is based on the proposed multi-agent object detection framework which tightly integrates different AI agents within the same task of providing object detection and tracking capabilities. The proposed design principles highlight the fast prototyping approach that is characteristic for transformational potential of generative AI systems, which are applied during both development and implementation stages. Instead of specialized communication and control interface, the system is made by using Slack channel chatbot agent and accompanying Ollama LLM reporting agent, which are both run locally on the same Raspberry Pi platform, alongside the dedicated YOLO based computer vision agent performing real time object detection and tracking. Agent orchestration is implemented through a specially designed event based message exchange subsystem, which represents an alternative to completely autonomous agent orchestration and control characteristic for contemporary LLM based frameworks like the recently proposed OpenClaw. Conducted experimental investigation provides valuable insights into limitations of the low cost testbed platforms in the design of completely centralized multi-agent AI systems. The paper also discusses comparative differences between presented approach and the solution that would require additional cloud based external resources.
Recent advancements in the field of natural language processing (NLP) and especially large language models (LLMs) and their numerous applications have brought research attention to design of different document processing tools and enhancements in the process of document archiving, search and retrieval. Domain of official, legal documents is especially interesting due to vast amount of data generated on the daily basis, as well as the significant community of interested practitioners (lawyers, law offices, administrative workers, state institutions and citizens). Providing efficient ways for automation of everyday work involving legal documents is therefore expected to have significant impact in different fields. In this work we present one LLM based solution for Named Entity Recognition (NER) in the case of legal documents written in Serbian language. It leverages on the pre-trained bidirectional encoder representations from transformers (BERT), which had been carefully adapted to the specific task of identifying and classifying specific data points from textual content. Besides novel dataset development for Serbian language (involving public court rulings), presented system design and applied methodology, the paper also discusses achieved performance metrics and their implications for objective assessment of the proposed solution. Performed cross-validation tests on the created manually labeled dataset with mean F_1 score of 0.96 and additional results on the examples of intentionally modified text inputs confirm applicability of the proposed system design and robustness of the developed NER solution.
Large-scale habitat degradation and destruction have been marked as the most important causes of biodiversity loss in the last 100 years. To combat this problem, proper monitoring and, if deemed necessary, restoration measures, must be taken. The economic value of restoration actions in Europe alone has exceeded one billion euros but despite this fact, many projects have failed to restore ecosystem functions of targeted areas. The lack of proper monitoring and steering of restoration actions, i.e. translating scientific findings into practical, easy-to-do actions for practitioners is probably one of the causes for this failure. „Sunčani salaš“ eLTER site (https://deims.org/5f5c850d-0036-49ac-97be-f9b314898607) was established in 2021 to monitor the progress of the revitalization of this former arable land parcel into the original sandy grassland habitat. The site is located within the Subotica sands protected landscape in northern Serbia, along the Hungarian border. Prior to the field campaign, we divided the study area into zones based on visual differences inferred from drone imagery. We performed a classical plot-based botanical survey and used hand-held GPS to determine the spatial distribution of finely differentiated vegetation types. Based on fieldwork data, we classified Sunčani salaš vegetation into appropriate EUNIS (European nature information system) classes and assessed its conservation status. These classes were then translated into site-specific key habitat types for site management purposes. In parallel, we used a commercially available and easy to operate DJI Inspire UAV (unmanned aerial vehicle), equipped with a Zenmuse X3 RGB camera to assess whether the commercially available RGB sensor and a relatively high flight altitude (100m) of the UAV have discriminative capacity to aid site managers by mapping identified steppe development stages. Both campaigns were performed monthly throughout the vegetative season (April-September). Based on the presence of characteristic species, we identified four main habitat types according to the EUNIS classification: Pannonic loess steppe grassland (E1.2C1 and E1.2C2), Pannonic sandy steppes (E1.2F4), Broadleaved deciduous woodland (G1.4) and Bare tilled land (I1.51). These were then translated into site-specific categories, termed: Young steppe I, Young steppe II, Forest steppe and Fallow land, respectively, which were all quantified across the habitat-specific zones. For image classification purposes, these categories were translated into categories: Class C0 or “Steppe”, encompassing Young steppe I and II, Class C1 or “Shrubs”, Class C2 or “Forest–Steppe”, Class C3 or “Bare/fallow land”. Of all detected habitat types, Pannonic loess steppe grassland (E1.2C1 and E1.2C2) and Pannonic sandy steppes (E1.2F4) are of conservation importance and are listed in the Annex I of the Habitat Directive. UAV vegetation maps show that the estimated extent of steppe habitat, class C0, dominates in most of the identified observation zones. In centrally positioned zones, steppe cover varies between 58% and 68% between the zones. Shrub cover across these zones is low (<5%), while the cover of bare soil is even lower, and stays under 2%. The extent of the western-marginal zone characterizes the presence of steppe cover C0 with high percentages of shrubs C1 (47% and 18%, respectively). The extent of southern-marginal zone characterizes the presence of forest C2 area range of approximately 60%. Both these zones are characterized by a significant percentage of pixels with a low confidence score. We consider 60% of the open habitat to be characterized by young steppe vegetation. According to the results of the UAV mapping, the extent of the young steppe is the largest in centrally positioned zones, where it reaches almost 70%. Results of the UAV campaign show that the proposed flight characteristics allow for more generalized classification of habitat types, but fine scale classification was not possible (e.g., it was not possible to distinguish Pannonic loess steppe from Pannonic sandy steppe). For this, higher spectral resolutions instead of RGB images would ease the classification problem and probably enable solitary UAV acquisitions, instead of repeated ones. Nevertheless, protected area managers can benefit from the results of the study, offering cost-effective and efficient tool for assessing habitat diversity and detecting broader ecological trends within protected areas.
Face detection and face recognition have been in the focus of vision community since the very beginnings. Inspired by the success of the original Videoface digitizer, a pioneering device that allowed users to capture video signals from any source, we have designed an advanced video analytics tool to efficiently create structured video stories, i.e. identity-based information catalogs. VideoFace2.0 is the name of the developed system for spatial and temporal localization of each unique face in the input video, i.e. face re-identification (ReID), which also allows their cataloging, characterization and creation of structured video outputs for later downstream tasks. Developed near real-time solution is primarily designed to be utilized in application scenarios involving TV production, media analysis, and as an efficient tool for creating large video datasets necessary for training machine learning (ML) models in challenging vision tasks such as lip reading and multimodal speech recognition. Conducted experiments confirm applicability of the proposed face ReID algorithm that is combining the concepts of face detection, face recognition and passive tracking-by-detection in order to achieve robust and efficient face ReID. The system is envisioned as a compact and modular extensions of the existing video production equipment. Presented results are based on test implementation that achieves between 18-25 fps on consumer type notebook. Ablation experiments also confirmed that the proposed algorithm brings relative gain in the reduction of number of false identities in the range of 73%-93%.
Due to the large-scale disappearance of grasslands there is an urgent need for revitalization. It calls for consistent and accessible monitoring and mapping plans, and an integrated management approach. However, revitalization efforts often focus solely on the vegetation component, and skip the link to other animal species that perform vital functions as ecosystem engineers and umbrella species. In this study, we combine an in-situ standard phytocoenological survey with an UAV-based technology in the effort to improve the monitoring and mapping of the sandy steppe habitat of the European ground squirrel (Spermophilus citellus; EGS), undergoing revitalization in the northern Serbia. It is a model organism of an animal species that enables identifying habitat quality and quantity indicators to understand the broader implications of the ecosystem revitalization efforts on the wildlife populations. The proposed approach tested whether the commercially available RGB sensor and a relatively high flight height of the UAV have discriminative capacity to aid site managers by mapping identified steppe development stages (specific plant assemblages, reflecting different habitat types). Thus, a novel set of high-resolution image descriptors that are capable of discriminating plant mixtures corresponding to Fallow land, Forest steppe and shrubs, Young steppe I and II, was proposed. Despite high resolution imaging, the method solves a challenging problem of UAV vegetation mapping in the case of limited spectral and spatial information in the image (by using only RGB camera and multitemporal approach). Although the lack of visual information that would allow identification of individual plant parts and shapes prevented the use of usual object-based image analysis, proposed pixel-based descriptors and feature selection were able to provide the extent of the targeted areas and their compositional carriers. Presented holistic approach enables implementation of effective management strategies that support the entire ecological community.
In this paper, a low-cost, Raspberry Pi based imaging system is proposed as compact standalone interrogation unit for analysis of fiber specklegram sensors. Standard methods for specklegram analysis are based on image correlation. Proposed imaging system is used for both capturing specklegram images at the output of the standard telecommunication optical fiber and for correlation analysis. Experimental setup for controlled mechanical deformation of the optical fiber is designed and zero-normalized cross-correlation, structural similarity and normalized mutual information score correlation methods are implemented and compared in order to verify proposed Raspberry Pi based imaging system functionally. A statistical method for detection of region of interest is used for dynamic output range extension. Additionally, to further extend dynamic range and increase linearity, correlation output is provided as difference of correlation coefficient for two reference samples located at the ends of measurement range.
A novel similarity measure between Gaussian mixture models (GMMs), based on similarities between the low-dimensional representations of individual GMM components and obtained using deep autoencoder architectures, is proposed in this paper. Two different approaches built upon these architectures are explored and utilized to obtain low-dimensional representations of Gaussian components in GMMs. The first approach relies on a classical autoencoder, utilizing the Euclidean norm cost function. Vectorized upper-diagonal symmetric positive definite (SPD) matrices corresponding to Gaussian components in particular GMMs are used as inputs to the autoencoder. Low-dimensional Euclidean vectors obtained from the autoencoder’s middle layer are then used to calculate distances among the original GMMs. The second approach relies on a deep convolutional neural network (CNN) autoencoder, using SPD representatives to generate embeddings corresponding to multivariate GMM components given as inputs. As the autoencoder training cost function, the Frobenious norm between the input and output layers of such network is used and combined with regularizer terms in the form of various pieces of information, as well as the Riemannian manifold-based distances between SPD representatives corresponding to the computed autoencoder feature maps. This is performed assuming that the underlying probability density functions (PDFs) of feature-map observations are multivariate Gaussians. By employing the proposed method, a significantly better trade-off between the recognition accuracy and the computational complexity is achieved when compared with other measures calculating distances among the SPD representatives of the original Gaussian components. The proposed method is much more efficient in machine learning tasks employing GMMs and operating on large datasets that require a large overall number of Gaussian components.
CycleGAN domain transfer architectures use cycle consistency loss mechanisms to enforce the bijectivity of highly underconstrained domain transfer mapping. In this paper, in order to further constrain the mapping problem and reinforce the cycle consistency between two domains, we also introduce a novel regularization method based on the alignment of feature maps probability distributions. This type of optimization constraint, expressed via an additional loss function, allows for further reducing the size of the regions that are mapped from the source domain into the same image in the target domain, which leads to mapping closer to the bijective and thus better performance. By selecting feature maps of the network layers with the same depth d in the encoder of the direct generative adversarial networks (GANs), and the decoder of the inverse GAN, it is possible to describe their d-dimensional probability distributions and, through novel regularization term, enforce similarity between representations of the same image in both domains during the mapping cycle. We introduce several ground distances between Gaussian distributions of the corresponding feature maps used in the regularization. In the experiments conducted on several real datasets, we achieved better performance in the unsupervised image transfer task in comparison to the baseline CycleGAN, and obtained results that were much closer to the fully supervised pix2pix method for all used datasets. The PSNR measure of the proposed method was, on average, 4.7% closer to the results of the pix2pix method in comparison to the baseline CycleGAN over all datasets. This also held for SSIM, where the described percentage was 8.3% on average over all datasets.
Motivated by the recent trends in the field of em-bedded vision platforms, we discuss potential of such solutions in providing foundations for the next generation of Cyber-Physical Systems (CPS). Improved capabilities and reduced price of these platforms will have profound effect on their everyday usage and applications. In comparison to speech and natural language processing, which have established speech recognition and machine translation applications as indispensable in many contemporary CPSs, the vision community is still searching for an application that would be so necessary and desirable to make most of the consumers buy specific vision hardware just to run it. That would be the ultimate proof of the core value of the technology in the market. Thus, also vision problems come with a longstanding tradition and history of numerous solutions, it is still hard to point out a single application that would incorporate many specific vision tasks into one device, and which would be ubiquitously useful and affordable to all (e.g. like smartphone has done in the fields of communication and personal computing). However, with development of new miniaturization technologies and spatial AI it is reasonable to expect that there will be more possibilities for designing CPS with capabilities of visual understanding of outdoor, dynamic and uncontrolled environments. One step in such direction are embedded vision platforms that besides powerful computing capabilities also provide multimodal perception, and thus improve the algorithm performance. As an example, we will discuss stereo depth perception in the context of new spatial AI platforms like OAK-D lite, and point out some possibilities for its improvement and integration into future CPS.
Most CycleGAN domain transfer architectures require a large amount of data belonging to domains on which the domain transfer task is to be applied. Nevertheless, in many real-world applications one of the domains is reduced, i.e., scarce. This means that it has much less training data available in comparison to the other domain, which is fully observable. In order to tackle the problem of using CycleGAN framework in such unfavorable application scenarios, we propose and invoke a novel Bootstrapped SSL CycleGAN architecture (BTS-SSL), where the mentioned problem is overcome using two strategies. Firstly, by using a relatively small percentage of available labelled training data from the reduced or scarce domain and a Semi-Supervised Learning (SSL) approach, we prevent overfitting of the discriminator belonging to the reduced domain, which would otherwise occur during initial training iterations due to the small amount of available training data in the scarce domain. Secondly, after initial learning guided by the described SSL strategy, additional bootstrapping (BTS) of the reduced data domain is performed by inserting artifically generated training examples into the training poll of the data discriminator belonging to the scarce domain. Bootstrapped samples are generated by the already trained neural network that performs transferring from the fully observable to the scarce domain. The described procedure is periodically repeated during the training process several times and results in significantly improved performance of the final model in comparison to the original unsupervised CycleGAN approach. The same also holds in comparison to the solutions that are exclusively based either on the described SSL, or on the bootstrapping strategy, i.e., when these are applied separately. Moreover, in the considered scarce scenarios it also shows competitive results in comparison to the fully supervised solution based on the pix2pix method. In that sense, it is directly applicable to many domain transfer tasks that are relying on the CycleGAN architecture.
Varietal classification of rice seeds is a crucial task in the process of rice crop production, management, and quality control. Traditionally, classification is performed manually which gives slow and inconsistent results. Machine vision technology provides an automated, real-time, non-destructive and cost-effective solution to this problem. Methods that combine RGB and hyperspectral imaging have shown very good results in rice seed classification. In this paper, we demonstrate the significance of morphological and border related features used in addition to spectral information and propose a feature set that provides a substantial improvement in classification results. The proposed approach was successfully tested on a publicly available dataset of 8640 seed samples corresponding to 90 different rice seed varieties, contained in 180 hyperspectral and RGB image pairs, and resulted in an average F1 score of 85.65%.
Remote sensing applications have gained in popularity in recent years, which has resulted in vast amounts of data being produced on a daily basis. Managing and delivering large sets of data becomes extremely difficult and resource demanding for the data vendors, but even more for individual users and third party stakeholders. Hence, research in the field of efficient remote sensing data handling and manipulation has become a very active research topic (from both storage and communication perspectives). Driven by the rapid growth in the volume of optical satellite measurements, in this work we explore the lossy compression technique for multispectral satellite images. We give a comprehensive analysis of the High Efficiency Video Coding (HEVC) still-image intra coding part applied to the multispectral image data. Thereafter, we analyze the impact of the distortions introduced by the HEVC’s intra compression in the general case, as well as in the specific context of crop classification application. Results show that HEVC’s intra coding achieves better trade-off between compression gain and image quality, as compared to standard JPEG 2000 solution. On the other hand, this also reflects in the better performance of the designed pixel-based classifier in the analyzed crop classification task. We show that HEVC can obtain up to 150:1 compression ratio, when observing compression in the context of specific application, without significantly losing on classification performance compared to classifier trained and applied on raw data. In comparison, in order to maintain the same performance, JPEG 2000 allows compression ratio up to 70:1.
Sparse representation of structured signals requires modelling strategies that maintain specific signal properties, in addition to preserving original information content and achieving simpler signal representation. Therefore, the major design challenge is to introduce adequate problem formulations and offer solutions that will efficiently lead to desired representations. In this context, sparse representation of covariance and precision matrices, which appear as feature descriptors or mixture model parameters, respectively, will be in the main focus of this paper.
The paper presents a novel precision matrix modeling technique for Gaussian Mixture Models (GMMs), which is based on the concept of sparse representation. Representation coefficients of each precision matrix (inverse covariance), as well as an accompanying overcomplete matrix dictionary, are learned by minimizing an appropriate functional, the first component of which corresponds to the sum of Kullback-Leibler (KL) divergences between the initial and the target GMM, and the second represents the sparse regularizer of the coefficients. Compared to the existing, alternative approaches for approximate GMM modeling, like popular subspace-based representation methods, the proposed model results in notably better trade-off between the representation error and the computational (memory) complexity. This is achieved under assumption that the training data in the recognition system utilizing GMM have an inherent sparseness property, which enables application of the proposed model and approximate representation using only one dictionary and a significantly smaller number of coefficients. Proposed model is experimentally compared with the Subspace Precision and Mean (SPAM) model, a state of the art instance of subspace-based representation models, using both the data from a real Automatic Speech Recognition (ASR) system, and specially designed sets of artificially created/synthetic data.
An approach to automatic hoverfly species discrimination based on detection and extraction of vein junctions in wing venation patterns of insects is presented in the paper. The dataset used in our experiments consists of high resolution microscopic wing images of several hoverfly species collected over a relatively long period of time at different geographic locations. Junctions are detected using the combination of the well known HOG (histograms of oriented gradients) and the robust version of recently proposed CLBP (complete local binary pattern). These features are used to train an SVM classifier to detect junctions in wing images. Once the junctions are identified they are used to extract statistics characterizing the constellations of these points. Such simple features can be used to automatically discriminate four selected hoverfly species with polynomial kernel SVM and achieve high classification accuracy.
A pixel-based cropland classification study based on the fusion of data from satellite images with different resolutions is presented. It is based on a time series of multispectral images acquired at different resolutions by different imaging instruments, Landsat-8 and RapidEye. The proposed data fusion method capabilities are explored with the aim of overcoming the shortcomings of different instruments in the particular cropland classification scenario characterized by the very small size of crop fields over the chosen agricultural region situated in the plains of Vojvodina in northern Serbia. This paper proposes a data fusion method that is successfully utilized in combination with arobust random forest classifier in improving the overall classification performance, as well as in enabling application of satellite imagery with a coarser spatial resolution in the given specific cropland classification task. The developed method effectively exploits available data and provides an improvement over the existing pixel-based classification approaches through the combination of different data sources. Another contribution of this paper is the employment of crowdsourcing in the process of reference data collection via dedicated smartphone application. (C) The Authors. Published by SPIE under a Creative Commons Attribution 3.0 Unported License. Distribution or reproduction of this work in whole or in part requires full attribution of the original publication, including its DOI.
Darko Stefanović合作论文数Department of Computer Science
University of New Mexico1