Financial time series are noisy, non-stationary, and volatile, which makes accurate forecasting particularly challenging. State-of-the-art time series foundation models achieve strong performance on benchmark problems with clear seasonality or trend, such as energy demand or traffic forecasting, yet struggle to produce reliable forecasts in low signal-to-noise environments such as financial markets. To address this limitation, we introduce a two-stage learning architecture that reformulates short-term trend forecasting as a selective classification problem. The first stage consists of a primary predictive model that outputs directional forecasts (up/down), while the second stage introduces a learned reliability estimator that models the error characteristics of the primary model. This explicit decoupling of prediction and reliability estimation enables the system to learn when the primary model is likely to produce an incorrect forecast. The proposed architecture supports selective inference, where decisions are executed only for instances classified as reliable by the second stage. We formalize this mechanism within a risk–coverage optimization framework, which provides explicit control over the trade-off between prediction precision and execution frequency. The proposed approach is model-agnostic and can be seamlessly integrated with modern neural forecasting models, including time series foundation models. Extensive experiments on cryptocurrencies and equity markets demonstrate consistent improvements in precision (+14% on average) and risk reduction (−14%) compared to non-selective baselines. Historical backtests with a simple trading strategy confirm the practical viability of the approach.
This paper proposes and evaluates a token-based incentive and coordination mechanism for decentralized Advanced Air Mobility (AAM) systems. Building on prior work on privacy-preserving Unmanned Aerial Vehicle (UAV) flight path verification, we introduce a tokenomic framework aligned with Requirements Engineering (RE) principles. The system incentivizes compliance, penalizes misbehavior, and leverages economic design to fulfill stakeholder requirements such as safety, transparency, efficiency, fairness, scalability, and resilience. A simulation environment allows the validation of system dynamics under varying assumptions about UAV operator behavior, flight slot demand, and token issuance models. We position our framework within the context of socio-technical RE, demonstrating how tokenomic mechanisms can encode and enforce stakeholder values within an AAM system. The simulator used in this study is publicly available at http://52.68.19.251/tokenomic-simulator to support reproducibility.
In this article, we explore the potential of zero-shot Large Multimodal Models (LMMs) in the domain of drone perception. We focus on person detection and action recognition tasks and evaluate two prominent LMMs, namely YOLO-World and GPT-4V(ision) using a publicly available dataset captured from aerial views. Traditional deep learning approaches rely heavily on large and high-quality training datasets. However, in certain robotic settings, acquiring such datasets can be resource-intensive or impractical within a reasonable timeframe. The flexibility of prompt-based Large Multimodal Models (LMMs) and their exceptional generalization capabilities have the potential to revolutionize robotics applications in these scenarios. Our findings suggest that YOLO-World demonstrates good detection performance. GPT-4V struggles with accurately classifying action classes but delivers promising results in filtering out unwanted region proposals and in providing a general description of the scenery. This research represents an initial step in leveraging LMMs for drone perception and establishes a foundation for future investigations in this area.
Data from different modalities, such as infrared and visible images, can offer complementary information, and integrating such information can significantly enhance the capabilities of a system to perceive and recognize its surroundings. Thus, multi-modal object detection has widespread applications, particularly in challenging weather conditions like low-light scenarios. The core of multi-modal fusion lies in developing a reasonable fusion strategy, which can fully exploit the complementary features of different modalities while preventing a significant increase in model complexity. To this end, this paper proposes a novel lightweight cross-fusion module named Channel-Patch Cross Fusion (CPCF), which leverages Channel-wise Cross-Attention (CCA), Patch-wise Cross-Attention (PCA) and Adaptive Gating (AG) to encourage mutual rectification among different modalities. This process simultaneously explores commonalities across modalities while maintaining the uniqueness of each modality. Furthermore, we design a versatile intermediate fusion framework that can leverage CPCF to enhance the performance of multi-modal object detection. The proposed method is extensively evaluated on multiple public multi-modal datasets, namely FLIR, LLVIP, and DroneVehicle. The experiments indicate that our method yields consistent performance gains across various benchmarks and can be extended to different types of detectors, further demonstrating its robustness and generalizability. Our codes are available at https://github.com/Superjie13/CPCF_Multispectral.
This article introduces the DeepDream approach for object detection, which allows us to visualize how objects are represented in single stage object detector networks like YOLO. Such networks work by predicting objects for thousands of fixed image positions in parallel which makes them even more opaque compared to classification CNNs. While there has been much work on feature visualization for classification, this study examines how visualization methods can deal with the multitude of possible object positions in detection tasks and investigates the necessary adaptions of the DeepDream method. Our experiments suggest that YOLO detects objects relative to the scene composition. YOLO does not only recognize single objects but it also has a clear representation of scene context, object sub-types, positions, and orientations. We visualize our findings with interactive, web-based demo applications, which are available on our webpage: https://limchr.github.io/yolo dreaming/. This research broadens the understanding of how objects are represented in object detection networks.
Time Series forecasting has been approached by a multiplicity of techniques including deep learning methods of various degrees of sophistication, showcasing notable advancements and improved performance over the past few years. More recently, there has been a sustained interest in the study of Transformers, a class of models renowned for their remarkable capacity to capture intricate long-range dependencies and interactions. This ability is perceived as particularly relevant and impactful in the context of time series modeling, reflecting a growing recognition of their potential in enhancing forecasting accuracy and understanding of complex temporal patterns. However, taking advantage of this principle to deploy successful forecasting methods is not yet clearly understood, and requires significant experimentation or engineering. Therefore, in this paper, we compare multiple variations of the Transformer model (standard Transformer, Autoformer, Informer), coupled with diverse combinations of embedding data. In particular, as the emphasis of our work is on forecasting, we investigate the relationship between Transformers’ input segment length and prediction performance in a multi-step time intervals framework. Our results suggest that the Autoformer outperforms both standard Transformer and Informer across various prediction steps. We also observe that shorter input lengths and shorter prediction lengths generally produce better model performance.
In engineering, prognostics can be defined as the estimation of the remaining useful life of a system given current and past health conditions. This field has drawn attention from research, industry, and government as this kind of technology can help improve efficiency and lower the costs of maintenance in a variety of technical applications. An approach to prognostics that has gained increasing attention is the use of data-driven methods. These methods typically use pattern recognition and machine learning to estimate the residual life of equipment based on historical data. Despite their promising results, a major disadvantage is that it is difficult to interpret this kind of methodologies, that is, to understand why a certain prediction of remaining useful life was made at a certain point in time. Nevertheless, the interpretability of these models could facilitate the use of data-driven prognostics in different domains such as aeronautics, manufacturing, and energy, areas where certification is critical. To help address this issue, we use Local Interpretable Model-agnostic Explanations (LIME) from the field of eXplainable Artificial Intelligence (XAI) to analyze the prognostics of a Gated Recurrent Unit (GRU) on the C-MAPSS data. We select the GRU as this is a deep learning model that a) has an explicit temporal dimension and b) has shown promising results in the field of prognostics and c) is of simplified nature compared to other recurrent networks. Our results suggest that it is possible to infer the feature importance for the GRU both globally (for the entire model) and locally (for a given RUL prediction) with LIME.
Prognostics and health management is an engineering discipline that aims to support system operation while ensuring maximum safety and performance. Prognostics is a key step of this framework, focusing on developing effective maintenance policies based on predictive methods. Traditionally, prognostics models forecast the degradation process using regression techniques that approximate a mapping function from input to continuous remaining useful life estimates. These models are typically of high complexity and low interpretability. Classification approaches are an alternative solution to these types of models. We propose a predictive classification model that translates the input into discrete output variables instead of mapping the input to a single remaining useful life estimate. Each discrete output variable corresponds to a range of remaining useful life values. In other words, each output class variable represents the likelihood or risk of failure within a specific time range. We apply this model to a real-world case study involving the unscheduled and scheduled removals of a set of engine bleed valves from a fleet of Boeing 737 aircraft. The model can reach an area under the (micro-average) receiver operating characteristic curve of 72%. Our results suggest that the proposed multiclass gated recurrent unit network can provide valuable information about the different fault stages (corresponding to intervals of residual lives) of the studied valves.
This article describes the artificial intelligence (AI) component of a drone for monitoring and patrolling tasks associated with disaster relief missions in specific restricted disaster scenarios, as specified by the Advanced Robotics Foundation in Japan. The AI component uses deep learning models for environment recognition and object detection. For environment recognition, we use semantic segmentation, or pixel-wise labeling, based on RGB images. Object detection is key for detecting and locating people in need. Since people are relatively small objects from the drone perspective, we use both RGB and thermal images. To train our models, we created a novel multispectral and publicly available data set of people. We used a geo-location method to locate people on the ground. The semantic segmentation models were extensively tested using different feature extractors. We created two dedicated data sets, which we have made publicly available. Compared with the baseline model, the best-performing model could increase the mean intersection over union (IoU) by 1.3%. Furthermore, we compared two types of person detection models. The first one is an ensemble model that combines RGB and thermal information via "late fusion"; the second one is a 4-channel model that combines these two types of information in an "early fusion" manner. The results suggest that the 4-channel model had a 40.6% increase of average precision for stricter IoU values (0.75) compared with the ensemble model and a 5.8% increase in the average precision compared with the thermal model. All models were deployed and tested on the NVIDIA AGX Xavier platform. To the best of our knowledge, this study was the first to use both RGB and thermal data from the perspective of a drone for monitoring tasks.
The development of a real-world Unmanned Aircraft System (UAS) Traffic Management (UTM) system to ensure the safe integration of Unmanned Aerial Vehicles (UAVs) in low altitude airspace, has recently generated novel research challenges. A key problem is the development of Pre-Flight Conflict Detection and Resolution (CDR) methods that provide collision-free flight paths to all UAVs before their takeoff. Such problem can be represented as a Multi-Agent Path Finding (MAPF) problem. Currently, most MAPF methods assume that the UTM system is a centralized entity in charge of CDR. However, recent discussions on UTM suggest that such centralized control might not be practical or desirable. Therefore, we explore Pre-Flight CDR methods where independent UAS Service Providers (UASSPs) with their own interests, communicate with each other to resolve conflicts among their UAV operations—without centralized UTM directives. We propose a novel MAPF model that supports the decentralized resolution of conflicts, whereby different ‘agents’, here UASSPs, manage their UAV operations. We present two approaches: (1) a prioritization approach and (2) a simple yet practical pairwise negotiation approach where UASSPs agents determine an agreement to solve conflicts between their UAV operations. We evaluate the performance of our proposed approaches with simulation scenarios based on a consultancy study of predicted UAV traffic for delivery services in Sendai, Japan, 2030. We demonstrate that our negotiation approach improves the “fairness” between UASSPs, i.e. the distribution of costs between UASSPs in terms of total delays and rejected operations due to replanning is more balanced when compared to the prioritization approach.
Human action recognition and detection from unmanned aerial vehicles (UAVs), or drones, has emerged as a popular technical challenge in recent years, since it is related to many use case scenarios from environmental monitoring to search and rescue. It faces a number of difficulties mainly due to image acquisition and contents, and processing constraints. Since drones’ flying conditions constrain image acquisition, human subjects may appear in images at variable scales, orientations, and occlusion, which makes action recognition more difficult. We explore low-resource methods for ML (machine learning)-based action recognition using a previously collected real-world dataset (the “Okutama-Action” dataset). This dataset contains representative situations for action recognition, yet is controlled for image acquisition parameters such as camera angle or flight altitude. We investigate a combination of object recognition and classifier techniques to support single-image action identification. Our architecture integrates YoloV5 with a gradient boosting classifier; the rationale is to use a scalable and efficient object recognition system coupled with a classifier that is able to incorporate samples of variable difficulty. In an ablation study, we test different architectures of YoloV5 and evaluate the performance of our method on Okutama-Action dataset. Our approach outperformed previous architectures applied to the Okutama dataset, which differed by their object identification and classification pipeline: we hypothesize that this is a consequence of both YoloV5 performance and the overall adequacy of our pipeline to the specificities of the Okutama dataset in terms of bias–variance tradeoff.
In this paper, we present our approach for drone detection which we submitted for the Drone-Vs-Bird Detection Challenge. In our work, we used the Fully Convolutional One-Stage Object Detection (FCOS) approach tuned to detect drones. Throughout our experiments, we opted for a simple data augmentation technique to reduce the amount of False Positives (FPs). Upon observing the results of our early experiments, our technique for data augmentation incorporates adding extra samples to the training sets including the object which generated the most number of FPs, namely other flying objects, leaves and objects with sharp edges. With the newly introduced data to the training set, our results for drone detection on the validation set are as follows: AP scores of 0.16, 0.34 and 0.65 for small-sized, medium-sized and large drones respectively.
This paper reports the results of the 5th edition of the "Drone-vs-Bird" detection challenge, organized within the 21st International Conference on Image Analysis and Processing (ICIAP). By taking as input video samples recorded by common cameras, the aim of the challenge is to devise advanced approaches aimed at spotlighting the presence of drones flying in the monitored area, while limiting the number of wrong alarms raised when similar flying entities such as birds suddenly appear in the scene. To this end, a number of important issues such as the dynamic variations in the scene and the background/foreground motion effects should be carefully considered, so as to allow the proposed solutions to correctly identify drones only when they are actually present. The paper summarizes the novel algorithms proposed by the four participating teams that succeeded in providing satisfactory detection performance on the 2022 challenge dataset.
We describe iCO2, a simulation platform for collecting driving behavior data. It is designed as the first massively multiplayer online game for mobile devices to practice eco-friendly driving. It facilitates the collection of large-scale data on driving behavior to better understand compliance and incentive mechanisms for eco-driving and users' preferences. We present the results of a campaign with iCO2 that used a game promoter to attract 2455 users. The results are described from three angles: (1) types of drivers are identified by clustering driving behavior; (2) types of players are identified by relating players' interaction with game elements and their driving behavior; (3) by looking at longer sessions, we demonstrate that players who show eco-unfriendly behavior at the beginning of the session improve their eco-driving behavior throughout their playtime.
Unmanned aerial vehicles (UAVs) operate with large degrees of autonomy when executing tasks from a ground-based controller. In this work, we design an (immersive) interface for drone management. Our interface supports operators in mission execution of one or several (automated) UAVs within a fleet. This allows operators to assign tasks to UAVs, while safe operations are ensured through in-flight conflict detection and prevention. Our interface fulfills both functional (e.g., reliability) and non-functional requirements (e.g., perceived cognitive load during use). A key design challenge is the latency in information exchanges, which introduces information asymmetries between the humans and UAVs. We then validate the effectiveness of our interface design across several experiments. We find that our design reduces the perceived cognitive load and improves task performance. Altogether, this work has implications for designing interfaces in human-machine collaboration, so humans can effectively control, interact, or collaborate with automated machines, such as UAVs.
The development of an unmanned aircraft system (UAS) traffic management (UTM) system for the safe integration of unmanned aerial vehicles (UAVs) requires pre-flight conflict detection and resolution (CDR) methods to provide collision-free flight paths for all UAVs before takeoff. A popular solution consists in adapting multi-agent path finding (MAPF) techniques. However, standard MAPF solvers consider only a fixed takeoff time and fixed uniform speed for each UAV flight path, which can lead to inefficiencies in the resolution of instances. Therefore, in this article, we propose incorporating scheduling elements into MAPF solvers, which allows us to adjust the takeoff times and speeds of each UAV to solve conflicts. We introduce two time-related resolution techniques: 1) takeoff scheduling, whereby the start time of a UAV agent is delayed, and 2) speed adjustment, wherein the speed of a UAV agent is decreased over a segment on its flight path. Importantly, we present a distinction of conflict types, which enables us to combine replanning resolution to the aforementioned temporal resolution techniques. We evaluate our proposed approaches on a realistic, high-density UAV delivery scenario in Tokyo, Japan. We show that the combination of takeoff scheduling, replanning, and speed-adjustment resolution techniques improves the efficiency of route planning by reducing the average delay per flight path and the number of rejected flight paths.
Traditionally, prognostics approaches to predictive maintenance have focused on estimating the remaining useful life of the equipment. However, from an industrial point of view, the goal is often not to predict the residual life but to determine the need for a maintenance action at a given time window. This approach allows us to frame the data-driven prognostics problem as a binary classification task rather than a regression one. To address this problem, we propose in this paper to explore the relative strengths and limitations of a set of classifier approaches such as random forests, support vector machines, nearest neighbors, and deep learning techniques. We evaluate the models using metrics such as sensitivity, specificity, accuracy, receiver operating characteristic curve, and F-score. This work's novelty lies in adopting a modeling approach with a natural probabilistic interpretation of the prognostics exercise. The comparison of an extensive range of classifier models is performed on two real-world datasets from the aeronautics sector. Results indicate that deep learning classifier methods are well suited for this kind of prognostics and can outperform by a significant margin the traditional classification techniques. Importantly, the proposed modeling approach aims to generate an alternative prognostics representation that goes in line with the expectations of aeronautical engineers.
Structural Health Monitoring (SHM) has greatly benefited from computer vision. Recently, deep learning approaches are widely used to accurately estimate the state of deterioration of infrastructure. In this work, we focus on the problem of bridge surface structural damage detection, such as delamination and rebar exposure. It is well known that the quality of a deep learning model is highly dependent on the quality of the training dataset. Bridge damage detection, our application domain, has the following main challenges: (i) labeling the damages requires knowledgeable civil engineering professionals, which makes it difficult to collect a large annotated dataset; (ii) the damage area could be very small, whereas the background area is large, which creates an unbalanced training environment; (iii) due to the difficulty to exactly determine the extension of the damage, there is often a variation among different labelers who perform pixel-wise labeling. In this paper, we propose a novel model for bridge structural damage detection to address the first two challenges. This paper follows the idea of an atrous spatial pyramid pooling (ASPP) module that is designed as a novel network for bridge damage detection. Further, we introduce the weight balanced Intersection over Union (IoU) loss function to achieve accurate segmentation on a highly unbalanced small dataset. The experimental results show that (i) the IoU loss function improves the overall performance of damage detection, as compared to cross entropy loss or focal loss, and (ii) the proposed model has a better ability to detect a minority class than other light segmentation networks.
Recurrent neural networks (RNNs) such as LSTM and GRU are not new to the field of prognostics. However, the performance of neural networks strongly depends on their architectural structure. In this work, we investigate a hybrid network architecture that is a combination of recurrent and feed-forward (conditional) layers. Two networks, one recurrent and another feed-forward, are chained together, with inference and weight gradients being learned using the standard back-propagation learning procedure. To better tune the network, instead of using raw sensor data, we do some preprocessing on the data, using mostly simple but effective statistics (researched in previous work). This helps the feature extraction phase and eases the problem of finding a suitable network configuration among the immense set of possible ones. This is not the first proposal of a hybrid network in prognostics but our work is novel in the sense that it performs a more comprehensive comparison of this type of architecture for different RNN layers and number of layers. Also, we compare our work with other classical machine learning methods. Evaluation is performed on two real-world case studies from the aero-engine industry: one involving a critical valve subsystem of the jet engine and another the whole reliability of the jet engine. Our goal here is to compare two cases contrasting micro (valve) and macro (whole engine) prognostics. Our results indicate that the performance of the LSTM and GRU deep networks are significantly better than that of other models.
Kugamoorthy Gajananan合作论文数National Institute of Informatics, The Graduate University for Advanced Studies (SOKENDAI)11
Pedro Santos合作论文数Instituto Superior Tecnico;Departamento de Matematica5