Recent deep learning-based remote sensing analysis models often struggle with performance degradation due to domain shifts caused by illumination variations (clear to overcast), changing atmospheric conditions (clear to foggy, dusty), and physical scene changes (clear to snowy). Addressing domain shift in aerial image segmentation is challenging due to limited training data availability, including costly data collection and annotation. We propose Multi-Weather DomainShifter, a comprehensive multi-weather domain transfer system that augments single-domain images into various weather conditions without additional laborious annotation, coordinated by a large language model (LLM) agent. Specifically, we utilize Unreal Engine to construct a synthetic dataset featuring images captured under diverse conditions such as overcast, foggy, and dusty settings. We then propose a latent space style transfer model that generates alternate domain versions based on real aerial datasets. Additionally, we present a multi-modal snowy scene diffusion model with LLM-assisted scene descriptors to add snowy elements into scenes. Multi-weather DomainShifter integrates these two approaches into a tool library and leverages the agent for tool selection and execution. Extensive experiments on the ISPRS Vaihingen and Potsdam dataset demonstrate that domain shift caused by weather change in aerial image-leads to significant performance drops, then verify our proposal’s capacity to adapt models to perform well in shifted domains while maintaining their effectiveness in the original domain.
Egocentric pose estimation is the technique for estimating the positions of all the joints of the human from images obtained from the camera attached to the body. In this paper we propose a method for egocentric pose estimation using images acquired from an omni-directional camera, which is attached on the human-body chest. The method consists of two steps; In the first step, the image acquired by omni-directional camera is converted into Cubemap, which is the set of the six images which capture front, back, left, right, top and bottom of view. The six images are mapped to the sub-body-parts which correspond to arms, lower, and torso, respectively. In the second step, the pose estimation is conducted for each sub-body-part. For estimating the pose of the lower, torso and upper, the Cubemap images which are assigned to sub-body-parts are input to and processed via the trained Convolutional Neural Network respectively. Finally, the prediction from each network is marginalized for estimating the pose of full body. To confirm the validity of our method, we made the dataset by under the simulation-environment which simulates the human-model based on human motion dataset[17]. Experimental results show that ours outperforms the baseline which processes equirectangular image by pre-trained large-scale CNN in most of many joints in each coordinate.
To enable the practical use of quadcopters equipped with microphone arrays in disaster sites for locating survivors’ voices, this paper proposes a comprehensive method for modeling and simulating complex acoustic environments using PyRoomAcoustics, and for locating sound sources using variants of the MUSIC algorithms. By comparing impulse responses in PyRoomAcoustics simulations with those in real environments, we observed a high degree of correlation, indicating the simulations’ suitability for real-world applications. Utilizing these simulations, we identified the optimal microphone array configuration for sound source localization (SSL) and examined the relationship between flight altitude and SSL performance. Key insights include minimizing ground reflection impacts at higher altitudes and enhancing SSL performance at lower altitudes with reduced ground reflectivity. Additionally, we found that power variations among multiple sound sources significantly affect the SSL performance of weaker sources. Among the MUSIC algorithm variants, iGEVD-MUSIC achieved the highest SSL performance, successfully locating multiple sound sources, including human voices. These findings demonstrate that the proposed simulation method is a valuable tool for developing and optimizing SSL techniques and parameters for real-world quadcopter applications. Furthermore, the insights gained from these simulations can be directly applied to the practical deployment of microphone array-equipped quadcopters in disaster response scenarios, aiding in the precise localization of survivors. This research significantly advances SSL methods and the practical realization of quadcopters for disaster site applications, ultimately enhancing the effectiveness and reliability of search and rescue operations.
In this paper, we propose a unified framework that leverages a single pretrained LLM for Motion-related Multimodal Generation, referred to as MoMug. MoMug integrates diffusion-based continuous motion generation with the model's inherent autoregressive discrete text prediction capabilities by fine-tuning a pretrained LLM. This enables seamless switching between continuous motion output and discrete text token prediction within a single model architecture, effectively combining the strengths of both diffusion- and LLM-based approaches. Experimental results show that, compared to the most recent LLM-based baseline, MoMug improves FID by 38 metrics by 16.61 accuracy across eight metrics by 8.44 of our knowledge, this is the first approach to integrate diffusion- and LLM-based generation within a single model for motion-related multimodal tasks while maintaining low training costs. This establishes a foundation for future advancements in motion-related generation, paving the way for high-quality yet cost-efficient motion synthesis.
This paper proposes an automatic method to identify important positions and their color features in retinal fundus images for gender classification using deep learning. The proposed method consists of MALCC (Model Analysis by Local Color Characteristics) and U-test. By partitioning each YCbCr image in the dataset into blocks and randomly masking them, the deep learning model predicts gender probabilities. Multiple regression analysis for the probabilities and masked blocks identifies significant blocks and colors, which are visualized in a Total Colormap. The Cr and Cb values of these blocks are then analyzed using a U-test to check if significant differences exist between classes. As a result of experiments, important blocks (positions) and their colors in the images are identified. Our code(RGB version) is available at https://github.com/tsutsu-22/MALCC.
While many unsupervised learning models focus on one family of tasks, either generative or discriminative, we explore the possibility of a unified representation learner: a model which addresses both families of tasks simultaneously. We identify diffusion models, a state-of-the-art method for generative tasks, as a prime candidate. Such models involve training a U-Net to iteratively predict and remove noise, and the resulting model can synthesize high-fidelity, diverse, novel images. We find that the intermediate feature maps of the U-Net are diverse, discriminative feature representations. We propose a novel attention mechanism for pooling feature maps and further leverage this mechanism as DifFormer, a transformer feature fusion of features from different diffusion U-Net blocks and noise steps. We also develop DifFeed, a novel feedback mechanism tailored to diffusion. We find that diffusion models are better than GANs, and, with our fusion and feedback mechanisms, can compete with state-of-the-art unsupervised image representation learning methods for discriminative tasks - image classification with full and semi-supervision, transfer for fine-grained classification, object detection and segmentation, and semantic segmentation. Our project website (https://mgwillia.github.io/diffssl/) and code (https://github.com/soumik-kanad/diffssl) are available publicly.
Compared to traditional agricultural environments, the high density and diversity of vegetation layouts in Synecoculture farms present significant challenges in locating and harvesting occluded fruits and pedicels (cutting points). To address this challenge, this study proposes a Mask R-CNN-based method for locating fruits (tomatoes, yellow bell peppers, etc.) and estimating the pedicels from RGBD images acquired by a camera moved along fixed paths. After obtaining masks of all fruits and pedicels, this method judges the matching relationship between the located fruit and pedicel according to the 3D distance between the fruit and pedicel. Subsequently, this research determines the least occluded best viewpoint for harvesting based on the visible real areas of located fruits in images acquired under the fixed paths, and harvesting is then completed from this best viewpoint following a straight path. Experimental results show this method effectively identifies occluded targets and their cutting positions in both Gazebo simulation environments and real-world farms. This method can select the least occluded viewpoint for a high harvesting success rate.
Cable tendency is the potential shape or characteristic that a cable may possess while being manipulated, of which some are considered erroneous and should be identified as a part of anomaly detection during an automatic manipulation. This research explores the ability of deep-learning models in learning the cable tendencies that, contrary to typical classification tasks of multi-object scenarios, is to differentiate the multiple states displayable by the same object - in this case, cables. By training multiple models with different combinations of self-collected real-world data and self-generated simulation data, a comparative study is carried out to compare the performance of each approach. In conclusion, the effectiveness of detecting three abnormal states and shapes of cables, and using simulation data is certificated in experiments.
This paper proposes a ski training system using VR (Virtual Reality) that enables beginners to acquire skiing skills without going to an actual ski ground. The proposed system obtains the speed of skiing based on the center of pressure (COP) of each player's foot. The first-person perspective of skiing at the obtained speed down a ski slope is fed back to the player as a VR image. Experiments were conducted to evaluate the effectiveness of the proposed system and the VR interface. Specifically, beginner skiers were categorized into three groups: “a group trained with the proposed VR system”, “a group trained with a system that provides feedback of the skiing speed calculated from the COP by increasing or decreasing the gauge (a bar-shaped graph representing changes in numerical values), instead of VR”, and “a group that does not train with the system”. After training under each of these conditions, a sliding test was conducted on an actual ski slope to check the degree of skill acquisition. The results show that subjects trained with the proposed system acquired more skiing skills than subjects who did not use the system on actual ski slopes. Furthermore, there was no clear difference in the result of the sliding test between subjects trained by the VR interface and those trained by the gauge interface, but the VR interface yields better deceleration postures.
This paper proposes a comprehensive method for estimating thrombus formation factors in the left atrial appendage (LAA). First, using 3D CT (Computer Tomography) image data as input, classification of thrombus presence/absence is learned using 3D ResNet. Besides, 3D Grad-CAM is applied to the prediction results to visualize regions of interest in thrombus formation. Second, features are extracted based on the visualization of regions of interest. Using the extracted features and numerical data obtained from the hospital as input, a regression analysis is performed to predict the presence/absence of thrombus using LightGBM. Visualization of regions of interest using 3D ResNet and 3D Grad-CAM shows that the right inferior pulmonary vein and the LAA were particularly correlated with thrombus formation. Estimation of important factors for thrombus formation using LightGBM shows that the LAA ostium area has the greatest influence on thrombus formation.Clinical Relevance—This paper shows the factors that contribute to thrombus formation in the LAA from the viewpoint of three-dimensional structure. In addition, the features considered important in thrombus formation were identified by comparing a variety of features.
In order to promote the widespread use of personal mobility (PM) in society, it is necessary to develop a driver support system that can accurately recognize the environment and improve the sense of security, safety, and comfort. This paper proposes a method for supporting the driver by recognizing obstacles and estimating the width of traffic from 3D point cloud obtained by the 3D lidar attached to the PM. To verify the effectiveness of the proposed method, we created a driving course with obstacles and conducted a driving test with 20 subjects. The results of this test show that the proposed method is effective in improving safety and comfort indices.