Radar offers unique advantages for localization in unstructured environments, including robustness to weather, lighting, and airborne particulates. While most prior work has studied radar odometry in urban, largely planar settings, its performance in off-road environments remains less understood. In this paper, we investigate the potential of radar for off-road odometry estimation and identify key challenges that arise from full SE(3) vehicle motion, terrain-induced ground returns, and sparse or unstable features. To address these issues, we introduce two simple baselines: Radar-KISSICP, which applies motion compensation to generate 3D-aware radar pointclouds, and Radar-IMU, which leverages IMU preintegration to stabilize scan matching. Experiments on the Great Outdoors (GO) dataset demonstrate that these baselines improve trajectory estimation in challenging routes and provide a reference point for future development of radar odometry in off-road robotics.
Autonomous navigation across large off-road environments remains a challenging problem. Onboard sensors perceive only the immediate surroundings, yet safe and efficient routes depend on terrain features that extend well beyond the sensor horizon. Geo-spatial data sources such as satellite imagery, aerial LiDAR, and vector maps can close this gap, but learning traversability from them is difficult: dense labels are unavailable at scale, and existing methods rely on short-range sensing. We propose an efficient formulation that learns a continuous traversability map from overhead data, supervised directly by human-driven GPS trajectories and shaped by self-supervised geometric priors from LiDAR. Alongside the model, we release a public dataset of 299 scenes spanning $\sim\!1{,}244\,\mathrm{km}^{2}$ of diverse terrain, paired with $1{,}130\,\mathrm{km}$ of human driving. In field trials on a Clearpath Warthog across seven routes at two sites, our method achieves trajectories within $3.66\%$ of human path length and reduces operator interventions by $\sim\!85\%$ compared to local-planner-only autonomy.
The Great Outdoors (GO) dataset is a multi-modal annotated data resource aimed at advancing ground robotics research in unstructured environments. Existing off-road datasets often lack sensor diversity and exclude vital modalities like thermal and radar that are critical for operation in degraded conditions (e.g., low visibility or adverse weather). To address these gaps, we introduce a large-scale multimodal off-road dataset with six complementary sensor modalities, along with semantic annotations and GPS traces, to support tasks such as semantic segmentation, object detection, and SLAM. The diverse environmental conditions represented in the dataset present significant real-world challenges, which provide opportunities to develop more robust solutions to support the continued advancement of field robotics, autonomous exploration, and perception systems in natural environments. The dataset can be downloaded at: https://www.unmannedlab.org/the-great-outdoors-dataset/
We present Stylos, a single-forward 3D Gaussian framework for 3D style transfer that operates on unposed content, from a single image to a multi- view collection, conditioned on a separate reference style image. Stylos synthesizes a stylized 3D Gaussian scene without per-scene optimization or precomputed poses, achieving geometry-aware, view-consistent stylization that generalizes to unseen categories, scenes, and styles. At its core, Stylos adopts a Transformer backbone with two pathways: geometry predictions retain self-attention to preserve geometric fidelity, while style is injected via global cross-attention to enforce visual consistency across views. With the addition of a voxel-based 3D style loss that aligns aggregated scene features to style statistics, Stylos enforces view-consistent stylization while preserving geometry. Experiments across multiple datasets demonstrate that Stylos delivers high-quality zero-shot stylization, highlighting the ef- fectiveness of global style–content coupling, the proposed 3D style loss, and the scalability of our framework from single view to large-scale multi-view settings. Our codes will be fully open-sourced soon.
Pedestrian behavior prediction is one of the most critical tasks in urban driving scenarios, playing a key role in ensuring road safety. Traditional learning-based methods have relied on vision models for pedestrian behavior prediction. However, fully understanding pedestrians’ behaviors in advance is very challenging due to the complex driving environments and the multifaceted interactions between pedestrians and road elements. Additionally, these methods often show a limited understanding of driving environments not included in the training. The emergence of Multimodal Large Language Models (MLLMs) provides an innovative approach to addressing these challenges through advanced reasoning capabilities. This paper presents OmniPredict, the first study to apply GPT-4o(mni), a state-of-the-art MLLM, for pedestrian behavior prediction in urban driving scenarios. We assessed the model using the JAAD and WiDEVIEW datasets, which are widely used for pedestrian behavior analysis. Our method utilized multiple contextual modalities and achieved 67% accuracy in a zero-shot setting without any task-specific training, surpassing the performance of the latest MLLM baselines by 10%. Furthermore, when incorporating additional contextual information, the experimental results demonstrated a significant increase in prediction accuracy across four behavior types (crossing, occlusion, action, and look). We also validated the model s generalization ability by comparing its responses across various road environment scenarios. OmniPredict exhibits strong generalization capabilities, demonstrating robust decision-making in diverse and unseen driving rare scenarios. These findings highlight the potential of MLLMs to enhance pedestrian behavior prediction, paving the way for safer and more informed decision-making in road environments.
LiDAR semantic segmentation frameworks predominantly use geometry-based features to differentiate objects within a scan. These methods excel with clear boundaries but struggle in ambiguous environments, especially off-road. Recent 3D segmentation advances use raw LiDAR intensity for better prediction accuracy. Nonetheless, existing models face challenges in relating raw intensity to distance, angle, reflectivity, and atmospheric conditions. Building on prior work Viswanath et al. (2024), we examine the benefits of calibrated intensity (reflectivity) in learning-based LiDAR segmentation. Adding reflectivity as input enhances data representation, resulting in a 4% mIoU improvement on the Rellis-3d off-road dataset. We also explore calibrated intensity benefits for urban segmentation (SemanticKITTI) and cross-sensor adaptation. Testing the Segment Anything Model (SAM) Kirillov et al. (2023) with reflectivity led to improved masks for LiDAR images.
Recent advances have improved autonomous navigation and mapping under payload constraints, but current multi-robot inspection algorithms are unsuitable for nano-drones, due to their need for heavy sensors and high computational resources. To address these challenges, we introduce ExploreBug , a novel hybrid frontier range-bug algorithm designed to handle limited sensing capabilities for a swarm of nano-drones. This system includes three primary components: a mapping subsystem, an exploration subsystem, and a navigation subsystem. Additionally, an intra-swarm collision avoidance system is integrated to prevent collisions between drones. We validate the efficacy of our approach through extensive simulations and real-world exploration experiments, involving up to seven drones in simulations and three in real-world settings, across various obstacle configurations and with a maximum navigation speed of 0.75 m/s. Our tests prove that the algorithm efficiently completes exploration tasks, even with minimal sensing, across different swarm sizes and obstacle densities. Furthermore, our frontier allocation heuristic ensures an equal distribution of explored areas and paths traveled by each drone in the swarm. We publicly release the source code of the proposed system to foster further developments in mapping and exploration using autonomous nano drones.
A crucial technology in autonomous driving is the ability to predict whether a pedestrian will cross the road in the future, allowing autonomous vehicles to respond accordingly. Traditional methods have employed visual networks to predict pedestrian crossing intentions. However, these methods rely on trained datasets, resulting in a lack of generalization when faced with previously unseen driving scenarios. The advent of Multimodal Large Language Models (MLLMs), proficient in processing and reasoning with both text and images, offers a breakthrough new approach to overcome these challenges. In this paper, we propose LLaMAPed, the first study to apply the open-source MLLM, VideoLLaMA2, to predict pedestrian crossing intentions. We evaluated the performance of our method on the widely used JAAD dataset for pedestrian behavior prediction and compared its performance to traditional visual benchmark models and the closed-source GPT-4V. VideoLLaMA2, designed to enhance the understanding of spatial-temporal modeling, has been utilized in LLaMAPed to predict pedestrian crossing intentions in a zero-shot manner. LLaMAPed achieved a prediction accuracy of 58
Autonomous navigation in off-road environments remains a significant challenge in field robotics, particularly for Unmanned Ground Vehicles (UGVs) tasked with search and rescue, exploration, and surveillance. Effective long-range planning relies on the integration of onboard perception systems with prior environmental knowledge, such as satellite imagery and LiDAR data. This work introduces Trailblazer, a novel framework that automates the conversion of multi-modal sensor data into costmaps, enabling efficient path planning without manual tuning. Unlike traditional approaches, Trailblazer leverages imitation learning and a differentiable A* planner to learn costmaps directly from expert demonstrations, enhancing adaptability across diverse terrains. The proposed methodology was validated through extensive real-world testing, achieving robust performance in dynamic and complex environments, demonstrating Trailblazer's potential for scalable, efficient autonomous navigation.
Off-road traversability segmentation enables autonomous navigation with applications in search-and-rescue, military operations, wildlife exploration, and agriculture. Current frameworks struggle due to significant variations in unstructured environments and uncertain scene changes, and are not adaptive to be used for different robot types. We present AnyTraverse, a framework combining natural language-based prompts with human-operator assistance to determine navigable regions for diverse robotic vehicles. The system segments scenes for a given set of prompts and calls the operator only when encountering previously unexplored scenery or unknown class not part of the prompt in its region-of-interest, thus reducing active supervision load while adapting to varying outdoor scenes. Our zero-shot learning approach eliminates the need for extensive data collection or retraining. Our experimental validation includes testing on RELLIS-3D, Freiburg Forest, and RUGD datasets and demonstrate real-world deployment on multiple robot platforms. The results show that AnyTraverse performs better than GA-NAV and Off-seg while offering a vehicle-agnostic approach to off-road traversability that balances automation with targeted human supervision.
In the area of autonomous driving, navigating off-road terrains presents a unique set of challenges, from unpredictable surfaces like grass and dirt to unexpected obstacles such as bushes and puddles. In this work, we present a novel learning-based local planner that addresses these challenges by directly capturing human driving nuances from real-world demonstrations using only a monocular camera. The key features of our planner are its ability to navigate in challenging off-road environments with various terrain types and its fast learning capabilities. By utilizing minimal human demonstration data (5-10 mins), it quickly learns to navigate in a wide array of off-road conditions. The local planner significantly reduces the real world data required to learn human driving preferences. This allows the planner to apply learned behaviors to real-world scenarios without the need for manual fine-tuning, demonstrating quick adjustment and adaptability in off-road autonomous driving technology.
This paper presents a novel system designed for 3D mapping and visual relocalization using 3D Gaussian Splatting. Our proposed method uses LiDAR and camera data to create accurate and visually plausible representations of the environment. By leveraging LiDAR data to initiate the training of the 3D Gaussian Splatting map, our system constructs maps that are both detailed and geometrically accurate. To mitigate excessive GPU memory usage and facilitate rapid spatial queries, we employ a combination of a 2D voxel map and a KD-tree. This preparation makes our method well-suited for visual localization tasks, enabling efficient identification of correspondences between the query image and the rendered image from the Gaussian Splatting map via normalized cross-correlation (NCC). Additionally, we refine the camera pose of the query image using feature-based matching and the Perspective-n-Point (PnP) technique. The effectiveness, adaptability, and precision of our system are demonstrated through extensive evaluation on the KITTI360 dataset.
Object detection and subsequent perception of the environment surrounding a vehicle play a very important role in autonomous driving applications. Existing perception algorithms do not generalize well since most algorithms are trained in well-structured driving environment datasets. To deploy self-driving cars on the road, they should have a reliable and robust perception system to handle all corner cases. This paper introduces a multimodal dataset, TIAND * (TiHAN-IITH Autonomous Navigation Dataset), collected from structured and unstructured environments seen in and around the city of Hyderabad, India, as an aid to further research in the generalization of object detection algorithms. The sensor suite contains four cameras, six radars, one Lidar, and GPS and IMU. TIAND comprises 150 scenes, each spanning a duration ranging from 2 minutes to 4 minutes. Subsequently, we present the object detection model’s performance using camera, radar, and Lidar data. Additionally, we offer insights into projecting data from Lidar to camera and from radar to camera.
In the realm of sustainable autonomous driving, this study investigates the impact of network disturbances on autonomous vehicle behavior and maneuvering, along with the influence of vehicle maneuvering profiles on battery and fuel consumption. Firstly, we present preliminary data and initial findings regarding the impact of autonomous vehicle brake and speed profiles on battery consumption. Secondly, we provide a detailed evaluation of network disturbances, particularly latency, on autonomous vehicle behaviors and profiles, leveraging CARLA simulations and NETEM to emulate network disruptions. Key metrics such as Time to Collision and Safe Stop Distance serve as primary indicators to assess AV behavior under constant and variable latency conditions. Early results indicate that sharp speed reductions and abrupt braking correlate with increased battery consumption in autonomous electric shuttles, with a 1.9% overall increase in battery consumption in case of frequent stops with hard braking in autonomous mode. Additionally, our simulations underscore the adverse effects of network disturbances on the vehicle’s ability to maintain safe stopping distances and time to collision in waypoint following and car following behaviors respectively. Notably, even a modest latency of 5ms extends the duration of collision risk zones by 28%, compelling vehicles to frequently apply brakes in car following behaviour. Overall, our investigations underscore that constant latency exerts a more pronounced effect on desired vehicle stop behavior, while variable latency presents a more formidable challenge by adversely impacting time to collision in car following behavior.
Autonomous Vehicles face significant safety challenges in complex urban environments, particularly in detecting and tracking vulnerable road users (VRUs) like pedestrians and cyclists, who are at higher risk of fatal accidents. Traditional sensor fusion techniques that combine LiDAR with Vision struggle with occlusions and adverse weather or lighting conditions. This paper explores the potential of Ultra-Wideband (UWB) technology as an additional sensing modality, known for its high ranging accuracy and robustness in challenging environments. Through real-world experiments, we compare UWB's performance to that of common vision sensors, demonstrating its effectiveness in improving VRUs' detection and tracking in urban driving scenarios. Finally, we demonstrate the potential applications of the UWB technology in VRU localization scenarios.
LiDAR is used in autonomous driving to provide 3D spatial information and enable accurate perception in off-road environments, aiding in obstacle detection, mapping, and path planning. Learning-based LiDAR semantic segmentation utilizes machine learning techniques to automatically classify objects and regions in LiDAR point clouds. Learning-based models struggle in off-road environments due to the presence of diverse objects with varying colors, textures, and undefined boundaries, which can lead to difficulties in accurately classifying and segmenting objects using traditional geometric-based features. In this paper, we address this problem by harnessing the LiDAR intensity parameter to enhance object segmentation in off-road environments. Our approach was evaluated in the RELLIS-3D data set and yielded promising results as a preliminary analysis with improved mIoU for classes “puddle” and “grass” compared to more complex deep learning-based benchmarks ( https://github.com/MOONLABIISERB/lidar-intensity-predictor/tree/main ). The methodology was evaluated for compatibility across both Velodyne and Ouster LiDAR systems, assuring its cross-platform applicability. This analysis advocates for the incorporation of calibrated intensity as a supplementary input, aiming to enhance the prediction accuracy of learning based semantic segmentation frameworks.
Predicting pedestrian behavior is the key to ensure safety and reliability of autonomous vehicles. While deep learning methods have been promising by learning from annotated video frame sequences, they often fail to fully grasp the dynamic interactions between pedestrians and traffic, crucial for accurate predictions. These models also lack nuanced common sense reasoning. Moreover, the manual annotation of datasets for these models is expensive and challenging to adapt to new situations. The advent of Vision Language Models (VLMs) introduces promising alternatives to these issues, thanks to their advanced visual and causal reasoning skills. To our knowledge, this research is the first to conduct both quantitative and qualitative evaluations of VLMs in the context of pedestrian behavior prediction for autonomous driving. We evaluate GPT-4V(ision) on publicly available pedestrian datasets: JAAD and WiDEVIEW. Our quantitative analysis focuses on GPT-4V's ability to predict pedestrian behavior in current and future frames. The model achieves a 57% accuracy in a zero-shot manner, which, while impressive, is still behind the state-of-the-art domain-specific models (70%) in predicting pedestrian crossing actions. Qualitatively, GPT-4V shows an impressive ability to process and interpret complex traffic scenarios, differentiate between various pedestrian behaviors, and detect and analyze groups. However, it faces challenges, such as difficulty in detecting smaller pedestrians and assessing the relative motion between pedestrians and the ego vehicle.