• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    德尔福汽车

    德尔福汽车

    Aptiv Inc.
    企业
    899论文总数
    1.8万引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Chantal S. Parenteau
    Chantal S. Parenteau
    Department of Injury Prevention, Chalmers University of Technology
    论文:11引用:0H-index:0
    Juliano Fujioka Mologni
    Juliano Fujioka Mologni
    ansys
    论文:9引用:0H-index:0
    Galen Fisher
    Galen Fisher
    Delphi Corporation;College of Engineering, University of Michigan
    论文:9引用:0H-index:0
    mark sellnau
    mark sellnau
    Automot Engn, Clemson Univ
    论文:9引用:0H-index:0
    R. K. Shah
    R. K. Shah
    Arts Sci & RA Patel Commerce Coll
    论文:8引用:0H-index:0
    david m bliesner
    david m bliesner
    delphi automotive
    论文:8引用:0H-index:0
    Minoo Shah
    Minoo Shah
    Delphi Automotive Systems
    论文:7引用:0H-index:0
    Laci Jalics
    Laci Jalics
    Aptiv
    论文:7引用:0H-index:0
    Joseph G. D'Ambrosio
    Joseph G. D'Ambrosio
    General Motors R&D Center
    论文:7引用:0H-index:0

    论文(899)

    年份
    起
    –
    止
    排序
    1Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
    Vignesh Gopinathan, Urs Zimmermann, Michael Arnold,Matthias Rottmann

    Video captioning models have seen notable advancements in recent years, especially with regard to their ability to capture temporal information. While many research efforts have focused on architectural advancements, such as temporal attention mechanisms, there remains a notable gap in understanding how models capture and utilize temporal semantics for effective temporal feature extraction, especially in the context of Advanced Driver Assistance Systems. We propose an automated LiDAR-based captioning procedure that focuses on the temporal dynamics of traffic participants. Our approach uses a rule-based system to extract essential details such as lane position and relative motion from object tracks, followed by a template-based caption generation. Our findings show that training SwinBERT, a video captioning model, using only front camera images and supervised with our template-based captions, specifically designed to encapsulate fine-grained temporal behavior, leads to improved temporal understanding consistently across three datasets. In conclusion, our results clearly demonstrate that integrating LiDAR-based caption supervision significantly enhances temporal understanding, effectively addressing and reducing the inherent visual/static biases prevalent in current state-of-the-art model architectures.

    20262026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)(2026)引用:1
    引用
    AI阅读
    加入学术空间
    2FutrTrack: A Camera-LiDAR Fusion Transformer for 3D Multiple Object Tracking
    Martha Teiko Teye, Ori Maoz, Matthias Rottmann

    We propose FutrTrack, a modular camera-LiDAR multi-object tracking framework that builds on existing 3D detectors by introducing a transformer-based smoother and a fusion-driven tracker. Inspired by query-based tracking frameworks, FutrTrack employs a multimodal two-stage transformer refinement and tracking pipeline. Our fusion tracker integrates bounding boxes with multimodal bird's-eye-view (BEV) fusion features from multiple cameras and LiDAR without the need for an explicit motion model. The tracker assigns and propagates identities across frames, leveraging both geometric and semantic cues for robust re-identification under occlusion and viewpoint changes. Prior to tracking, we refine sequences of bounding boxes with a temporal smoother over a moving window to refine trajectories, reduce jitter, and improve spatial consistency. Evaluated on nuScenes and KITTI, FutrTrack demonstrates that query-based transformer tracking methods benefit significantly from multimodal sensor features compared with previous single-sensor approaches. With an aMOTA of 74.7 on the nuScenes test set, FutrTrack achieves strong performance on 3D MOT benchmarks, reducing identity switches while maintaining competitive accuracy. Our approach provides an efficient framework for improving transformer-based trackers to compete with other neural-network-based methods even with limited data and without pretraining.

    2026International Conference on Computer Vision Theory and Applications(2026)引用:1
    引用
    AI阅读
    加入学术空间
    3ControlMap: Controllable High-Definition Map Generation for Traffic Scenario Simulation
    Marwan Farag, Steffen Wäldele, Yu Yao

    Simulation is central to validating autonomous driving systems, yet current pipelines are limited by insufficient scenario diversity due to costly High Definition (HD) map creation. Scaling HD maps requires expensive data collection and manual processing. Moreover, existing generative models lack the fine-grained control necessary to target specific road topologies during generation. This paper presents a data-driven pipeline for controllable HD map generation using latent diffusion and ControlNet for spatial conditioning. To our knowledge, we are the first to inject spatial guidance signals into a diffusion model for HD map synthesis. Furthermore, our model supports adjustable conditioning strength through classifier-free guidance and city-level style transfer via city label conditioning. To complement existing metrics, we introduce two novel metrics to evaluate adherence to the control signal and similarity to ground-truth maps. Experiments demonstrate that our model generates realistic HD maps that faithfully follow input road topologies while accurately preserving city-specific details.

    2026引用:1
    引用
    AI阅读
    加入学术空间
    4AdaFuse-Det: Adaptive Cross-Modal Fusion of Event Cameras for Robust Object Detection in Low-Light RGB Imagery
    Raju Imandi, Chethana B,Bharatesh Chakravarthi,Yong-Guk Kim, Manipriya S, Pavan Kumar B N

    Detecting objects reliably under extreme low-light conditions is an open problem in computer vision, with practical urgency in applications ranging from nighttime surveillance to search-and-rescue robotics. Conventional RGB cameras degrade sharply at low photon flux, while event cameras which record asynchronous per-pixel brightness changes at microsecond resolution and high dynamic range provide complementary structural cues that are largely illumination-invariant. We present AdaFuse-Det, a dual-stream framework that fuses CLAHE-enhanced RGB frames with voxelized event tensors through an Adaptive Cross-Modal Fusion (ACMF) module grounded in minimum-variance linear estimation theory. We formally show that the learned attention map asymptotically recovers the Gauss-Markov optimal fusion weights, and establish event conservation and temporal resolution bounds for the voxelization stage. On the LLE-VOS benchmark, AdaFuse-Det achieves a Recall of 65.54%, Precision of 53.85%, and F1-Score of 59.12% under severe illumination degradation, outperforming single-modality detectors in recall by a margin that reflects the theoretically predicted illumination-adaptation behavior.

    2026
    引用
    AI阅读
    加入学术空间
    5Texture-Shape Bias Balancing for Robust Synthetic-to-Real Semantic Segmentation in Automotive NIR Imagery
    Felix Stillger, Ben Hamscher, Lukas Hahn, Annika Mütze,Tobias Meisen, Kira Maag

    Semantic segmentation is a fundamental component of visual perception in modern automotive systems, enabling pixel-level scene understanding. Near-Infrared imaging (NIR) offers stable detection under difficult illumination conditions, but the development of domain-specific semantic segmentation models remains challenging due to the lack of high-quality annotated data from real-world scenarios. Synthetic datasets offer a scalable alternative, but models trained on synthetic images often suffer performance degradation when transferred to real domains. We present the first systematic study on synthetic to real domain adaptation for semantic segmentation in NIR images in the automotive domain. We propose a generative augmentation framework that transforms synthetic images into realistic NIR-style variants via our introduced target style adaptation (TSA). TSA fine-tunes a latent diffusion model via low-rank adaptation on a small curated set of real NIR images and applies it to synthetic training data using structure-preserving multi-signal conditioning. To reduce texture bias and improve segmentation robustness, we further apply a Voronoi-based style diversification strategy (VSD) that modifies the original textures while preserving scene geometry. Experiments with multiple model architectures on NIR data from vehicle interiors and street scenes show that balancing inductive bias during training leads to noticeably more robust semantic segmentation and effectively reduces the domain gap in our real-world scenarios by up to 63.6 GitHub .

    2026Machine Learning and Knowledge Discovery in Databases Applied Data Science Track, Demo Track and Ind...(2026)
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 899 篇论文

    合作机构(100)

    通用汽车合作论文 22
    伍伯塔尔大学合作论文 14
    韦恩州立大学合作论文 11
    浙江大学合作论文 10
    密歇根大学合作论文 9
    福特汽车公司合作论文 9
    麻省理工学院合作论文 7
    俄亥俄州立大学合作论文 7
    肯塔基大学合作论文 7
    卢森堡大学合作论文 6

    机构统计