• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    奔

    奔驰

    Mercedes-Benz
    企业EST. 1926
    609论文总数
    9,289引用总数

    梅赛德斯是西班牙语,是优雅的意思,代表着戴姆勒生产的汽车的理念;“梅赛德斯”的三叉星做标志,象征着陆上、水上和空中的机械化。自从奔驰制造了第一辆世界公认的汽车后,一百多年过去了,汽车早已度过了他的百岁寿辰,而在这一百多年来,随着汽车工业的蓬勃发展,曾涌现出很多的汽车厂家,也有显赫一时的,但最终不过是昙花一现。到如今,能够经历风风雨雨而最终保存下来,不过三四家、而百年老店,却仅只奔驰公司一家。

    论文量&引用量时间轴

    机构学者

    排序
    Felix Heide
    Felix Heide
    Computational Imaging Lab, Department of Computer Science, School of Engineering and Applied Science, Princeton University;Torc Robotics;Cephia AI
    论文:14引用:0H-index:0
    Michael Sedlmair
    Michael Sedlmair
    Institute for Visualization and Interactive Systems, University of Stuttgart
    论文:10引用:0H-index:0
    Markus Enzweiler
    Markus Enzweiler
    Daimler AG Research & Development Heßbruhlstraße
    论文:10引用:0H-index:0
    Wolfgang Minker
    Wolfgang Minker
    Faculty of Engineering and Computer Science,University of Ulm Institute of Information Technology
    论文:9引用:0H-index:0
    Klaus Dietmayer
    Klaus Dietmayer
    Institute of Measurement, Control and Microtechnology, University of Ulm
    论文:8引用:0H-index:0
    Luca Delgrossi
    Luca Delgrossi
    mercedes benz
    论文:8引用:0H-index:0
    Juergen Dickmann
    Juergen Dickmann
    Mercedes-Benz AG
    论文:8引用:0H-index:0
    Pascal Hirmer
    Pascal Hirmer
    Institute for Parallel and Distributed Systems, University of Stuttgart
    论文:6引用:0H-index:0
    Mario Bijelic
    Mario Bijelic
    Daimler AG
    论文:6引用:0H-index:0

    论文(609)

    年份
    起
    –
    止
    排序
    1SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
    Peizheng Li, Zhenghao Zhang, David Holtz, Hang Yu, Yutong Yang, Yuzhi Lai, Rui Song,Andreas Geiger,Andreas Zell

    End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-grained 3D spatial relationships which is a fundamental requirement for systems interacting with the physical world. To address this issue, we propose SpaceDrive, a spatial-aware VLM-based driving framework that treats spatial information as explicit positional encodings (PEs) instead of textual digit tokens, enabling joint reasoning over semantic and spatial representations. SpaceDrive employs a universal positional encoder to all 3D coordinates derived from multi-view depth estimation, historical ego-states, and text prompts. These 3D PEs are first superimposed to augment the corresponding 2D visual tokens. Meanwhile, they serve as a task-agnostic coordinate representation, replacing the digit-wise numerical tokens as both inputs and outputs for the VLM. This mechanism enables the model to better index specific visual semantics in spatial reasoning and directly regress trajectory coordinates rather than generating digit-by-digit, thereby enhancing planning accuracy. Extensive experiments validate that SpaceDrive achieves state-of-the-art open-loop performance on the nuScenes dataset and the second-best Driving Score of 78.02 on the Bench2Drive closed-loop benchmark over existing VLM-based methods. Code will be released upon acceptance.

    2026CVPR 2026(2026)引用:21
    引用
    AI阅读
    加入学术空间
    2CAIM: Development and Evaluation of a Cognitive AI Memory Framework for Long-Term Interaction with Intelligent Agents
    Rebecca Westhäußer, Frederik Berenz,Wolfgang Minker, Sebastian Zepf

    Large language models (LLMs) have advanced the field of artificial intelligence (AI) and are a powerful enabler for interactive systems. However, they still face challenges in long-term interactions that require adaptation towards the user as well as contextual knowledge and understanding of the ever-changing environment. To overcome these challenges, holistic memory modeling is required to efficiently retrieve and store relevant information across interaction sessions for suitable responses. Cognitive AI, which aims to simulate the human thought process in a computerized model, highlights interesting aspects, such as thoughts, memory mechanisms, and decision-making, that can contribute towards improved memory modeling for LLMs. Inspired by these cognitive AI principles, we propose our memory framework CAIM. CAIM consists of three modules: 1.) The Memory Controller as the central decision unit; 2.) the Memory Retrieval, which filters relevant data for interaction upon request; and 3.) the Post-Thinking, which maintains the memory storage. We compare CAIM against existing approaches, focusing on metrics such as retrieval accuracy, response correctness, contextual coherence, and memory storage. The results demonstrate that CAIM outperforms baseline frameworks across different metrics, highlighting its context-awareness and potential to improve long-term human-AI interactions.

    2026International Conference on Intelligent User Interfaces(2026)引用:7
    引用
    AI阅读
    加入学术空间
    3LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
    Julian Ost,Andrea Ramazzina, Amogh Joshi, Maximilian Bömer, Mario Bijelic,Felix Heide

    Large-scale scene data is essential for training and testing in robot learning. Neural reconstruction methods have promised the capability of reconstructing large physically-grounded outdoor scenes from captured sensor data. However, these methods have baked-in static environments and only allow for limited scene control -- they are functionally constrained in scene and trajectory diversity by the captures from which they are reconstructed. In contrast, generating driving data with recent image or video diffusion models offers control, however, at the cost of geometry grounding and causality. In this work, we aim to bridge this gap and present a method that directly generates large-scale 3D driving scenes with accurate geometry, allowing for causal novel view synthesis with object permanence and explicit 3D geometry estimation. The proposed method combines the generation of a proxy geometry and environment representation with score distillation from learned 2D image priors. We find that this approach allows for high controllability, enabling the prompt-guided geometry and high-fidelity texture and structure that can be conditioned on map layouts -- producing realistic and geometrically consistent 3D generations of complex driving scenes.

    2026AAAI 2026(2026)引用:5
    引用
    AI阅读
    加入学术空间
    4Radar Misalignment Calibration Using Vehicle Dynamics Model
    Luis Diener, Jens Kalkkuhl,Markus Enzweiler

    In this article, we explore a novel approach for the extrinsic misalignment calibration of automotive radar sensors. This article usually neglects the lateral velocity at the vehicle's rear axis when calibrating the radar misalignment. This limits the achievable accuracy and availability of their methods. Thus, given our previous success in estimating the side-slip gradient online; we propose an algorithm that estimates both the vehicle ego-motion states and the misalignment of multiple radar sensors, along with the sideslip gradient. Our findings demonstrate that this integrated approach allows for accurate and consistent calibration of radar misalignments while accurately estimating the side-slip gradient. Empirical validation using multiple datasets shows that the algorithm provides consistent accuracy and can track changes of parameters. The approach can, therefore, be applied to consumer-grade vehicles.

    2026IEEE SENSORS JOURNAL(2026)引用:3
    引用
    AI阅读
    加入学术空间
    5What Matters for Scalable and Robust Learning in End-to-End Driving Planners?
    David Holtz,Niklas Hanselmann,Simon Doll,Marius Cordts,Bernt Schiele

    End-to-end autonomous driving has gained significant attention for its potential to learn robust behavior in interactive scenarios and scale with data. Popular architectures often build on separate modules for perception and planning connected through latent representations, such as bird's eye view feature grids, to maintain end-to-end differentiability. This paradigm emerged mostly on open-loop datasets, with evaluation focusing not only on driving performance, but also intermediate perception tasks. Unfortunately, architectural advances that excel in open-loop often fail to translate to scalable learning of robust closed-loop driving. In this paper, we systematically re-examine the impact of common architectural patterns on closed-loop performance: (1) high-resolution perceptual representations, (2) disentangled trajectory representations, and (3) generative planning. Crucially, our analysis evaluates the combined impact of these patterns, revealing both unexpected limitations as well as underexplored synergies. Building on these insights, we introduce BevAD, a novel lightweight and highly scalable end-to-end driving architecture. BevAD achieves 72.7% success rate on the Bench2Drive benchmark and demonstrates strong data-scaling behavior using pure imitation learning. Our code and models are publicly available here: https://dmholtz.github.io/bevad/

    2026引用:3
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 609 篇论文

    合作机构(100)

    斯图加特大学合作论文 55
    乌尔姆大学合作论文 21
    卡尔斯鲁厄理工学院合作论文 20
    戴姆勒股份公司合作论文 15
    达姆施塔特工业大学合作论文 14
    图宾根大学合作论文 13
    普林斯顿大学合作论文 12
    Esslingen University of Applied Sciences合作论文 12
    慕尼黑工业大学合作论文 10
    思科系统合作论文 10

    机构统计