Autonomous driving depends on successful interactions among humans, vehicles, and roads. However, people often lack an understanding of autonomous vehicle (AV) behaviours and decisions. Moreover, AVs have difficulty aligning with human intentions in their interactions. To overcome the obstacles associated with the absence of interactive intelligence, especially in complex and uncertain environments, we introduce the concept of embodied interactive intelligence towards autonomous driving (EIIAD), which establishes representation and learning methods aligned with the physical world, enhancing human–machine integration. Building on this concept, we propose an end-to-end unified constrained vehicle environment interaction (UniCVE) model, which involves the construction of an end-to-end perception–cognition–behaviour closed-loop feedback paradigm and continuous learning through accumulated split driving scenarios. This model realizes interaction cognition through networks designed for pedestrians and vehicles, and it unifies the cognition as a value network of AVs to generate socially compatible behaviours. The UniCVE model is implemented on Dongfeng autonomous buses, which have successfully travelled 22 thousand kilometres and completed 45 thousand navigation tasks in Xiong’an New Area, China, demonstrating its general applicability in various driving scenarios. In addition, we highlight the high-level interactive intelligence of the UniCVE model in selected simulated complex interaction scenarios, demonstrating that it makes AVs more intelligent, more reliable, and more attuned to human relationships. Furthermore, the UniCVE model’s capacity for self-learning and self-growth allows it to infinitely approximate true intelligence, even with limited experience.
This study explores the effect of takeover time (TOT) on decision-making for Level-3 autonomous vehicles (L3-AVs). The existing research on L3-AV lacks an in-depth analysis of the mechanisms affecting TOT, ignores the importance of spatial and temporal variations in features for TOT prediction, and also lacks consideration of TOT in downstream trajectory planning tasks. This study proposed an exponential smoothing transformers (ETS) former model for TOT prediction, and then, the spatial-temporal predictive transformer (ST-Preformer) was employed to forecast the trajectories of surrounding vehicles, assess lane availability, and determine lane-changing probabilities. Ultimately, these evaluations contribute to the decision-making process of L3-AVs. The findings showed that the ETSformer was able to explain more than 83% of the characteristics of the TOT distribution in the TOT prediction task, effectively reducing the absolute percentage error by 0.7%, based on which the decision-making framework was able to make safe and comfortable optimal decisions. Decision-making is closely related to driving conditions and the surrounding traffic state, and TOT has a critical impact on the safety and stability of decision-making. A comprehensive understanding the impact of TOT on decision-making can help improve the safety of autonomous driving and provide guidance for improving decision-making techniques.
This paper tackles the challenges of sparse and unevenly distributed of ring-like mechanical LiDAR point clouds in road environment perception for autonomous driving by proposing the LPSF-LiDARNet framework. The framework enhances fine-grained 3D semantic segmentation through inter-frame spatiotemporal feature enhancement and balanced voxel sampling. Temporal window fusion is achieved via multi-frame stacking, integrating complementary temporal features to mitigate single-frame incompleteness. Spatially, log-polar coordinate voxel sampling leverages spatial distribution patterns to improve feature consistency. Additionally, adaptive heterogeneous convolution kernels with dynamic attention mechanisms are introduced, combined with semantic-level data augmentation, to optimize feature extraction for sparse points and rare samples. The framework demonstrates superior performance over existing models in experiments using SemanticKITTI and local datasets, validating the effectiveness of its spatiotemporal fusion strategy and sampling mechanism. Ultimately, the model achieves a segmentation accuracy of 96.4
Active Disturbance Rejection Control (ADRC) possesses robust disturbance rejection capabilities, making it well‐suited for longitudinal velocity control. However, the conventional Extended State Observer (ESO) in ADRC fails to fully exploit feedback from first‐order and higher‐order estimation errors and tracking error simultaneously, thereby diminishing the control performance of ADRC. To address this limitation, an enhanced car‐following algorithm utilising ADRC is proposed, which integrates the improved ESO with a feedback controller. In comparison to the conventional ESO, the enhanced version effectively utilises multi‐order estimation and tracking errors. Specifically, it enhances convergence rates by incorporating feedback from higher‐order estimation errors and ensures the estimated value converges to the reference value by utilising tracking error feedback. The improved ESO significantly enhances the disturbance rejection performance of ADRC. Finally, the effectiveness of the proposed algorithm is validated through the Lyapunov approach and experiments.
ABSTRACT The widespread use of ChatGPT has normalized the dialogue Turing test. To meet this challenge, China's major national development strategy suggests that for a new generation of artificial intelligence, it is first necessary to answer the big questions raised by Turing in 1950 from the perspective of cognitive physics: Can machines think? How do machines think? How do machines cognize? Whether it is carbon-based human cognition or silicon-based machine cognition, it is an interaction between complex constructs composed of the four most basic elements: matter, energy, structure, and time. Both humans and machines depend on negative entropy for living, and time is the cornerstone of cognition. Structure and time are parasitic on matter and energy in physical space, forming hard-structured ware. The soft-structured ware in cognitive space is mind, which is parasitic on the hard-structured ware or other existing soft-structured ware, and constitutes a rich hierarchy of multi-scale feelings, concepts, information, and knowledge. Extending “abstraction” from the symbolic school of artificial intelligence, “association” from the connectionist school, and “interaction” from the behaviorist school, the core of cognition is established on the shoulders of such scientific giants such as Schrödinger, Turing and Wiener. Soft and hard-structured ware interact. Cognitive machine can comprise heterogeneous hard-structured ware, such as field programmable gate arrays (FPGAs), data processing units (DPUs), central processing units (CPUs), graphics processing units (GPUs), tensor processing units (TPUs), and memory. It can also be implanted with the “Baby Cognitive Nucleus” which is hard-structured ware genetically inherited and naturally evolved to form the embodied machine. Then the hard-structured ware is parasitized by rich, multi-scale soft-structured ware. By regulating matter and energy through soft-structured ware, machines produce orderly events, form coordinated and orderly thinking activities. The heterogeneous sensors configured by the machine and the speed of thinking will no longer be trapped by the extreme values of biochemical parameters of carbon-based organisms but will be able to perceive through multi-channel cross-modal means, carry out intense thinking, and maintain cognitive continuity with memory. To generate computational and memory intelligence in cognitive space which can bootstrap, self-reuse and self-replicate, imagination and creativity are improved through memory-constrained computing. The new generation of artificial intelligence will leap beyond mechanized mathematics to automation in thinking and self-driven growth of cognition, and the thinking in the cognitive space and the behavior in the physical space verify each other, from the dialogue Turing test to the embodied Turing test. Humans have entered the intelligent era of human-machine co-creation with cognitive machines iteratively inventing, discovering, and creating alongside scientists, engineers, and skilled craftsmen, each wise in its way, improving thinking ability, and amplifying human energy.
This paper discusses how intelligent machines have replaced humans in tasks requiring, heavy, and repetitive labor, whilst being better suited to the requirements of these jobs. The increased capacity for brute force computation has facilitated increased collaborative innovation between man and machines. For example, the intelligent farming machines have overcome the confines of computational power, algorithms, and data, and the next generation of intelligent farming machines is expected to interact, learn, and grow autonomously. In the future, in addition to self enhancement, humans are expected to teach machines to learn and work. Scientists and engineers will collaborate with machines to accomplish invention, discovery, and creation. For "embodied intelligence" in the farming machine context, we propose (1) deep learning should be performed iteratively via real-time interactions with the external world; (2) embodied control and self-regulation can ensure coordination between behaviors of machines and their environment; (3) intelligent farming machines are characterized by the ability to interact, learn, and grow autonomously.
Intelligent driving has become one of the most compelling topics of interest.Nevertheless,current intelligent driving technologies still face challenges,such as missed detection due to large vehicle occlusion and false detection caused by sensor accuracy degradation in sudden light changes.Multimodal perception technology for intelligent vehicles has emerged to ensure the safety of vehicle perception in complex scenarios.However,the existing multimodal fusion methods are still limited to the improvement of detection accuracy,lack of interpretability of the perception process and lack of evaluation indexes for the model perception process.In this paper,from the information theory perspective,we design the perception model according to the communication model.We propose a multimodal fusion perception model based on joint source-channel coding theory to explain the perception process of the model theoretically.At the same time,we propose a new evaluation index,average entropy variation(AEV),which is used to reflect the stability of the model during its perceptual interaction with the outside world in real time.Further,the perceptual process is quantified and analyzed to increase the interpretability of the model.Finally,we compare the evaluation results with other advanced perceptual models in the KITTI dataset,and our model decreases the average entropy variation to 0.5904,which better ensures the perceptual safety of the detection task.
城市立体交通是未来智慧出行发展的热点方向,近年来受到了广泛的关注.作为城市立体交通的载体,智能飞行汽车融合了飞机与汽车两种运动模态,能够灵活地在空中与地面进行切换.本文介绍了智能飞行汽车的背景、历史与现状,阐述了其与城市空中交通载具的区别,分析与讨论了飞行汽车的系统设计,并介绍了智能飞行汽车的关键技术创新,包括动力技术和机电总体设计、多模态切换、模块复用与飞车脑认知等.重点讨论了飞行汽车的智能化技术,包括近地感知、决策与规划、智能控制与智能通信系统的关键技术与瓶颈.最后,结合现有技术,对智能飞行汽车的技术进行了剖析,并讨论了潜在的解决方案与发展趋势.
The mechanical LiDAR sensor is crucial in autonomous vehicles. After projecting a 3D point cloud onto a 2D plane and employing a deep learning model for computation, accurate environmental perception information can be supplied to autonomous vehicles. Nevertheless, the vertical angular resolution of inexpensive multi-beam LiDAR is limited, constraining the perceptual and mobility range of mobile entities. To address this problem, we propose a point cloud super-resolution model in this paper. This model enhances the density of sparse point clouds acquired by LiDAR, consequently offering more precise environmental information for autonomous vehicles. Firstly, we collect two datasets for point cloud super-resolution, encompassing CARLA32-128in simulated environments and Ruby32-128 in real-world scenarios. Secondly, we propose a novel temporal and spatial feature-enhanced point cloud super-resolution model. This model leverages temporal feature attention aggregation modules and spatial feature enhancement modules to fully exploit point cloud features from adjacent timestamps, enhancing super-resolution accuracy. Ultimately, we validate the effectiveness of the proposed method through comparison experiments, ablation studies, and qualitative visualization experiments conducted on the CARLA32-128 and Ruby32-128 datasets. Notably, our method achieves a PSNR of 27.52 on CARLA32-128 and a PSNR of 24.82 on Ruby32-128, both of which are better than previous methods.
Human intelligence begins with language, and artificial intelligence begins with words. The greatest human intelligence is the invention of education. Intelligence is rooted in education. Education has changed the development of human intelligence from the “present continuous tense” to the “present perfect tense”. The matter and energy in the intelligent machine are the real existence at the physical level, and the structure and time are the abstract thinking at the cognitive level. Structure and time are parasitic on matter and energy, forming a hard-structure ware. The information in the machine is a large number of soft-structured ware, reflecting the spiritual world, and it can be parasitized on the hard-structured ware or other soft-structured ware. There are both virtual and real, and the combination of virtual and real, and it can be bootstrapped and reused by itself. There is always at least the next time cycle, so that the machine can “think” again. Its order shows the ability to maintain its own thinking and produce orderly events. Human thinking and machine thinking are isomorphic in mathematics and homologous in physics, supported by energy and living on negative entropy. The birth of intelligent machines has impacted the all-round and all-around elements of education, from teaching “books” to teaching “learning”, and then to teaching “education”, from how to acquire knowledge, to how to use knowledge, and then to how to create knowledge. The essence of education in the era of intelligence is to cultivate the imagination and creativity of thinking. The core of human thinking is abstraction, association and interaction, and so is the machine. Machines use soft-structured ware to extend and expand human thinking. More importantly, machines can think with violence. People and machines can teach and learn from each other, complement each other, and form iterative intelligence. The problem of education reform in the era of intelligence has been put in front of all mankind urgently. Reading changes one person and education changes all mankind. Let us welcome the revolution of learning!
新一代人工智能如何从传统人工智能中脱颖而出? 当前中国人工智能的现状,一是"基础研究弱",我国获得图灵奖的只有姚期智院士一人,还是在美国取得的,中国人在自己国土上还没有拿到过这个奖;二是"辐射市场大",许多人、许多行业都在这上面忙碌;"叶茂枝不壮,树大根不深,叫好难叫座",叫好的人多,投资的人少.
The sequential recommendation (also known as the next-item recommendation), which aims to predict the following item to recommend in a session according to users’ historical behavior, plays a critical role in improving session-based recommender systems. Most of the existing deep learning-based approaches utilize the recurrent neural network architecture or self-attention to model the sequential patterns and temporal influence among a user's historical behavior and learn the user's preference at a specific time. However, these methods have two main drawbacks. First, they focus on modeling users’ dynamic states from a user-centric perspective and always neglect the dynamics of items over time. Second, most of them deal with only the first-order user-item interactions and do not consider the high-order connectivity between users and items, which has recently been proved helpful for the sequential recommendation. To address the above problems, in this article, we attempt to model user-item interactions by a bipartite graph structure and propose a new recommendation approach based on a Position-enhanced and Time-aware Graph Convolutional Network (PTGCN) for the sequential recommendation. PTGCN models the sequential patterns and temporal dynamics between user-item interactions by defining a position-enhanced and time-aware graph convolution operation and learning the dynamic representations of users and items simultaneously on the bipartite graph with a self-attention aggregator. Also, it realizes the high-order connectivity between users and items by stacking multi-layer graph convolutions. To demonstrate the effectiveness of PTGCN, we carried out a comprehensive evaluation of PTGCN on three real-world datasets of different sizes compared with a few competitive baselines. Experimental results indicate that PTGCN outperforms several state-of-the-art sequential recommendation models in terms of two commonly-used evaluation metrics for ranking. In particular, it can make a better trade-off between recommendation performance and model training efficiency, which holds great potential for online session-based recommendation scenarios in the future.
Knee osteoarthritis (Knee OA) is a degenerative disease that often perplexes the elderly and the whole society, and its timely recognition receives interest worldwide. However, traditional imaging examinations cannot reflect dynamic function nor implement long-term monitoring. To address this issue, this article suggests a piezoresistive-based gait monitoring method to recognize patients with Knee OA by assessing the plantar pressure signals during the subjects' walking, which is mobile, wearable, low-cost, and convenient. Eighteen subjects diagnosed with Knee OA and twenty-two control subjects participated in the experiment. Considering the asymmetric pressure distribution in feet and the landing habits of Knee OA patients, the plantar surface was split into eight areas, calculating the contact time and maximum force of each area in a gait cycle. Using these characteristics to train, the support vector machine (SVM) reached an accuracy of 93.15%, a precision of 92.39%, and a recall of 92.79%. Furthermore, a prediction model was proposed for the application that aggregates all the results in one test and gives a more accurate result, and the classification accuracy for individuals in the ensemble model is 90.90%. Our technique fills the vacancy of the recognition of patients with Knee OA based on wearable instruments. It provides ideas for intelligent healthcare, which benefits potential Knee OA patients' early diagnosis and treatment.
鉴于采用边缘云进行集中式车辆队列控制时,通信时延将会降低队列控制性能指标甚至导致队列失稳,本文在考虑通信时延和车辆纵向非线性动力学特性前提下,从包括队列能效的多目标优化出发,提出了一种基于边缘云的队列集中式模型预测控制算法,并设计了一种时延补偿方法.首先分析了控制算法的渐进稳定性;然后通过不同时延下的仿真试验对控制算法的串稳定性和随机时延补偿方法在一定时延范围的有效性进行了验证;最后,分析了时延对队列稳定性和燃料消耗的影响.结果表明,随着时延增大,队列稳定性和燃料经济性变差,当时延均值为250 ms且时延抖动20%时,队列处于失稳的边缘.
Mechanical LiDAR is one of the most crucial perception sensors for autonomous vehicles. However, the vertical angular resolution of low-cost multi-beam LiDAR is small, limiting the perception and movement range of mobile agents. This paper presents a novel temporal convolutional (TC)-based U-Net model for point cloud super-resolution, which can optimize the point cloud of low-cost LiDAR based on fusing spatiotemporal features of the point cloud. We project the 3D point cloud on a 2D image plane and extend a U-Net convolutional neural network model with a temporal convolutional (TC) module for processing consecutive frames. Each time the model generates one dense/up-sampled image from low-end LiDAR consecutive frames and projects it back into the 3D space as the final result. Considering the intrinsic noise of LiDAR, the structural similarity index measure (SSIM) is introduced as the loss function. Experiments are carried out on both datasets generated by the CARLA simulator and a small-scale dataset collected from actual road conditions with a local vehicle platform. Results show that the proposed model achieves a high peak signal to noise ratio (PSNR). It means the T-UNet model can effectively upsample the sparse point cloud of low-cost LiDAR to a dense point cloud which is almost indistinguishable from the high-end LiDAR point cloud. The source code can be accessed at https://github.com/donkeyofking/lidar-sr.git
基础研究崇尚想象力和创造力的完全自由,依赖独立学者的兴趣和自由合作,它可以不限研究者的身份,不设完成的时限,不以落地应用为目的,也不一定要组织大团队"攻关",不搞群众运动,允许试错,宽容失败,更不以获得自然科学奖为目的;需要研究者有深厚的人文艺术素养,耐得住寂寞,沉得下心来,虽然研究结果和产出时间无法被精确预测,但一旦出现原始创新,对引领技术进步必然会有长期且深刻的影响.阿兰·图灵的研究就是一例.
在我国当前助力乡村振兴的大背景下,从广袤的东北平原到美丽的西南山地,已经大面积解决了农业机械化的问题.东北地区人少地多,一望无际的黑土地,需要的农机主要是大马力;西南地区山地和丘陵、人多地少,更需要微型农机的精细作业.生态特色各有不同,我们要从根本上改变农民面朝黄土背朝天的艰苦劳作状态,都需要用人工智能打造有温度的农机,研发有感知、有认知、有行为、可交互、会学习、自成长的农田作业机器人.
As one of the important signs of the third wave of artificial intelligence, wheeled robots not only inherit knowledge but also learn independently, which brings about to learnable wheeled robots that use a driving brain to achieve data-driven control and learning. Presently, most existing technologies for self-driving vehicles can learn positively from the benchmark drivers to guarantee safe driving. However, in many unpredicted situations, such as rollover, human drivers often cause the behavior of irrational subconscious on account of human emotions like panic. In this paper, we propose a learnable wheeled robot using the driving brain by taking the rollover as an example, which is the most serious and dangerous situation in dynamic vehicle operations. Then, based on the analysis of rollover accidents, we utilize the driving brain reversely and conduct negative learning, materializing, and condensing the group intelligence of accident experts, to solve the problem of the lack of individual intelligence in emergencies and further promote real-time response to other dangerous conditions, such as puncture for self-driving vehicles.
Herein, we present a novel approach for monocular dual quadric initialization that combines three-dimensional (3D) map points with two-dimensional (2D) object detection for forward-translating camera movements. The traditional approach using 2D detection bounding boxes in multiple views fails in straight vehicle motion scenarios as object observation is limited to few frames. Although single image initialization is possible when multiple constraints are introduced, such initialization is based on strong assumptions. In this letter, we incorporate constraints from 3D map points with single-view 2D object detection to robustly initialize the dual quadric. Constraints from 3D map points are converted to planar constraints from their convex hull. Together with the projective planar constraints from bounding boxes, the proposed method can infer accurate dual quadric parameters. Further, comparison studies with the state of the art (SOTA) show that the proposed approach achieves the same accuracy of center localization but outperforms the existing methods in shape estimation and success ratio of initialization. The proposed method dose not rely on assumptions of dimension and pose of 3D objects; hence, it is more generic and accurate. Based on the KITTI raw dataset, the initialization success ratio is up to 97.7% with an average position error of 1.58 m, and 2D IoU of 80% when the number of map points per object accumulates to 60. When applied to the TUM RGB-D dataset, the proposed approach yields an initialization success ratio of 92.7% when the number of map points per object accumulates to 30, revealing a 16.2% increment compared with the SOTA using an RGB-D camera. Finally, we integrate the initialization method into a simultaneous localization and mapping system.