An accurate fuel consumption prediction system for transportation units is crucial for efficient fuel management, offering both cost reduction and emission savings. While extensive research has been conducted on fuel prediction for modes like airplanes, trucks, and vehicles, studies on cargo ships are scarce and often rely on traditional machine learning models. The complexity of real-world factors, such as data collection challenges and varying weather conditions, adds to the difficulty of accurate prediction. This paper addresses these challenges by comparing traditional machine learning algorithms with advanced deep learning models for predicting fuel consumption in ship engines. Our comparative study shows that LSTM-GRU hybrid models emerge as particularly effective, capturing the intricate dependencies and variabilities inherent in fuel consumption forecasting. The results underscore the superior capability of deep learning models, particularly LSTM-GRU, over traditional regression techniques in managing the complexities of fuel consumption in cargo ships.
Autonomous pavement compaction requires centimeter-level localization, yet RTK-GNSS and pre-surveyed maps add complexity and become unreliable under signal occlusion. This paper presents a mapless framework for autonomous rollers that combines LiDAR geometric understanding with relative localization. Road boundaries are extracted in real time from LiDAR point clouds using a lightweight bird's-eye-view segmentation network. Parametric curve fitting estimates lateral position and heading, while ultra-wideband ranging provides longitudinal distance to the paver. These relative states support a multi-roller collaboration strategy for coordinated compaction without global maps. A dedicated simulation platform enables large-scale evaluation across diverse road geometries, and the system is validated in simulation and field tests with multiple double-drum rollers. Results show centimeter-level localization comparable to RTK-GNSS, robust operation in GNSS-degraded environments, and uniform coverage through cooperative planning. The framework offers a practical and scalable foundation for autonomous collaborative roller operation on real construction sites.
Jointly recovering explicit surface geometry and high-quality appearance from multi-view images remains challenging. This capability is essential for maintaining high-fidelity real-to-sim environments for embodied intelligence, where local changes should be incorporated without complete reconstruction. Existing neural surface reconstruction and 3DGS-to-mesh pipelines often learn geometry indirectly or separate geometry construction from appearance modeling. This separation introduces optimization redundancy and makes local geometry or appearance updates expensive. We propose an end-to-end mesh-Gaussian scene representation that binds 3D Gaussians to mesh faces and uses differentiable 3DGS rendering for photometric supervision. This design provides a direct information pathway for jointly learning explicit geometry and renderable appearance. Experiments on indoor and outdoor scenes demonstrate improved efficiency and rendering quality while preserving high-quality surface reconstruction. The explicit mesh also enables mesh-based manipulation, and the coupled representation adapts efficiently to local scene modifications. These properties support scalable visual scene modeling and the efficient maintenance of real-to-sim environments for embodied-agent training and evaluation.
Natural mathematical objects for representing spatially distributed physical attributes are 3D field functions, which are prevalent in applied sciences and engineering, including areas such as fluid dynamics and computational geometry. The representations of these objects are task-oriented, which are achieved using various techniques that are suitable for specific areas. A recent breakthrough involves using flexible parameterized representations, particularly through neural networks, to model a range of field functions. This technique aims to uncover fields for computational vision tasks, such as representing light-scattering fields. Its effectiveness has led to rapid advancements, enabling the modeling of time dependence in various applications. This survey provides an informative taxonomy of the recent literature in the field of learnable field representation, as well as a comprehensive summary in the application field of visual computing. Open problems in field representation and learning are also discussed, which help shed light on future research.
Human–machine interaction is a critical component in robotic rehabilitation systems. A mutual learning strategy involving both machine- and human-oriented learning has shown improvements in learning efficiency and receptiveness. Despite these advancements, a theoretical framework that encompasses high-level human responses during robot-assisted rehabilitation is still needed. This paper introduces a novel human–machine interface that uses a Co-adaptive Markov Decision Process (CaMDP) model based on cooperative multi-agent reinforcement learning. The CaMDP model effectively measures user adaptation to machines, treating the entire rehabilitation process as a collaborative learning experience. It quantifies learning rates at a higher system abstraction level. Policy Iteration in Reinforcement Learning is employed for the cooperative adjustment of Policy Improvement between the human and machine. Simulation studies demonstrate that the proposed new Policy Improvement approach has great potential to address non-stationarity issues and significantly reduce the switching frequency of patients — from 16.6% to 2% in a sample of 120,000 cases. The CaMDP model provides valuable insights into rehabilitation effect prediction and risk avoidance through dual-agent simulation, thereby enhancing the overall performance of the rehabilitation process.
In the electronic nose (e-nose), a stable feature representation of the gas sensor’s response is a key step to realize subsequent odor identification algorithms. However, the noises in gas sensors hinder the acquisition of such features. In order to solve this problem, this article proposes a stable feature extraction algorithm which takes the impulse response of the e-nose system as the feature. The impulse response is estimated from a nonparametric model constrained by a multiscale wavelet kernel regularization matrix. The kernel regularization matrix equips the proposed feature extraction method with an ability in resistance to random noise. A numerical experiment proves that compared with single-scale kernel regularization, the use of multiscale wavelet kernel helps to achieve more stable and accurate impulse response estimation. Then, a field experiment is conducted to demonstrate the performance of the proposed features. This experiment aims to identify four different whiskies measured by a self-designed e-nose with four commercial gas sensors. Under the framework of transfer learning, the classification result based on the proposed features outperforms those using other considered features. The accuracy of whisky identification reaches 92.00%, showing a good potential of applying the proposed feature representations in the area of e-noses.
High-quality estimation of surface normal can help reduce ambiguity in many geometry understanding problems, such as collision avoidance and occlusion inference. This paper presents a technique for estimating the normal from 3D point clouds and 2D colour images. We have developed a transformer neural network that learns to utilise the hybrid information of visual semantic and 3D geometric data, as well as effective learning strategies. Compared to existing methods, the information fusion of the proposed method is more effective, which is supported by experiments. We have also built a simulation environment of outdoor traffic scenes in a 3D rendering engine to obtain annotated data to train the normal estimator. The model trained on synthetic data is tested on the real scenes in the KITTI dataset. And subsequent tasks built upon the estimated normal directions in the KITTI dataset show that the proposed estimator has advantage over existing methods.
Rapid economic development of any country will usually lead to negative environmental impacts. A free market economy cannot fundamentally solve this issue, which requires the guidance and control of the government. The environmental policies of governments can effectively improve the ecological conditions in a region. This study quantifies environmental regulation policies and takes four urban agglomerations in eastern China as the research object to explore the influence of environmental regulation on regional ecological efficiency. First, policies can be divided into policy control, pollution control, ecological protection, and social adjustment by building a Latent Dirichlet Allocation (LDA) model of a policy text library to measure differences in urban policy bias. Next, a slack-based measure (SBM) model was used to measure urban ecological efficiency. Finally, using the qualitative comparison analysis method, the time effect was considered and the ecological efficiency configuration of the entire region was obtained. In addition, contrast analysis was conducted for different path configurations and the reasons of the statuses of the four urban agglomerations were obtained. The results show that policy control, pollution control, ecological protection, and other mandatory policies can significantly reduce the negative environmental effects, especially when the policy controls involve an explicit form of punishment, which is a necessary condition for achieving a high level of ecological efficiency. However, social adjustment means having higher requirements for regional total factor development, which requires improving the environmental awareness of residents and improving corporate social responsibility. Regional differences will mean that areas with a lower level of economic development have a higher intensity of policy control, and the means of policy control should be matched with the level of economic development. In addition, this study proposes some policy suggestions, such as strengthening policy control and pollution prevention, strengthening social regulation, and promulgating environmental laws and regulations according to local conditions.
Obtaining a high-quality frontal face image from a low-resolution (LR) non-frontal face image is primarily important for many facial analysis applications. However, mainstreams either focus on super-resolving near-frontal LR faces or frontalizing non-frontal high-resolution (HR) faces. It is desirable to perform both tasks seamlessly for daily-life unconstrained face images. In this paper, we present a novel Vivid Face Hallucination Generative Adversarial Network (VividGAN) for simultaneously super-resolving and frontalizing tiny non-frontal face images. VividGAN consists of coarse-level and fine-level Face Hallucination Networks (FHnet) and two discriminators, i.e., Coarse-D and Fine-D. The coarse-level FHnet generates a frontal coarse HR face and then the fine-level FHnet makes use of the facial component appearance prior, i.e., fine-grained facial components, to attain a frontal HR face image with authentic details. In the fine-level FHnet, we also design a facial component-aware module that adopts the facial geometry guidance as clues to accurately align and merge the frontal coarse HR face and prior information. Meanwhile, two-level discriminators are designed to capture both the global outline of a face image as well as detailed facial characteristics. The Coarse-D enforces the coarsely hallucinated faces to be upright and complete while the Fine-D focuses on the fine hallucinated ones for sharper details. Extensive experiments demonstrate that our VividGAN achieves photo-realistic frontal HR faces, reaching superior performance in downstream tasks, i.e., face recognition and expression classification, compared with other state-of-the-art methods.
With the continuous development of science and technology, self-driving vehicles will surely change the nature of transportation and realize the automotive industry's transformation in the future. Compared with self-driving cars, self-driving buses are more efficient in carrying passengers and more environmentally friendly in terms of energy consumption. Therefore, it is speculated that in the future, self-driving buses will become more and more important. As a simulator for autonomous driving research, the CARLA simulator can help people accumulate experience in autonomous driving technology faster and safer. However, a shortcoming is that there is no modern bus model in the CARLA simulator. Consequently, people cannot simulate autonomous driving on buses or the scenarios interacting with buses. Therefore, we built a bus model in 3ds Max software and imported it into the CARLA to fill this gap. Our model, namely KIT bus, is proven to work in the CARLA by testing it with the autopilot simulation. The video demo is shown on our Youtube.
Capsule networks (CapsNet) are recently proposed neural network models containing newly introduced processing layer, which are specialized in entity representation and discovery in images. CapsNet is motivated by a view of parse tree-like information processing mechanism and employs an iterative routing operation dynamically determining connections between layers composed of capsule units, in which the information ascends through different levels of interpretations, from raw sensory observation to semantically meaningful entities represented by active capsules. The CapsNet architecture is plausible and has been proven to be effective in some image data processing tasks, the newly introduced routing operation is mainly required for determining the capsules’ activation status during the forward pass. However, its influence on model fitting and the resulted representation is barely understood. In this work, we investigate the following: 1) how the routing affects the CapsNet model fitting; 2) how the representation using capsules helps discover global structures in data distribution, and; 3) how the learned data representation adapts and generalizes to new tasks. Our investigation yielded the results some of which have been mentioned in the original paper of CapsNet, they are: 1) the routing operation determines the certainty with which a layer of capsules pass information to the layer above and the appropriate level of certainty is related to the model fitness; 2) in a designed experiment using data with a known 2D structure, capsule representations enable a more meaningful 2D manifold embedding than neurons do in a standard convolutional neural network (CNN), and; 3) compared with neurons of the standard CNN, capsules of successive layers are less coupled and more adaptive to new data distribution.
In multipath video streaming transmission, the selection of the best vehicle for video packet forwarding considering the junction area is a challenging task due to the several diversions in the junction area. The vehicles in the junction area change direction based on the different diversions, which lead to video packet drop. In the existing works, the explicit consideration of different positions in the junction areas has not been considered for forwarding vehicle selection. To address the aforementioned challenges, a Junction-Aware vehicle selection for Multipath Video Streaming (JA-MVS) scheme has been proposed. The JA-MVS scheme considers three different cases in the junction area including the vehicle after the junction, before the junction and inside the junction area, with an evaluation of the vehicle signal strength based on the signal to interference plus noise ratio (SINR), which is based on the multipath data forwarding concept using greedy-based geographic routing. The performance of the proposed scheme is evaluated based on the Packet Loss Ratio (PLR), Structural Similarity Index (SSIM) and End-to-End Delay (E2ED) metrics. The JA-MVS is compared against two baseline schemes, Junction-Based Multipath Source Routing (JMSR) and the Adaptive Multipath geographic routing for Video Transmission (AMVT), in urban Vehicular Ad-Hoc Networks (VANETs).
This work presents a novel architecture of deep neural networks to generate meshes approximating the surface of a 3D object from a single image. Compared to existing learning-based 3D reconstruction models, our architecture is characterized by (1) deep mesh deformation stacks with residual network design, where a simple mesh is transformed to approximate the target surface and undergoes multiple deformation steps to progressively refine the result and reduce the residuals, and (2) parallel paths per deformation step, which can exponentially enrich the generated meshes using deeper structure and more model parameters. We also propose novel regularization scheme that encourages the meshes to be both globally complementary to cover the target surface and locally consistent with each other. Empirical evaluation on benchmark datasets show advantage of the proposed architecture over existing methods.
Generative adversarial nets (GANs) are effective framework for constructing data models and enjoys desirable theoretical justification. On the other hand, realizing GANs for practical complex data distribution often requires careful configuration of the generator, discriminator, objective function and training method and can involve much non-trivial effort. We propose an novel family of generative adversarial nets (GANs), where we employ both continuous noise and random binary codes in the generating process. The binary codes in the new GAN model (named BGANs) play the role of categorical latent variables helps improve the model capability and training stability when dealing with complex data distributions. BGAN has been evaluated and compared with existing GANs trained with the state-of-the-art method on both synthetic and practical data. The empirical evaluation shows effectiveness of BGAN.
It is common that practical data has multiple attributes of interest. For example, a picture can be characterized in terms of its content, e.g. the categories of the objects in the picture, and in the meanwhile the image style such as photo-realistic or artistic is also relevant. This work is motivated by taking advantage of all available sources of information about the data, including those not directly related to the target of analytics. We propose an explicit and effective knowledge representation and transfer architecture for image analytics by employing Capsules for deep neural network training based on the generative adversarial nets (GAN). The adversarial scheme help discover capsule-representation of data with different semantic meanings in respective dimensions of the capsules. The data representation includes one subset of variables that are particularly specialized for the target task – by eliminating information about the irrelevant aspects. We theoretically show the elimination by mixing conditional distributions of the represented data. Empirical evaluations show the propose method is effective for both standard transfer-domain recognition tasks and zero-shot transfer.
Capsule Networks (CapsNet) are recently proposed multi-stage computational models specialized for entity representation and discovery in image data. CapsNet employs iterative routing that shapes how the information cascades through different levels of interpretations. In this work, we investigate i) how the routing affects the CapsNet model fitting, ii) how the representation by capsules helps discover global structures in data distribution and iii) how learned data representation adapts and generalizes to new tasks. Our investigation shows: i) routing operation determines the certainty with which one layer of capsules pass information to the layer above, and the appropriate level of certainty is related to the model fitness, ii) in a designed experiment using data with a known 2D structure, capsule representations allow more meaningful 2D manifold embedding than neurons in a standard CNN do and iii) compared to neurons of standard CNN, capsules of successive layers are less coupled and more adaptive to new data distribution.
In this research, we propose deep networks that discover Granger causes from multivariate temporal data generated in financial markets. We introduce a Deep Neural Network (DNN) and a Recurrent Neural Network (RNN) that discover Granger-causal features for bivariate regression on bivariate time series data distributions. These features are subsequently used to discover Granger-causal graphs for multivariate regression on multivariate time series data distributions. Our supervised feature learning process in proposed deep regression networks has favourable F-tests for feature selection and t-tests for model comparisons. The experiments, minimizing root mean squared errors in the regression analysis on real stock market data obtained from Yahoo Finance, demonstrate that our causal features significantly improve the existing deep learning regression models.
Recent years have witnessed the success of deep neural networks in dealing with a plenty of practical problems. The invention of effective training techniques largely contributes to this success. The so-called "Dropout" training scheme is one of the most powerful tool to reduce over-fitting. From the statistic point of view, Dropout works by implicitly imposing an L2 regularizer on the weights. In this paper, we present a new training scheme: Shakeout. Instead of randomly discarding units as Dropout does at the training stage, our method randomly chooses to enhance or inverse the contributions of each unit to the next layer. We show that our scheme leads to a combination of L1 regularization and L2 regularization imposed on the weights, which has been proved effective by the Elastic Net models in practice.We have empirically evaluated the Shakeout scheme and demonstrated that sparse network weights are obtained via Shakeout training. Our classification experiments on real-life image datasets MNIST and CIFAR-10 show that Shakeout deals with over-fitting effectively.
Hierarchical neural networks have been shown to be effective in learning representative image features and recognizing object classes. However, most existing networks combine the low/middle level cues for classification without accounting for any spatial structures. For applications such as understanding a scene, how the visual cues are spatially distributed in an image becomes essential for successful analysis. This paper extends the framework of deep neural networks by accounting for the structural cues in the visual signals. In particular, two kinds of neural networks have been proposed. First, we develop a multitask deep convolutional network, which simultaneously detects the presence of the target and the geometric attributes (location and orientation) of the target with respect to the region of interest. Second, a recurrent neuron layer is adopted for structured visual detection. The recurrent neurons can deal with the spatial distribution of visible cues belonging to an object whose shape or structure is difficult to explicitly define. Both the networks are demonstrated by the practical task of detecting lane boundaries in traffic scenes. The multitask convolutional neural network provides auxiliary geometric information to help the subsequent modeling of the given lane structures. The recurrent neural network automatically detects lane boundaries, including those areas containing no marks, without any explicit prior knowledge or secondary modeling.
Distance metric learning (DML) is successful in discovering intrinsic relations in data. However, most algorithms are computationally demanding when the problem size becomes large. In this paper, we propose a discriminative metric learning algorithm, develop a distributed scheme learning metrics on moderate-sized subsets of data, and aggregate the results into a global solution. The technique leverages the power of parallel computation. The algorithm of the aggregated DML (ADML) scales well with the data size and can be controlled by the partition. We theoretically analyze and provide bounds for the error induced by the distributed treatment. We have conducted experimental evaluation of the ADML, both on specially designed tests and on practical image annotation tasks. Those tests have shown that the ADML achieves the state-of-the-art performance at only a fraction of the cost incurred by most existing methods.
Hesham El-Sayed合作论文数United Arab Emirates University;College of Information Technology1