To increase the realism of fluid simulations, interactions between different fluids (such as water and air, or water and oil) must be considered. This allows for the simulation of phenomena such as the ‘glugging’ effect seen when a filled bottle is turned upside-down and air must fill the space left by the liquid. A common approach to fluid animation is to use Particle-in-Cell (PIC) methods. However, such methods are known to lose fluid volume over time due to accumulating numerical errors, an issue that is exacerbated when simulating multiple interacting fluids. Implicit Density Projection (IDP) is used to overcome volume preservation issues for PIC methods. However, it is formulated for single fluids. To address this, we present two novel extensions to IDP: generalised IDP and D1-IDP. Generalised IDP extends IDP to multiple fluids. D1-IDP then further improves on the volume preservation capabilities. We show that D1-IDP is particularly good in multiphase fluid simulations with complex fluid interfaces. We also demonstrate its applicability to multiphase simulations involving variable density fluids. D1-IDP is able to achieve a maximum volume error of <1.0% for the majority of presented scenarios, while having a negligible impact on computation performance compared to generalised IDP. The code used to generate the results presented in this paper can be found at https://github.com/robden820/Multiphase_Fluids .
Recent progress in text-to-3D object generation enables the synthesis of detailed geometry from text input by leveraging 2D diffusion models and differentiable 3D representations. However, the approaches often suffer from limited controllability and texture ambiguity due to the limitation of the text modality. To address this, we present SIC3D, a controllable image-conditioned text-to-3D generation pipeline with 3D Gaussian Splatting (3DGS). There are two stages in SIC3D. The first stage generates the 3D object content from text with a text-to-3DGS generation model. The second stage transfers style from a reference image to the 3DGS. Within this stylization stage, we introduce a novel Variational Stylized Score Distillation (VSSD) loss to effectively capture both global and local texture patterns while mitigating conflicts between geometry and appearance. A scaling regularization is further applied to prevent the emergence of artifacts and preserve the pattern from the style image. Extensive experiments demonstrate that SIC3D enhances geometric fidelity and style adherence, outperforming prior approaches in both qualitative and quantitative evaluations.
Artistic style transfer is concerned with the generation of imagery that combines the content of an image with the style of an artwork. In the realm of computer games, most work has focused on post-processing video frames. Some recent work has integrated style transfer into the game pipeline, but it is limited to single styles. Integrating an arbitrary style transfer method into the game pipeline is challenging due to the memory and speed requirements of games. We present PQDAST, the first solution to address this. We use a perceptual quality-guided knowledge distillation framework and train a compressed model using the FLIP evaluator, which substantially reduces both memory usage and processing time with limited impact on stylisation quality. For better preservation of depth and fine details, we utilise a synthetic dataset with depth and temporal considerations during training. The developed model is injected into the rendering pipeline to further enforce temporal stability and avoid diminishing post-process effects. Quantitative and qualitative experiments demonstrate that our approach achieves superior performance in temporal consistency, with comparable style transfer quality, to state-of-the-art image, video and in-game methods.
The digitization of natural history specimens has unlocked opportunities for large-scale phenotypic trait analysis. In recent years, deep learning has shown significant results in accurately predicting annotations on 2D specimen photographs. However, it can be challenging for biologists without extensive related expertise to easily use deep learning. Here, we introduce PhenoLearn, a toolkit developed for biologists to generate annotations on 2D specimen images using deep learning. PhenoLearn integrates graphical user interfaces (GUIs) within its two main modules, PhenoLabel for image annotation and PhenoTrain for model training and prediction. GUIs increase accessibility and reduce the need for computational expertise, allowing biologists to intuitively go through a workflow of labelling training sets, using deep learning, and reviewing predictions in the same tool. We demonstrate PhenoLearn's capabilities through a case study involving the segmentation of plumage areas on bird images, showcasing prediction accuracy and the running time with and without graphics processing unit, highlighting its potential to generate annotations with minimal computational cost and time. The toolkit's modular design and flexibility ensure adaptability, allowing for integration with other tools amidst rapidly evolving deep learning approaches. PhenoLearn bridges the gap between specimen digitization and downstream analysis, providing biologists with broader access to deep learning. The source code, installation guides, tutorials with screenshots, and a small demo dataset for PhenoLearn can be found at https://github.com/echanhe/phenolearn.
Artistic Neural Style Transfer (NST) has achieved remarkable success for images. However, this is not the case for dynamic 3D environments, such as computer games, where temporal coherence remains a challenge. Our paper presents an approach that uses the G-buffer information available in a game pipeline to generate robust and temporally consistent in-game artistic stylizations based on a style reference image. We use a synthetic dataset created from open-source computer games and demonstrate that the utilization of depth, normals, and edge information enables the stylization process to be more aware of the geometric and semantic aspects of a game scene. The proposed approach builds on previous work by injecting style transfer in the rendering pipeline, while also utilizing G-buffer information during inference time to improve upon the stability of the stylizations, offering a controllable way to stylize computer games in terms of temporal coherence and content preservation. Qualitative and quantitative evaluations of our in-game stylization network demonstrate significantly higher temporal stability compared to existing style transfer approaches when stylizing 3D computer games.
The development of automatic methods for early cognitive impairment (CI) detection has a crucial role to play in helping people obtain suitable treatment and care. Video-based analysis offers a promising, low-cost alternative to resource-intensive clinical assessments. This paper investigates visual features (eye blink rate (EBR), head turn rate (HTR), and head movement statistical features (HMSFs)) for distinguishing between neurodegenerative disorders (NDs), mild cognitive impairment (MCI), functional memory disorders (FMDs), and healthy controls (HCs). Following prior work, we improve the multiple thresholds (MTs) approach specifically for EBR calculation to enhance performance and robustness, while the HTR and HMSFs are extracted using methods from previous work. The EBR, HTR, and HMSFs are evaluated using an in-the-wild video dataset captured in challenging environments. This method leverages clinically validated cues and automatically extracts features to enable classification. Experiments show that the proposed approach achieves competitive performance in distinguishing between ND, MCI, FMD, and HCs on in-the-wild datasets, with results comparable to audiovisual-based methods conducted in a lab-controlled environment. The findings highlight the potential of visual-based approaches to complement existing diagnostic tools and provide an efficient home-based monitoring system. This work advances the field by addressing traditional limitations and offering a scalable, cost-effective solution for early detection.
The field of Neural Style Transfer (NST) has witnessed remarkable progress in the past few years, with approaches being able to synthesize artistic and photorealistic images and videos of exceptional quality. To evaluate such results, a diverse landscape of evaluation methods and metrics is used, including authors' opinions based on side-by-side comparisons, human evaluation studies that quantify the subjective judgements of participants, and a multitude of quantitative computational metrics which objectively assess the different aspects of an algorithm's performance. However, there is no consensus regarding the most suitable and effective evaluation procedure that can guarantee the reliability of the results. In this review, we provide an in-depth analysis of existing evaluation techniques, identify the inconsistencies and limitations of current evaluation methods, and give recommendations for standardized evaluation practices. We believe that the development of a robust evaluation framework will not only enable more meaningful and fairer comparisons among NST methods but will also enhance the comprehension and interpretation of research findings in the field.
Lip motion accuracy is important for speech intelligibility, especially for users who are hard of hearing or second language learners. A high level of realism in lip movements is also required for the game and film production industries. 3D morphable models (3DMMs) have been widely used for facial analysis and animation. However, factors that could influence their use in facial animation, such as the differences in facial features between recorded real faces and animated synthetic faces, have not been given adequate attention. This paper investigates the mapping between real speakers and similar and non-similar 3DMMs and the impact on the resulting 3D lip motion. Mouth height and mouth width are used to determine face similarity. The results show that mapping 2D videos of real speakers with low mouth heights to 3D heads that correspond to real speakers with high mouth heights, or vice versa, generates less good 3D lip motion. It is thus important that such a mismatch is considered when using a 2D recording of a real actor's lip movements to control a 3D synthetic character.
Liquid-fabric interaction simulations using particle-in-cell (PIC) based models have been used to simulate a wide variety of phenomena and yield impressive visual results. However, these models suffer from numerical damping due to the data interpolation between the particles and grid. Our paper addresses this by using the polynomial PIC (PolyPIC) model instead of the affine PIC (APIC) model that is used in current state-of-the-art wet cloth models. The affine transfers of the APIC model are replaced by the higher order polynomials of PolyPIC, thus reducing numerical dissipation and improving resolution of vorticial details. This improved energy preservation enables more dynamic simulations to be generated although this is at an increased computational cost.
Neural Style Transfer (NST) research has been applied to images, videos, 3D meshes and radiance fields, but its application to 3D computer games remains relatively unexplored. Whilst image and video NST systems can be used as a post-processing effect for a computer game, this results in undesired artefacts and diminished post-processing effects. Here, we present an approach for injecting depth-aware NST as part of the 3D rendering pipeline. Qualitative and quantitative experiments are used to validate our in-game stylisation framework. We demonstrate temporally consistent results of artistically stylised game scenes, outperforming state-of-the-art image and video NST methods.
Close interaction with robots in Human-Robot Collaboration (HRC) can increase worker productivity in production, but cages around the robot often limit this. Our research aims to visualise virtual safety zones around a real robot arm with Augmented Reality (AR), thereby replacing the cages. We tested our system with a collaborative pick-and-place application which mimics a real manufacturing scenario in an industrial robot cell. The shape, size and visualisation of the AR safety zones were tested with 19 participants. The overwhelming preference was for a visualisation that used cylindrical AR safety zones together with a virtual cage bars effect.
A key challenge in mobilising growing numbers of digitised biological specimens for scientific research is finding high-throughput methods to extract phenotypic measurements on these datasets. In this paper, we test a pose estimation approach based on Deep Learning capable of accurately placing point labels to identify key locations on specimen images. We then apply the approach to two distinct challenges that each requires identification of key features in a 2D image: (i) identifying body region-specific plumage colouration on avian specimens and (ii) measuring morphometric shape variation in Littorina snail shells. For the avian dataset, 95% of images are correctly labelled and colour measurements derived from these predicted points are highly correlated with human-based measurements. For the Littorina dataset, more than 95% of landmarks were accurately placed relative to expert-labelled landmarks and predicted landmarks reliably captured shape variation between two distinct shell ecotypes (‘crab’ vs ‘wave’). Overall, our study shows that pose estimation based on Deep Learning can generate high-quality and high-throughput point-based measurements for digitised image-based biodiversity datasets and could mark a step change in the mobilisation of such data. We also provide general guidelines for using pose estimation methods on large-scale biological datasets.
The use of robot arms in various industrial settings has changed the way tasks are completed. However, safety concerns for both humans and robots in these collaborative environments remain a critical challenge. Traditional approaches to visualising safety zones, including physical barriers and warning signs, may not always be effective in dynamic environments or where multiple robots and humans are working simultaneously. Mixed reality technologies offer dynamic and intuitive visualisations of safety zones in real time, with the potential to overcome these limitations. In this study, we compare the effectiveness of safety zone visualisations in virtual and real robot arm environments using the Microsoft HoloLens 2. We tested our system with a collaborative pick-and-place application that mimics a real manufacturing scenario in an industrial robot cell. We investigated the impact of safety zone shape, size, and appearance in this application. Visualisations that used virtual cage bars were found to be the most preferred safety zone configuration for a real robot arm. However, the results for this aspect were mixed for a virtual robot arm experiment. These results raise the question of whether or not safety visualisations can initially be tested in a virtual scenario and the results transferred to a real robot arm scenario, which has implications for the testing of trust and safety in human–robot collaboration environments.
Early detection of dementia has attracted much research interest due to its crucial role in helping people get suitable treatment or care. Video analysis may provide an effective approach for detection, with low cost and effort compared to current expensive and intensive clinical assessments. This paper investigates the use of a range of visual features - eye blink rate (EBR), head turn rate (HTR) and head movement statistical features (HMSF) - for identifying neurodegenerative disorder (ND), mild cognitive impairment (MCI) and functional memory disorder (FMD). These features are used in a noval multiple thresholds approach, which is applied to an in-the-wild video dataset which includes data recorded in a range of challenging environments. A combination of EBR and HTR gives 78 % accuracy in a three-way classification task (ND/MCI/FMD) and 83%, 83% and 92%, respectively, for the two-way classifications ND/MCI, ND/FMD and MCI/FMD. These results are comparable to related work that uses more features from different modalities. They also provide evidence to support the possibility of an in-the-home detection process for dementia or cognitive impairment.
Temporal consistency and content preservation are the prominent challenges in artistic video style transfer. To address these challenges, we present a technique that utilizes depth data and we demonstrate this on real-world videos from the web, as well as on a standard video dataset of three-dimensional computer-generated content. Our algorithm employs an image-transformation network combined with a depth encoder network for stylizing video sequences. For improved global structure preservation and temporal stability, the depth encoder network encodes ground-truth depth information which is fused into the stylization network. To further enforce temporal coherence, we employ ConvLSTM layers in the encoder, and a loss function based on calculated depth information for the output frames is also used. We show that our approach is capable of producing stylized videos with improved temporal consistency compared to state-of-the-art methods whilst also successfully transferring the artistic style of a target painting.
Railway networks systems are by design open and accessible to people, but this presents challenges in the prevention of events such as terrorism, trespass, and suicide fatalities. With the rapid advancement of machine learning, numerous computer vision methods have been developed in closed-circuit television (CCTV) surveillance systems for the purposes of managing public spaces. These methods are built based on multiple types of sensors and are designed to automatically detect static objects and unexpected events, monitor people, and prevent potential dangers. This survey focuses on recently developed CCTV surveillance methods for rail networks, discusses the challenges they face, their advantages and disadvantages and a vision for future railway surveillance systems. State-of-the-art methods for object detection and behaviour recognition applied to rail network surveillance systems are introduced, and the ethics of handling personal data and the use of automated systems are also considered.
Neural Style Transfer (NST) is concerned with the artistic stylization of visual media. It can be described as the process of transferring the style of an artistic image onto an ordinary photograph. Recently, a number of studies have considered the enhancement of the depth-preserving capabilities of the NST algorithms to address the undesired effects that occur when the input content images include numerous objects at various depths. Our approach uses a deep residual convolutional network with instance normalization layers that utilizes an advanced depth prediction network to integrate depth preservation as an additional loss function to content and style. We demonstrate results that are effective in retaining the depth and global structure of content images. Three different evaluation processes show that our system is capable of preserving the structure of the stylized results while exhibiting style-capture capabilities and aesthetic qualities comparable or superior to state-of-the-art methods. Project page: https://ioannoue.github.io/depth-aware-nst-using-in.html.
Ultraviolet colouration is thought to be an important form of signalling in many bird species, yet broad insights regarding the prevalence of ultraviolet plumage colouration and the factors promoting its evolution are currently lacking. In this paper, we develop a image segmentation pipeline based on deep learning that considerably outperforms classical (i.e. non deep learning) segmentation methods, and use this to extract accurate information on whole-body plumage colouration from photographs of >24,000 museum specimens covering >4500 species of passerine birds. Our results demonstrate that ultraviolet reflectance, particularly as a component of other colours, is widespread across the passerine radiation but is strongly phylogenetically conserved. We also find clear evidence in support of the role of light environment in promoting the evolution of ultraviolet plumage colouration, and a weak trend towards higher ultraviolet plumage reflectance among bird species with ultraviolet rather than violet-sensitive visual systems. Overall, our study provides important broad-scale insight into an enigmatic component of avian colouration, as well as demonstrating that deep learning has considerable promise for allowing new data to be brought to bear on long-standing questions in ecology and evolution.
Revealing the functioning of compound eyes is of interest to biologists and engineers alike who wish to understand how visually complex behaviours (e.g. detection, tracking, and navigation) arise in nature, and to abstract concepts to develop novel artificial sensory systems. A key investigative method is to replicate the sensory apparatus using artificial systems, allowing for investigation of the visual information that drives animal behaviour when exposed to environmental cues. To date, 'compound eye models' (CEMs) have largely explored features such as field of view and angular resolution, but the role of shape and overall structure have been largely overlooked due to modelling complexity. Modern real-time ray-tracing technologies are enabling the construction of a new generation of computationally fast, high-fidelity CEMs. This work introduces a new open-source CEM software (CompoundRay) that is capable of accurately rendering the visual perspective of bees (6000 individual ommatidia arranged on 2 realistic eye surfaces) at over 3000 frames per second. We show how the speed and accuracy facilitated by this software can be used to investigate pressing research questions (e.g. how low resolution compound eyes can localise small objects) using modern methods (e.g. machine learning-based information exploration).
Investigating automatic methods for the early detection of dementia and related conditions that cause cognitive impairment is an area of growing interest. Video processing could play a role by providing a non-invasive and low-cost alternative to current expensive assessments. For this to be successful it is crucial that approaches are robust to in-the-wild challenges. In this paper, visual cues, related to the eye blink rate (EBR), are investigated to quantify the early phase of neurodegenerative disorder (ND) and mild cognitive impairment (MCI) as well as functional memory disorder (FMD; problems with memory not related to neurodegenerative disorder). This paper aims to improve the detection of ND and MCI by investigating a novel approach to calculating the EBR that is more robust to in-the-wild challenges. An in-house dataset with 18 participants is used. The EBR is calculated from eye landmarks extracted using two libraries (Dlib and Openface). To mitigate issues observed in the noisy, in-the-wild recordings, a multiple threshold approach for EBR detection is proposed. It involves generating multiple thresholds for identifying a blink, where a threshold is used to determine whether an eye is open or closed. Several supervised machine learning approaches are used for automatic classification. The results show that accuracy measures of 89% and 78% are achieved using Dlib and OpenFace data, respectively, when distinguishing between three conditions with ND, MCI and FMD.