Electroencephalography (EEG) is a pivotal tool for exploring brain functions. However, the low amplitude of EEG signals renders them inherently susceptible to contamination from diverse physiological and environmental artifacts, including electromyogram artifacts, electrocardiogram interference, and electrical noise from power lines. These contaminants significantly hinder the analysis and interpretation of EEG data, posing substantial challenges for signal processing. Recently, deep learning paradigms have catalyzed significant progress in EEG denoising, with many studies reporting competitive reconstruction fidelity and artifact suppression in benchmarked settings. Despite this progress, there remains a notable gap in the literature regarding comprehensive reviews of deep learning-based EEG denoising strategies. To bridge this gap, we use the end-to-end denoising pipeline as an analytical framework, examining how data/target construction, input representation, modular architecture, objective design, and evaluation strategies influence model assumptions, the interpretation of model performance, and practical utility. We further discuss selective and multi-task denoising strategies, downstream validation, and model deployment as key issues for translating reconstruction performance into usable EEG applications. Finally, we identify future research directions aimed at developing more reliable, interpretable, and practically useful EEG denoising systems, thereby enhancing the utility of EEG technologies in broader applications.
3D Gaussian Splatting (3DGS) has enabled photorealistic and real-time rendering of 3D head avatars. Existing 3DGS-based avatars typically rely on tens of thousands of 3D Gaussian points (Gaussians), with the number of Gaussians fixed after training. However, many practical applications require adjustable levels of detail (LOD) to balance rendering efficiency and visual quality. In this work, we propose "ArchitectHead", the first framework for creating 3D Gaussian head avatars that support continuous control over LOD. Our key idea is to parameterize the Gaussians in a 2D UV feature space and propose a UV feature field composed of multi-level learnable feature maps to encode their latent features. A lightweight neural network-based decoder then transforms these latent features into 3D Gaussian attributes for rendering. ArchitectHead controls the number of Gaussians by dynamically resampling feature maps from the UV feature field at the desired resolutions. This method enables efficient and continuous control of LOD without retraining. Experimental results show that ArchitectHead achieves state-of-the-art (SOTA) quality in self and cross-identity reenactment tasks at the highest LOD, while maintaining near SOTA performance at lower LODs. At the lowest LOD, our method uses only 6.2% of the Gaussians while the quality degrades moderately (L1 Loss +7.9%, PSNR −0.97%, SSIM −0.6%, LPIPS Loss +24.1%), and the rendering speed nearly doubles. Project homepage: https://peizhiyan.github.io/docs/architect/.
Recent advancements in 3D Gaussian Splatting (3DGS) have unlocked significant potential for modeling 3D head avatars, providing greater flexibility than mesh-based methods and more efficient rendering compared to NeRF-based approaches. Despite these advancements, the creation of controllable 3DGS-based head avatars remains time-intensive, often requiring tens of minutes to hours. To expedite this process, we here introduce the "Gaussian Deja-vu" framework, which first obtains a generalized model of the head avatar and then personalizes the result. The generalized model is trained on large 2D (synthetic and real) image datasets. This model provides a well-initialized 3D Gaussian head that is further refined using a monocular video to achieve the personalized head avatar. For personalizing, we propose learnable expression-aware rectification blendmaps to correct the initial 3D Gaussians, ensuring rapid convergence without the reliance on neural networks. Experiments demonstrate that the proposed method meets its objectives. It outperforms state-of-the-art 3D Gaussian head avatars in terms of photorealistic quality as well as reduces training time consumption to at least a quarter of the existing methods, producing the avatar in minutes. Project homepage: https://peizhiyan.github.io/docs/dejavu
For 3D face modeling, the recently developed 3D-aware neural rendering methods are able to render photorealistic face images with arbitrary viewing directions. The training of the parametric controllable 3D-aware face models, however, still relies on a large-scale dataset that is lab-collected. To address this issue, this paper introduces "StyleMorpheus", the first style-based neural 3D Morphable Face Model (3DMM) that is trained on in-the-wild images. It inherits 3DMM's disentangled controllability (over face identity, expression, and appearance) but without the need for accurately reconstructed explicit 3D shapes. StyleMorpheus employs an auto-encoder structure. The encoder aims at learning a representative disentangled parametric code space and the decoder improves the disentanglement using shape and appearance-related style codes in the different sub-modules of the network. Furthermore, we fine-tune the decoder through style-based generative adversarial learning to achieve photorealistic 3D rendering quality. The proposed style-based design enables StyleMorpheus to achieve state-of-the-art 3D-aware face reconstruction results, while also allowing disentangled control of the reconstructed face. Our model achieves real-time rendering speed, allowing its use in virtual reality applications. We also demonstrate the capability of the proposed style-based design in face editing applications such as style mixing and color editing. Project homepage: https://github.com/ubc-3d-vision-lab/StyleMorpheus.
3D Face shape stylization refers to transforming a realistic 3D face shape into a different style, such as a cartoon face style. To solve this problem, this paper proposes modeling this task as a deformation transfer problem. This approach significantly reduces labor costs, as the artists would only need to create a single template for each face style. Realistic facial features of the original 3D face e.g. the nose or chin shape, would thus be automatically transferred to those in the style template. Deformation transfer methods, however, have two drawbacks. They are slow and they require re-optimization for every new input face. To address these weaknesses, we propose a neural network-based 3D face shape stylization method. This method is trained through weakly supervised learning, and its template's structure is preserved using our novel template-guided mesh smoothing regularization. Our method is the first learning-based deformation transfer method for 3D face shape stylization. Its employment offers the useful and practical benefit of not requiring paired training data. The experiments show that the quality of the stylized faces obtained by our method is comparable to that of the traditional deformation transfer method, achieving an average Chamfer Distance of approximately 0.01 mm. However, our approach significantly boosts the processing speed, achieving a rate approximately 3,000 times faster than the traditional deformation transfer.
Laryngeal hemiplegia (LH) is a major upper respiratory tract (URT) complication in racehorses. Endoscopy imaging of horse throat is a gold standard for URT assessment. However, current manual assessment faces several challenges, stemming from the poor quality of endoscopy videos and subjectivity of manual grading. To overcome such limitations, we propose an explainable machine learning (ML)-based solution for efficient URT assessment. Specifically, a cascaded YOLOv8 architecture is utilized to segment the key semantic regions and landmarks per frame. Several spatiotemporal features are then extracted from key landmarks points and fed to a decision tree (DT) model to classify LH as Grade 1,2,3 or 4 denoting absence of LH, mild, moderate, and severe LH, respectively. The proposed method, validated through 5-fold cross-validation on 107 videos, showed promising performance in classifying different LH grades with 100%, 91.18%, 94.74% and 100% sensitivity values for Grade 1 to 4, respectively. Further validation on an external dataset of 72 cases confirmed its generalization capability with 90%, 80.95%, 100%, and 100% sensitivity values for Grade 1 to 4, respectively. We introduced several explainability related assessment functions, including: (i) visualization of YOLOv8 output to detect landmark estimation errors which can affect the final classification, (ii) time-series visualization to assess video quality, and (iii) backtracking of the DT output to identify borderline cases. We incorporated domain knowledge (e.g., veterinarian diagnostic procedures) into the proposed ML framework. This provides an assistive tool with clinical-relevance and explainability that can ease and speed up the URT assessment by veterinarians.
This paper proposes an end-to-end framework for generating 3D human pose datasets using Neural Radiance Fields (NeRF). Public datasets generally have limited diversity in terms of human poses and camera viewpoints, largely due to the resource-intensive nature of collecting 3D human pose data. As a result, pose estimators trained on public datasets significantly underperform when applied to unseen out-of-distribution samples. Previous works proposed augmenting public datasets by generating 2D-3D pose pairs or rendering a large amount of random data. Such approaches either overlook image rendering or result in suboptimal datasets for pre-trained models. Here we propose PoseGen, which learns to generate a dataset (human 3D poses and images) with a feedback loss from a given pre-trained pose estimator. In contrast to prior art, our generated data is optimized to improve the robustness of the pre-trained model. The objective of PoseGen is to learn a distribution of data that maximizes the prediction error of a given pre-trained model. As the learned data distribution contains OOD samples of the pre-trained model, sampling data from such a distribution for further fine-tuning a pre-trained model improves the generalizability of the model. This is the first work that proposes NeRFs for 3D human data generation. NeRFs are data-driven and do not require 3D scans of humans. Therefore, using NeRF for data generation is a new direction for convenient user-specific data generation. Our extensive experiments show that the proposed PoseGen improves two baseline models (SPIN and HybrIK) on four datasets with an average 6% relative improvement.
The increasing demand for medical imaging has surpassed the capacity of available radiologists, leading to diagnostic delays and potential misdiagnoses. Artificial intelligence (AI) techniques, particularly in automatic medical report generation (AMRG), offer a promising solution to this dilemma. This review comprehensively examines AMRG methods from 2021 to 2024. It (i) presents solutions to primary challenges in this field, (ii) explores AMRG applications across various imaging modalities, (iii) introduces publicly available datasets, (iv) outlines evaluation metrics, (v) identifies techniques that significantly enhance model performance, and (vi) discusses unresolved issues and potential future research directions. This paper aims to provide a comprehensive understanding of the existing literature and inspire valuable future research.
The increasing demand for medical imaging has surpassed the capacity of available radiologists, leading to diagnostic delays and potential misdiagnoses. Artificial intelligence (AI) techniques, particularly in automatic medical report generation (AMRG), offer a promising solution to this dilemma. This review comprehensively examines AMRG methods from 2021 to 2024. It (i) presents solutions to primary challenges in this field, (ii) explores AMRG applications across various imaging modalities, (iii) introduces publicly available datasets, (iv) outlines evaluation metrics, (v) identifies techniques that significantly enhance model performance, and (vi) discusses unresolved issues and potential future research directions. This paper aims to provide a comprehensive understanding of the existing literature and inspire valuable future research.
is an organization within the framework of the IEEE, with professional interest
Brain–computer interfaces (BCIs) employ neurophysiological signals derived from the brain to control computers or external devices. By enhancing or replacing human peripheral functioning capacity, BCIs offer supplementary degrees of freedom, significantly improving individuals’ quality of life, particularly offering hope for those with locked-in syndrome (LIS). Moreover, BCI applications have expanded across medical and nonmedical domains, including rehabilitation, clinical diagnosis, cognitive and affective computing, and gaming. Over the past decades, with a wealth of brain signals captured invasively or noninvasively, BCI has made spectacular progress. However, this also poses new challenges for signal processing techniques, such as characterization and classification. In this review, we first introduce signal enhancement and characterization methods to mine inherent patterns of nonstationary and time-varying brain signals. Then, we highlight widely adopted classification methods in BCI and the challenges they face. This article aims to comprehensively overview crucial signal processing techniques in BCI and provide suggestions for future directions.
The Signal Processing Society is an organization, within the framework of the IEEE, of members with principal professional interest in the technology of transmission, recording, reproduction, processing, and measurement of speech; other audio-frequency waves and other signals by digital electronic, electrical, acoustic, mechanical, and optical means; the components and systems to accomplish these and related aims; and the environmental, psychological, and physiological factors of these technologies.For membership and subscription information and pricing, please visit www.
Unlike 2D face images, obtaining a 3D face is not easy. Existing methods, therefore, create a 3D face from a 2D face image (3D face reconstruction). A user might wish to edit the reconstructed 3D face, but 3D face editing has seldom been studied. This paper presents such method and shows that reconstruction and editing can help each other. In the presented framework named NEO-3DF, the 3D face model we propose has independent sub-models corresponding to semantic face parts. It allows us to achieve both local intuitive editing and better 3D-to-2D alignment. Each face part in our model has a set of controllers designed to allow users to edit the corresponding features (e.g., nose height). In addition, we propose a differentiable module for blending the face parts and making it possible to automatically adjust the face parts (both the shapes and the locations) so that they are better aligned with the original 2D image. Experiments show that the results of NEO-3DF outperform existing methods in intuitive face editing and have better 3D-to-2D alignment accuracy (14% higher IoU) than global face model-based reconstruction. Code available at https://github.com/ubc-3d-vision-lab/NEO-3DF .
The 3D-aware parametric face model named HeadNeRF achieved advantages in rendering photo-realistic face images. However, it has two limitations: (1) it uses single-image fitting reconstruction that is slow and prone to overfitting; (2) it lacks explicit 3D geometry information, making using semantic facial-parts-based loss challenging. This paper presents a 3D-aware face reconstruction learning framework tailored for HeadNeRF to address the limitations. We train a face encoder network that can directly learn the disentangled features for facial reconstruction to address the first limitation. For the second limitation, we introduce a lightweight semantic face segmentation network and facial-parts-based loss function to improve the reconstruction accuracy and quality. Our experiments show that the proposed method achieves a low reconstruction time consumption and enhanced reconstruction accuracy. Project page: https://peizhiyan.github.io/docs/headnerf+
Shahram Shirani合作论文数Professional Engineers Ontario;The Institute of Electrical and Electronics Engineers (IEEE);UBC Alumni Association10
Maher Ahmed合作论文数Wilfrid Laurier University, Waterloo, Canada N2L 3C57