Motion Capture (MoCap) is still dominated by optical MoCap as it remains the gold standard. However, the raw captured data even from such systems suffer from high-frequency noise and errors sourced from ghost or occluded markers. To that end, a post-processing step is often required to clean up the data, which is typically a tedious and time-consuming process. Some studies tried to address these issues in a data-driven manner, leveraging the availability of MoCap data. However, there is a high-level data redundancy in such data, as the motion cycle is usually comprised of similar poses (e.g. standing still). Such redundancies affect the performance of those methods, especially in the rarer poses. In this work, we address the issue of long-tailed data distribution by leveraging representation learning. We introduce a novel technique for imbalanced regression that does not require additional data or labels. Our approach uses a Mahalanobis distance-based method for automatically identifying rare samples and properly reweighting them during training, while at the same time, we employ high-order interpolation algorithms to effectively sample the latent space of a Variational Autoencoder (VAE) to generate new tail samples. We prove that the proposed approach can significantly improve the results, especially in the tail samples, while at the same time is a model-agnostic method and can be applied across various architectures.
Radiance field-based methods have recently been used to reconstruct human avatars, showing that we can significantly downscale the systems needed for creating animated human avatars. Although this progress has been initiated by neural radiance fields, their slow rendering and backward mapping from the observation space to the canonical space have been the main challenges. With Gaussian splatting overcoming both challenges, a new family of approaches has emerged that are faster to train and render, while also straightforward to implement using forward skinning from the canonical to the observation space. However, the linear blend skinning required for the deformation of the Gaussians does not provide valid results for their non-linear rotation properties. To address such artifacts, recent works use mesh properties to rotate the non-linear Gaussian properties or train models to predict corrective offsets. Instead, we propose a weighted rotation blending approach that leverages quaternion averaging. This leads to simpler vertex-based Gaussians that can be efficiently animated and integrated in any engine by only modifying the linear blend skinning technique, and using any Gaussian rasterizer.
We present a novel marker-based motion capture (MoCap) technology. Instead of leveraging initialization and tracking for marker labeling as traditional solutions do, the present system is built upon real-time and low-latency data-driven models and optimization techniques, offering new possibilities and overcoming limitations currently present in the MoCap landscape. Even though we similarly begin with unlabeled markers captured with optical sensing within a capturing area, our approach diverges as we follow a data-driven and optimization pipeline to simultaneously denoise the markers and robustly and accurately solve the skeleton per frame. Similarly to traditional marker-based options, our work demonstrates higher stability and accuracy than inertial and/or markerless optical MoCap. Inertial MoCap lacks absolute positioning and suffers from drifting, therefore, it is almost impossible to achieve comparable positional accuracy. Markerless solutions lack the existence of a strong prior (i.e., markers) to increase the capturing precision, while, due to the heavier workload, the capturing frequency cannot easily scale, resulting in inaccuracies in fast movements. On the other hand, traditional marker-based motion capture heavily relies on high-quality marker data, assuming precise localization, outlier elimination and consistent marker tracking. In contrast, our innovative approach operates without such assumptions, effectively mitigating input noise, including ghost markers, occlusions, marker swaps, misplacement and mispositioning. This noise tolerance enables our system to function seamlessly with cameras with lower cost and specifications. Our method introduces body structure invariance, empowering automatic marker layout configuration by selecting from a diverse pool of models trained with different marker layouts. Our proposed MoCap technology integrates various consumer-grade optical sensors, leverages efficient data acquisition, succeeds in precise marker position estimation and allows for spatio-temporal alignment of multi-view streams. Sequentially, by incorporating data-driven models, our system achieves low latency and real-time rate performances. Finally, efficient body optimization techniques further improve the final MoCap solving, enabling seamless integration into various applications requiring real-time, accurate and robust motion capture. Concluding, real-time communities can be benefited from our MoCap which is a) affordable; with the use of low-cost equipment, b) scalable; with processing on the edge, c) portable; with easy setup and spatial calibration, d) robust; on heavy occlusions, marker removal and camera coverage and e) flexible; no need for super precise marker placement, super precise camera calibration, body calibration per actor or manual marker configuration.
Real-time optical Motion Capture (MoCap) systems have not benefited from the advances in modern data-driven modeling. In this work we apply machine learning to solve noisy unstructured marker estimates in real-time and deliver robust marker-based MoCap even when using sparse affordable sensors. To achieve this we focus on a number of challenges related to model training, namely the sourcing of training data and their long-tailed distribution. Leveraging representation learning we design a technique for imbalanced regression that requires no additional data or labels and improves the performance of our model in rare and challenging poses. By relying on a unified representation, we show that training such a model is not bound to high-end MoCap training data acquisition, and exploit the advances in marker-less MoCap to acquire the necessary data. Finally, we take a step towards richer and affordable MoCap by adapting a body model-based inverse kinematics solution to account for measurement and inference uncertainty, further improving performance and robustness. Project page: moverseai.github.io/noise-tail.
Human motion capture and perception without the need for complex systems with specialized cameras or wearable equipment is the holy grail for many human-centric applications. Here, we present a scalable markerless motion capture method that estimates 3D human poses in real-time using low-cost hardware. We do so by replacing the inefficient 3D joint reconstruction techniques, such as learnable triangulation and feature splatting, with a novel uncertainty-driven approach that exploits the available depth information and the edge sensors’ spatial alignment to fuse the per viewpoint estimates into final 3D joint positions.
HUMAN4D: A Human-Centric Multimodal Dataset for Motions & Immersive Media (Subject #2) The dataset was captured with the use of VCL Volumetric Capture free software (https://github.com/VCL3D/VolumetricCapture) device_repository.json includes the camera instrinsic parameters. pose.zip includes the camera extrinsic calibration parameters. offsets.zip include the frame offset between the pose ids (name_of_file==id) and the group frame ids of the RGBD data (first number before underscore in the filename of each file) S2_activities.txt files that maps the zip filenames with data for specific activities. HUMAN4D is a large and multimodal 4D dataset that contains a variety of human activities simultaneously captured by a professional marker-based MoCap, a volumetric capture and an audio recording system. By capturing 2 female and 2 male professional actors performing various full-body movements and expressions, HUMAN4D provides a diverse set of motions and poses encountered as part of single- and multi-person daily, physical and social activities (jumping, dancing, etc.), along with multi-RGBD (mRGBD), volumetric and audio data. Despite the existence of multi-view color datasets captured with the use of hardware (HW) synchronization, to the best of our knowledge, HUMAN4D is the first and only public resource that provides volumetric depth maps with high synchronization precision due to the use of intra- and inter-sensor HW-SYNC.
A series of 2D (and 3D) keypoint estimation tasks are built upon heatmap coordinate representation, i.e. a probability map that allows for learnable and spatially aware encoding and decoding of keypoint coordinates on grids, even allowing for sub-pixel coordinate accuracy. In this report, we aim to reproduce the findings of DARK that investigated the 2D heatmap representation by highlighting the importance of the encoding of the ground truth heatmap and the decoding of the predicted heatmap to keypoint coordinates. The authors claim that a) a more principled distribution-aware coordinate decoding method overcomes the limitations of the standard techniques widely used in the literature, and b), that the reconstruction of heatmaps from ground-truth coordinates by generating accurate and continuous heatmap distributions lead to unbiased model training, contrary to the standard coordinate encoding process that quantizes the keypoint coordinates on the resolution of the input image grid.
Since December 2019, the world has been devastated by the Coronavirus Disease 2019 (COVID-19) pandemic. Emergency Departments have been experiencing situations of urgency where clinical experts, without long experience and mature means in the fight against COVID-19, have to rapidly decide the most proper patient treatment. In this context, we introduce an artificially intelligent tool for effective and efficient Computed Tomography (CT)-based risk assessment to improve treatment and patient care. In this paper, we introduce a data-driven approach built on top of volume-of-interest aware deep neural networks for automatic COVID-19 patient risk assessment (discharged, hospitalized, intensive care unit) based on lung infection quantization through segmentation and, subsequently, CT classification. We tackle the high and varying dimensionality of the CT input by detecting and analyzing only a sub-volume of the CT, the Volume-of-Interest (VoI). Differently from recent strategies that consider infected CT slices without requiring any spatial coherency between them, or use the whole lung volume by applying abrupt and lossy volume down-sampling, we assess only the "most infected volume" composed of slices at its original spatial resolution. To achieve the above, we create, present and publish a new labeled and annotated CT dataset with 626 CT samples from COVID-19 patients. The comparison against such strategies proves the effectiveness of our VoI-based approach. We achieve remarkable performance on patient risk assessment evaluated on balanced data by reaching 88.88%, 89.77%, 94.73% and 88.88% accuracy, sensitivity, specificity and F1-score, respectively.
We introduce HUMAN4D, a large and multimodal 4D dataset that contains a variety of human activities simultaneously captured by a professional marker-based MoCap, a volumetric capture and an audio recording system. By capturing 2 female and 2 male professional actors performing various full-body movements and expressions, HUMAN4D provides a diverse set of motions and poses encountered as part of single- and multi-person daily, physical and social activities (jumping, dancing, etc.), along with multi-RGBD (mRGBD), volumetric and audio data. Despite the existence of multi-view color datasets captured with the use of hardware (HW) synchronization, to the best of our knowledge, HUMAN4D is the first and only public resource that provides volumetric depth maps with high synchronization precision due to the use of intra- and inter-sensor HW-SYNC. Moreover, a spatio-temporally aligned scanned and rigged 3D character complements HUMAN4D to enable joint research on time-varying and high-quality dynamic meshes. We provide evaluation baselines by benchmarking HUMAN4D with state-of-the-art human pose estimation and 3D compression methods. We apply OpenPose and AlphaPose reaching 70.02% and 82.95% mAP(PCKh-0.5) on single- and 68.48% and 73.94% mAP(PCKh-0.5) on two-person 2D pose estimation, respectively. In 3D pose, a recent multi-view approach named Learnable Triangulation, achieves 80.26% mAP(PCK3D-10cm). For 3D compression, we benchmark Draco, Corto and CWIPC open-source 3D codecs, respecting online encoding and steady bit-rates between 7-155 and 2-90 Mbps for mesh- and point-based volumetric video, respectively. Qualitative and quantitative visual comparison between mesh-based volumetric data reconstructed in different qualities and captured RGB, showcases the available options with respect to 4D representations. HUMAN4D is introduced to enable joint research on spatio-temporally aligned pose, volumetric, mRGBD and audio data cues. The dataset and its code are available online.
The recent advances in real-time volumetric capturing enable new interaction paradigms in virtual environments. Volumetric representations can elevate the feeling of presence, and therefore increase the user immersion by emplacing them within virtual environments allowing them to occupy space. This research demonstration aims to showcase the possibilities offered by real-time volumetric capturing in an Augmented VR context. More specifically, a novel single-player obstacle avoidance game is presented where the users try to avoid collision with incoming carved walls by moving their actual bodies.
HUMAN4D constitutes a large and multimodal 4D dataset that contains a variety of human activities simultaneously captured by a professional marker-based MoCap, a volumetric capture and an audio recording system.
In this paper, a marker-based, single-person optical motion capture method (DeepMoCap) is proposed using multiple spatio-temporally aligned infrared-depth sensors and retro-reflective straps and patches (reflectors). DeepMoCap explores motion capture by automatically localizing and labeling reflectors on depth images and, subsequently, on 3D space. Introducing a non-parametric representation to encode the temporal correlation among pairs of colorized depthmaps and 3D optical flow frames, a multi-stage Fully Convolutional Network (FCN) architecture is proposed to jointly learn reflector locations and their temporal dependency among sequential frames. The extracted reflector 2D locations are spatially mapped in 3D space, resulting in robust 3D optical data extraction. The subject’s motion is efficiently captured by applying a template-based fitting technique on the extracted optical data. Two datasets have been created and made publicly available for evaluation purposes; one comprising multi-view depth and 3D optical flow annotated images (DMC2.5D), and a second, consisting of spatio-temporally aligned multi-view depth images along with skeleton, inertial and ground truth MoCap data (DMC3D). The FCN model outperforms its competitors on the DMC2.5D dataset using 2D Percentage of Correct Keypoints (PCK) metric, while the motion capture outcome is evaluated against RGB-D and inertial data fusion approaches on DMC3D, outperforming the next best method by 4 . 5 % in total 3D PCK accuracy.
In this supplementary material we complement our original manuscript with additional quantitative and qualitative results, which better showcase the advantages of the proposed self-supervised denoising model over traditional filtering and supervised CNN-based approaches. In particular, we present the adopted implementation details used for training our model, as well as additional qualitative results for the two 3D application experiments presented in the original manuscript, namely 3D scanning with KinectFusion [8] and full-body 3D reconstructions using Poisson 3D surface reconstruction [4]. A comparative evaluation with the learning-based state-of-the-art methods on InteriorNet (IN) [7] follows, while an ablation study concludes the document. The aforementioned results based on all methods presented in the originally manuscript, namely Bilateral Filter (BF [10]), Joint Bilateral Filter (JBF [6]), Rolling Guidance (RGF [12]), and data-driven approaches (DRR [3], DDRNet [11]). Note that for the DRR and DDRNet methods, additional results aim to highlight the over-smoothing effect of the former and the weakness of the latter to denoise depth maps captured by the Intel RealSense D415 sensors. In more detail, DRR is trained on static scenes that contain dominant planar surfaces and, thus tends to flatten (i.e. oversmooth) the input data. On the other hand, the available DDRNet model 1 that we used, produces high levels of flying pixels (i.e. spraying, see Fig. 1) which can be attributed to background (zero depth values) and foreground blending, even though its predictions are appropriately masked. While the authors have not provided the necessary information, it is our speculation that the available model is trained using Kinect 1 data, which is partly supported by the suboptimal results it produces on Kinect 2 data. Qualitatively, the remaining traditional filters (BF, JBF, RGF) perform similarly, with RGF showcasing the most competitive results to our method. However, note that
Depth perception is considered an invaluable source of information for various vision tasks. However, depth maps acquired using consumer-level sensors still suffer from non-negligible noise. This fact has recently motivated researchers to exploit traditional filters, as well as the deep learning paradigm, in order to suppress the aforementioned non-uniform noise, while preserving geometric details. Despite the effort, deep depth denoising is still an open challenge mainly due to the lack of clean data that could be used as ground truth. In this paper, we propose a fully convolutional deep autoencoder that learns to denoise depth maps, surpassing the lack of ground truth data. Specifically, the proposed autoencoder exploits multiple views of the same scene from different points of view in order to learn to suppress noise in a self-supervised end-to-end manner using depth and color information during training, yet only depth during inference. To enforce selfsupervision, we leverage a differentiable rendering technique to exploit photometric supervision, which is further regularized using geometric and surface priors. As the proposed approach relies on raw data acquisition, a large RGB-D corpus is collected using Intel RealSense sensors. Complementary to a quantitative evaluation, we demonstrate the effectiveness of the proposed self-supervised denoising approach on established 3D reconstruction applications. Code is avalable at https://github.com/VCL3D/DeepDepthDenoising
Proteins are macromolecules central to biological processes that display a dynamic and complex surface. They display multiple conformations differing by local (residue side-chain) or global (loop or domain) structural changes which can impact drastically their global and local shape. Since the structure of proteins is linked to their function and the disruption of their interactions can lead to a disease state, it is of major importance to characterize their shape. In the present work, we report the performance in enrichment of six shape-retrieval methods (3D-FusionNet, GSGW, HAPT, DEM, SIWKS and WKS) on a 2 267 protein structures dataset generated for this protein shape retrieval track of SHREC'18.
Proteins are macromolecules central to biological processes that display a dynamic and complex surface. They display multiple conformations differing by local (residue side-chain) or global (loop or domain) structural changes which can impact drastically their global and local shape. Since the structure of proteins is linked to their function and the disruption of their interactions can lead to a disease state, it is of major importance to characterize their shape. In the present work, we report the performance in enrichment of six shape-retrieval methods (3D-FusionNet, GSGW, HAPT, DEM, SIWKS and WKS) on a 2 267 protein structures dataset generated for this protein shape retrieval track of SHREC’18.
BACKGROUND:Exercise-based rehabilitation plays a key role in improving the health and quality of life of patients with Cardiovascular Disease (CVD). Home-based computer-assisted rehabilitation programs have the potential to facilitate and support physical activity interventions and improve health outcomes.OBJECTIVES:We present the development and evaluation of a computerized Decision Support System (DSS) for unsupervised exercise rehabilitation at home, aiming to show the feasibility and potential of such systems toward maximizing the benefits of rehabilitation programs.METHODS:The development of the DSS was based on rules encapsulating the logic according to which an exercise program can be executed beneficially according to international guidelines and expert knowledge. The DSS considered data from a prescribed exercise program, heart rate from a wristband device, and motion accuracy from a depth camera, and subsequently generated personalized, performance-driven adaptations to the exercise program. Communication interfaces in the form of RESTful web service operations were developed enabling interoperation with other computer systems.RESULTS:The DSS was deployed in a computer-assisted platform for exercise-based cardiac rehabilitation at home, and it was evaluated in simulation and real-world studies with CVD patients. The simulation study based on data provided from 10 CVD patients performing 45 exercise sessions in total, showed that patients can be trained within or above their beneficial HR zones for 67.1 ± 22.1% of the exercise duration in the main phase, when they are guided with the DSS. The real-world study with 3 CVD patients performing 43 exercise sessions through the computer-assisted platform, showed that patients can be trained within or above their beneficial heart rate zones for 87.9 ± 8.0% of the exercise duration in the main phase, with DSS guidance.CONCLUSIONS:Computerized decision support systems can guide patients to the beneficial execution of their exercise-based rehabilitation program, and they are feasible.
•A novel framework for real-time motion analysis is proposed.•The devised framework can perform action detection/recognition and evaluation based on motion capture data.•Automatic and dynamic motion data weighting is introduced, altering joint data significance based on action involvement aiming for more efficient action detection and recognition.•Action evaluation is performed and, exploiting fuzzy logic, semantic feedback is automatically retrieved proposing ways of action execution improvement.
Cardiac Rehabilitation (CR) can significantly improve mortality and morbidity rates from Cardiovascular Diseases (CVD). Nevertheless, traditional CR is diminished by low subsequent adherence rates. Thus, in this paper, an e-Health technological module for human motion analysis and user modelling is proposed, in order to address the requirements of unsupervised, tele-rehabilitation systems for CVD, by evaluating and personalizing prescribed physical CR programs. The proposed module consists of a) an exercise capturing and evaluation component, and b) a user modelling and decision support system for personalization of cardiac rehabilitation programs. In particular, the module monitors and analyses the body movements of the patient when exercising in real-time, while based on this analysis and the heart-rate measurements, it is capable of short-term and long-term CR session adaptation. The proposed module constitutes a significant tool for internet-enabled sensor-based home exercise platforms.