Splatting-based algorithms reconstruct photorealistic, real-time-renderable, and mesh-exportable 3D scenes from regular images, but they represent a scene as a single monolithic field. Therefore, the reconstruction has no object-level structure, leaving it infeasible for downstream editing or interaction. Moreover, regions that are never directly observed in the input scans are contaminated by the surrounding texture and left uncorrected, capping both mesh fidelity and novel-view synthesis. We propose a decompose-before-reconstruct approach: we segment the instances out of every frame, consider the remaining as background and inpaint it, reconstruct each instance and the background independently with mesh splatting, and compose them into a single scene. Our method significantly improves mesh fidelity (over a 5% gain in F-score) and novel-view synthesis, while supporting object-wise modifiability and interactivity. The code will be made publicly available.
When developing an application for production, end-user questionnaires are a rapid way of acquiring user evaluations about the product’s usability. These questionnaires include topics such as the base of the application, user interface, and various ailments related to the application, to say the least. Multiple types of questionnaires and methods have been developed to assist developers with feedback acquisition. However, with such a wide range of disciplines, there is not a standard used by all developers in the research community. Furthermore, as applications are delivered using various devices, such as mobile, desktop, and now XR, new questionnaires have arisen. Compiling these evaluations leads to a large number of questions (100+), a multitude of redundancies, and difficulty comparing related applications due to lack of a solidified system. Furthermore, XR development requires unique feedback that typical presentations of data and engagement do not contain. In this work, we create a concise subset of questions, on a 5-point Likert scale, that remove redundancies and allows for comprehensive statistical analysis, allowing developers to measure differences in populations more readily.
With the rise of AI-generated content (AIGC) and advanced techniques for 3D human representation, the task of generating 3D dance movements has become an exciting area of research. Despite significant advancements, current methods often fail to provide comprehensive and distinct control over various multimodal inputs from users, such as music or specific descriptions of desired movements. As a result, the generated motions may be statistically plausible and technically correct, but they often lack depth, expressiveness, and alignment with the user's creative vision. To address this issue, we present CustomDance, a coarse-to-fine interactive system designed for customized 3D dance generation. Inspired by the workflows of expert choreographers, CustomDance introduces a novel paradigm to AI-assisted choreography through three interconnected stages. First, a multimodal Large Language Model (MLLM) analyzes the music and a high-level text prompt to identify key temporal anchors and creative cues for the piece. Next, for each anchor, a multimodal retriever suggests high-quality motion clips from a dance library based on local music and text, empowering the user with concrete and predictable options. Finally, a custom music-conditioned diffusion in-painter seamlessly connects the selected phrases, allowing for iterative, user-guided refinement of the final composition, supported by visualizations of motion dynamics. Our evaluations demonstrate that CustomDance not only highlights the significant creative utility and empowering potential of our AI-assisted choreography paradigm, but also outperforms competitive baselines across quantitative and qualitative comparisons. Project page: https://github.com/XulongT/CustomDance
Remote collaboration systems based on physical environments face several critical challenges, including data-heavy virtual representations and high latencies during data acquisition, reconstruction, rendering, and transmission. Existing approaches often suffer from significant latency, making them unsuitable for real-time collaboration, rely on static scenes that limit interaction, and require multiple specialized hardware, restricting accessibility. To address these challenges, we present Collaborative Interactive Dynamic Environments for co-located users and Virtual Reality (VR) for remote participants through a fully automated pipeline that replicates entire physical environments. CIDER dynamically transforms a user's physical space into an interactive virtual environment, shareable with remote collaborators within seconds. It employs an efficient approach to represent, render, distribute, and synchronize virtual scenes, achieving interaction latencies of 0.22 seconds, about 10 times lower than comparable systems (2.4 seconds). We evaluate CIDER's performance quantitatively with collaboration-oriented metrics in scenarios where participants are separated by up to 12,000 km. We also conducted a questionnaire-based user study with 17 participants to evaluate usability and overall user experience. Furthermore, CIDER allows collaborators to participate using a broad range of devices, including personal computers (via Unity emulators, functioning similarly to a MR/VR device), MR devices (e.g., HoloLens 2), and VR devices (e.g., Meta Quest 2 and 3), enhancing accessibility and usability for diverse user groups.
Point cloud stands as the most widely adopted format for representing 3D shapes and scenes due to its simplicity and geometric fidelity. However, its inherent unordered and irregular nature, exacerbated by sensor noise and occlusions, introduces unique challenges for machine learning based methodologies. To combat these issues, diverse strategies have been developed, including converting to a format that has orderliness, extracting local geometry, and permutation-invariant or self-attention-based processing. In this paper, our focus is directed towards deep learning models for three fundamental tasks in 3D vision: point cloud classification, part segmentation, and semantic segmentation. We begin by formally defining point cloud data, followed by an in-depth discussion on its structural characteristics. Then, we categorize notable works based on their backbone structure and evaluate their performance on popular benchmarks. Beyond empirical comparison, we offer insights into architectural innovations and limitations. We also outline open challenges and promising future directions for 3D point cloud understanding.
The growing presence of large language models (LLMs) and connected sensors or the Internet of Things (IoT) offers substantial potential across various tasks and fields, albeit with accompanying challenges. On the other hand, as advances in spatial computing redefine immersive interactions, ambient intelligence becomes essential for enabling environments that intuitively respond to user behavior without explicit commands. By fusing real-time sensory input with LLM-driven reasoning, ambient intelligence transforms spatially embedded systems into adaptive, context-aware, and human-centered digital experiences. In fact, according to reports,1 two prominent technological trends of 2025 are spatial computing and ambient invisible intelligence.
Abstract BackgroundCurrent telemedicine technologies are not fully optimized for conducting physical examinations. The Virtual Remote Tele-Physical Examination (VIRTEPEX) system, a novel proprietary technology platform using a Microsoft Kinect-based augmented reality game system to track motion and estimate force, has the potential to assist with conducting asynchronous, remote musculoskeletal examinations. ObjectiveThis pilot study evaluated the feasibility of the VIRTEPEX system as a supplement to telehealth musculoskeletal strength assessments. MethodsIn this cross-sectional pilot study, 12 study participants with upper extremity pain and/or weakness underwent strength evaluations for four upper extremity movements using in-person, telehealth, VIRTEPEX, and composite (telehealth plus VIRTEPEX) assessments. The evaluators were blinded to each other’s assessments. The primary outcome was feasibility, as determined by participant recruitment, study completion, and safety. The secondary outcome was preliminary evaluation of inter-rater agreement between in-person, telehealth, and VIRTEPEX strength assessments, including κ statistics. ResultsThis pilot study had an 80% recruitment rate, a 100% completion rate, and reported no adverse events. In-person and telehealth evaluations achieved highest overall agreement (85.71%), followed by agreements between in-person and composite (75%), in-person and VIRTEPEX (62.5%), and telehealth and VIRTEPEX (62.5%) evaluations. However, for shoulder flexion, agreement between in-person and VIRTEPEX evaluations (78.57%; κ=0.571, 95% CI 0.183 to 0.960) and in-person and composite evaluations (78.57%; κ=0.571, 95% CI 0.183 to 0.960) was higher than that between in-person and telehealth evaluations (71.43%; κ=0.429, 95% CI −0.025 to 0.882). ConclusionsThis study demonstrates the feasibility of asynchronous VIRTEPEX examinations and supports the potential for VIRTEPEX to supplement and add value to standard telehealth platforms. Further studies with an additional development of VIRTEPEX and larger sample sizes for adequate power are warranted.
Precise and effective processing of cardiac imaging data is critical for the identification and management of the cardiovascular diseases. We introduce IntelliCardiac, a comprehensive, web-based medical image processing platform for the automatic segmentation of 4D cardiac images and disease classification, utilizing an AI model trained on the publicly accessible ACDC dataset. The system, intended for patients, cardiologists, and healthcare professionals, offers an intuitive interface and uses deep learning models to identify essential heart structures and categorize cardiac diseases. The system supports analysis of both the right and left ventricles as well as myocardium, and then classifies patient's cardiac images into five diagnostic categories: dilated cardiomyopathy, myocardial infarction, hypertrophic cardiomyopathy, right ventricular abnormality, and no disease. IntelliCardiac combines a deep learning-based segmentation model with a two-step classification pipeline. The segmentation module gains an overall accuracy of 92.6%. The classification module, trained on characteristics taken from segmented heart structures, achieves 98% accuracy in five categories. These results exceed the performance of the existing state-of-the-art methods that integrate both segmentation and classification models. IntelliCardiac, which supports real-time visualization, workflow integration, and AI-assisted diagnostics, has great potential as a scalable, accurate tool for clinical decision assistance in cardiac imaging and diagnosis.
We introduce a novel grasp representation named the Unified Gripper Coordinate Space (UGCS) for grasp synthesis and grasp transfer. Our representation leverages spherical coordinates to create a shared coordinate space across different robot grippers, enabling it to synthesize and transfer grasps for both novel objects and previously unseen grippers. The strength of this representation lies in the ability to map palm and fingers of a gripper in the unified coordinate space. Grasp synthesis is formulated as predicting the unified spherical coordinates on object surface points via a conditional variational autoencoder. The predicted unified gripper coordinates establish exact correspondences between the gripper and object points, which is used to optimize grasp pose and joint values. Grasp transfer is facilitated through the point-to-point correspondence between any two (potentially unseen) grippers and solved via a similar optimization. Extensive simulation and real-world experiments showcase the efficacy of the unified grasp representation for grasp synthesis in generating stable and diverse grasps. Similarly, we showcase real-world grasp transfer from human demonstrations across different objects.(1)
To effectively display visual information in Mixed Reality (MR), it is essential to understand how various representations of virtual objects influence users' perceptions and decision-making. The perceived depth of virtual objects is crucial for understanding their spatial relationships with one another and with the physical environment. In contrast to real-world objects, which typically present multiple depth cues, virtual objects often have limited depth cues. The current research reported an experiment with 59 participants that investigated how color, luminance, and rendered distance of virtual objects in a HoloLens 2 MR environment affected human depth perception of virtual objects. Results indicated that objects with high luminance tended to be perceived closer than low luminance ones, and objects with cool colors (green and blue) tended to be perceived closer than those with warm colors (red and yellow), which contrasts significantly with the results of previous studies.
GANs are a class of machine learning framework that are used to generate new data instances that resemble the training data. First proposed by Goodfellow et al.,1 the GAN architecture (Figure 1) typically consists of two separate competing adversarial neural networks that learn from each other. The two neural networks in a GAN model are a generator model, which creates new synthetic data samples, and a discriminator model, which evaluates the synthetic samples generated by the generator model against real data samples. The evaluation of synthetic data by the discriminator helps the generator create better and more realistic, accurate samples. As the generator improves, so does the discriminator, enabling effective learning with the aim of producing synthetic data that are so realistic that the discriminator cannot tell if they are real or fake.
Motion parallax refers to the perceived difference in motion of objects at varying distances from the observer; objects that are closer appear to move faster, and vice versa. This depth cue is crucial for humans in judging distance in physical surroundings. However, in virtual environments, the limited field of view of display devices makes it difficult for humans to perceive the same level of motion parallax as they can in the physical environment. Recent studies have aimed to improve users gaining motion parallax for a more realistic virtual environment, but there is currently no standardized method to conclude a large amount of data for evaluating and comparing their proposed techniques. In this paper, we first propose an improved formula with normalization as a motion parallax evaluation standard and evaluate it by conducting motion parallax experiments using a Microsoft HoloLens 2 device. The experiments consisted of a standard one with different relative moving distances and distance reporting methods (blind walking and oral reporting), and another using an X-ray vision system that helps users see through physical obstacles, potentially providing motion parallax. The aim was to assess the effectiveness of our proposed standard and evaluate whether the proposed conditions influence users' ability to estimate distance by utilizing motion parallax.
We present CIS2VR (CNN-based Indoor Scan to VR), an authoring framework designed to transform input RGB- D scans captured by conventional sensors into an interactive VR environment. Existing state-of-the-art 3D instance segmentation algorithms are employed to extract object instances from RGB- D scans. A novel 3D Convolutional Neural Network (3D CNN) architecture is used to learn 3D shape features common to both classification and 3D pose estimation problems, enabling rapid shape encoding and pose estimation of objects detected in the scan. The generated embedding vector and predicted pose are then used to retrieve and align a matching 3D CAD (Computer-Aided-Design) model. The aligned models, along with the estimated layout of the scene, are transferred to Unity, a 3D game engine, to create a VR scene. An optional human-in-the- loop system allows users to validate results at various steps of the pipeline, improving the quality of the final VR scene. We evaluate and compare our approach to existing semantic reconstruction methods on key metrics. The proposed approach outperforms several existing methods in object alignment, coming close to the state-of-the-art, while speeding up the process an order of magnitude. CIS2VR takes an average of 0.68 seconds for the entire conversion across our test dataset of 312 scenes. The code for the proposed framework will be made publicly available on GitHub.
Mixed Reality (MR) [ 19 , 20 ] has developed rapidly in recent years and is used to potentially improve human living environments (such as life-related and entertainment applications) and work efficiency. Microsoft HoloLens [ 17 ] has played an essential role in the progress of MR as a state-of-the-art head-mounted device (HMD) from the first generation to the second generation. By incorporating multiple sensors, such as depth and RGB cameras, along with the Inertial Measurement Unit (IMU), the tracking and sensing accuracy has significantly improved from the previous generation. However, currently, there are no studies evaluating the accuracy of HoloLens 2 sensors, which makes it difficult for researchers to compare the device with others and improve it. In this paper, we systematically evaluate the sensors utilized in the HoloLens 2, including RGB, eye, depth cameras, and microphone array to show the progress compared to the previous HoloLens generation HMD. Based on the evaluation results for most of the available features in the HoloLens 2, we provide discussion and suggestions for future researchers to design applications (such as entertainment [ 5 , 7 ], education [ 12 , 25 , 27 ] and training [ 23 ]) or research works with existing capabilities.
We present a new reproducible benchmark for evaluating robot manipulation in the real world, specifically focusing on pick-and-place. Our benchmark uses the YCB objects, a commonly used dataset in the robotics community, to ensure that our results are comparable to other studies. Additionally, the benchmark is designed to be easily reproducible in the real world, making it accessible to researchers and practitioners. We also provide our experimental results and analyzes for model-based and model-free 6D robotic grasping on the benchmark, where representative algorithms are evaluated for object perception, grasping planning, and motion planning. We believe that our benchmark will be a valuable tool for advancing the field of robot manipulation. By providing a standardized evaluation framework, researchers can more easily compare different techniques and algorithms, leading to faster progress in developing robot manipulation methods.
With the advent of 3D humanoid reconstruction techniques, using a realistic 3D human avatar in a serious game has become popular. This realistic representation in the virtual environment could be achieved using a single low-cost RGB-D camera such as Kinect. However, properly setting up such a camera system for high-quality rendering can be challenging due to the relatively restricted in-home environment. In this paper, we address the challenge of finding optimized camera setup guidelines for an in-home first-person perspective mixed reality (MR) gaming system. We use an MR system with personalized humanoids to simulate the texture reconstruction for a user under a specific camera configuration. Then, a derivative-free optimization is leveraged as a black box approach to search for the optimized camera setup through iterative simulation. We also introduce a novel skeleton-based calibration to address the effects of physically varying the camera's position. For evaluation, two experiments are carried out to evaluate the correctness and effectiveness of the proposed calibration. Furthermore, we conduct a case study using simulation-based optimization for reconstructing lower limb amputees. This work can potentially help locate a proper camera setup for an MR system within the constraints of an in-home environment.
We introduce a large-scale dataset named MultiGripperGrasp for robotic grasping. Our dataset contains 30.4M grasps from 11 grippers for 345 objects. These grippers range from two-finger grippers to five-finger grippers, including a human hand. All grasps in the dataset are verified in the robot simulator Isaac Sim to classify them as successful and unsuccessful grasps. Additionally, the object fall-off time for each grasp is recorded as a grasp quality measurement. Furthermore, the grippers in our dataset are aligned according to the orientation and position of their palms, allowing us to transfer grasps from one gripper to another. The grasp transfer significantly increases the number of successful grasps for each gripper in the dataset. Our dataset is useful to study generalized grasp planning and grasp transfer across different grippers. 1
Motion analysis is used for several applications related to exercise, movements, entertainment, and therapy. In-person clinical evaluations limit the flexibility of both the physician and the patients in terms of time and location. On the other hand, remote healthcare, such as telehealth options, provides patients the ability to receive healthcare without the need for in-person evaluation, increasing availability. Past research has been conducted as proof of concept of using virtual, joint tracking tools, such as RGB-D and RGB equipment, to assist with portability and low-cost solutions for computational motion analysis. Our work aims to ensure that a physician can record custom, lightweight abstract representation of movement data that can be stored in a library and later be retrieved by a patient for reference. To demonstrate that our system is device agnostic given the compact nature of data representation, we choose to present a prototype using Kinect V2 for RGB-D input and Google's MediaPipe for RGB input. We utilize the system to capture motions by both a physician and patient, and calculate the Range of Motion for multiple exercises using either KinectV2 or a web cam based on the nature of the input. Our project acts as an easy-to-use system, allowing customization of motion plans and virtual range of motion feedback.