We introduce HOT3D, a publicly available dataset for egocentric hand and object tracking in 3D. The dataset offers over 833 minutes (3.7M+ images) of recordings that feature 19 subjects interacting with 33 diverse rigid objects. In addition to simple pick-up, observe, and put-down actions, the subjects perform actions typical for a kitchen, office, and living room environment. The recordings include multiple synchronized data streams containing egocentric multi-view RGB/monochrome images, eye gaze signal, scene point clouds, and 3D poses of cameras, hands, and objects. The dataset is recorded with two headsets from Meta: Project Aria, which is a research prototype of AI glasses, and Quest 3, a virtual-reality headset that has shipped millions of units. Ground-truth poses were obtained by a motion-capture system using small optical markers attached to hands and objects. Hand annotations are provided in the UmeTrack and MANO formats, and objects are represented by 3D meshes with PBR materials obtained by an in-house scanner. In our experiments, we demonstrate the effectiveness of multi-view egocentric data for three popular tasks: 3D hand tracking, model-based 6DoF object pose estimation, and 3D lifting of unknown in-hand objects. The evaluated multi-view methods, whose benchmarking is uniquely enabled by HOT3D, significantly outperform their single-view counterparts.
Accurate tracking of a user's body pose while wearing a virtual reality (VR), augmented reality (AR) or mixed reality (MR) headset is a prerequisite for authentic self-expression, natural social presence, and intuitive user interfaces. Existing body tracking approaches on VR/AR devices are either under-constrained, e.g., attempting to infer full body pose from only headset and controller pose, or require impractical hardware setups that place cameras far from a user's face to improve body visibility. In this paper, we present the first controller-less egocentric body tracking solution that runs on an actual VR device using the same cameras that are used for SLAM tracking. We propose a novel egocentric tracking architecture that models the temporal history of body motion using multi-view latent features. Furthermore, we release the first large-scale real-image dataset for egocentric body tracking, EgoBody3M, with a realistic VR headset configuration and diverse subjects and motions. Benchmarks on the dataset shows that our approach outperforms other state-of-the-art methods in both accuracy and smoothness of the resulting motion. We perform ablation studies on our model choices and demonstrate the method running in realtime on a VR headset. Our dataset with more than 30 h of recordings and 3 million frames will be made publicly available.
Atherosclerosis is an inflammatory disease of the arteries associated with alterations in lipid and other metabolism and is a major cause of cardiovascular disease (CVD). LDL consists of several subclasses with different sizes, densities, and physicochemical compositions. Small dense LDL (sd-LDL) is a subclass of LDL. There is growing evidence that sd-LDL-C is associated with CVD risk, metabolic dysregulation, and several pathophysiological processes. In this study, we present a straightforward membrane device filtration method that can be performed with simple laboratory methods to directly determine sd-LDL in serum without the need for specialized equipment. The method consists of three steps: first, the precipitation of lipoproteins with magnesium harpin; second, the collection of effluent from a 100 nm filter; and third, the quantification of sd-LDL-ApoB in the effluent with an SH-SAW biosensor. There was a good correlation between ApoB values obtained using the centrifugation (y = 1.0411x + 12.96, r = 0.82, n = 20) and filtration (y = 1.0633x + 15.13, r = 0.88, n = 20) methods and commercially available sd-LDL-C assay values. In addition to the filtrate method, there was also a close correlation between sd-LDL-C and ELISA assay values (y = 1.0483x - 4489, r = 0.88, n = 20). The filtration treatment method also showed a high correlation with LDL subfractions and NMR spectra ApoB measurements (y = 2.4846x + 4.637, r = 0.89, n = 20). The presence of sd-LDL-ApoB in the effluent was also confirmed by ELISA assay. These results suggest that this filtration method is a simple and promising pretreatment for use with the SH-SAW biosensor as a rapid in vitro diagnostic (IVD) method for predicting sd-LDL concentrations. Overall, we propose a very sensitive and specific SH-SAW biosensor with the ApoB antibody in its sensitive region to monitor sd-LDL levels by employing a simple delay-time phase shifted SH-SAW device. In conclusion, based on the demonstration of our study, the SH-SAW biosensor could be a strong candidate for the future measurement of sd-LDL.
Hands are the primary means through which humans interact with the world. Reliable and always-available hand pose inference could yield new and intuitive control schemes for human-computer interactions, particularly in virtual and augmented reality. Computer vision is effective but requires one or multiple cameras and can struggle with occlusions, limited field of view, and poor lighting. Wearable wrist-based surface electromyography (sEMG) presents a promising alternative as an always-available modality sensing muscle activities that drive hand motion. However, sEMG signals are strongly dependent on user anatomy and sensor placement, and existing sEMG models have required hundreds of users and device placements to effectively generalize. To facilitate progress on sEMG pose inference, we introduce the emg2pose benchmark, the largest publicly available dataset of high-quality hand pose labels and wrist sEMG recordings. emg2pose contains 2kHz, 16 channel sEMG and pose labels from a 26-camera motion capture rig for 193 users, 370 hours, and 29 stages with diverse gestures - a scale comparable to vision-based hand pose datasets. We provide competitive baselines and challenging tasks evaluating real-world generalization scenarios: held-out users, sensor placements, and stages. emg2pose provides the machine learning community a platform for exploring complex generalization problems, holding potential to significantly enhance the development of sEMG-based human-computer interactions.
Sepsis is a life-threatening condition that can progress to septic shock as the body's extreme response to pathogenesis damages its own vital organs. Staphylococcus aureus (S. aureus) accounts for 50% of nosocomial infections, which are clinically treated with antibiotics. However, methicillin-resistant strains (MRSA) have emerged and can withstand harsh antibiotic treatment. To address this problem, curcumin (CCM) is employed to prepare carbonized polymer dots (CPDs) through mild pyrolysis. Contrary to curcumin, the as-formed CCM-CPDs are highly biocompatible and soluble in aqueous solution. Most importantly, the CCM-CPDs induce the release of neutrophil extracellular traps (NETs) from the neutrophils, which entrap and eliminate microbes. In an MRSA-induced septic mouse model, it is observed that CCM-CPDs efficiently suppress bacterial colonization. Moreover, the intrinsic antioxidative, anti-inflammatory, and anticoagulation activities resulting from the preserved functional groups of the precursor molecule on the CCM-CPDs prevent progression to severe sepsis. As a result, infected mice treated with CCM-CPDs show a significant decrease in mortality even through oral administration. Histological staining indicates negligible organ damage in the MRSA-infected mice treated with CCM-CPDs. It is believed that the in vivo studies presented herein demonstrate that multifunctional therapeutic CPDs hold great potential against life-threatening infectious diseases. The biocompatible curcumin-derived carbonized polymer dots (CCM-CPDs) stimulate neutrophils to release traps that capture and eradicate harmful microbes. Tested on mice with MRSA-induced sepsis, CCM-CPDs notably curtailed bacterial growth, prevents severe septic developments, and significantly reduces mortality rates. Histological analysis reveals minimal organ damage in treated mice. Harnessing the power of natural phytochemicals to create carbonized nanomaterials to boost immune responses, offers promising potential against severe infectious diseases, marking a significant stride in combating antibiotic-resistant pathogens.image
The task of inputting text within virtual reality has attracted significant research attention over the last five years. Less well explored is the related task of correcting inputted text when errors are made. This is despite the fact that considerable time and frustration stems from efforts to correct text. In this paper, we bridge this gap in prior research and explore efficient methods for supporting text input correction in virtual reality. We present a characterization of the types and frequencies of errors encountered when inputting text in virtual reality and an analysis of effective editing strategies. We also present the results of a user study evaluating the performance and usability trade-offs for several interaction methods leveraging the unique capabilities of modern head-mounted displays.
While passive surfaces offer numerous benefits for interaction in mixed reality, reliably detecting touch input solely from head-mounted cameras has been a long-standing challenge. Camera specifics, hand self-occlusion, and rapid movements of both head and fingers introduce considerable uncertainty about the exact location of touch events. Existing methods have thus not been capable of achieving the performance needed for robust interaction. In this paper, we present a real-time pipeline that detects touch input from all ten fingers on any physical surface, purely based on egocentric hand tracking. Our method TouchInsight comprises a neural network to predict the moment of a touch event, the finger making contact, and the touch location. TouchInsight represents locations through a bivariate Gaussian distribution to account for uncertainties due to sensing inaccuracies, which we resolve through contextual priors to accurately infer intended user input. We first evaluated our method offline and found that it locates input events with a mean error of 6.3 mm, and accurately detects touch events (F1 = 0.99) and identifies the finger used (F1 = 0.96). In an online evaluation, we then demonstrate the effectiveness of our approach for a core application of dexterous touch input: two-handed text entry. In our study, participants typed 37.0 words per minute with an uncorrected error rate of 2.9% on average.
We introduce HOT3D, a publicly available dataset for egocentric hand and object tracking in 3D. The dataset offers over 833 minutes (more than 3.7M images) of multi-view RGB/monochrome image streams showing 19 subjects interacting with 33 diverse rigid objects, multi-modal signals such as eye gaze or scene point clouds, as well as comprehensive ground truth annotations including 3D poses of objects, hands, and cameras, and 3D models of hands and objects. In addition to simple pick-up/observe/put-down actions, HOT3D contains scenarios resembling typical actions in a kitchen, office, and living room environment. The dataset is recorded by two head-mounted devices from Meta: Project Aria, a research prototype of light-weight AR/AI glasses, and Quest 3, a production VR headset sold in millions of units. Ground-truth poses were obtained by a professional motion-capture system using small optical markers attached to hands and objects. Hand annotations are provided in the UmeTrack and MANO formats and objects are represented by 3D meshes with PBR materials obtained by an in-house scanner. We aim to accelerate research on egocentric hand-object interaction by making the HOT3D dataset publicly available and by co-organizing public challenges on the dataset at ECCV 2024. The dataset can be downloaded from the project website: https://facebookresearch.github.io/hot3d/.
Text input is a critical component of any general purpose computing system, yet efficient and natural text input remains a challenge in AR and VR. Headset based hand-tracking has recently become pervasive among consumer VR devices and affords the opportunity to enable touch typing on virtual keyboards. We present an approach for decoding touch typing on uninstrumented flat surfaces using only egocentric camera-based hand-tracking as input. While egocentric hand-tracking accuracy is limited by issues like self occlusion and image fidelity, we show that a sufficiently diverse training set of hand motions paired with typed text can enable a deep learning model to extract signal from this noisy input. Furthermore, by carefully designing a closed-loop data collection process, we can train an end-to-end text decoder that accounts for natural sloppy typing on virtual keyboards. We evaluate our work with a user study (n=18) showing a mean online throughput of 42.4 WPM with an uncorrected error rate (UER) of 7% with our method compared to a physical keyboard baseline of 74.5 WPM at 0.8% UER, showing progress towards unlocking productivity and high throughput use cases in AR/VR.
AR/VR devices have started to adopt hand tracking, in lieu of controllers, to support user interaction. However, today’s hand input rely primarily on one gesture: pinch. Moreover, current mappings of hand motion to use cases like VR locomotion and content scrolling involve more complex and larger arm motions than joystick or trackpad usage. STMG increases the gesture space by recognizing additional small thumb-based microgestures from skeletal tracking running on a headset. We take a machine learning approach and achieve a 95.1% recognition accuracy across seven thumb gestures performed on the index finger surface: four directional thumb swipes (left, right, forward, backward), thumb tap, and fingertip pinch start and pinch end. We detail the components to our machine learning pipeline and highlight our design decisions and lessons learned in producing a well generalized model. We then demonstrate how these microgestures simplify and reduce arm motions for hand-based locomotion and scrolling interactions.
The fight against hand, foot, and mouth disease (HFMD) remains an arduous challenge without existing point-of-care (POC) diagnostic platforms for accurate diagnosis and prompt case quarantine. Hence, the purpose of this salivary biomarker discovery study is to set the fundamentals for the realization of POC diagnostics for HFMD. Whole salivary proteome profiling was performed on the saliva obtained from children with HFMD and healthy children, using a reductive dimethylation chemical labeling method coupled with high-resolution mass spectrometry-based quantitative proteomics technology. We identified 19 upregulated (fold change = 1.5-5.8) and 51 downregulated proteins (fold change = 0.1-0.6) in the saliva samples of HFMD patients in comparison to that of healthy volunteers. Four upregulated protein candidates were selected for dot blot-based validation assay, based on novelty as biomarkers and exclusions in oral diseases and cancers. Salivary legumain was validated in the Singapore (n = 43 healthy, 28 HFMD cases) and Taiwan (n = 60 healthy, 47 HFMD cases) cohorts with an area under the receiver operating characteristic curve of 0.7583 and 0.8028, respectively. This study demonstrates the feasibility of a broad-spectrum HFMD POC diagnostic test based on legumain, a virus-specific host systemic signature, in saliva.
Convolutional neural network inference on video input is computationally expensive and requires high memory bandwidth. Recently, DeltaCNN [26] managed to reduce the cost by only processing pixels with significant updates over the previous frame. However, DeltaCNN relies on static camera input. Moving cameras add new challenges in how to fuse newly unveiled image regions with already processed regions efficiently to minimize the update rate - without increasing memory overhead and without knowing the camera extrinsics of future frames. In this work, we propose MotionDeltaCNN, a sparse CNN inference framework that supports moving cameras. We introduce spherical buffers and padded convolutions to enable seamless fusion of newly unveiled regions and previously processed regions – without increasing memory footprint. Our evaluation shows that we outperform DeltaCNN by up to 90% for moving camera videos.
Point-of-care testing (POCT), also known as on-site or near-patient testing, has been exploding in the last 20 years. A favorable POCT device requires minimal sample handling (e.g., finger-prick samples, but plasma for analysis), minimal sample volume (e.g., one drop of blood), and very fast results. Shear horizontal surface acoustic wave (SH-SAW) biosensors have attracted a lot of attention as one of the effective solutions to complete whole blood measurements in less than 3 min, while providing a low-cost and small-sized device. This review provides an overview of the SH-SAW biosensor system that has been successfully commercialized for medical use. Three unique features of the system are a disposable test cartridge with an SH-SAW sensor chip, a mass-produced bio-coating, and a palm-sized reader. This paper first discusses the characteristics and performance of the SH-SAW sensor system. Subsequently, the method of cross-linking biomaterials and the analysis of SH-SAW real-time signals are investigated, and the detection range and detection limit are presented.
Recent studies have suggested that BK polyomavirus (BKPyV) may be associated with the development of urothelial carcinoma. In Merkel cell carcinoma, TAg and tAg are the major viral proteins of Merkel cell polyomavirus with oncogenic potential. In this study, we aimed to distinguish the role of TAg and tAg in cell migration. Our result demonstrated that ERK was phosphorylated in human renal tubular cells expressing its TAg and tAg after BKPyV infection. Treatment with the ERK inhibitor U0126 suppressed BKPyV gene expression and reduced BKPyV replication. Both TAg and tAg induced cell migration via ERK-dependent signaling. Furthermore, the expression of TAg and tAg had a significant regulatory effect on focal adhesion molecules in renal proximal tubular cells, which strongly suggests that alterations in the focal adhesion complexes are critically involved in TAg and tAg-induced cell migration. Gelatin zymography profiling revealed that TAg regulates the expression and activity of MMP-2 and MMP-9, but not tAg. Interestingly, TAg regulates the expression and activity of MMP-9 through ERK signaling, whereas MMP-2 is regulated through an ERK-independent pathway. Unbalanced ERK pathway activity is frequently observed in many cancers, while MMP proteins are usually overexpressed in aggressive tumors. These findings support the view that BKPyV is an oncogenic virus.
Enterovirus A71 (EV-A71) is transmitted through the respiratory tract, gastrointestinal system, and fecal-oral routes. The main symptoms caused by EV-A71 are hand, foot, and mouth disease (HFMD) or vesicular sore throat. Upf1 (Up-frameshift protein 1) was reported to degrade mRNA containing early stop codons, known as nonsense-mediated decay (NMD). Upf1 is also involved in the NMD mechanism as a host factor detrimental to viral replication. In this study, we dissected the potential roles of Upf1 in the EV-A71-infected cells. Upf1 was virulently down-regulated in three different EV-A71-infected cells, RD, Hela, and 293T, implying that Upf1 is a host protein unfavorable for EV-A71 replication. Knockdown of Upf1 protein resulted in increased viral RNA expression and production of progeny virus, and conversely, overexpression of Upf1 protein resulted in decreased viral RNA expression and production of progeny virus. Importantly, we observed increased RNA levels of asparagine synthetase (ASNS), one of the indicator substrates for the NMD mechanism, which indirectly suggests that EV-A71 infection of cells suppresses NMD activity in the host. The results shown in this study are useful for subsequent analysis of the relationship between the NMD/Upf1 mechanism and other picornaviruses, which may lead to the development of anti-picornavirus drugs.
Integrated hand-tracking on modern virtual reality (VR) headsets can be readily exploited to deliver mid-air virtual input surfaces for text entry. These virtual input surfaces can closely replicate the experience of typing on a Qwerty keyboard on a physical touchscreen, thereby allowing users to leverage their pre-existing typing skills. However, the lack of passive haptic feedback, unconstrained user motion, and potential tracking inaccuracies or observability issues encountered in this interaction setting typically degrades the accuracy of user articulations. We present a comprehensive exploration of error-tolerant probabilistic hand-based input methods to support effective text input on a mid-air virtual Qwerty keyboard. Over three user studies we examine the performance potential of hand-based text input under both gesture and touch typing paradigms. We demonstrate typical entry rates in the range of 20 to 30 wpm and average peak entry rates of 40 to 45 wpm.
Real-time tracking of 3D hand pose in world space is a challenging problem and plays an important role in VR interaction. Existing work in this space are limited to either producing root-relative (versus world space) 3D pose or rely on multiple stages such as generating heatmaps and kinematic optimization to obtain 3D pose. Moreover, the typical VR scenario, which involves multi-view tracking from wide field of view (FOV) cameras is seldom addressed by these methods. In this paper, we present a unified end-to-end differentiable framework for multi-view, multi-frame hand tracking that directly predicts 3D hand pose in world space. We demonstrate the benefits of end-to-end differentiabilty by extending our framework with downstream tasks such as jitter reduction and pinch prediction. To demonstrate the efficacy of our model, we further present a new large-scale egocentric hand pose dataset that consists of both real and synthetic data. Experiments show that our system trained on this dataset handles various challenging interactive motions, and has been successfully applied to real-time VR applications.
To prevent the COVID-19 pandemic that threatens human health, vaccination has become a useful and necessary tool in the response to the pandemic. The vaccine not only induces antibodies in the body, but may also cause adverse effects such as fatigue, muscle pain, blood clots, and myocarditis, especially in patients with chronic disease. To reduce unnecessary vaccinations, it is becoming increasingly important to monitor the amount of anti-SARS-CoV-2 S protein antibodies prior to vaccination. A novel SH-SAW biosensor, coated with SARS-CoV-2 spike protein, can help quantify the amount of anti-SARS-CoV-2 S protein antibodies with 5 μL of finger blood within 40 s. The LoD of the spike-protein-coated SAW biosensor was determined to be 41.91 BAU/mL, and the cut-off point was determined to be 50 BAU/mL (Youden’s J statistic = 0.94733). By using the SH-SAW biosensor, we found that the total anti-SARS-CoV-2 S protein antibody concentrations spiked 10–14 days after the first vaccination (p = 0.0002) and 7–9 days after the second vaccination (p = 0.0116). Furthermore, mRNA vaccines, such as Moderna or BNT, could achieve higher concentrations of total anti-SARS-CoV-2 S protein antibodies compared with adenovirus vaccine, AZ (p < 0.0001). SH-SAW sensors in vitro diagnostic systems are a simple and powerful technology to investigate the local prevalence of COVID-19.
Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101 is the first multi-view action dataset, with simultaneous static (8) and egocentric (4) recordings. Sequences are annotated with more than 100K coarse and 1M fine-grained action segments, and 18M 3D hand poses. We benchmark on three action understanding tasks: recognition, anticipation and temporal segmentation. Additionally, we propose a novel task of detecting mistakes. The unique recording format and rich set of annotations allow us to investigate generalization to new toys, cross-view transfer, long-tailed distributions, and pose vs. appearance. We envision that Assembly101 will serve as a new challenge to investigate various activity understanding problems.
We propose a method for estimating the 6DoF pose of a rigid object with an available 3D model from a single RGB image. Unlike classical correspondence-based methods which predict 3D object coordinates at pixels of the input image, the proposed method predicts 3D object coordinates at 3D query points sampled in the camera frustum. The move from pixels to 3D points, which is inspired by recent PIFu-style methods for 3D reconstruction, enables reasoning about the whole object, including its (self-)occluded parts. For a 3D query point associated with a pixel-aligned image feature, we train a fully-connected neural network to predict: (i) the corresponding 3D object coordinates, and (ii) the signed distance to the object surface, with the first defined only for query points in the surface vicinity. We call the mapping realized by this network as Neural Correspondence Field. The object pose is then robustly estimated from the predicted 3D-3D correspondences by the Kabsch-RANSAC algorithm. The proposed method achieves state-of-the-art results on three BOP datasets and is shown superior especially in challenging cases with occlusion. The project website is at: linhuang17.github.io/NCF .