Recent advancements in deep learning, high-fidelity simulation, and robotic hardware have propelled significant progress in Physical Artificial Intelligence (AI). This field marks a revolutionary step in the evolution of AI by combining the precision of physical laws with the adaptability of machine learning. In this paper, we review the development of Physical AI and its taxonomy by examining the relevant literature, categorizing it into three sub-domains: Physical-Informed AI, Generative Physical AI, and Embodied AI. These sub-domains primarily tackle scientific and engineering challenges, create physics-plausible scenarios, and enable robots or autonomous vehicles to interact with the physical world. This approach also addresses the questions of how to perceive, generate, and interact with the physical world by integrating physics with AI algorithms. Additionally, we discuss related benchmarks and datasets. Finally, we outline the current challenges and propose potential opportunities for future research.
With the rapid advancements in wearable device technologies, there is a growing interest in learning useful digital biomarkers from wearable device data as objective, low-cost, real-time alternatives to use in healthcare settings. They have the potential to facilitate disease progression monitoring, medication tailoring, and supplementing clinical trial endpoints. For example, triaxial accelerometer sensor data is promising for monitoring symptoms of movement-related diseases, such as tremors in Parkinson's disease (PD). However, existing methods for accelerometer studies based on hidden Markov models (HMM) often analyze each individual's activity data separately, leading to inefficiency and limited generalizability. This paper proposes a joint nonparametric Bayesian method that extends the hierarchical Dirichlet process autoregressive HMM (HDP-AR-HMM) to incorporate subject-specific transition parameters. This approach allows for simultaneous estimation across multiple subjects and repeated measurements, accounts for between-subject variability, and provides consistent hidden state estimation without pre-specifying the number of states. We validate our method on simulated data and show that it can achieve higher accuracy in detecting the true hidden states compared to alternative methods. We apply the method to a free-living study, the Biomarker & Endpoint Assessment to Track Parkinson's disease (BEAT-PD) DREAM Challenge CIS-PD study, to demonstrate its utility in monitoring disease symptoms in PD patients.
Edge-preserving filters, as known as bilateral filters, are fundamental to graphics rendering techniques, providing greater generality and capability of edge preservation than pure convolution filters. However, sampling with a large kernel per pixel for these filters can be computationally intensive in real-time rendering. Existing acceleration methods for approximating edge-preserving filters still struggle to balance blur controllability, edge clarity, and runtime efficiency. In this paper, we propose a novel scheme for approximating edge-preserving filters with large anisotropic kernels by recursively reconstructing them from multi-image pyramid (MIP) layers that are weightedly filtered in a dual 3x3 kernel space. Our approach introduces a concise unified processing pipeline independent of kernel size, which includes upsampling and downsampling on MIP layers and enables the integration of custom edge-stopping functions. We also derive the implicit relations of the sampling weights and formulate a weight template model for inference. Furthermore, we convert the pipeline into a lightweight neural network for numerical solutions through data training. Consequently, our image post-processors achieve high-quality and high-performance edge-preserving filters in real-time, using the same control parameters as the original bilateral filters. These filters are applicable for depth-of-fields, global illumination denoising, and screen-space particle rendering. The simplicity of the reconstruction process in our pipeline makes it user-friendly and cost-effective, saving both runtime and implementation costs.
Creating visualizations of multiple volumetric density fields is demanding in virtual reality (VR) applications, which often include divergent volumetric density distributions mixed with geometric models and physics-based simulations. Real-time rendering of such complex environments poses significant challenges for rendering quality and performance. This article presents a novel scheme for efficient real-time rendering of varying translucent volumetric density fields with global illumination (GI) effects on high-resolution binocular VR displays. Our scheme proposes creative solutions to address three challenges involved in the target problem. First, to tackle the doubled heavy workloads of binocular ray marching, we explore the anti-aliasing principles and more advanced potentials of ray marching on interior cube-map faces, and propose a coupled ray-marching technique that converges to multi-resolution cube maps with interleaved adaptive sampling. Second, we devise a fully dynamic ambient GI approximation method that leverages spherical-harmonics (SH) transform information of the phase function to reduce the huge amount of ray sampling required for GI while ensuring fidelity. The method catalyzes spatial ray-marching reuse and adaptive temporal accumulation. Third, we deploy a two-phase ray-tracing algorithm with a tiled k-buffer to achieve fast processing of order-independent transparency (OIT) for multiple volume instances. Consequently, high-quality and high-performance real-time dynamic volume rendering can be achieved under constrained budgets controlled by developers. As our solution supports mixed mesh-volume rendering, the test results prove the practical usefulness of our approach for high-resolution binocular VR rendering on hybrid multi-volumetric and geometric environments.
Multivariate longitudinal data are frequently encountered in practice such as in our motivating longitudinal microbiome study. It is of general interest to associate such high-dimensional, longitudinal measures with some univariate continuous outcome. However, incomplete observations are common in a regular study design, as not all samples are measured at every time point, giving rise to the so-called blockwise missing values. Such missing structure imposes significant challenges for association analysis and defies many existing methods that require complete samples. In this paper we propose to represent multivariate longitudinal data as a three-way tensor array (i.e., sample-by-feature-by-time) and exploit a parsimonious scalar-on-tensor regression model for association analysis. We develop a regularized covariance-based estimation procedure that effectively leverages all available observations without imputation. The method achieves variable selection and smooth estimation of time-varying effects. The application to the motivating microbiome study reveals interesting links between the preterm infant's gut microbiome dynamics and their neurodevelopment. Additional numerical studies on synthetic data and a longitudinal aging study further demonstrate the efficacy of the proposed method.
Digital technologies (e.g., mobile phones) can be used to obtain objective, frequent, and real-world digital phenotypes from individuals. However, modeling these data poses substantial challenges since observational data are subject to confounding and various sources of variabilities. For example, signals on patients' underlying health status and treatment effects are mixed with variation due to the living environment and measurement noises. The digital phenotype data thus shows extensive variabilities between- and within-patient as well as across different health domains (e.g., motor, cognitive, and speaking). Motivated by a mobile health study of Parkinson's disease (PD), we develop a mixed-response state-space (MRSS) model to jointly capture multi-dimensional, multi-modal digital phenotypes and their measurement processes by a finite number of latent state time series. These latent states reflect the dynamic health status and personalized time-varying treatment effects and can be used to adjust for informative measurements. For computation, we use the Kalman filter for Gaussian phenotypes and importance sampling with Laplace approximation for non-Gaussian phenotypes. We conduct comprehensive simulation studies and demonstrate the advantage of MRSS in modeling a mobile health study that remotely collects real-time digital phenotypes from PD patients.
Chinese calligraphy is one of the excellent expressions of Chinese traditional art. But people without domain knowledge of calligraphy can hardly read, appreciate, or learn this art form, due to it contains many brush strokes with unique shapes and complicate structural topological relationship. In this paper, we explore the solution of text sequence recognition of calligraphy, which is a challenging task because traditional algorithms of text recognition can rarely obtain the satisfied results for the varied styles of calligraphy. Therefore, based on a trainable neural network, this paper proposes an easy recognition method, which combines feature sequence extraction based on DenseNet (Dense Convolutional Network) model, sequence modeling and transcription into a consolidated architecture. Compared with previous algorithms for text recognition, it has two distinctive properties: One is to acquire the artistic features on shapes and structures of different styles of calligraphic characters and the other is to handel sequences in arbitrary lengths without character segmentation. Our experimental results prove that in contrast with several common recognition methods, our method of Chinese calligraphic character in diverse styles demonstrates greater robustness, and the recognition accuracy rate reaches 84.70%.
Electronic health records (EHRs) collected from large‐scale health systems provide rich subject‐specific information on a broad patient population at a lower cost compared to randomized controlled trials. Thus, EHRs may serve as a complementary resource to provide real‐world data to construct individualized treatment rules (ITRs) and achieve precision medicine. However, in the absence of randomization, inferring treatment rules from EHR data may suffer from unmeasured confounding. In this article, we propose a self‐matched learning method inspired by the self‐controlled case series (SCCS) design to mitigate this challenge. We alleviate unmeasured time‐invariant confounding between patients by matching different periods of treatments within the same patient (self‐controlled matching) to infer the optimal ITRs. The proposed method constructs a within‐subject matched value function for optimizing ITRs and bears similarity to the SCCS design. We examine assumptions that ensure Fisher consistency, and show that our method requires weaker assumptions on unmeasured confounding than alternative methods. Through extensive simulation studies, we demonstrate that self‐matched learning has comparable performance to other existing methods when there are no unmeasured confounders, but performs markedly better when unobserved time‐invariant confounders are present, which is often the case for EHRs. Sensitivity analyses show that the proposed method is robust under different scenarios. Finally, we apply self‐matched learning to estimate the optimal ITRs from type 2 diabetes patient EHRs, which shows our estimated decision rules lead to greater advantages in reducing patients' diabetes‐related complications.
We compare two deletion-based methods for dealing with the problem of missing observations in linear regression analysis. One is the complete-case analysis (CC, or listwise deletion) that discards all incomplete observations and only uses common samples for ordinary least-squares estimation. The other is the available-case analysis (AC, or pairwise deletion) that utilizes all available data to estimate the covariance matrices and applies these matrices to construct the normal equation. We show that the estimates from both methods are asymptotically unbiased under missing completely at random (MCAR) and further compare their asymptotic variances in some typical situations. Surprisingly, using more data (i.e., AC) does not necessarily lead to better asymptotic efficiency in many scenarios. Missing patterns, covariance structure and true regression coefficient values all play a role in determining which is better. We further conduct simulation studies to corroborate the findings and demystify what has been missed or misinterpreted in the literature. Some detailed proofs and simulation results are available in the online supplemental materials.
OBJECTIVE:There is long-standing interest in how best to define stages of illness for anorexia nervosa, including remission and recovery. The authors used data from a previously published study to examine the time course of relapse over the year following full weight restoration.METHODS:Following weight restoration in an acute care setting, 93 women with anorexia nervosa were randomly assigned to receive fluoxetine or placebo and were discharged to outpatient care, where they also received cognitive-behavioral therapy for up to 1 year. Relapse was defined on the basis of a priori clinical criteria. Fluoxetine had no impact on the time to relapse. In the present analysis, for each day after entry into the study, the risk of relapse over the following 60 days and the following 90 days was calculated and a parametric function was fitted to approximate the Kaplan-Meier estimator.RESULTS:The risk of relapse rose immediately after entry into the study, reached a peak after approximately 60 days, and then gradually declined. There was no indication of an inflection point at which the risk of relapse fell precipitously after the initial peak.CONCLUSIONS:This analysis highlights the fact that adult patients with anorexia nervosa are at increased risk of relapse in the first months following discharge from acute care, suggesting a need for frequent follow-up and relapse prevention-focused treatment during this period. After approximately 2 months, the risk of relapse progressively decreases over time.
This paper presents a novel approach to anti-aliased ray marching by indirect shading in cube-map space. Our volume renderer firstly performs ray marching on each visible interior pixel of a maximum-resolution-limited cube map, and then resamples (usually up-scales) the cube imposter in viewport space. By this viewport-resolution-independent strategy, developers can improve both ray-marching performance and its quality of anti-aliasing when allowing larger marching strides. Moreover, our solution also covers depth-occlusion anti-aliasing for mixed mesh-volume rendering, cube-map level-of-details (LOD) optimization for a further performance boost, and multiple-volume rendering by leveraging the GPU inline ray tracing. Besides, our implementation is developer-friendly and the performance-quality tradeoff determined by the parameter configuration is easily controllable.
为实现电子化创作中国水墨画,提出一个绘画工具,通过对在照片上粗略勾勒的线条进行风格化,生成具有花卉水墨画风格的艺术笔画.首先用户勾勒线条并自动修正,它们分别作为风格模板的候选笔画和照片上风格化目标的基本路径;然后为每个基本路径选择最佳的风格模板笔画;通过形状计算和纹理合成,将最佳的风格模板映射至基本路径.实验结果表明,与已有的以整幅画为目标的风格迁移方法相比,文中提出的方法可充分展现艺术特征的细节并捕捉用户意图,且生成的绘画作品接近指定的艺术风格,满足普通用户的审美需求.在应用价值方面,该工具支持触摸式设备,可用于艺术教育与实践.
Dimension reduction of high-dimensional microbiome data facilitates subsequent analysis such as regression and clustering. Most existing reduction methods cannot fully accommodate the special features of the data such as count-valued and excessive zero reads. We propose a zero-inflated Poisson factor analysis model in this paper. The model assumes that microbiome read counts follow zero-inflated Poisson distributions with library size as offset and Poisson rates negatively related to the inflated zero occurrences. The latent parameters of the model form a low-rank matrix consisting of interpretable loadings and low-dimensional scores that can be used for further analyses. We develop an efficient and robust expectation-maximization algorithm for parameter estimation. We demonstrate the efficacy of the proposed method using comprehensive simulation studies. The application to the Oral Infections, Glucose Intolerance, and Insulin Resistance Study provides valuable insights into the relation between subgingival microbiome and periodontal disease.
Generally, it is regarded as challenge work to grasp the drawing style of an ancient masterpiece in Chinese painting learning. This paper presents a novel approach to the generation of a Chinese ink painting in a certain style and animating its brushwork process with expert skills. In order to demonstrate the techniques of brush and ink inside a stroke, a serials of geometric properties of a brush stroke, are first extracted, then through rational deformation calculation, the best stroke source is mapped onto the stroking path, which is sketched by the user, and finally a new Chinese painting can be synthesized by style migration and natural stroke composition. So with the generated strokes, the lifelike brushwork process of the new painting can be represented dramatically. Actually, by showing the authentic painting process, the tool we implemented helps the learners, who have no profound skills and knowledge in domain of Chinese painting, master the essence of a great painting style, and also provides an easy way to art creation and the comprehension of mysterious Chinese traditional art.
For mental disorders, patients' underlying mental states are non-observed latent constructs which have to be inferred from observed multi-domain measurements such as diagnostic symptoms and patient functioning scores. Additionally, substantial heterogeneity in the disease diagnosis between patients needs to be addressed for optimizing individualized treatment policy in order to achieve precision medicine. To address these challenges, we propose an integrated learning framework that can simultaneously learn patients' underlying mental states and recommend optimal treatments for each individual. This learning framework is based on the measurement theory in psychiatry for modeling multiple disease diagnostic measures as arising from the underlying causes (true mental states). It allows incorporation of the multivariate pre- and post-treatment outcomes as well as biological measures while preserving the invariant structure for representing patients' latent mental states. A multi-layer neural network is used to allow complex treatment effect heterogeneity. Optimal treatment policy can be inferred for future patients by comparing their potential mental states under different treatments given the observed multi-domain pre-treatment measurements. Experiments on simulated data and a real-world clinical trial data show that the learned treatment polices compare favorably to alternative methods on heterogeneous treatment effects, and have broad utilities which lead to better patient outcomes on multiple domains.
This paper presents a novel solution for approximations to some large convolution kernels by leveraging a weighted box-filtered image pyramid set. Convolution filters are widely used, but still compute-intensive for real-time rendering when the kernel size is large. Our algorithm approximates the convolution kernels, such as Gaussian and cosine filters, by two phases of down and up sampling on a GPU. The computational complexity only depends on the input image resolution and is independent of the kernel size. Therefore, our method can be applied to nonuniform blurs, irradiance probe generations, and ray-traced glossy global illuminations in real time, and runs in effective and efficient performance.
To address substantial heterogeneity in patient response to treatment of chronic disorders and achieve the promise of precision medicine, individualized treatment rules (ITRs) are estimated to tailor treatments according to patient-specific characteristics. Randomized controlled trials (RCTs) provide gold standard data for learning ITRs not subject to confounding bias. However, RCTs are often conducted under stringent inclusion/exclusion criteria, and participants in RCTs may not reflect the general patient population. Thus, ITRs learned from RCTs lack generalizability to the broader real world patient population. Real world databases such as electronic health records (EHRs) provide new resources as complements to RCTs to facilitate evidence-based research for personalized medicine. However, to ensure the validity of ITRs learned from EHRs, a number of challenges including confounding bias and selection bias must be addressed. In this work, we propose a matching-based machine learning method to estimate optimal individualized treatment rules from EHRs using interpretable features extracted from EHR documentation of medications and ICD diagnoses codes. We use a latent Dirichlet allocation (LDA) model to extract latent topics and weights as features for learning ITRs. Our method achieves confounding reduction in observational studies through matching treated and untreated individuals and improves treatment optimization by augmenting feature space with clinically meaningful LDA-based features. We apply the method to EHR data collected at New York Presbyterian Hospital clinical data warehouse in studying optimal second-line treatment for type 2 diabetes (T2D) patients. We use cross validation to show that ITRs outperforms uniform treatment strategies (i.e., assigning same treatment to all individuals), and including topic modeling features leads to more reduction of post-treatment complications.
Dimension reduction of high-dimensional microbiome data facilitates subsequent analysis such as regression and clustering. Most existing reduction methods cannot fully accommodate the special features of the data such as count-valued and excessive zero reads. We propose a zero-inflated Poisson factor analysis (ZIPFA) model in this article. The model assumes that microbiome absolute abundance data follow zero-inflated Poisson distributions with library size as offset and Poisson rates negatively related to the inflated zero occurrences. The latent parameters of the model form a low-rank matrix consisting of interpretable loadings and low-dimensional scores which can be used for further analyses. We develop an efficient and robust expectation-maximization (EM) algorithm for parameter estimation. We demonstrate the efficacy of the proposed method using comprehensive simulation studies. The application to the Oral Infections, Glucose Intolerance and Insulin Resistance Study (ORIGINS) provides valuable insights into the relation between subgingival microbiome and periodontal disease.
We propose a new photo-guided drawing tool for the easy generation of artistic strokes in the style of flower painting, which is the most representative one in Chinese paintings, through stylizing rough sketch lines over a photoimage. While existing style migration methods aim at processing the entire image, we depict that a stroke-based stylization framework better performs detailed style features and catches the user's intent. The steps in our technique are first to correct user inaccurate strokes that mark candidates of style patterns on a painting and skeletal paths to be stylized over a photograph, then to select the optimal pattern for every skeletal stroke, and finally, to map the pattern onto skeletal stroke by shape calculation and texture synthesis. This tool can tolerate errors from user inputs and work well for touch devices. Our experiments prove that the paintings our technique helped to generate show close to the specified style and are aesthetically satisfied by normal users.
The easy acquisition of various information of a Chinese calligraphic character can greatly enhance comprehension and learning efficiency of Chinese calligraphy. This paper presents a handy tool for fast searching all kinds of information of a written calligraphic character. Taking the characters written in regular style for example, which are the popular and representative style in calligraphy, first we establish an information library of over 6,700 standard Chinese calligraphy including Pinyin, structure, components and explanation. Then in order to fast identify a given image of a written calligraphic character in the library, we build a neural network with a feature set that largely determines the shape and spatial structure of a character and for the character that is provided more than one identifying results, we turn to image-based matching algorithm to find out the one with the highest similarity. Finally, users can easily obtain the whole information of the input character.