Generating interactive 3D scenes from text requires not only synthesizing assets but arranging them with spatial intelligence—support, affordances, and plausibility. However, training data for interactive scenes is dominated by a few indoor datasets, so learning-based methods overfit to in-distribution layouts and struggle to compose diverse arrangements (e.g., outdoor settings and small-on-large relations). Meanwhile, LLM-based layout planners can propose diverse arrangements, but the lack of visual grounding often yields implausible placements that violate commonsense physics. We propose Scenethesis, a training-free, agentic framework that couples LLM-based scene planning with vision-guided layout refinement. Given a text prompt, Scenethesis first drafts a coarse layout with an LLM; a vision module refines the layout and extracts scene structure to capture inter-object relations. A novel optimization stage enforces pose alignment and physical plausibility, and a final judge verifies spatial coherence and triggers targeted repair when needed. Across indoor and outdoor prompts, Scenethesis produces realistic, relation-rich, and physically plausible 3D interactive scenes, reducing collisions and stability failures compared to SOTA methods, making it practical for virtual content creation, simulation, and embodied AI.
Generalization remains the central challenge for interactive 3D scene generation. Existing learning‑based approaches ground spatial understanding in limited scene dataset, restricting generalization to new layouts.We instead reprogram a pre‑trained 3D instance generator to act as a scene‑level learner via, replacing dataset-bounded supervision with model-centric spatial supervision.This reprogramming unlocks the generator's transferable spatial knowledge, enabling generalization to unseen layouts and novel object compositions.Remarkably, spatial reasoning still emerges even when the training scenes are randomly composed objects. This demonstrates that the generator’s transferable scene prior provides a rich learning signal for inferring proximity, support, and symmetry from purely geometric cues.Replacing widely used canonical space, we instantiate this insight with a view‑centric formulation of the scene space, yielding a fully feed‑forward, generalizable scene generator that learns spatial relations directly from the instance model.Quantitative and qualitative results show that a 3D instance generator is an implicit spatial learner and reasoner, pointing toward foundation models for interactive 3D scene understanding and generation.
Synthesizing interactive 3D scenes from text is essential for gaming, virtual reality, and embodied AI. However, existing methods face several challenges. Learning-based approaches depend on small-scale indoor datasets, limiting the scene diversity and layout complexity. While large language models (LLMs) can leverage diverse text-domain knowledge, they struggle with spatial realism, often producing unnatural object placements that fail to respect common sense. Our key insight is that vision perception can bridge this gap by providing realistic spatial guidance that LLMs lack. To this end, we introduce Scenethesis, a training-free agentic framework that integrates LLM-based scene planning with vision-guided layout refinement. Given a text prompt, Scenethesis first employs an LLM to draft a coarse layout. A vision module then refines it by generating an image guidance and extracting scene structure to capture inter-object relations. Next, an optimization module iteratively enforces accurate pose alignment and physical plausibility, preventing artifacts like object penetration and instability. Finally, a judge module verifies spatial coherence. Comprehensive experiments show that Scenethesis generates diverse, realistic, and physically plausible 3D interactive scenes, making it valuable for virtual content creation, simulation environments, and embodied AI research.
This study explores the vaccine prioritization strategy to reduce the overall burden of the pandemic when the supply is limited. Existing vaccine distribution methods focus on macro-level or simplified micro-level assuming homogeneous behavior within populations without considering mobility patterns. Directly applying these models for micro-level vaccine allocation leads to sub-optimal solutions. To address the issue, we first proposed a Trans-vaccine-SEIR model to incorporate mobility heterogeneity in disease propagation. Then we develop a novel deep reinforcement learning to seek the optimal vaccine allocation strategy for the disease evolution system. The graph neural network is used to effectively capture the structural properties of the mobility network and extract disease features. In our evaluation, the proposed framework reduces 7% - 10% of infections and deaths compared to the baseline strategies. Extensive evaluation shows that the proposed framework is robust to seek the optimal vaccine allocation with diverse mobility patterns. In particular, we find transit usage restriction is significantly more effective than restricting cross-zone mobility for the top 10% age-based and income-based zones under optimal vaccine allocation strategy. These results provide valuable insights for areas with limited vaccines and low logistic efficacy.
We have witnessed significant progress in deep learning-based 3D vision, ranging from neural radiance field (NeRF) based 3D representation learning to applications in novel view synthesis (NVS). However, existing scene-level datasets for deep learning-based 3D vision, limited to ei-ther synthetic environments or a narrow selection of real-world scenes, are quite insufficient. This insufficiency not only hinders a comprehensive benchmark of existing methods but also caps what could be explored in deep learning-based 3D analysis. To address this critical gap, we present DL3DV-10K, a large-scale scene dataset, featuring 51.2 million frames from 10,510 videos captured from 65 types of point- of-interest (POI) locations, covering both bounded and unbounded scenes, with different levels of reflection, transparency, and lighting. We conducted a comprehensive benchmark of recent NVS methods on DL3DV-10K, which revealed valuable insights for future research in NVS. In addition, we have obtained encouraging results in a pilot study to learn generalizable NeRF from DL3DV-10K, which manifests the necessity of a large-scale scene-level dataset to forge a path toward a foundation model for learning 3D representation. Our DL3DV-10K dataset, benchmark results, and models will be publicly accessible.
Bokeh is widely used in photography to draw attention to the subject while effectively isolating distractions in the background. Computational methods simulate bokeh effects without relying on a physical camera lens. However, in the realm of digital bokeh synthesis, the two main challenges for bokeh synthesis are color bleeding and partial occlusion at object boundaries. Our primary goal is to overcome these two major challenges using physics principles that define bokeh formation. To achieve this, we propose a novel and accurate filtering-based bokeh rendering equation and a physically-based occlusion-aware bokeh renderer, dubbed Dr.Bokeh, which addresses the aforementioned challenges during the rendering stage without the need of post-processing or data-driven approaches. Our rendering algorithm first preprocesses the input RGBD to obtain a layered scene representation. Dr.Bokeh then takes the layered representation and user-defined lens parameters to render photo-realistic lens blur. By softening non-differentiable operations, we make Dr.Bokeh differentiable such that it can be plugged into a machine-learning framework. We perform quantitative and qualitative evaluations on synthetic and real-world images to validate the effectiveness of the rendering quality and the differentiability of our method. We show Dr.Bokeh not only outperforms state-of-the-art bokeh rendering algorithms in terms of photo-realism but also improves the depth quality from depth-from-defocus.
While the growth of TNCs took a substantial part of ridership and asset value away from the traditional taxi industry, existing taxi market policy regulations and planning models remain to be reexamined, which requires reliable estimates of the sensitivity of labor supply and income levels in the taxi industry. This study aims to investigate the impact of TNCs on the labor supply of the taxi industry, estimate wage elasticity, and understand the changes in taxi drivers' work preferences. We introduce the wage decomposition method to quantify the effects of TNC trips on taxi drivers' work hours over time, based on taxi and TNC trip record data from 2013 to 2018 in New York City. The data are analyzed to evaluate the changes in overall market performances and taxi drivers' work behavior through statistical analyses, and our results show that the increase in TNC trips not only decreases the income level of taxi drivers but also discourages their willingness to work. We find that 1 of the yellow taxi industry and 0.68 green taxi industry in recent years. More importantly, we report that the work behavior of taxi drivers shifts from the widely accepted neoclassical standard behavior to the reference-dependent preference (RDP) behavior, which signifies a persistent trend of loss in labor supply for the taxi market and hints at the collapse of taxi industry if the growth of TNCs continues. In addition, we observe that yellow and green taxi drivers present different work preferences over time. Consistently increasing RDP behavior is found among yellow taxi drivers. Green taxi drivers were initially revenue maximizers but later turned into income targeting strategy
The causal impact of COVID-19 vaccine coverage on effective reproduction number R(t) under the disease control measures in the real-world scenario is understudied, making the optimal reopening strategy (e.g., when and which control measures are supposed to be conducted) during the recovery phase difficult to design. In this study, we examine the demographic heterogeneity and time variation of the vaccine effect on disease propagation based on the Bayesian structural time series analysis. Furthermore, we explore the role of non-pharmaceutical interventions (NPIs) and the entrance of the Delta variant of COVID-19 in the vaccine effect for U.S. counties. The analysis highlights several important findings: First, vaccine effects vary among the age-specific population and population densities. The vaccine effect for areas with high population density or core airport hubs is 2 times higher than for areas with low population density. Besides, areas with more older people need a high vaccine coverage to help them against the more contagious variants (e.g., the Delta variant). Second, the business restriction policy and mask requirement are more effective in preventing COVID-19 infections than other NPI measures (e.g., bar closure, gather ban, and restaurant restrictions) for areas with high population density and core airport hubs. Furthermore, the mask requirement consistently amplifies the vaccine effects against disease propagation after the presence of contagious variants. Third, areas with a high percentage of older people are suggested to postpone relaxing the restaurant restriction or gather ban since they amplify the vaccine effect against disease infections. Such empirical insights assist recovery phases of the pandemic in designing more efficient reopening strategies, vaccine prioritization, and allocation policies.
Lighting effects such as shadows or reflections are key in making synthetic images realistic and visually appealing. To generate such effects, traditional computer graphics uses a physically-based renderer along with 3D geometry. To compensate for the lack of geometry in 2D Image compositing, recent deep learning-based approaches introduced a pixel height representation to generate soft shadows and reflections. However, the lack of geometry limits the quality of the generated soft shadows and constrains reflections to pure specular ones. We introduce PixHt-Lab, a system leveraging an explicit mapping from pixel height representation to 3D space. Using this mapping, PixHt-Lab reconstructs both the cutout and background geometry and renders realistic, diverse lighting effects for image compositing. Given a surface with physically-based materials, we can render reflections with varying glossiness. To generate more realistic soft shadows, we further propose using 3D-aware buffer channels to guide a neural renderer. Both quantitative and qualitative evaluations demonstrate that PixHt-Lab significantly improves soft shadow generation. Project: https://shengcn.github.io/PixHtLab/
Introduction: Right-turn lane (RTL) crashes are among the key contributors to intersection crashes in the US. Unfortunately, the lack of deep insights into understanding the effects of RTL geometric design factors on crash frequency impedes improving RTL safety performance. Method: Taking the crash data in ten counties in Indiana state from 2013 to 2016 as a case study, this study investigates the safety performance of RTL geometric configuration based on multi-sources. We introduce the geographically and temporally weighted negative binomial model (GTWNBR) to capture the space and time instability in crashes. Results: The results show that the impacts of RTL geometric design factors on crash frequency vary significantly among space and time. Several key insights can be obtained from the state-wide and multi-years crash analysis by associating the estimated parameters with road classes, localities, and counties. Conclusions: First, the RTL’s length, width, turning radius, and the installments of traffic roundabouts present higher spatiotemporal heterogeneity than other factors in modeling the crash frequency. Second, the effects of RTL’s geometric factors vary significantly across space and time. The presence of bicycle and pedestrian lanes is more likely to increase crashes in urban areas than in rural ones, especially at nighttime. Third, while exclusive RTLs decrease the crash frequency compared to the shared RTLs, the exclusive RTLs are more likely to increase the crashes for RTLs on the county road than on other road classes. Increasing RTL’s turning radius and decreasing RTL’s length is more likely to promote crashes for RTLs on county roads than on other road classes. Practical Applications: The insights provide vital guidance to improve the safety performance of geometric configuration for RTLs and intersections.
User study In addition to the quantitative and qualitative evaluations in the paper, we further applied a user study to quantitatively evaluate the soft shadow quality generated by SSG++ from perception perspective. The user study is based on the cross-comparison of images generated by GSSN, SSG, and SSN. Besides, we prepared a reference image for participants to identify the shadow and rank the images from the most similar to the least. In total, the user study contains 25 questions. To be fair with SSN, 9 of the 25 questions are experiment1, in which the shadow receiver is always ground plane. Experiment 2 is composed of the other 16 questions, in which the shadow receiver is general shadow receiver. In experiment 2, we just compare SSG and SSG++. We show the image group in random order and positions to 50 participants. The gender statistical distribution of participants is: 81% male, 17% female, 2% not identified. The age distribution of participants is: 71% from the class of 18−30 years older, 25% from the class of > 30 years older, and the rest prefer not to tell. The average similarity rank score from 1-3 (the lower the score, the more similar the image is) and standard deviation of each model is presented in Fig 1. In experiment 1, SSG++ has an average similarity rank score 1.8; SSG has an average similarity rank score 2.2; SSN has an average similarity rank score 2.0. In experiment 2, SSG++ has an average similarity rank score 1.3; SSG has an average similarity rank score 1.7. In 75% of the questions, the users prefer SSG++, while only in 25% of the questions, the users prefer SSG. Thus, SSG++ performs the best in all the three methods. The T-test for the similarity rank score is significant at 0.001 level, which indicates that images generated from GSSN are significantly more similar to the reference image than images generated from SSG and SSN. SSG++ SSG SSN 0.0 0.5 1.0 1.5 2.0 2.5 3.0
Electric bikes (e-bikes), including lightweight e-bikes with pedals and e-bikes in scooter form, are gaining popularity around the world because of their convenience and affordability. At the same time, e-bike-related accidents are also on the rise and many policymakers and practitioners are debating the feasibility of building e-bike lanes in their communities. By collecting e-bikes and bikes data in Shanghai City, the study first recalibrates the capacity of the conventional bike lane based on the traffic movement characteristics of the mixed bikes flow. Then, the study evaluates the traffic safety performance of the mixed bike flow in the conventional bike lane by the observed passing events. Finally, this study proposes a comprehensive model for evaluating the feasibility of building an e-bike lane by integrating the Analytic Hierarchy Process and fuzzy mathematics by considering the three objectives: capacity, safety, and budget constraint. The proposed model, one of the first of its kind, can be used to (i) evaluate the existing road capacity and safety performance improvement of a mixed bike flow with e-bikes and human-powered bikes by analyzing the mixed bike flow arrival rate and passing maneuvers, and (ii) quantify the changes to the road capacity and safety performance if a new e-bike lane is constructed. Numerical experiments are performed to calibrate the proposed model and evaluate its performance using non-motorized vehicles' trajectories in Shanghai, China. The numerical experiment results suggest that the proposed model can be used by policymakers and practitioners to evaluate the feasibility of building e-bike lanes.
Improving information efficiency is vital in guiding better protective actions for managers. During hurricane events, individuals assess certainty through received information and translate the certainty of evacuation decision into final decision outcome. While the literature has investigated the correlation between information and the evacuation decision, little is known about the causality in the chain of information, the certainty of evacuation decision, and the final evacuation decision. This study addresses the gap by establishing a two-layer structure to capture the causality chain using data from a representative hurricane. The upper layer infers certainty assessment based on the amount of received information and the certainty of evacuation decision, and models certainty assessment based on the multinomial regression using covariates such as individual and household characteristics. The lower layer applies logistic regression to model the evacuation decision using the joint effect of covariates, the amount of received information, and certainty assessment. The Bayesian approach is adopted to impute two layers simultaneously and systematically capture the uncertainty of the model parameter. Finally, the association rule investigates the association pattern among variables and provides additional interpretation for the Bayesian estimates. The results imply the heterogeneity of the causal impact of warning information on the evacuation decision. The insights offer practical implications for improving information efficiency in managing protective actions.
BACKGROUND:Understanding non-epidemiological factors is essential for the surveillance and prevention of infectious diseases, and the factors are likely to vary spatially and temporally as the disease progresses. However, the impacts of these influencing factors were primarily assumed to be stationary over time and space in the existing literature. The spatiotemporal impacts of mobility-related and social-demographic factors on disease dynamics remain to be explored. METHODS:Taking daily cases data during the coronavirus disease 2019 (COVID-19) outbreak in the US as a case study, we develop a mobility-augmented geographically and temporally weighted regression (M-GTWR) model to quantify the spatiotemporal impacts of social-demographic factors and human activities on the COVID-19 dynamics. Different from the base GTWR model, the proposed M-GTWR model incorporates a mobility-adjusted distance weight matrix where travel mobility is used in addition to the spatial adjacency to capture the correlations among local observations. RESULTS:The results reveal that the impacts of social-demographic and human activity variables present significant spatiotemporal heterogeneity. In particular, a 1% increase in population density may lead to 0.63% more daily cases, and a 1% increase in the mean commuting time may result in 0.22% increases in daily cases. Although increased human activities will, in general, intensify the disease outbreak, we report that the effects of grocery and pharmacy-related activities are insignificant in areas with high population density. And activities at the workplace and public transit are found to either increase or decrease the number of cases, depending on particular locations. CONCLUSIONS:Through a mobility-augmented spatiotemporal modeling approach, we could quantify the time and space varying impacts of non-epidemiological factors on COVID-19 cases. The results suggest that the effects of population density, socio-demographic attributes, and travel-related attributes will differ significantly depending on the time of the pandemic and the underlying location. Moreover, policy restrictions on human contact are not universally effective in preventing the spread of diseases.
The causal impact of COVID-19 vaccine coverage on effective reproduction number R(t) under the disease control measures in the real-world scenario is understudied, making the optimal reopening strategy (e.g., when and which control measures are supposed to be conducted) during the recovery phase difficult to design. In this study, we examine the demographic heterogeneity and time variation of the vaccine effect on disease propagation based on the Bayesian structural time series analysis. Furthermore, we explore the role of non-pharmaceutical interventions (NPIs) and the entrance of the Delta variant of COVID-19 in the vaccine effect for U.S. counties. The analysis highlights several important findings: First, vaccine effects vary among the age-specific population and population densities. The vaccine effect for areas with high population density or core airport hubs is 2 times higher than for areas with low population density. Besides, areas with more older people need a high vaccine coverage to help them against the more contagious variants (e.g., the Delta variant). Second, the business restriction policy and mask requirement are more effective in preventing COVID-19 infections than other NPI measures (e.g., bar closure, gather ban, and restaurant restrictions) for areas with high population density and core airport hubs. Furthermore, the mask requirement consistently amplifies the vaccine effects against disease propagation after the presence of contagious variants. Third, areas with a high percentage of older people are suggested to postpone relaxing the restaurant restriction or gather ban since they amplify the vaccine effect against disease infections. Such empirical insights assist recovery phases of the pandemic in designing more efficient reopening strategies, vaccine prioritization, and allocation policies.
Shadow evacuation and non-compliance are among undesirable behaviors during hurricane events. Based on a post-Hurricane Matthew household survey, this study aims to understand the combined effects of information source and uncertainty on individual-level evacuate–stay decisions from the Jacksonville, Florida metropolitan area. A random parameter logit model is developed to capture the heterogeneous effects among individuals. The combined effects are substantial and distinct in the two cases and the study reveals several significant findings. First, consistent and sufficient warning information leads to better alignment with recommended actions. Second, larger social networks encourage compliance and induce shadow evacuation. Third, a highly diverse network induces compliance with the recommended actions, and isolated individuals should be provided with more information. Fourth, risk information from traditional media should be clearly delineated between the high-risk and low-risk areas for effective compliance. However, information from social media has a nonsignificant effect.
Right-turn lane (RTL) crashes are among the most key contributors to intersection crashes in the US. Different right turn lanes based on their design, traffic volume, and location have varying levels of crash risk. Therefore, engineers and researchers have been looking for alternative ways to improve the safety and operations for right-turn traffic. This study investigates the traffic safety performance of the RTL in Indiana state based on multi-sources, including official crash reports, official database, and field study. To understand the RTL crashes' influencing factors, we introduce a random effect negative binomial model and log-linear model to estimate the impact of influencing factors on the crash frequency and severity and adopt the robustness test to verify the reliability of estimations. In addition to the environmental factors, spatial and temporal factors, intersection, and RTL geometric factors, we propose build environment factors such as the RTL geometrics and intersection characteristics to address the endogeneity issues, which is rarely addressed in the accident-related research literature. Last, we develop a case study with the help of the Indiana Department of Transportation (INDOT). The empirical analyses indicate that RTL crash frequency and severity is mainly influenced by turn radius, traffic control, and other intersection related factors such as right-turn type and speed limit, channelized type, and AADT, acceleration lane and AADT. In particular, the effects of these factors are different among counties and right turn lane roadway types.
Public transport system needs to serve passengers continuously, accurately and effectively. Service quality in public transportation has been widely researched since it can evaluate public transportation systems within the passengers' aspect. At present, however, the study of public transport system service quality is usually based on survey, which requires a lot of time and economic cost. Besides, the result is insufficient in some cases since few passengers take the survey. Commuter traffic in peak hours is always a hot topic in public transport research. Based on field experiments, this paper proposes an evaluation method for public transportation service quality based on the energy cost of passengers. This method utilizes the heart rate, acceleration and speed data automatically collected by the experimenters when they are walking in subway transfer stations, fits these data to Physical Activity Intensity and uses it as the index of travel energy cost. Subsequently, the accuracy, theoretical and practical prospect of this method are verified by the transfer passenger data of Beijing Subway Line 1 and Line 2 in May, 2017. The results show that the service quality evaluation method can accurately perceive the change of system service efficiency and its recovery ability according to different travel demands of the passengers. At the same time, this method uses automatic data collection to analyze, improves its accuracy and analysis adaptability compared with the traditional methods.
Providing convenient transit services at reasonable cost is important for transit agencies. Timed transfers that schedule vehicles from various routes to arrive at some transfer stations simultaneously (or nearly so) can significantly reduce wait times in transit networks, while stochastic passenger flows and complex operating environments may reduce this improvement. Although transit priority methods have been applied in some high-density cities, operating delays may cause priority failures. This paper proposes a resilient schedule coordination method for a bus transit corridor, which analyzes link travel time, passenger loading delay, and priority signal intersection delay. It maximizes resilience based on realistic passenger flow volume, whether or not transit priority is provided. The data accuracy and result validity are improved with automatically collected data from multiple bus routes in a corridor. The Yan’an Road transit corridor in Shanghai is used as a case study. The results show that the proposed method can increase the system resilience by balancing operation cost and passenger-based cost. It also provides a guideline for realistic bus schedule coordination.
Recently, taxi play an increasingly important role in transit mode due to its accessibility and convenient. However, vacant e-hailing vehicles occupy the road capacity, thus, aggravated traffic congestions. With the availability of real-time data, one way to deal with these concerns is to improve the forecast accuracy of demand and supply thus help improving the dispatching efficiency. This paper proposes a two-stage forecast model based on big data to real-time predict the gap between demand and supply in large scale of network. The framework includes four steps, GPS data dimensionality reduction based on Principle Component Analysis, pattern analysis, the proposed methodology of two-stage forecast model and model verification. The methodology combines both non-linear Support Vector Machine and Backpropagation neural network. In the cast study of Beijing city, the model is testified and the results show that two-stage forecast model fastens responsive performance and improves the prediction accuracy. The proposed framework not only reveals the mobility pattern, it also improves the prediction accuracy for the gap between demand and supply of taxis thus helps to improve the taxi utilizations.