Three-dimensional laser gated range-intensity correlation imaging (GRICI) can obtain long-range 3D images in real time with high resolution by suppressing backscatter and background noise outside the 3D depth of field. Range precision is crucial for evaluating the performance of 3D GRICI. However, traditional range precision models can only estimate it from experimental results, or cannot accurately predict it theoretically due to over-simplified analysis or insufficient theoretical modeling capability. In this paper, we present a range precision prediction model for 3D GRICI based on the law of error propagation and stochastic process analysis. Particularly, we take into account the photo-electronic process and the temporal jitter of a 3D GRICI system. Based on the classic laser range-gated imaging model, the signal-to-noise model, and the compound random distribution, the mean intensity value and the shot noise can be solved. Thus, the range precision within the 3D depth of field can be accurately predicted. We conducted experiments to validate this model. The results demonstrate that its predictions are more consistent with the experimental results compared to the traditional models. The relative error of the traditional models at 13.45 m is 76.9%, while the relative error of the proposed prediction model is only 3.1%. This research can support the design of 3D gated imaging systems.
We propose ESSC-RM, a plug-and-play Enhancing framework for Semantic Scene Completion with a Refinement Module, which can be seamlessly integrated into existing semantic scene completion (SSC) models. ESSC-RM operates in two phases: a baseline SSC network first produces a coarse voxel prediction, which is subsequently refined by a 3D U-Net-based Prediction Noise-Aware Module (PNAM) and Voxel-level Local Geometry Module (VLGM) under multiscale supervision. Experiments on SemanticKITTI show that ESSC-RM consistently improves semantic prediction performance. When integrated into CGFormer and MonoScene, the mean IoU increases from 16.87 to 17.27% and from 11.08 to 11.51%, respectively. These results demonstrate that ESSC-RM serves as a general refinement framework applicable to a wide range of SSC models. Project page: https://github.com/LuckyMax0722/ESSC-RM and https://github.com/LuckyMax0722/VLGSSC.
Sea surface reflectance is a key underlying-surface parameter for remote sensing retrievals of sea fog, sea wind, and marine aerosols, and it exhibits significant variability across different times and ocean-current-controlled regions, with maximum variations of up to approximately 15
A dual-front-camera system with wide and narrow field-of-view (FOV) coverage is crucial for perception performance in autonomous driving. To address the inefficiencies and potential risks associated with long-tail data collection challenges in training perception models, we present CGDD, a Consistency-Guided Diffusion model for image generation that produces synthetic data in Dual-front cameras. To leverage the narrow-angle view’s zoomed-in representation of the overlapping area in the wide-angle view, we propose four key components in our diffusion method to enhance consistency. First, the diffusion model shares noise between corresponding areas of the two views. Second, a cross-view attention mechanism is applied to capture contextual consistency in the overlapping regions. Third, a self-supervised consistency loss ensures both geometric and semantic alignment. Finally, during inference, we introduce an adaptive blending and stack strategy that dynamically fuses images from both views, enhancing consistency and reducing denoising steps. We evaluate our diffusion-based framework for dual-front-camera system with wide and narrow angles on a real-world dataset. The experimental results demonstrate that our method achieves promising performance. The effectiveness of each component is validated through ablation studies.
Topology reasoning is crucial for autonomous driving. Current methods primarily focus on instance-level learning for centerline detection, followed by a sequential module for topology reasoning that relies on simplified MLP layers. Moreover, they often neglect the importance of point-to-instance (P2I) relationships in topology reasoning. To address these limitations, we present TopoHR (Topological Hierarchical Representation), a novel end-to-end framework that establishes cyclic interaction between centerline detection and topology reasoning, allowing them to iteratively enhance each other. Specifically, we introduce a hierarchical centerline representation including point queries, instance queries, and semantic representations. These multi-level features are seamlessly integrated and fused within a hierarchical centerline decoder. Furthermore, we design a hierarchical topology reasoning module that captures both fine-grained P2I relationships and global instance-to-instance (I2I) connections within a unified architecture. With these novel components, TopoHR ensures accurate and robust topology reasoning. On the OpenLane-V2 benchmark, TopoHR refreshes state-of-the-art performance with significant improvements. Notably, compared with previous best results, TopoHR achieves +3.8 in DET_l, +5.4 in TOP_ll on subset_A and +11.0 in DET_l, +7.9 in TOP_ll on subset_B, validating the effectiveness of the proposed components. The code will be shared publicly at https://github.com/Yifeng-Bai/TopoHR.git.
Objective Coral reef ecosystems are facing dual threats from global climate change and human activities, necessitating urgent support from cross-scale, multi-dimensional in situ observation technologies for effective monitoring. Coral reefs are structurally complex, multi-scale ecosystems whose health changes may appear simultaneously at multiple levels, from symbiont condition and individual macro-organisms to habitat patches and community/landscape patterns. Diver-based transect surveys remain accurate, but they cover limited areas, operate with low efficiency, and are constrained by working depth. Remote sensing and conventional underwater optical imaging methods are also limited by spatial resolution, water attenuation, and viewing point, which together preclude continuous monitoring from shallow to deep reefs and from microscopic processes to macroscopic patterns. To meet the demand for cross-scale in situ intelligent reef monitoring, this research develops a cross-scale multi-dimensional optical in situ integrated observation system of Shuijing for coral reef ecosystems. The system covers the full workflow from data acquisition, image enhancement, 3D reconstruction, and intelligent recognition to quantitative assessment. Methods The system was organized along two dimensions: the spatial scale of observed targets and their level of ecological organization. It partitioned coral reef observation into four scales: community/landscape, habitat patch, individual macro-organism, and microorganism/tissue. Five complementary imaging units were developed for these scales: Shuijing-FlyCam, Shuijing-360, Shuijing-Trap, Shuijing-LiRAI, and Shuijing-Micro. At the community/landscape scale, Shuijing-FlyCam performed optical cruising imaging along seabed transects to acquire large-area, high-resolution benthic imagery, from which coral cover, community composition, and the spatial distribution of bleaching and disease could be extracted. At the habitat-patch scale, Shuijing-360 captured circumferential panoramic imagery to record reef-associated fish together with their surrounding habitat context. At the individual macro-organism scale, Shuijing-Trap and Shuijing-LiRAI jointly supported 3D morphological measurement of fish, corals, and other targets, and could operate under challenging conditions such as low light, turbid water, and transparent organisms such as jellyfish. At the microorganism/tissue scale, Shuijing-Micro performed in situ optical microscopy of coral symbiotic zooxanthellae and plankton, enabling early warning of coral bleaching and the collection of baseline food-web data. Data from the four scales were organized within a unified spatiotemporal reference framework. Transect segments served as the spatial backbone, while position logging, timestamp alignment, and synchronized multi-device triggering established correspondences among macroscopic transects, local patches, individual targets, and microscopic fields of view. This formed a nested, transect-anchored data structure in which macroscopic transect imagery anchored large-area spatial positioning, panoramic imagery refined local habitat characterization, 3D imaging contributed structural and morphological parameters, and microscopic imagery supplied tissue-and symbiont-level early indicators. To address the color cast, haze, low contrast, and loss of structural information commonly found in underwater optical imagery, image enhancement and biological quantification based on artificial intelligence (AI) were embedded into the processing pipeline. Two-dimensional image enhancement improved color fidelity, edge sharpness, and texture visibility in transect and habitat imagery, providing more stable inputs for downstream target recognition, image mosaicking, and 3D reconstruction. Scene geometry was recovered by monocular depth estimation, binocular stereo matching, or multi-view reconstruction. A comparison based on representative stereo image pairs showed that monocular depth estimation agreed well with binocular stereo reconstruction in capturing the overall trend of depth variation, while binocular stereo reconstruction remained more reliable for local geometric details and absolute-scale measurement. For large-area transect scenes, multi-view reconstruction further yielded continuous 3D representations that facilitated reef structural analysis and spatially explicit quantification. Results and Discussions Initial deployments during coral reef surveys in the Xisha Islands, South China Sea, from 2023 to 2025 verified the feasibility of the system, indicating that Shuijing provides a relatively complete technical framework for coral reef ecosystem monitoring, assessment, and conservation research. For AI-based visual tasks, an AI-oriented Detection-Recognition-Identification (AI-DRI) standard was established, adapted from classical optical DRI criteria. AI-DRI reframed the traditional three levels as presence detection, target-category classification, and fine-grained discrimination. Using fish targets as a case study, two lightweight networks were used to examine how target pixel count related to task accuracy. Accuracy across all three task levels increased with target pixel count and gradually approached saturation, while finer-grained tasks required more pixels. Under a common usable-accuracy threshold, detection required the fewest pixels, recognition required more, and identification required the most. AI-DRI thus provides a quantifiable and comparable link among data, algorithms, and hardware, and can inform training-data screening, model evaluation, and sensor-parameter design. Given the depth-driven ecological zonation of coral reefs, this study further proposed a collaborative survey scheme based on cross-domain heterogeneous group (CDHG). Shallow zones are mainly covered by unmanned surface vehicles and aerial-aquatic vehicles capable of switching between flight and underwater operation. These platforms can perform large-area transect imaging, non-contact observation in extremely shallow areas, rapid relocation among discrete sites, and vertical-profile sampling. Mesophotic and deeper zones are covered by autonomous underwater vehicles, remotely operated vehicles, and benthic landers, which together form an observation chain of large-area cruising, fine-scale verification, and long-term fixed-point monitoring. Human-occupied vehicles and diver propulsion vehicles serve as external scientific validation nodes, supporting on-site judgment, ground-truth verification, and sample collection in priority areas. Corresponding to this hardware-side architecture, coral reef health assessment follows a strategy of independent quantification within each scale followed by unified integration. Each scale first produces standardized intermediate products, including coral cover, classified patches, 3D morphological parameters, fish abundance and size distribution, plankton counts, and zooxanthellae status. Results from transects, patches, individuals, and microscopic fields of view are then aligned through transect-based spatial units. Finally, compositional, structural, and microscopic evidence is integrated within common assessment units to yield interpretable and traceable reef health assessments. Conclusions Traditional coral reef ecosystem monitoring methods, constrained by single-scale designs, struggle to simultaneously capture multi-scale ecological processes ranging from microscopic to macroscopic levels. To address this challenge, this paper proposes an underwater cross-scale multi-dimensional optical in situ integrated observation technology system of Shuijing for coral reef ecosystems. This system integrates specialized optical observation technologies across four scales: community/landscape, habitat patch, individual macro-organism, and microorganism/tissue. Addressing the issues of underwater image degradation and the loss of three-dimensional information, AI-powered image visualization enhancement technology is introduced to effectively improve image clarity and restore the three-dimensional structure of scenes, providing high-quality input for target recognition and geometric quantitative analysis. Building upon this foundation, the AI-DRI standard is proposed, mapping the three-level tasks of detection, recognition, and identification to AI vision tasks. It establishes a quantitative relationship between target pixel count and recognition accuracy, providing a unified framework for training data screening, model performance evaluation, and sensor selection. Integrating the ecological depth stratification of coral reefs, a collaborative survey concept involving CDHG is proposed. A stratified survey scheme is constructed using core platforms such as unmanned surface vehicles, aerial-aquatic vehicles, autonomous underwater vehicles, remotely operated vehicles, and benthic landers. This work provides a systematic technical framework for efficient monitoring and scientific conservation of coral reef ecosystems, with the potential to propel this field towards systematic and intelligent development.
Magnetic microrobots (MMs) have emerged as promising tools for targeted therapies, including non-invasive in vivo treatments and precise drug delivery, owing to their untethered controllability and biocompatibility. Current actuation strategies for MMs primarily rely on two magnetic field (MF) generation approaches: gradient-based and rotational methods. Unlike the gradient method, rotational actuation enables efficient manipulation of MMs under significantly weaker magnetic fields. To fully leverage the potential of rotationally driven MMs, a comprehensive understanding of their fundamental spin motility is essential. Achieving accurate characterization of these MMs necessitates the development of an MF generation system equipped with rapid motion-tracking and broad-range measurement capabilities. This study proposes a high-speed rotating states observation scheme by developing a tracking-based optimal local imaging and estimation scheme, simultaneously meeting the broad-range observation capability and the high imaging speed requirement. Specifically, the CSR-DCF tracking method is adopted to detect the MM’s location, and based on this, the observation system adjusts the imaging region optimally. An estimation scheme based on the asynchronous rectification method is derived to measure the MM rotating states consistently using measured MF data and local optical images of the target. Experimental studies are carried out to validate the effectiveness of the proposed scheme.
Efficient motion planning with the error tolerance is crucial for dynamic robotic tasks, particularly robotic table tennis. This task demands simultaneous high efficiency and error tolerance. First, the incoming ball's high speed allows only tens of milliseconds for motion planning. Second, two types of errors, namely, ball-paddle motion uncertainty and joint limit violation, must be tolerated to ensure a high success rate of planning (SRP) and striking. This article proposes an advanced joint planning framework designed for high efficiency and error tolerance. To tolerate the error of ball-paddle motion uncertainty, this work introduces a joint classification criterion according to the joint motion characteristics. To tolerate the error of joint limit violation, based on the classification criterion, this study also develops a robust reference trajectory generation scheme, named error-tolerant-variable-sigmoid-based motion template (ETVSMT), to fully consider the motion capabilities of different joints. The ETVSMT approach utilizes the variable-sigmoid-based motion template (VSMT) as the backbone and designs its basic and remedial portions to tolerate the error of hard joint limit violation. The implementation of the ETVSMT scheme results in an average SRP of 98% and a ball-striking success rate of up to 90%, with execution time on an Intel Xeon CPU as low as 15 ms. Furthermore, the proposed method ensures the outgoing ball lands on the opposite side of the table with an average landing error of 19.62 cm and a net-passing height error of 12.68 cm. The proposed method can also benefit other dynamic tasks like human-robot interaction to improve the planning efficiency.
3D lane detection is an integral part of autonomous driving systems. Previous CNN and Transformer-based methods usually first generate a bird's-eye-view (BEV) feature map from the front view image, and then use a sub-network with BEV feature map as input to predict 3D lanes. Such approaches require an explicit view transformation between BEV and front view, which itself is still a challenging problem. In this paper, we propose CurveFormer, a single-stage Transformer-based method that directly calculates 3D lane parameters and can circumvent the difficult view transformation step. Specifically, we formulate 3D lane detection as a curve propagation problem by using curve queries. A 3D lane query is represented by a dynamic and ordered anchor point set. In this way, queries with curve representation in Transformer decoder iteratively refine the 3D lane detection results. Moreover, a curve cross-attention module is introduced to compute the similarities between curve queries and image features. Additionally, a context sampling module that can capture more relative image features of a curve query is provided to further boost the 3D lane detection performance. We evaluate our method for 3D lane detection on both synthetic and real-world datasets, and the experimental results show that our method achieves promising performance compared with the state-of-the-art approaches. The effectiveness of each component is validated via ablation studies as well.
The effectiveness of Vision Transformer (ViT)-based feature encoding network has been demonstrated in medical image analysis tasks. However, the complexity growing quadratically with the token number limits its application in dense prediction. To accelerate ViT, we propose an efficient and accurate token halting and reconstruction encoder framework, termed HRViT, designed for precise medical image semantic segmentation. Our approach is motivated by the observation that background and internal tokens can be easily identified and halted in early layers, while complex and ambiguous edge regions require deeper computational processing for accurate segmentation. HRViT leverages this insight by incorporating an edge-aware token halting module, which dynamically identifies edge patches and halts non-edge tokens. The preserved edge tokens are propagated to deeper layers and further refined through edge reinforcement. After encoding, all tokens are restored to their original positions, and auxiliary supervision is also introduced to strengthen the encoder's representation power. We evaluate the segmentation performance of our method using two public medical image datasets and the experimental results show that our method achieves promising performance compared with the state-of-the-art approaches. Our code is released at https://github.com/guoyh6/hrvit.
Accurate spin estimation is crucial for assisting table tennis robots in defeating high-level human players. Currently, most researchers use methods based on physical models or identify logos on the flying table tennis ball to estimate spin. We try a new approach, considering the use of a data-driven method to estimate spin. However, directly using such models does not produce satisfactory results, as these methods do not consider that table tennis spin estimation is a coarse-to-fine process. To address this problem, we develop a hierarchical spin estimation network to estimate spin progressively. Furthermore, we introduce a Mix Conv-Attn Block to enhance feature extraction from table tennis trajectories. This block can capture both short-term and long-term features to improve the estimation accuracy. Comparing our approach with physical-based and non-hierarchical neural network methods, experimental results show that our method achieves superior performance.
Semi-supervised medical image segmentation aims to leverage a limited set of labeled images alongside a substantial volume of unlabeled images to train semantic segmentation models. Existing studies often employ consistency regularization to maximize the utilization of unlabeled data, thereby enhancing the model's robustness and accuracy. However, the methods for constructing perturbations at image-level on unlabeled data are typically simplistic, involving techniques such as color transformations, additive noise, which do not adequately leverage the precise and reliable supervisory information available from labeled images. To address this limitation, in addition to image perturbation, we propose a cross-image feature perturbation approach for semi-supervised medical image segmentation. This method utilizes feature information from labeled images to guide the refinement of ambiguous semantic representations in unlabeled images, thereby expanding the perturbation space more effectively. Moreover, recognizing the limitations of existing consistency regularization frameworks that rely on confidence thresholds to filter pseudo-labels, we introduce an uncertainty-based pseudo-label fusion strategy. This strategy mitigates the effects of unreliable predictions caused by perturbations by calculating the uncertainty and using it as a weight during pseudo-label fusion. We have conducted extensive experiments on the 2D ACDC and 3D LA datasets. The results demonstrate that our approach achieves performance comparable to the current state-of-the-art (SOTA) methods.
X-ray machines are vital in medical imaging for viewing internal body structures. However, in lumbar vertebrae surgery, the X-ray machine has to move and position frequently and manually, which risks infection, misalignment, and overexposure to radiation. There’s a need for mobile X-ray machines with autonomous recognition and positioning. Although visual servoing suitable for X-ray positioning naturally, it still experiencing challenges like limited Field-of-View (FOV) and the structural similarity among vertebrae. In this paper, a Bayesian Estimation based approach is developed to improve the detection accuracy in case of vertebrae positions outside the FOV. Furthermore, a visual servoing method is proposed for X-ray positioning, considering their mechanical structure and imaging characteristics. The experiment results indicates that the proposed approach improves lumbar vertebrae detection accuracy significantly. This approach facilitates precise positioning in X-ray imaging and enhances the safety and effectiveness of lumbar vertebrae surgery.
Magnetic microrobots (MMs) enable precise, minimally invasive manipulation via external magnetic field (MF), allowing remote operation in complex or inaccessible environments. Common actuation methods for MMs include gradient-based and rotational MF strategies, with the latter requiring significantly weaker MF to achieve efficient propulsion compared to gradient-driven approaches. To optimize the functionality of rotationally actuated MMs, their fundamental spin dynamics must be systematically characterized. This necessitates advanced MF generation systems capable of rapid motion-tracking and wide-range measurement to ensure precise analysis. This study introduces a real-time planar spin velocity measurement scheme, combining a tracking-based local imaging framework with a practical estimation algorithm to simultaneously achieve high-speed and large-scale measurement. The proposed method tracks one target MM to focus on the small local region, enabling the imaging system to dramatically enhance its capturing speed. Based on the fast imaging strategy, a singular value decomposition (SVD)-based orientation detection algorithm is then applied to calculate spin velocity of the target MM in real-time. Experimental study was carried out to demonstrate the capability of the proposed measuring scheme.
Motion generation for robots in highly dynamic environments is challenging due to the need for high efficiency and error tolerance. Robots must generate motion efficiently within milliseconds to react to environmental changes swiftly. Furthermore, the limitations of perceptual systems in detecting fast-moving objects necessitate robust motion generation that can tolerate perceptual errors. This paper presents a motion generation framework that ensures both high efficiency and error tolerance for the dynamic task of robotic table tennis. To achieve high efficiency, this study introduces a novel analytical solution to the ball flight equation using Chebyshev approximation, offering a time complexity of O (1). For robotic motion generation, this study proposes a dimension-reduced task planner and employs piecewise cubic polynomial and Variable-Sigmoid-Based motion templates for individual joint motions, with coefficients determined analytically and rapidly. To enhance perceptual error tolerance, this study develops a policy optimizer based on a striking-at-the-center principle. The proposed task planner accelerates paddle motion determination by an average of 9.27 times. The proposed ball motion prediction method enhances the efficiency of determining the ball-paddle collision point by up to a factor of two. Additionally, the proposed framework improves the success rate of planning and striking by 16.0% and 18.3%, respectively, compared to methods that do not account for error tolerance. It also reduces the landing position error in the x and y directions by 32.5% and 35.6%, while improving the accuracy of passing-net height by 24.0%.
In table tennis, developing a precise ball-racket rebound model is crucial for predicting the trajectory and spin of a ball after it hits the racket, which is instrumental in racket design and enhancing the capabilities of table tennis robots. To this end, accuracy and computational efficiency are two challenges to overcome, which has not been perfectly handled using existing methods such as finite element and simplified rigid-body models. This paper introduces a new model that calculates ball-racket rebounds in two orthogonal directions. Vertically, the collision dynamics are analyzed with the Kelvin-Voigt model, revealing that the contact duration is independent of the incoming ball’s velocity. Horizontally, we establish that the restitution coefficient varies as a function of the incident velocity, based on a continuous contact force model and momentum conservation principles. High-speed camera data corroborate these findings and confirm the model’s efficacy across diverse conditions. Compared to an established representative model, our method not only maintains high computational efficiency but also improves the accuracy of predicting the ball’s linear and angular velocities by an average of 42.72% and 33.77%, respectively, as evidenced by our experimental data.
Sport and game industry has grown rapidly in recent years due to the application of novel sensors and algorithms for quantitative analysis. For example, flying speed and spin estimation is essential to help players to improve their skills in table tennis. However, the spin estimation for a table tennis ball is challenging, as it is difficult to observe using cameras and model the aerodynamics of ball flight with spin. This article proposes a generalized aerodynamic model with variable aerodynamic coefficients to accurately represent the flying state of a table tennis ball. Analytical solutions for the aerodynamic coefficients and the acceleration due to the Magnus force are also developed for accurate ball spin estimation using pre- and postrebounding flight trajectories. The experimental results showed that compared to current state-of-the-art methods, the proposed method has achieved the best performance in angular velocity magnitude estimation for topspinning and backspinning balls. It also achieved an error of below 10° in angular velocity amplitude estimation. Using the proposed spin estimation method, our table tennis robot could strike balls with either topspin or backspin with a high success rate of up to 84.6 $%$ . Besides, the experimental results also demonstrated the potential of the proposed method in the area of table tennis training and sports-broadcasting.
Object detection represents a fundamental and pivotal task within the domain of computer vision, which has attracted considerable interest in approaches that directly utilize Vision Transformer (ViT) to perform region-level recognition. However, despite the efforts of early pioneers in exploring vanilla ViT detectors, a significant performance disparity remains. Addressing this limitation, in this work, we propose an ViT-based detector to further enhance the ability of plain ViT in object detection by incorporating a frozen vision-language large model. Specifically, to boost the ViT detector with the frozen CLIP model, we construct ViT-based Side Prompt-Adapter Tuning, which align and fine-tune CLIP features without requiring gradients to flow through CLIP to provide additional rich semantic information for the ViT detector. Furthermore, a CLIP-based visual token selection module is proposed to leverage fine-tuned CLIP features to filter out irrelevant background visual tokens, resulting in decreased computational complexity and memory usage of the ViT-based detector. Additionally, we introduce query denoising training and adapt the position embeddings to further enhance training efficiency. Compared to the latest ViT-based detector, experimental results show that our method converges 3× faster and achieves promising performance.
Underwater fishing net recognition plays an indispensable role in applications such as safe navigation of unmanned underwater vehicles, protection of marine ecology and marine ranching. However, the performance of underwater fishing net recognition usually degrades seriously due to noise interference in underwater environments. In this paper, we use range gated imaging as the detection device, and propose a semantic fishing net recognition network (SFNR-Net) for underwater fishing net recognition at long distance. The proposed SFNR-Net introduces an auxiliary semantic segmentation module (ASSM) to introduce extra semantic information and enhance feature representation under noisy conditions. Besides, to address the problem of unbalanced training data, we employ semantic regulated cycle-consistent generative adversarial network (CycleGAN) as a data augmentation approach. To improve the quality of generated data, we propose a semantic loss to regulate the training of CycleGAN. Comprehensive experiments on the test data show that SFNR-Net can effectively solve noise interference and achieve the best recognition accuracy of 96.28% compared with existing methods. Field experiments in underwater environments with different turbidity further validate the advantages of our method.
Dyna-style Model-based reinforcement learning (MBRL) methods have demonstrated superior sample efficiency compared to their model-free counterparts, largely attributable to the leverage of learned models. Despite these advancements, the effective application of these learned models remains challenging, largely due to the intricate interdependence between model learning and policy optimization, which presents a significant theoretical gap in this field. This paper bridges this gap by providing a comprehensive theoretical analysis of Dyna-style MBRL for the first time and establishing a return bound in deterministic environments. Building upon this analysis, we propose a novel schema called Model-Based Reinforcement Learning with Model-Free Policy Optimization (MBMFPO). Compared to existing MBRL methods, the proposed schema integrates model-free policy optimization into the MBRL framework, along with some additional techniques. Experimental results on various continuous control tasks demonstrate that MBMFPO can significantly enhance sample efficiency and final performance compared to baseline methods. Furthermore, extensive ablation studies provide robust evidence for the effectiveness of each individual component within the MBMFPO schema. This work advances both the theoretical analysis and practical application of Dyna-style MBRL, paving the way for more efficient reinforcement learning methods.