Brain MRI underpins a wide range of neuroscientific and clinical applications, yet most learning-based methods remain task-specific and require substantial labeled data. Here we show that a single self-supervised representation can generalize across heterogeneous brain MRI endpoints. We trained BrainDINO, a self-distilled foundation model, on approximately 6.6 million unlabeled axial slices from 20 datasets encompassing broad variation in population, disease, and acquisition setting. Using a frozen encoder with lightweight task heads, BrainDINO supported transfer across tumor segmentation, neurodegenerative and neurodevelopmental conditions classification, brain age estimation, post-stroke temporal prediction, molecular status prediction, MRI sequence classification, and survival modeling. Across tasks and supervision regimes, BrainDINO consistently equaled or exceeded natural-image and MRI-specific self-supervised baselines, with particularly strong advantages under label scarcity. Representation analyses further showed anatomically organized and pathology-sensitive feature structure in the absence of task-specific supervision. Our findings indicate that large-scale slice-wise self-supervised learning can yield a unified brain MRI representation that supports diverse neuroimaging tasks without volumetric pretraining or full-network fine-tuning, establishing a scalable foundation for robust and data-efficient brain imaging analysis.
Photon-counting CT (PCCT) provides superior image quality with higher spatial resolution and lower noise compared to conventional energy-integrating CT (EICT), but its limited clinical availability restricts large-scale research and clinical deployment. To bridge this gap, we propose SUMI, a simulated degradation-to-enhancement method that learns to reverse realistic acquisition artifacts in low-quality EICT by leveraging high-quality PCCT as reference. Our central insight is to explicitly model realistic acquisition degradations, transforming PCCT into clinically plausible lower-quality counterparts and learning to invert this process. The simulated degradations were validated for clinical realism by board-certified radiologists, enabling faithful supervision without requiring paired acquisitions at scale. As outcomes of this technical contribution, we: (1) train a latent diffusion model on 1,046 PCCTs, using an autoencoder first pre-trained on both these PCCTs and 405,379 EICTs from 145 hospitals to extract general CT latent features that we release for reuse in other generative medical imaging tasks; (2) construct a large-scale dataset of over 17,316 publicly available EICTs enhanced to PCCT-like quality, with radiologist-validated voxel-wise annotations of airway trees, arteries, veins, lungs, and lobes; and (3) demonstrate substantial improvements: across external data, SUMI outperforms state-of-the-art image translation methods by 15
Gadolinium-based contrast agents (GBCAs) are commonly employed with T1-weighted (T1w) MRI to enhance lesion visualization but are restricted in patients at risk of nephrogenic systemic fibrosis. In addition, variations in GBCA administration can introduce imaging inconsistencies. This study develops an efficient 3D deep-learning framework to generate T1-contrast enhanced images (T1C) from pre-contrast multiparametric MRI. We propose the 3D latent rectified flow (T1C-RFlow) model for generating high-quality T1C images. First, T1w and T2-FLAIR images are input into a pretrained autoencoder to acquire an efficient latent space representation. A rectified flow diffusion model is then trained in this latent space representation. The T1C-RFlow model was trained on a curated dataset comprised of the Brain Tumor Segmentation (BraTS) 2024 glioma (GLI; 1480 patients), meningioma (MEN; 1141 patients), and metastases (MET; 1475 patients) datasets. Selected patients were split into training (N = 2860), test (N = 614), and validation (N = 612) sets. Model performance was evaluated with the normalized mean squared error (NMSE) and structural similarity index measure (SSIM). Both qualitative and quantitative results demonstrate that the T1C-RFlow model outperforms benchmark 3D models (pix2pix, denoising diffusion probability models (DDPM), Diffusion Transformers (DiT-3D) trained in the same latent space. T1C-RFlow achieved the following metrics - GLI: NMSE 0.044 ± 0.047, SSIM 0.935 ± 0.025; MEN: NMSE 0.046 ± 0.029, SSIM 0.937 ± 0.021; MET: NMSE 0.098 ± 0.088, SSIM 0.905 ± 0.082. In a blinded reader study of 15 patients (5 GLI, 5 MEN, 5 MET), T1C-RFlow received the highest diagnostic-quality scores across all tumor types (3.80 ± 0.45, 3.20 ± 0.45, and 2.60 ± 0.89 on a 5-point Likert scale), significantly outperforming all baseline methods (p < .05). Further studies showed T1C-RFlow to have the best tumor reconstruction performance and significantly faster denoising times (6.9 s volume-1, 200 steps) than conventional DDPM models in both latent space (37.7 s, 1000 steps) and patch-based in image space (4.3 h volume-1). Our proposed method generates synthetic T1C images that closely resemble radiological features of ground truth T1C in much less time than previous diffusion models. Further development may permit a practical method for contrast-agent-free MRI for brain tumors. Code is made available at https://github.com/zacheidex/An-Efficient-3D-Latent-Diffusion-Model-for-T1-contrast-Enhanced-MRI-Generation.
The traditional grid-following (GFL) and grid-forming (GFM) control for modular multilevel converters (MMC) face potential oscillation risks under wide grid strength, where the grid strength is quantified by the short-circuit ratio (SCR). The existing improved methods have a limited expansion of stability margin, which cannot meet the operation requirements when SCR changes in wide range. Therefore, a multi-mode fusion control (MFC) strategy is proposed for MMC. In the proposed control strategy, the fusion point of GFL and GFM control is designed at the input end of the multiplexed current control loop, which helps reduce the number of controllers. The control ratios of GFL and GFM are determined by the fusion factor associated with SCR, which is designed by analogy with the Logistic equation. Then, the proposed control strategy comprises five operation modes in MMC, among which the equal-proportion fusion control is a residence design thereby preventing interference resulting from the frequent variations of the fusion factor. And under the adaptive fusion mode, the proportion of GFL and GFM is adjusted through fusion factor. The operation range is expanded with higher combination flexibility. Additionally, the MFC strategy includes pure GFL and GFM control modes, which makes MMC maintain stability in extremely strong grid or weak grid. Impedance model is also established to analyze the small-signal stability of the proposed control strategy. Both simulation and experimental results verify the effectiveness and superiority of the proposed MFC strategy.
The accurate and fast short-circuit ratio (SCR) online estimation (SOE) technology is critical to the stable operation of the grid-connected inverters (GCI), particularly in a wide range of grid strength. However, the existing SOE methods cannot balance the indicators of grid-friendliness, accuracy, and rapidity. Furthermore, most strategies are designed exclusively for either grid following (GFL) or grid forming (GFM) control, which cannot be directly applied to other control structures. To solve the above-mentioned questions, a grid-friendly universal short-circuit ratio online estimation (FUSOE) strategy is proposed in this paper. FUSOE strategy has grid friendliness, whose estimation function is derived from the steady-state variables of the system. Therefore, FUSOE has no additional disturbance to the grid, making the continuous SCR online estimation possible. Moreover, the inner current loop reference serves as a middle bridge, breaking through the application limitations of the traditional method between GFL and GFM modes while ensuring estimation rapidity. Therefore, the proposed FUSOE strategy can adapt to multi control structures without second design. The proposed method is verified in simulations and experiments under grid strength fluctuations and different controls. The simulation and experimental results verify the correctness of the theoretical analysis and the advantage of FUSOE strategy in accuracy, rapidity, and universality.
Background:High-resolution MRI is essential for accurate diagnosis and treatment planning, but its clinical acquisition is often constrained by long scanning times, which increase patient discomfort and reduce scanner throughput. While super-resolution (SR) techniques offer a post-acquisition solution to enhance resolution, existing deep learning approaches face trade-offs between reconstruction fidelity and computational efficiency, limiting their clinical applicability. Purpose:This study aims to develop an efficient and accurate deep learning framework for MRI super-resolution that preserves fine anatomical detail while maintaining low computational overhead, enabling practical integration into clinical workflows. Materials and Methods:We propose a novel SR framework based on multi-head selective state-space models (MHSSM) integrated with a lightweight channel multilayer perceptron (MLP). The model employs 2D patch extraction with hybrid scanning strategies (vertical, horizontal, and diagonal) to capture long-range dependencies while mitigating pixel forgetting. Each MambaFormer block combines MHSSM, depthwise convolutions, and gated channel mixing to balance local and global feature representation. The framework was trained and evaluated on two distinct datasets: 7T brain T1 MP2RAGE maps (142 subjects) and 1.5T prostate T2w MRI (334 subjects). Performance was compared against multiple baselines including Bicubic interpolation, GAN-based (CycleGAN, Pix2pix, SPSR), transformer-based (SwinIR), Mamba-based (MambaIR), and diffusion-based (I2SB, Res-SRDiff) methods. Results:The proposed model demonstrated superior performance across all evaluation metrics while maintaining exceptional computational efficiency. On the 7T brain dataset, our method achieved the highest structural similarity (SSIM: 0.951±0.021) and peak signal-to-noise ratio (PSNR: 26.90±1.41 dB), along with the best perceptual quality scores (LPIPS: 0.076±0.022; GMSD: 0.083±0.017). These results represented statistically significant improvements over all baselines (p < 0.001), including a 2.1% SSIM gain over SPSR and a 2.4% PSNR improvement over Res-SRDiff. For the prostate dataset, the model similarly outperformed competing approaches, achieving SSIM of 0.770±0.049, PSNR of 27.15±2.19 dB, LPIPS of 0.190±0.095, and GMSD of 0.087±0.013. Notably, our framework accomplished these results with only 0.9 million parameters and 57 GFLOPs, representing reductions of 99.8% in parameters and 97.5% in computational operations compared to Res-SRDiff, while also substantially outperforming SwinIR and MambaIR in both accuracy and efficiency metrics. Conclusion:The proposed framework provides a computationally efficient yet accurate solution for MRI super-resolution, delivering well-defined anatomical details and improved perceptual fidelity across anatomically distinct datasets. By significantly reducing computational demands while maintaining state-of-the-art performance, the model offers strong potential for feasibility toward clinical translation and scalable integration into future imaging workflows.
PURPOSE:We investigated the effect of scanning speed, beam configuration, and dose-rate modeling on the FLASH effect in postmastectomy proton transmission beams planning and evaluated the potential of spot scanning path optimization for enhancing the FLASH effect. METHODS AND MATERIALS:Five patients with left-sided postmastectomy breast cancer (32 Gy/5 fractions) were retrospectively replanned with single-energy (249 MeV) tangential transmission beams, supplemented by a clinical en face beam for dose homogenization. FLASH evaluation employed 2 models: Krieger's FLASH effectiveness model (FEM) and Folkerts' average dose-rate (ADR) framework. Plans were simulated under conventional pencil beam scanning, split-field, and optimized spot sequences (using genetic algorithm [GA]), with vertical scan speeds varied from 10 to 20 mm/ms. FLASH effect in normal tissues was quantified by the percentage of voxels meeting the threshold (≥4 Gy at ≥40 Gy/s). A dose adjustment factor of 0.67 was applied to voxels meeting FLASH criteria to compute FLASH-weighted dose metrics in normal tissues within the chest wall target, whereas the physical dose to tumor cells remained unchanged. RESULTS:The FLASH effect showed high sensitivity to scanning patterns and model selection. Increasing vertical scan speed from 10 to 20 mm/ms increased the FLASH in clinical target volume (CTV) by 22% (ADR) and 12% (FEM), whereas in skin, it rose from 41.4% to 58.8% (ADR) and 8.4% to 13.1% (FEM). Split-field delivery improved the temporal distance between the vertical columns of the spot scanning pattern, yielding a superior FLASH effect, which is up to a 9.2 Gy reduction in CTV (FLASH-corrected) Dmean(mean dose) with the ADR model. GA-based optimization shortened scan time and provided FLASH comparable with split-field delivery, with a CTV Dmean reduction of 7.87 Gy (ADR GA), with skin Dmean reductions of 2 to 3 Gy. CONCLUSION:This study demonstrates that FLASH outcomes are highly sensitive to scanning trajectory, scan speed, and model selection. Beyond these parameters, optimizing spot delivery using a path minimizer, such as GA, can further improve the dose-rate distribution in healthy voxels across all scenarios.
Deep learning-based registration methods have made significant progress in recent years, but their performance is often limited by the size and diversity of training data. This paper explores the effectiveness of leveraging a Vision Foundation Model pre-trained on a very large 3D magnetic resonance imaging (MRI) dataset to improve the performance of downstream registration tasks. We employ two leading registration architectures (TransMorph and SwinUNETR) and systematically compare the performance of various parameter initialization strategies (training from scratch, VoCo-SSL, and our model's pre-trained weights) on three publicly available registration datasets (IXI, OASIS, and ACDC). Experimental results demonstrate that the pre-training strategy significantly impacts registration performance, particularly when the network architecture fully leverages the encoder's pretrained features. Specifically, a SwinUNETR model initialized with our model's pretrained weights achieved a 6.90% improvement in Dice Similarity Coefficient (DSC) on the OASIS brain MRI dataset compared to training from scratch. The study found that incomplete encoder weight loading can limit the benefits of pretraining, while pretraining data that is highly consistent with the downstream task modality and organ can maximize performance.
Objective. Accurate segmentation of the prostate and dominant intraprostatic lesions (DILs) on magnetic resonance imaging (MRI) is important for prostate cancer radiation therapy treatment planning and targeted dose escalation. However, DIL segmentation remains challenging due to small datasets, institutional bias, and variable imaging protocols. Although the segment anything model (SAM) has shown promise in medical image segmentation, most prior work depends on manual prompts. This study developed a fully automated pipeline that combines localization with a fine-tuned SAM model to segment the prostate and DIL.Approach. Two datasets were utilized: the PI-CAI dataset, comprising 1476 patients, and the cancer imaging archive dataset, comprising 803 patients. The pipeline consisted of two stages: (1) a reinforcement learning-based localization network predicted bounding boxes as segmentation inputs, and (2) a fine-tuned SAM model performed segmentation. Model performance was evaluated using the dice similarity coefficient (DSC), intersection over union (IoU), and detection rates, with additional analysis based on lesion volumes.Main results. The proposed method achieved a mean and median DSC of 0.896 ± 0.070 and 0.915, and an IoU of 0.818 ± 0.100 and 0.844 for prostate segmentation. For DIL segmentation, the mean and median DSC were 0.592 ± 0.192 and 0.636, IoU of 0.446 ± 0.190 and 0.466, with a detection rate of 89%. Four DIL groups were created based on lesion volume percentile. The mean/median DSC and IoU for each volume group are as follows: 0.5-1.0 cubic centimeters (cc): 0.555 ± 0.201/0.562 & 0.414 ± 0.205/0.391; 1.0-1.8 cc: 0.603 ± 0.185/0.660 & 0.454 ± 0.180/0.492; 1.8-4.0 cc: 0.588 ± 0.183/0.627 & 0.439 ± 0.174/0.456; >4.0 cc: 0.621 ± 0.197/0.669 & 0.477 ± 0.197/0.503.Significance. This study presented a fully automated prostate and DIL segmentation framework on MRI by integrating a localization network with fine-tuned SAM. The method achieved robust performance across large multi-institutional datasets and diverse lesion shapes. It shows strong potential for application to clinical workflows for prostate cancer radiation therapy planning and treatment.
Objective . This review investigates the use of cone-beam computed tomography (CBCT) in conjunction with radiomics for external beam radiation therapy (EBRT) in cancer treatment. CBCT, which provides high-resolution, volumetric images, offers a promising tool for precision treatment delivery. By integrating radiomics and quantitative features extracted from CBCT, this review explores potential advancements in tumor characterization, treatment planning, and monitoring treatment responses in personalized cancer therapy. Approach . We conducted this systematic review using the PRISMA (preferred reporting items for systematic reviews and meta-analyses) framework. This study focused on CBCT-only radiomics applications, examining publications in PubMed, Embase, and Scopus databases. The inclusion criteria were strictly peer-reviewed journal articles, resulting in 29 studies being selected for analysis. These studies were divided into two main categories: (1) method development for treatment outcome prediction; (2) verification, validation, and uncertainty quantification (VVUQ) for CBCT-based radiomics. Main Results . The literature encompasses a range of investigations into CBCT-based radiomics for EBRT, covering different cancer types such as head-and-neck squamous cell carcinoma, non-small cell lung cancer, esophageal squamous cell cancer, hepatocellular carcinoma, prostate cancer, and rectal cancer. These studies used radiomics to predict outcomes including tumor response, local failure, tissue toxicity, and patient survival. VVUQ studies addressed the robustness and reproducibility of radiomic features. Furthermore, the emerging field of 4D-CBCT radiomics shows potential in improving image quality. Significance . CBCT-based radiomics presents a promising advancement in personalized radiotherapy, allowing for enhanced cancer prognosis and treatment adaptation. However, challenges of imaging quality and acquisition need to be addressed to ensure consistency and reliability. Future research should focus on standardizing imaging protocols and incorporating multi-institutional collaborations to further validate the clinical applicability of CBCT-based radiomics. Integration of this technology can potentially induce a paradigm shift in personalized cancer radiotherapy. New technologies promise to make CBCT even more valuable in the future.
BACKGROUND:Ultra-high-field 7 Tesla (7T) magnetic resonance imaging (MRI) provides improved resolution and signal-to-noise ratio (SNR) over standard clinical field strengths (1.5T, 3T). However, 7T scanners are costly, scarce, and introduce additional challenges such as susceptibility artifacts. PURPOSE:We propose an efficient transformer-based model (7T-Restormer) to synthesize 7T-like MRI (quantitative T1 MP2RAGE maps) from routine 1.5T or 3T T1-weighted (T1W) images. METHODS:The proposed method leverages an efficient restoration transformer backbone and spatial attention layers to capture long-range dependencies and generate high-quality 7T-like images. Our model was validated on an institutional dataset, comprised of 35 1.5T and 108 3T T1w MRI paired with corresponding 7T T1 maps of patients with confirmed multiple sclerosis (MS). A total of 141 patient cases (32 128 slices) were randomly divided into 105 (25 1.5T and 80 3T) training cases (19 204 slices), 19 (5 1.5T and 14 3T) validation cases (3476 slices), and 17 (5 1.5T and 12 3T) test cases (3145 slices). The synthetic 7T T1 maps were evaluated by comparing their similarity to the ground truth volumes using peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and normalized mean squared error (NMSE), and were compared against ResViT and ResShift. In addition, a blinded reader study was performed in which three clinicians scored the diagnostic image quality of each method on a 5-point Likert scale (1 = non-acceptable, 5 = excellent). Statistical significance (α = 0.05) was assessed using two-sided paired t-tests with Holm correction for multiple comparisons and Wilcoxon signed-rank tests with Holm correction for reader scores. Effect sizes (Cohen's d z ${{d}_z}$ for paired data) were also computed to quantify practical significance. We interpreted ∣ d z ∣ ≈ 0.2 , 0.5 , $\mid {{d}_z}\mid \approx 0.2,\ 0.5,$ and 0.8 $0.8$ as small, medium, and large effects. RESULTS:For 1.5T inputs, 7T-Restormer achieved an NMSE of 0.018 ± 0.003, PSNR of 25.0 ± 0.6 dB, and SSIM of 0.870 ± 0.019; for 3T inputs, the corresponding values were 0.019 ± 0.006, 24.5 ± 1.2 dB, and 0.874 ± 0.027. Across both field strengths, 7T-Restormer provided the lowest NMSE and highest PSNR and SSIM among the three methods, reducing NMSE by approximately 5% and 25% relative to ResShift and ResViT at 1.5T and by 14% and 24% at 3T, respectively, and increasing PSNR by 0.5-1.5 dB (Holm-adjusted p < 0.05, paired Cohen's d z ${{d}_z}$ = 1.37-6.84). Training on mixed 1.5T+3T data improved performance for 1.5T inputs compared with 1.5T-only training (e.g., NMSE 0.018 vs. 0.019; p = 0.001, d z = - 4.29 ${{d}_z} = - 4.29$ ), without degrading performance for 3T inputs (p = 0.630, d z = - 0.14 ${{d}_z} = - 0.14$ for NMSE). In the blinded reader study, 7T-Restormer received higher diagnostic quality scores (3.50 ± 0.28) than ResShift (3.00 ± 0.27) and ResViT (2.33 ± 0.27), corresponding to mean improvements of 0.5-1.2 points on the 5-point scale (Holm-adjusted p = 0.014 and 0.004; paired Cohen's d z = 1.27 ${{d}_z} = 1.27$ and 3.24, respectively). CONCLUSION:We propose a novel method for predicting quantitative 7T MP2RAGE maps from 1.5T and 3T T1W scans with higher quality than existing state-of-the-art methods. Future development and application of our approach may enhance diagnostic accuracy, treatment planning, and facilitate downstream tasks by making the benefits of 7T MRI more accessible to standard clinical workflows.
CBCT is widely used to support adaptive image-guided radiotherapy (IGRT) in gynecologic (GYN) patients but is constrained by long acquisition times, additional radiation dose, and potential gantry-device collisions. Orthogonal Xray projections offer a practical alternative: they are low dose, fast, and already obtained in many IGRT workflows for setup verification. In proton therapy, anterior-posterior (AP) and lateral (LAT) projections are routinely acquired for daily positioning; in that setting, leveraging these images for CBCT reconstruction introduces no additional imaging protocol or dose burden. We propose a geometry-constrained dual-domain diffusion network (GCDD-DDPM) to reconstruct CBCT volumes directly from orthogonal X-rays. The framework consists of a Projection-DDPM, a deterministic Geometric Transformation Module (GTM), and a CBCT-DDPM, jointly optimized end-to-end to enforce cross-domain consistency. Patient-specific training was conducted on daily CBCTs from 12 GYN patients, with 24 scans per patient for training and 4 for testing. All projections were simulated under scanner-consistent geometry to enable a controlled feasibility evaluation. Quantitative results show that GCDD-DDPM achieved MAE = 25.89 +/- 3.82 HU, SSIM = 0.98 +/- 0.02, PSNR = 47.21 +/- 1.68 dB, significantly outperforming a conventional DDPM baseline (MAE = 37.43 +/- 12.35 HU, SSIM = 0.96 +/- 0.04, PSNR = 38.89 +/- 2.47 dB). Qualitative analysis further confirms improved anatomical fidelity in challenging regions. These findings indicate that GCDD-DDPM enables accurate, robust, and geometryconsistent CBCT reconstruction from orthogonal X-rays, providing a promising pathway toward low-dose, time-efficient volumetric imaging to support adaptive radiotherapy in GYN patients.
Metro stray current invades into the utility side, resulting in DC bias phenomenon in utility transformers. Existing mitigation measures exhibit limitations in suppressing the DC bias current. This paper proposes a rail-side source mitigation measure. The mechanism and calculation measures for both stray current and neutral DC current are discussed in detail. Simulation results indicate that reducing running rail resistance or increasing rail-to-ground transition resistance can effectively mitigate the transformer DC bias current magnitude. A sensitivity analysis is conducted to evaluate the marginal benefit and elasticity of the two key rail parameters, providing engineering guidance for prioritizing mitigation measures.
Computed tomography (CT) is extensively used for accurate visualization and segmentation of organs and lesions. While deep learning models such as convolutional neural networks (CNNs) and vision transformers (ViTs) have significantly improved CT image analysis, their performance often declines when applied to diverse, real-world clinical data. Although foundation models offer a broader and more adaptable solution, their potential is limited due to the challenge of obtaining large-scale, voxel-level annotations for medical images. In response to these challenges, prompting-based models using visual or text prompts have emerged. Visual-prompting methods, such as the Segment Anything Model (SAM), still require significant manual input and can introduce ambiguity when applied to clinical scenarios. Instead, foundation models that use text prompts offer a more versatile and clinically relevant approach. Notably, current text-prompt models, such as the CLIP-Driven Universal Model, are limited to text prompts already encountered during training and struggle to process the complex and diverse scenarios of real-world clinical applications. Instead of fine-tuning models trained from natural imaging, we propose OpenVocabCT, a vision-language model pretrained on large-scale 3D CT images for universal text-driven segmentation. Using the large-scale CT-RATE dataset, we decompose the diagnostic reports into fine-grained, organ-level descriptions using large language models for multi-granular contrastive learning. We evaluate our OpenVocabCT on downstream segmentation tasks across 14 public datasets and 1 institutional dataset for organ and tumor segmentation, demonstrating the superior performance of our model compared to existing methods. All code, datasets, and models will be publicly released at https://github.com/ricklisz/OpenVocabCT.
Recent advances in large language models (LLMs) have enabled general-purpose systems to perform increasingly complex domain-specific reasoning without extensive fine-tuning. In the medical domain, decision-making often requires integrating heterogeneous information sources, including patient narratives, structured data, and medical images. This study positions GPT-5 as a generalist multimodal reasoner for medical decision support and systematically evaluates its zero-shot chain-of-thought reasoning performance on both text-based question answering and visual question answering tasks under a unified protocol. We benchmark GPT-5, GPT-5-mini, GPT-5-nano, and GPT-4o-2024-11-20 against standardized splits of MedQA, MedXpertQA (text and multimodal), MMLU medical subsets, USMLE self-assessment exams, and VQA-RAD. Results show that GPT-5 consistently outperforms all baselines, achieving state-of-the-art accuracy across all QA benchmarks and delivering substantial gains in multimodal reasoning. On MedXpertQA MM, GPT-5 improves reasoning and understanding scores by +29.26% and +26.18% over GPT-4o, respectively, and surpasses pre-licensed human experts by +24.23% in reasoning and +29.40% in understanding. In contrast, GPT-4o remains below human expert performance in most dimensions. A representative case study demonstrates GPT-5's ability to integrate visual and textual cues into a coherent diagnostic reasoning chain, recommending appropriate high-stakes interventions. Our results show that, on these controlled multimodal reasoning benchmarks, GPT-5 moves from human-comparable to above human-expert performance. This improvement may substantially inform the design of future clinical decision-support systems. We make the code public at the GPT-5-Evaluation.
Medical vision–language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language. This paradigm is fundamentally limited in clinical scenarios, where accurate answers often depend on subtle, localized visual evidence that cannot be reliably preserved in static embeddings. We propose MedLVR, a latent visual reasoning framework that introduces an explicit visual evidence state into autoregressive decoding. Instead of relying solely on text-based intermediate reasoning, MedLVR interleaves a short latent reasoning segment within the decoder by reusing hidden states as continuous latent steps, enabling iterative preservation and refinement of query-relevant visual evidence before answer generation. To support effective visual supervision, we adopt a two-stage training strategy: region of interest (ROI)-supervised fine-tuning aligns latent states with clinically relevant image evidence, and Visual-Latent Policy Optimization (VLPO) further optimizes latent reasoning and answer generation under outcome-level rewards. Experiments on OmniMedVQA and five external medical VQA benchmarks show that MedLVR consistently outperforms recent reasoning baselines and improves the average score over the Qwen2.5-VL-7B backbone from 48.3% to 53.4%. These results show that latent visual reasoning provides an effective mechanism for preserving diagnostically relevant visual evidence and improving the reliability of medical VQA.
Computed tomography (CT) is a central to three-dimensional medical imaging, yet CT-based artificial intelligence remains fragmented across task-specific models for segmentation, classification, registration, and report analysis. Here we present FlexiCT, a family of CT foundation models trained by agglomerative continual pretraining on 266,227 CT volumes from 56 publicly available datasets, forming a large-scale public resource for CT representation learning. FlexiCT uses agglomerative pretraining across three stages: two-dimensional axial pretraining, three-dimensional anatomical pretraining and report-guided semantic alignment. This training strategy supports slice-level, volume-level and vision-language analysis. Across five downstream task families (segmentation, classification, registration, vision-language understanding and clinical retrieval), FlexiCT matches or exceeds prior task-specific approaches on multiple benchmarks. Its embeddings further organize CT scans along gradients associated with various tumor stages, suggesting that CT foundation models can capture imaging features relevant to disease phenotype characterization. Code is available at https://github.com/ricklisz/FlexiCT
The virtual-return-based DC traction power supply system (VRTPS) has been studied to solve the stray current and rail potential issues in urban rail transit. By utilizing the power electronic devices (PEDs) and return line, VRTPS constructs the zero-resistance virtual return path (VRP) by diverting traction current from the running rail into this newly established path. However, the mitigation performance of existing VRTPS control strategies is still suboptimal, leading to incomplete suppression of stray current and rail potential within the train-occupied section. Thus, the reverse running rail current control strategy of VRTPS is proposed to further reduce the stray current and rail potential. Firstly, the topology and modeling of VRTPS are investigated to analyze the inherent limitations of the ideal zero-resistance VRP. Secondly, the detailed analysis of the reverse running rail current control for VRTPS is presented, which optimizes the running rail current distribution with the variable-resistance VRP. Thirdly, the comparative researches on both the stray current and rail potential distribution are conducted between the conventional system and VRTPS with different control strategies. Finally, the correctness and effectiveness of VRTPS with proposed strategy on the stray current and rail potential mitigation are validated through simulation and experimental results.
BACKGROUND:Cone-beam computed tomography (CBCT) has gained considerable traction in medical imaging due to its ability to provide on-board three-dimensional (3D) volumetric imaging with relatively low radiation dose, aiding in tracking anatomical changes. However, artifacts such as scatter, noise, and respiratory motion, especially in thoracic imaging, complicate accurate segmentation of organs-at-risk (OARs). Additionally, patient-specific anatomical and tumor changes during treatment further challenge precise delineation, potentially affecting radiation therapy outcomes. PURPOSE:This study aims to address the challenges in accurate segmentation of OARs in CBCT imaging, considering the impact of imaging artifacts and patient-specific anatomical and tumor changes, which are critical for improving the efficacy of radiation therapy. METHODS:A 3D patient- and fraction-specific scalable and transferable U-Net (PFS-STU-Net) was developed using a large-scale pre-trained STU-Net with LoRA-based parameter-efficient fine-tuning. The framework first adapts a general STU-Net model (G-STU-Net) to the CBCT domain using paired planning CT (pCT) and deformed CBCT (dCBCT) images, then incrementally fine-tunes the model with patient-specific CBCT fractions to capture anatomical evolution throughout treatment. Segmentation accuracy was evaluated using dice similarity coefficient (DSC), 95th-percentile Hausdorff distance (HD95), and mean surface distance (MSD) on 33 lung SBRT patients (5 fractions each) and an external 10-patient cohort. Statistical comparisons were performed using paired Wilcoxon signed-rank tests with Benjamini-Hochberg false discovery rate (FDR) correction for multiple comparisons, with statistical significance set at p < 0.05 $p < 0.05$ . RESULTS:Compared with the pre-trained STU-Net, the proposed PFS-STU-Net achieved statistically significant improvements across all organs ( p < 0.05 $p < 0.05$ , FDR-corrected), with large effect sizes (Cliff's δ > 0.8 $\delta > 0.8$ for the esophagus and spinal cord). DSC increased from 0.60 to 0.87 for the esophagus, 0.79 to 0.92 for the heart, 0.83 to 0.98 for the left lung, 0.84 to 0.98 for the right lung, and 0.74 to 0.97 for the spinal cord. Average HD95 and MSD were reduced by more than 50%. External validation demonstrated consistent cross-institutional performance (average DSC = 0.91 ± 0.07). CONCLUSIONS:Integrating large-scale pretraining with LoRA-based patient- and fraction-specific adaptation significantly enhances CBCT multi-organ segmentation accuracy and clinical transferability. The PFS-STU-Net framework demonstrates robust performance across institutions and provides a scalable, efficient solution for real-time CBCT-guided adaptive radiotherapy.