
This survey presents a structured review of deep learning-based techniques for image data hiding, proposing a three-paradigm taxonomy organized by the method’s operational relationship to the carrier image. We classify existing methods into modification-based, synthesis-based, and logic-based approaches. In the modification-based tier, we trace the architectural progression from foundational Convolutional Neural Networks and Generative Adversarial Networks to high-capacity Invertible Neural Networks and Transformers, analyzing their distinct trade-offs between embedding capacity, imperceptibility, and robustness. In the synthesis-based tier, we examine how Diffusion Probabilistic Models and generative adversarial frameworks reframe data hiding as a carrier generation problem rather than a pixel editing task. This paradigm encompasses both generative steganography (where carriers are synthesized from scratch) and proactive watermarking (where provenance is embedded during AI content generation). In the logic-based tier, we review zero-watermarking and coverless steganography, where ownership is established through feature extraction and semantic mapping without modifying any image, a critical property for sensitive domains such as medical imaging. Finally, we identify four persistent infrastructure gaps: benchmarking fragmentation, narrow robustness evaluation, domain generalization failures, and computational infeasibility that prevent real-world deployment despite architectural progress, and we propose concrete research directions.
Pulmonary nodule malignancy classification requires effective integration of heterogeneous imaging features and contextual information among nodules. Existing models mainly analyze nodules independently and may overlook inter-nodule relationships. We propose RGGA-Net, a radiomics-guided graph attention network that integrates deep imaging features and radiomics representations through cross-modal interaction and graph-based reasoning. A learned edge gate is introduced to adaptively modulate graph message passing, and an anchor regularization strategy is used to improve representation stability. RGGA-Net was evaluated on the LUNA25 dataset using patient-level splitting with an internal held-out test set and further assessed on the LIDC-IDRI cohort under cross-dataset evaluation. On the internal test set, RGGA-Net achieved an AUC of 0.8910 and a PR-AUC of 0.5173. External evaluation on LIDC-IDRI demonstrated moderate discrimination (AUC = 0.7023). Gate–Anchor ablation analysis showed a favorable interaction pattern between the learned gate and anchor regularization, although statistical superiority was not established. These findings suggest that radiomics-guided graph attention provides a feasible framework for incorporating inter-nodule information into malignancy classification, while further multi-center validation remains necessary.
Small-scale pedestrians in UAV imagery often exhibit limited pixel coverage, weak texture, and severe background interference, while progressive network downsampling can further degrade their short-side structures and increase missed detections. To address this problem, we propose DGF-YOLO, a degradation-guided feature enhancement method built on YOLOv12n. The method introduces a degradation-level criterion to identify the feature stage at which a pedestrian first undergoes significant structural degradation and, based on the resulting statistics, incorporates a high-resolution P2 detection head. It further employs a Directional Structure-Aware module to enhance local, horizontal, and vertical structural cues through adaptive multi-branch fusion, a Degradation-Guided Attention module to learn a degradation guidance map under explicit supervision and reweight degradation-sensitive regions, and a Fine-Grained Structure Preservation module to retain local contours and contextual details using depthwise and dilated convolutions. On the single-class pedestrian detection task constructed from VisDrone2019-DET, DGF-YOLO achieves 70.4% precision, 51.8% recall, 59.7% mAP50, and 26.9% mAP50-95, improving the YOLOv12n baseline by 6.6, 7.4, 10.3, and 6.5 percentage points, respectively. The results suggest that the proposed feature enhancement strategy helps reduce missed detections associated with structural degradation in small-scale pedestrians.
Patients with implanted cardiac devices are a rapidly growing imaging population, and electrocardiogram-gated cardiac computed tomography (CT) is increasingly used to characterize device geometry, multi-device relationships, and dynamic behavior across the cardiac cycle. Cinematic rendering (CR) is a photorealistic three-dimensional (3D) visualization technique for cardiac CT whose established contribution in this population is communicative: it conveys 3D device geometry and material distinctions within a single rendered volume. We describe a demonstrative case series extending CR across the cardiac cycle—time-resolved “4D cine” CR—to depict dynamic device behavior and time-resolved multi-device interaction in a single volume; this is an illustrative technical experience rather than a systematic evaluation of diagnostic performance. Illustrative examples include an EVOQUE transcatheter tricuspid valve rendered together with concurrent surgical mitral and transcatheter aortic valves, a left atrial appendage occlusion device, a normally positioned Impella catheter, and a HeartMate 3 left ventricular assist device (LVAD). Across cases, 4D cine CR feasibility scaled inversely with metallic burden—the aggregate volume and radiodensity of metallic device components within the scan field—with renderings informative for low-metal nitinol and catheter devices but substantially degraded by streak artifact in high-metal LVAD housings. This relationship was observed qualitatively in a small selected series and is offered as an initial observation rather than an established characteristic of the technique. We discuss current limitations and emerging directions such as photon-counting detector CT, metal artifact reduction, and artificial-intelligence-assisted post-processing that may extend 4D cine CR in this population.
Due to the limited dynamic range of imaging sensors, most cameras can only capture low-dynamic-range (LDR) images. Multi-exposure fusion (MEF) is an effective technique for generating high-dynamic-range (HDR) images. However, to simultaneously preserve texture details and global exposure, most existing methods primarily rely on more complex models to improve performance, resulting in higher computational costs and longer processing times. To address this issue, we propose a real-time unsupervised MEF network driven by prior knowledge. To this end, a hierarchical feature extraction module is designed that utilizes filtering operations to decompose the source images into base layers and detail layers. Features are extracted from each layer separately to reduce the difficulty of extracting effective features. Then, the receptive field of feature maps is expanded by dilated convolutions, and a window-based self-attention mechanism is applied to perform context modeling, achieving effective contextual modeling with low computational cost. Subsequently, texture features and global features are extracted separately to enable the model to maintain both local texture clarity and global smoothness. In addition, a one-dimensional lookup table is utilized to accelerate the inference process. Comprehensive experiments are conducted to verify the effectiveness of the proposed method. The subjective evaluation results demonstrate that the fused images exhibit superior visual quality, while objective experiments further quantify its superior performance, demonstrating that the proposed method effectively reduces computation time.
Aviation ground safety and protective devices are critical for flight safety; however, their unintentional retention on aircraft after maintenance remains a persistent risk. Existing deep learning-based approaches for aviation safety have predominantly followed a reactive paradigm, detecting FOD on runways or inspecting the aircraft for inadvertently retained tools post-maintenance. In contrast, this paper advocates a proactive philosophy: using a neural network to recognize and inventory all ground safety and protective devices immediately after maintenance closure, thereby preventing retention incidents at their source. However, realizing this proactive verification is technically challenging—object detection for these devices often suffers from loss of fine-grained detail due to downsampling and inherently sparse semantic information of the targets. To this end, we propose an Information Retention and Feature Screening Synergistic Network (RS-Net) grounded in information bottleneck theory. The network comprises a main branch that enhances discriminative features through attention-guided screening, and an auxiliary branch, used only during training, that preserves fine-grained spatial details via information-retentive convolutions. A Dual-State Region Refinement Module (DRM) provides configurable support for both branches, decoupling the conflicting objectives of background compression and detail preservation. Experiments on a self-constructed dataset collected from real airline maintenance operations demonstrate that RS-Net substantially outperforms the strong YOLOv9 baseline, achieving gains of 4.531% in F1-score, 2.533% in mAP0.5, and 1.429% in mAP0.5:0.95. Cross-dataset experiments further validate its strong generalization capability.
Histopathology foundation models (FMs) have become widely used as patch-level feature extractors in computational pathology (CPath), where domain shift is a central challenge, yet their generalization ability across acquisition centers and scanning platforms remains insufficiently studied. In this work, we evaluate the domain generalization of ten state-of-the-art FMs on two multi-source datasets with different supervision settings: SemiCOL, a colorectal cancer cohort of 499 whole-slide images (WSIs) for weakly labeled slide-level binary tumor classification, and BEETLE, a breast cancer cohort of 583 WSIs for patch-level four-class tissue classification. In an ablation-style setting, FMs are used as patch-level feature extractors, with patch embeddings mean-pooled into slide-level representations for SemiCOL, and a lightweight multi-layer perceptron trained for slide-level and patch-level classification on SemiCOL and BEETLE, respectively. To test FM domain generalization, we use three evaluation protocols: a Baseline source-mixed 5-fold cross-validation (CV) and two leave-source-out CV settings that assess cross-center and cross-scanner performance. On SemiCOL, all FMs achieve near-saturated performance, indicating stable performance under acquisition-source domain shift for slide-level Tumor vs. Benign classification. In contrast, BEETLE reveals clear generalization gaps, with scanner-induced domain shift more challenging than center-induced domain shift, and class-wise results showing that performance losses concentrate in epithelial discrimination. Overall, Virchow2 shows the strongest robustness across all evaluated protocols. These findings show that standard source-mixed CV can overestimate domain generalization across centers and scanners, and that FM choice matters in multi-source cohorts, especially for more challenging CPath tasks, where cross-domain failures are more visible.
Systematic optimization of pediatric brain computed tomography (CT) protocols remains challenging because clinical data are limited for children younger than 5 years. Sixty-nine non-contrast brain CT examinations in children aged 5 years or younger were retrospectively analyzed. Multiple linear regression was used to assess associations between dose–length product (DLP) and age group, sex, body mass index (BMI), and scanner group. In univariable analyses, DLP differed significantly by age group (p < 0.001) and scanner group (p < 0.001). In the multivariable model, only age group remained significantly associated with DLP; examinations in children aged 1–5 years showed an adjusted DLP increase of 158.71 mGy·cm compared with examinations in children younger than 1 year (p < 0.001). BMI and scanner group were not independently associated with DLP. Model stability was supported by residual normality and absence of multicollinearity. This study found that age group was the only measured variable significantly associated with DLP after adjustment, whereas BMI, sex, and scanner group were not significant. These findings highlight the importance of considering age when interpreting dose variation in pediatric brain CT.
Non-rigid image registration is a fundamental problem in medical imaging and a representative example of continuous transformation modeling in image processing. Diffeomorphic registration methods, such as Large Deformation Diffeomorphic Metric Mapping (LDDMM) and its PDE-constrained variants (PDE-LDDMM), provide mathematically grounded formulations with strong geometric guarantees for transformation quality. However, existing approaches face persistent trade-offs between numerical stability, accuracy, and computational efficiency. Recent work has explored implicit neural representations (INRs) and neural ordinary differential equations (NODEs) as flexible neural representations for modeling continuous transformations. Despite their increasing adoption, their practical behavior and limitations in diffeomorphic registration remain insufficiently understood. In this paper, we present a unified formulation of INR- and NODE-based registration methods within LDDMM and PDE-LDDMM, enabling a systematic and controlled comparison across architectures, sampling strategies, and numerical solvers. Our analysis reveals fundamental trade-offs between these approaches. In particular, we show that MLP-based INR formulations introduce significant computational overhead and rely on sampling strategies that can degrade smoothness and lead to the increased occurrence of non-diffeomorphic transformations at higher resolutions. Moreover, these approximations do not fully alleviate the computational cost, with some variants exceeding the costs of expensive classical optimization-based methods. In contrast, NODE-based formulations and downsampling strategies consistently provide transformations with more controlled Jacobian extrema while maintaining competitive computational performance. Among the evaluated methods, the original NODE-LDDMM and NODE-PDE-LDDMM formulations achieve the most favorable trade-offs between registration accuracy, geometric consistency, and computational efficiency. These findings provide clear insights into the design of neural representations for continuous transformation modeling, with practical implications for diffeomorphic registration and computational anatomy applications.
Accurate and up-to-date knowledge of land use and land cover represents one of the central challenges in spatial planning and landscape sciences. In this context, the present work introduces GANCIU (Geospatial Analysis with Neural Classification and Image Understanding), an original hybrid pipeline for the automatic extraction of man-made infrastructure from high-resolution satellite imagery. The primary methodological contribution lies in the sequential integration of four technologically heterogeneous components: a per-pixel Random Forest classifier, a guided image modulation step, edge detection via the Mumford–Shah variational functional solved through the Ambrosio–Tortorelli approximation, and final object delineation via the Segment Anything Model (SAM). Each component does not operate independently but conditions and informs the next: The RF probability map guides the modulation, which in turn directs the sensitivity of the variational step exclusively towards regions of interest; the AT edges provide spatial prompts to SAM, for which its masks are finally filtered by the RF probability in an adaptive manner through a Gaussian Mixture Model. This progressive conditioning scheme constitutes the architectural core of GANCIU and distinguishes it from approaches that combine classification and segmentation in parallel or in purely sequential fashion with each stage conditioning the next but without any reverse correction between them. The Random Forest classifier was trained on 44 manually annotated scenes, geographically disjoint from the twelve independent scenes used for quantitative validation. This validation, based on an instance matching protocol (precision, recall, F1 score, and IoU), confirms the contribution of the full pipeline over a Random-Forest-only baseline: Pooled false positives fall by close to two orders of magnitude (from 8320 to 209), while true positives rise nearly twentyfold (from 5 to 95), with a mean IoU of 0.742 ± 0.060 on correctly matched objects. Notably, the entire pipeline—including SAM-based segmentation—runs end-to-end on a modest, GPU-free consumer laptop (four logical CPU cores, under 16 GB RAM), demonstrating that competitive infrastructure-extraction performance does not require specialised computing hardware.
Automated lumbar computed tomography (CT) Hounsfield unit (HU) measurement can support opportunistic osteoporosis screening, but heterogeneous Digital Imaging and Communications in Medicine (DICOM) inputs often disrupt automated workflows before measurement. We evaluated a local Qwen2.5-VL-7B vision-language model (VLM)-assisted front end for suitability routing, input planning, and HU-traceable DICOM-derived field-of-view/orientation adaptation upstream of an unchanged single-slice lumbar CT HU workflow. The VLM was restricted to routing and planning. In a 20-case FUJIFILM pilot cohort, observed HU output availability was higher with the front-end-assisted route than with direct processing (12/20 versus 5/20), while paired automatic-success cases showed excellent agreement with post hoc manual ImageJ (version 1.54g) measurements (ICC(A,1) = 0.997; MAE = 4.64 HU). Public DICOM stress testing further demonstrated improved workflow robustness across heterogeneous datasets. These findings support the use of a deployment-oriented front-end strategy to improve the auditability and robustness of automated lumbar CT HU measurement while preserving an unchanged downstream measurement workflow.
Pre-wash sorting of medical textiles is essential for hospital infection control, yet accurate stain detection remains challenging because visually apparent stains often have blurred boundaries, whereas dried urine stains lack distinguishable optical features. This study proposes a medical textile stain detection method integrating chemically enhanced visualization with deep semantic segmentation. Dimethylaminocinnamaldehyde (DMACA) was used to convert latent urine stains into chemically developed stains with orange–red visual features. Based on the spatial color difference ΔE in the L*a*b* color space, 0.0183 mol/L was selected as the most suitable DMACA concentration among those tested. A dataset of 1974 images was constructed, including blood stains, chemically developed urine stains, medication stains, and uncontaminated textiles. A cascaded preprocessing strategy was applied to enhance stain boundaries and suppress textile texture noise, after which an Enhanced semantic segmentation model incorporating residual feature extraction, multiscale feature fusion, and transfer learning was used for pixel-level recognition. The IoU values for blood stains, chemically developed urine stains, and medication stains were 88.11%, 82.67%, and 89.62%, respectively. The average time required for image preprocessing and network inference was 15.39 ms per image. An input-level ablation comparison showed that DMACA-based color development increased the urine-stain IoU from 3.07% to 86.23%, demonstrating its substantial contribution to latent urine-stain detection. These results support the feasibility of integrating front-end chemical feature enhancement with back-end semantic segmentation for multiclass medical textile stain recognition under the current experimental conditions.
Vision Transformers, especially Swin Transformer, have become default backbones for various vision tasks but suffer from high memory consumption and training costs. This letter proposes MoR–Swin, a novel architecture that integrates Mixture of Recursions (MoR) into Swin Transformer. An adaptive token-level recursion mechanism dynamically allocates computational depth based on semantic complexity. A recursive window attention module and a lightweight router with load balancing loss are introduced. Extensive experiments on ImageNet classification, COCO detection, and ADE20K segmentation show that MoR–Swin reduces parameters by about 50% and accelerates inference up to twofold at a modest accuracy cost (within about 0.5 points of Swin-B on ImageNet-1K). It provides a new technical pathway for optimizing Vision Transformer models, significantly enhancing their applicability in resource-constrained environments.
Breast cancer is the most common cancer among women, and early detection through mammography is essential for reducing mortality. Artificial intelligence can support radiologists by improving diagnostic accuracy. To develop and evaluate a two-stage ensemble machine learning pipeline for breast cancer diagnosis from digital mammograms. The proposed framework combines image preprocessing, multiple convolutional neural networks trained under different conditions, and a second-stage classifier that integrates the CNN outputs. Several machine learning models and feature selection techniques were evaluated using publicly available mammography datasets. Results: The ensemble approach consistently outperformed the individual CNN models. The MLP classifier achieved the best overall balance between precision and recall, while the heuristic fusion method provided the highest sensitivity. Feature selection reduced model complexity while maintaining comparable performance, and cross-validation confirmed the robustness of the proposed methodology. Combining complementary information from multiple CNNs with classical machine learning improves diagnostic performance and provides a robust framework for computer-aided breast cancer diagnosis. The proposed two-stage ensemble offers an effective and interpretable approach for mammographic breast cancer classification. A demonstration application incorporating Grad-CAM explainability further supports its potential use as a clinical decision-support tool.
Deep learning-based surgical vision systems achieve strong performance in semantic segmentation and phase recognition, but their black-box nature limits traceability in safety-critical clinical settings. Prototype-based networks offer an interpretable alternative by grounding predictions in learned visual exemplars, yet their suitability for surgical video understanding remains insufficiently characterized. We adapted a prototype-based architecture to two surgical datasets, laparoscopic cholecystectomy and robot-assisted minimally invasive esophagectomy (RAMIE), and benchmarked it against conventional baselines. We evaluated a segmentation-only setting, in which prototype size and capacity were ablated, and a multitask setting, in which three strategies for coupling prototype learning to semantic segmentation and surgical phase recognition were compared. Prototype-based models underperformed the conventional baselines across both tasks and datasets. In the segmentation-only setting, the selected prototype configurations achieved Dice scores of 72.23% on Cholecystectomy and 72.07% on RAMIE, compared with 74.35% and 74.02% for the corresponding conventional baselines, and showed weaker boundary agreement. In the multitask setting, the best prototype strategy recovered competitive segmentation performance but remained 6–12 F1 points below the conventional baseline for phase recognition. Qualitatively, prototype activation maps exposed intra-structure decompositions and contextual cues that are not directly available from black-box baselines. Prototype networks provide spatially traceable evidence for surgical scene understanding, but currently trade interpretability for reduced boundary precision and phase-recognition performance. These findings motivate future work on scene-level and temporally aware prototypes for explainable surgical AI.
Accurate segmentation of skin lesions and gastrointestinal polyps is essential for early diagnosis and treatment planning. Currently, Convolutional Neural Networks (CNNs) are limited by local receptive fields, missing small lesions. While Transformers model global context, their quadratic computational complexity incurs high costs. To address these limitations, we propose the Wavelet–Vision Mamba UNet (WVM-UNet), integrating State Space Models (SSMs) for linear-complexity long-range dependencies and wavelet transforms for fine-grained feature extraction. The network employs a Wavelet-based Residual State Space (WRSS) block, combining the multi-scale decomposition of discrete wavelet transforms with Vision Mamba to efficiently capture global features. A Fused Channel–Spatial Attention (FCSA) mechanism is incorporated to adaptively recalibrate feature representations. Additionally, we construct an Encoder–Decoder Semantic Connection (EDSC) to replace traditional skip connections, effectively bridging the semantic gap between cross-level features. Experimental results on multiple public datasets demonstrate the competitive performance of our method. Specifically, on the ISIC 2017 dataset, WVM-UNet achieves an mIoU of 82.94% and a DSC of 90.67%, outperforming the Mamba-based VM-UNet by 2.71% in mIoU. These results indicate our architecture effectively captures discriminative features for precise medical image segmentation.
Virtual reality-based subjective visual vertical (VR-SVV) has attracted attention as a potential solution to equipment-related limitations of conventional SVV testing. This study investigated the test–retest reliability and adequate trial number for stable assessment of vertical perception using VR-SVV in healthy adults. Participants performed 10 trials in a VR-SVV test and repeated the assessment after one week. SVV orientation and SVV variability were calculated. Test–retest reliability was evaluated using the intraclass correlation coefficient (ICC [1,2]), standard error of measurement (SEM), and minimal detectable change at the 95% confidence level (MDC95). The minimum number of trials required for stable assessment was examined by comparing results from fewer trials with those obtained from all 10 trials. SVV orientation and variability were −0.14° and 0.79°, respectively. ICC, SEM, and MDC95 were 0.62, 0.64°, and 1.76°, respectively. No adverse events occurred. SVV metrics derived from 6–9 trials differed by less than 10% from those obtained from 10 trials. VR-SVV is feasible and demonstrates moderate test–retest reliability. Six trials may be adequate for practical assessment of vertical perception.
Three-dimensional (3D) imaging systems, including depth cameras, LiDAR sensors, and multi-view scanning pipelines, often produce point clouds with noisy normals, outliers, sparse sampling, and non-uniform density, which can degrade downstream mesh reconstruction. Poisson surface reconstruction is lightweight and training-free, but its global implicit formulation is sensitive to unreliably oriented samples and fixed density-trimming thresholds. This paper presents RG-PSR, a reliability-guided enhancement framework for Poisson-family surface reconstruction from degraded 3D-imaging point clouds. RG-PSR estimates a deterministic per-point reliability score from local density regularity, spacing variation, and normal consistency, and propagates this score through conservative point filtering, reliability-guided normal refinement, adaptive density-reliability trimming, and structure-aware postprocessing. The main pipeline requires no manual labels, neural network training, or ground-truth meshes at inference time. Experiments on three groups of object meshes under five deterministic degradation types show that RG-PSR improves Poisson-family reconstruction under degraded inputs. Compared with fixed density-trimmed Poisson reconstruction, RG-PSR reduces the overall Chamfer-L1 from 0.0218 to 0.0172, improves F0.01 from 0.6618 to 0.6836, and reduces Artifact0.02 from 0.3090 to 0.2632. In the broader classical comparison, local triangulation methods achieve stronger point-wise accuracy, while RG-PSR yields the fewest connected components and the highest largest-component ratio. These results position RG-PSR as a practical reliability layer for coherent Poisson-family reconstruction rather than a universal replacement for all surface-reconstruction methods.
Accurate segmentation of non-small cell lung cancer (NSCLC) on positron emission tomography/computed tomography (PET/CT) is an essential prerequisite for automated metabolic tumor volume (MTV) quantification and staging. Although deep learning models achieve high performance on large-scale datasets, their generalization across different clinical domains is limited by variations in imaging protocols and patient demographics. This study aims to evaluate several deep learning architectures and investigate a transfer learning strategy to mitigate domain shift. Three architectures, including ResNet-backbone 3D U-Net, nnU-Net v2, and Swin UNETR, were benchmarked from scratch and compared with a fine-tuned nnU-Net initialized with AutoPET II weights. Results on the internal dataset showed that the fine-tuned nnU-Net achieved a Dice similarity coefficient (DSC) of 83.4 ± 6.5%, a 95% Hausdorff distance (HD95) of 5.1 ± 3.6 mm, and a precision of 89.6 ± 8.2%. Compared to the nnU-Net v2, the fine-tuned nnU-Net improved the absolute DSC by 6.8% while reducing local training time by 37.5% by bypassing the initial feature-learning phase. The fine-tuned nnU-Net model also demonstrated a high correlation between the MTV and the ground truth (Pearson r = 0.96, p < 0.001), indicating its potential as a reliable automated approach for quantitative MTV extraction and NSCLC prognostic-related analysis.
Synthesizing multi-phase contrast-enhanced CT images from non-contrast CT (NCCT) may provide complementary phase-specific cues for preliminary assessment. After rigorous clinical validation, such images could provide clinical decision support or aid diagnostic triage by identifying cases that warrant further acquired contrast-enhanced CT (CECT) work-up; they are not intended to replace acquired CECT. Arterial-phase (ART) and portal-venous-phase (PV) images share anatomical structures but exhibit distinct enhancement patterns and intensity distributions, which makes simultaneous multi-phase synthesis challenging for conventional single-decoder models. We propose a task-prompt-guided dual-decoder framework that combines a shared Swin Transformer encoder, two phase-specific decoders, learnable task prompts initialized from Qwen3-8B semantic representations, two phase-specific adversarial discriminators, and an independent ART/PV domain classifier. The shared encoder extracts phase-invariant anatomical features, whereas the two decoders independently model ART- and PV-specific enhancement. The Qwen3-derived prompt vectors provide phase-aware initialization and subsequently adapt through prompt–feature interaction at the bottleneck. Experiments on a single-center dataset of 86 patients show improved whole-image and lesion-focused PSNR, SSIM, MSE, and PCC relative to Pix2pix and MedGAN. Controlled ablation experiments indicate the benefits of Qwen3-based prompt initialization, subsequent prompt adaptation, and correct prompt–decoder correspondence; a reduced baseline additionally assesses the combined removal of LLTP and ART/PV domain classification. In a downstream slice-level four-class focal liver lesion classification experiment, synthetic multiphase input improved accuracy from 71.65% with NCCT alone to 85.04%, compared with 91.34% for acquired multiphase CT. These findings provide a proof of concept for LLM-guided multi-phase CT synthesis, although external validation, clinically oriented safety assessment, and reader studies remain necessary before clinical use for decision support or diagnostic triage.