Background: Accurate translation of radiology reports is important for multilingual research, clinical communication, and radiology education, but the validity of LLM-based evaluation remains unclear. Objective: To evaluate the educational suitability of LLM-generated Japanese translations of chest CT reports and compare radiologist assessments with LLM-as-a-judge evaluations. Methods: We analyzed 150 chest CT reports from the CT-RATE-JPN validation set. For each English report, a human-edited Japanese translation was compared with an LLM-generated translation by DeepSeek-V3.2. A board-certified radiologist and a radiology resident independently performed blinded pairwise evaluations across 4 criteria: terminology accuracy, readability, overall quality, and radiologist-style authenticity. In parallel, 3 LLM judges (DeepSeek-V3.2, Mistral Large 3, and GPT-5) evaluated the same pairs. Agreement was assessed using QWK and percentage agreement. Results: Agreement between radiologists and LLM judges was near zero (QWK=-0.04 to 0.15). Agreement between the 2 radiologists was also poor (QWK=0.01 to 0.06). Radiologist 1 rated terminology as equivalent in 59
Purpose: To develop and validate a deep learning ensemble for estimating adult sex, age, height, and weight from coronal digitally reconstructed radiographs (DRRs) generated from diagnostic CT. Materials and Methods: This retrospective study included 128,621 CT examinations from 80,004 adults at nine institutions in Japan. Three multitask models-ConvNeXt-Base, ViT-Base/16, and MaxViT-Base-were fine-tuned using coronal DRRs and combined by weighted averaging. Data were split by institution into training (114,147 examinations; seven institutions), tuning (4,305; one institution), and test (10,169; one institution) sets; generalizability was assessed on two non-Japanese datasets. Accuracy and mean absolute error (MAE) were used to evaluate sex classification and age, height, and weight regression, respectively. Body surface area (BSA)-corrected heart and liver volume trends were compared using true versus estimated height and weight. Results: In the test set (median age, 69.9 years; 4,899 of 10,169 [48.2
In this study, we propose a segmentation method for automatic estimation of lymph node maps (LNMs) from CT volumes. Esophageal cancer progresses rapidly with few symptoms, so early detection and treatment are essential. Because treatment decisions depend on nodal metastasis, level classification of lymph node for staging currently requires both specialist expertise and considerable time. In addition, each lymph node region is defined by its spatial relationship with the surrounding organs, so clear boundaries are absent on CT volumes, making segmentation difficult. This study therefore proposes constructing a segmentation model that can automatically estimate LNMs from CT volumes. First, we defined a small set of CT volumes with corresponding manual LNMs as templates and generated a large set of CT-map pairs by applying a B-spline transformation for nonrigid deformation to the templates, producing pseudo-LNMs. Second, to leverage the spatial relationships of each lymph node region to the surrounding organs, we constructed the training dataset by pairing each CT volume with its pseudo-LNM and with organ annotations automatically generated using TotalSegmentator. Third, we trained Swin UNETR to predict an LNM from a CT volume, using a loss function that includes a regularization term penalizing the distance between each lymph node region and the annotated organs, thereby improving generalization to unseen cases. Experimental results showed that the proposed method achieved a Dice coefficient of 0.69 for lymph node region extraction, indicating the potential to reduce radiologists ' workload and improve diagnostic objectivity. For future work, we will evaluate the quality of the pseudo-LNMs and apply additional data augmentation to enhance both transparency and accuracy.
Accurate segmentation of mediastinal lymph nodes (MLNs) from thoracic computed tomography (CT) volumes is essential for staging thoracic malignancies and radiotherapy planning. Manual MLN delineation is quite time-consuming, while supervised deep learning requires large annotated datasets that are rarely available. This paper proposes to utilize a foundation model to improve MLN segmentation from CT volumes, especially when training data is limited. First, we train a foundation model using a large unlabeled CT dataset with the masked autoencoder (MAE) method. The encoder of this pre-trained foundation model is then used to initialize the encoder of the segmentation model, which is fine-tuned for the MLN segmentation. Our results showed that, compared to the non-pre-trained method, the foundation model-based method achieved higher accuracy with limited training data, confirming its effectiveness in few-shot settings. The proposed method achieved an averaged Dice score of 0.621 when 80% of the dataset was used for training, compared to 0.605 by the model trained from scratch.
While foundation models in radiology are expected to be applied to various clinical tasks, computational cost constraints remain a major challenge when training on 3D-CT volumetric data. In this study, we propose TotalFM, a radiological foundation model that efficiently learns the correspondence between 3D-CT images and linguistic expressions based on the concept of organ separation, utilizing a large-scale dataset of 140,000 series. By automating the creation of organ volume and finding-sentence pairs through segmentation techniques and Large Language Model (LLM)-based radiology report processing, and by combining self-supervised pre-training via VideoMAE with contrastive learning using volume-text pairs, we aimed to balance computational efficiency and representation capability. In zero-shot organ-wise lesion classification tasks, the proposed model achieved higher F1 scores in 83
Background: Computer-aided detection (CAD) systems for chest radiographs are widely used; however, concurrent reader displays such as bounding-box (BB) highlights may influence interpretation. This pilot study used eye tracking to examine which aspects of visual search were affected by these factors. Methods: We sampled 180 chest radiographs from the VinDR-CXR dataset: 120 with solitary pulmonary nodules or masses and 60 without. BBs were configured for 80 % display sensitivity and specificity. Three radiologists (with 11, 5, and 1 years of experience) interpreted each case twice—once with BBs visible and once without—after a ≥ 2-week washout. Eye movements were recorded using an EyeTech VT3 Mini. Metrics included interpretation time, time to first fixation, lesion dwell time, total gaze-path length, and lung-field coverage. Outcomes were modeled using a linear mixed model with the reading condition set as a fixed effect and case and reader as random intercepts. Primary analysis was restricted to true positives (n = 96). Results: Concurrent BB display prolonged interpretation time by 4.9 s (p < 0.001) and increased lesion dwell time by 1.3 s (p < 0.001). Total gaze-path length increased by 2076 pixels (p < 0.001), and lung-field coverage increased by 10.5 % (p < 0.001). The time to first fixation was reduced by 1.3 s (p < 0.001). Conclusion: Eye tracking revealed measurable changes in search behavior associated with concurrent BB display during chest radiograph interpretation. These findings support this approach and highlight the need for larger studies across modalities and clinical contexts.
Abstract Objectives: To characterize longitudinal age-related changes in abdominal organ volumes using CT volumetry and to model nonlinear trajectories across multiple organs. Materials & Methods: This retrospective single-center study included adults who underwent whole-body screening low-dose CT between 2006 and 2017. Subjects with at least eight examinations during a follow-up period of at least 78 months were included. After applying exclusion criteria, 700 participants with 6,739 CT series were analyzed. Non-contrast CT images were processed using automated organ segmentation, and volumes of the liver, pancreas, spleen, and kidneys were quantified. Longitudinal changes were modeled using generalized additive mixed models with sex-specific smooth functions of age and subject-level random effects. Age-dependent rates of change were estimated from model derivatives. Results: A total of 700 participants (mean age, 56.9 ± 9.8 years, 29.6% women) were evaluated. Liver, pancreas, and kidney volumes showed mild increases or plateaued at approximately 40–60 years of age, depending on the organ, and were followed by gradual declines with advancing age, whereas splenic volume decreased more continuously across the age range. These patterns showed nonlinear age dependence. The transition from positive to negative change rates tended to occur earlier in women than in men for several organs, particularly the liver and kidneys. Conclusion: Longitudinal CT analysis demonstrated nonlinear age-related changes in abdominal organ volumes, with organ-specific trajectories and sex-related differences in the timing and magnitude of volume changes. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study did not receive any funding. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Research Ethics Committee of the Faculty of Medicine of the University of Tokyo gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The original clinical data cannot be publicly shared because of ethical and privacy restrictions.
Japanese language models for medical text classification face challenges with complex vocabulary and linguistic structures in radiology reports. This study compared three Japanese models—BERT Base, JMedRoBERTa, and ModernBERT—for multi-label classification of 18 chest CT findings. Using the CT-RATE-JPN dataset, all models were fine-tuned under identical conditions. ModernBERT showed clear efficiency advantages, producing substantially fewer tokens and achieving faster training and inference than the other models while maintaining comparable performance on the internal test dataset (exact match accuracy: 74.7% vs. 72.7% for BERT Base). To assess generalizability, we additionally constructed RR-Findings, an external dataset of 243 naturally written Japanese radiology reports annotated using the same schema. Under this domain-shifted setting, performance differences became pronounced: BERT Base outperformed both JMedRoBERTa and ModernBERT, whereas ModernBERT showed the largest decline in exact match accuracy. Average precision differences were smaller, indicating that ModernBERT retained reasonable ranking ability despite reduced calibration. Overall, ModernBERT offers substantial computational efficiency and strong in-domain performance but remains sensitive to real-world linguistic variability. These results highlight the need for more diverse natural-language training data and domain-specific calibration strategies to improve robustness when deploying modern transformer models in heterogeneous clinical environments.
BACKGROUND:Claudin-18 isoform 2 (CLDN18.2) is a tight junction protein expressed in gastric mucosa and gastric cancer (GC) cells. Although chemotherapeutic agents are suggested to increase CLDN18.2 expression in GC cells, their impact on GC tissues and their underlying mechanisms remain unclear. METHODS:We examined the effects of chemotherapy on CLDN18.2 expression in human GC tissues and cell lines, and investigated the role of cell-cycle regulation in this process. RESULTS:We found that CLDN18.2 expression was upregulated in GC tissues after chemotherapy. In GC cell lines (SNU-601, NUGC4, GSU), chemotherapeutic agents, including 5-fluorouracil, irinotecan, paclitaxel, and cisplatin, increased CLDN18.2 expression at the transcriptional level, although the response patterns varied among cell lines and agents. Since cell-cycle regulation appeared to be involved, we tested cyclin-dependent kinase (CDK) inhibitors and found that CDK1- and CDK4/6-specific inhibitors similarly enhanced CLDN18.2 expression. Importantly, both chemotherapeutic agents and CDK inhibitors significantly increased zolbetuximab-mediated antibody-dependent cellular cytotoxicity in GC cells. CONCLUSION:These results suggest that chemotherapeutic agents upregulate CLDN18.2 in GC cells at least in part through cell cycle arrest, and support combining zolbetuximab with chemotherapy and/or CDK inhibitors for GC treatment.
Large-scale CT-based reference standards for abdominal organ volume, incorporating age, sex, and body size, are limited. To establish sex- and age-specific reference distributions for major abdominal organ volumes on non-contrast abdominopelvic CT in a nationwide Japanese cohort to provide a foundation for automated clinical assessment and dose optimization. In this retrospective, multicenter study, using the Japan Medical Image Database, we identified all non-contrast abdominopelvic CT examinations performed in 2024. Unique adults with available data on age, sex, height, and weight were included in this study. The final sample comprised 49,764 examinations (26,456 men and 23,308 women) conducted at nine institutions. Automated segmentation (TotalSegmentator v2.10.0) was used to produce organ volumes, excluding hollow viscera. The sex-specific 10th, 25th, 50th, 75th, and 90th percentiles were calculated. Age–volume relationships of body surface area (BSA)-normalized volumes (mL/m 2 ) were modeled using natural cubic splines (four degrees of freedom) separately by sex. Median (mL) male vs female volumes were as follows: liver, 1194.7 vs 1024.0; pancreas, 63.6 vs 52.2; spleen, 118.1 vs 95.1; kidneys (total), 268.3 vs 221.2; adrenals (total), 6.6 vs 4.2; iliopsoas (total), 483.4 vs 317.7; prostate, 24.9 (men only). Age–volume relationships of BSA-normalized volumes showed convex patterns for the liver, pancreas, and kidneys in both sexes and for male adrenal glands; lower values in older age groups for the spleen and iliopsoas in both sexes; and higher values in older age groups for the prostate and female adrenal glands. This nationwide Japanese CT cohort provides sex- and age-resolved volumetric reference standards. These standards enable objective identification of abnormalities, support personalized medicine, and facilitate automated AI-based reporting to reduce radiologist workload and optimize radiation dose protocols. Median volumes (men vs women, mL): liver 1195/1024; pancreas 64/52; spleen 118/95; kidneys 268/221; adrenals 6.6/4.2; iliopsoas 483/318; prostate 25. Body surface area–normalized age–volume relationships were convex for liver, pancreas, and kidneys in both sexes and for male adrenal glands. Spleen and iliopsoas declined monotonically with age in both sexes, whereas prostate and female adrenal glands increased monotonically.
Purpose: To evaluate whether large language model (LLM)-assisted label cleaning can identify label-report discordance in CT-RATE, a large-scale public chest CT dataset. Materials and Methods: After report-level deduplication, 24,446 unique radiology reports were identified. Twelve reports were excluded from the primary GPT-5.4 analysis because of Microsoft Azure AI Foundry content-safety filtering, leaving 24,434 reports and 439,812 label instances across 18 abnormality categories. GPT-5.4-derived binary labels were generated from report text using structured JSON output and compared with existing CT-RATE labels. Discordant instances were adjudicated by radiologists. In addition, 100 randomly sampled reports were manually annotated to compare CT-RATE labels, individual LLM-derived labels, and multi-LLM majority-vote labels against radiologist-annotated reference labels. Results: Overall agreement between GPT-5.4-derived and CT-RATE labels was 96.4 Conclusion: LLM-assisted label cleaning identified clinically meaningful label-report discordance in CT-RATE and may support scalable quality improvement of public imaging datasets. The cleaned dataset will be made publicly available to support future research.
The purpose of this study is to develop and validate a deep learning model for automatic identification of acquisition sequences (including T1-weighted sequences with various dynamic contrast-enhanced phases and other auxiliary sequences) in gadolinium ethoxybenzyl diethylenetriamine pentaacetic acid (Gd-EOB-DTPA)–enhanced liver MRI, enabling automated examination-level data curation. This retrospective study included internal Gd-EOB-DTPA–enhanced liver MRI examinations acquired at our institution between June 2018 and May 2020, with independent external test datasets from three additional institutions. Each examination comprised multiple dynamic contrast-enhanced phases and auxiliary sequences, resulting in 13 predefined label categories. A deep learning pipeline was constructed using a convolutional neural network (ConvNeXt) for feature extraction, followed by sequential models (gated recurrent unit [GRU]-based or transformer-based) to model temporal relationships across series. Models were trained using series-level 3D image volumes without reliance on textual Digital Imaging and Communications in Medicine metadata. Model selection was performed on a validation set. Performance was evaluated on internal and external test sets using examination-level accuracy, complete-correct rate, and category-level accuracy. Sequential modeling substantially improved performance compared with a convolution-only baseline. The ConvNeXt + GRU model achieved the highest examination-level accuracy and was selected for final evaluation. Dynamic contrast-enhanced phases were identified with high accuracy across datasets. Reduced performance on the external test set was mainly observed for auxiliary sequences with high inter-institutional variability, particularly in T2-weighted imaging. The proposed framework enables accurate automatic identification of dynamic phases and auxiliary sequences in Gd-EOB-DTPA–enhanced liver MRI, supporting robust examination-level data organization in multicenter studies.
Background The integration of artificial intelligence (AI) in radiology has accelerated globally, with Japan's Pharmaceuticals and Medical Devices Agency (PMDA) approving numerous AI-based Software as a Medical Device (SaMD) products. However, the transparency and completeness of clinical evidence available to healthcare providers remain unclear. Purpose To systematically evaluate the availability and transparency of clinical evidence in package inserts of PMDA-approved AI-based radiology SaMD products, identifying gaps that may impact clinical implementation. Materials and methods We conducted a scoping review of all PMDA-approved SaMD products as of December 31, 2024. Products were included if they utilized AI technology and were classified for radiology applications. Data extraction focused on product characteristics, study designs, demographic information, and performance metrics. Results Of 151 approved SaMD products, 40 utilized AI technology, with 20 specifically designed for radiology applications. Critical gaps were identified in demographic reporting, with no products providing complete case demographic data. Performance metrics varied widely, with sensitivity ranging from 67.7% to 100% in standalone studies. Physician-assisted studies consistently demonstrated performance improvements but lacked stratified results by characteristics in all cases. Conclusion Current package insert requirements provide insufficient transparency for evidence-based clinical implementation of AI-based radiology SaMD. Enhanced regulatory frameworks and industry-led initiatives for comprehensive validation are essential for safe and effective AI deployment in Japanese healthcare.
Background: Automated structuring of radiology reports is essential for data utilization and the development of medical artificial intelligence models. However, manual annotation by experts is labor-intensive, and processing real clinical data through commercial large language models (LLMs) presents significant privacy risks. These challenges are particularly pronounced for non-English languages like Japanese, where specialized medical corpora are scarce. While synthetic data generation offers a potential privacy-preserving alternative, its effectiveness in capturing complex clinical nuances-such as negation and contextual dependencies-to train robust classification models without any real-world training data has not been fully established. Objective: This study aimed to develop a context-aware sentence classification model for Japanese radiology reports using an entirely synthetic training pipeline, thereby eliminating reliance on real-world clinical data during the development phase. Furthermore, we sought to evaluate the generalizability of this approach by validating the model's performance on diverse, multi-institutional, real-world reports. Methods: Japanese radiology reports (n=3104) were generated using GPT-4.1 and automatically annotated at the sentence level into 4 categories (background, positive finding, negative finding, and continuation) using GPT-4.1-mini. The synthetic data were partitioned into training (n=2670), validation (n=334), and test (n=100) sets. We fine-tuned several models, including lightweight local LLMs (Qwen3 and Llama 3.2 series) using low-rank adaptation and Japanese text classification models (Bidirectional Encoder Representations from Transformers [BERT]-base Japanese v3, Japanese Medical Robustly Optimized BERT Pretraining Approach [JMedRoBERTa]-base, and ModernBERT-Ja-130M). External validation was performed using 280 real-world reports (3477 sentences) from 7 institutions in the Japan Medical Image Database, with ground-truth labels established by board-certified radiologists. Evaluation metrics included accuracy, macro-averaged F1 (macro F1) score, and positive predictive value for positive findings (PPV_1). Results: All models achieved high performance on the synthetic test set (accuracy: 0.938-0.951; macro F1-score: 0.924- 0.940). Overall performance declined on the external validation dataset (accuracy: 0.783-0.813; macro F1-score: 0.761-0.790), reflecting distributional differences between synthetic and real-world reports; however, PPV_1 remained stable and high across datasets (eg, 0.957 on the synthetic test set vs 0.952 on the external validation dataset for Qwen3 [4B]). Parsing errors occurred in LLM-based approaches (19-260 sentences, 0.55%-7.48% in the external dataset). Conclusions: This study demonstrates the feasibility of developing context-aware sentence classification models for Japanese radiology reports using a training pipeline based entirely on synthetic data. The stability of PPV_1 indicates that the models successfully captured the essential clinical terminology and linguistic patterns required to identify positive findings in real-world reports, despite the observed performance degradation during external validation. This approach substantially reduces manual annotation requirements and privacy risks, providing a scalable foundation for constructing structured radiology datasets to support the development of clinically relevant medical artificial intelligence models.
PURPOSE:Large-scale CT-based reference distributions for abdominal organ volume that incorporate age, sex, and body size are limited. To establish sex- and age-specific reference distributions for major abdominal organ volumes on non-contrast abdominopelvic CT in a nationwide Japanese cohort to provide a foundation for automated clinical assessment. MATERIALS AND METHODS:In this retrospective, multicenter study, we used the Japan Medical Image Database to identify all non-contrast abdominopelvic CT examinations performed in 2024. Unique adults with available data on age, sex, height, and weight were included in this study. The final sample comprised 49,764 examinations (26,456 men and 23,308 women) conducted at nine different institutions. Automated segmentation (TotalSegmentator v2.10.0) was used to produce organ volumes, excluding the hollow viscera. Sex-specific 10th, 25th, 50th, 75th, and 90th percentiles were calculated. Age-volume relationships of body surface area (BSA)-normalized volumes (mL/m2) were modeled using natural cubic splines (four degrees of freedom) separately for each sex. RESULTS:The median male vs. female volumes (mL) were as follows: liver, 1,194.7 vs 1,024.0; pancreas, 63.7 vs 52.3; spleen, 118.7 vs 95.4; kidneys (total), 268.3 vs 221.2; adrenals (total), 6.7 vs 4.2; iliopsoas (total), 483.4 vs 317.7; and prostate, 25.5 (men only). The age-volume relationships of BSA-normalized volumes showed convex patterns for the liver, pancreas, and kidneys in both sexes and for male adrenal glands; lower values in older age groups for the spleen and iliopsoas in both sexes; and higher values in older age groups for the prostate and female adrenal glands. CONCLUSION:This nationwide Japanese CT cohort provides sex- and age-resolved organ volume reference distributions. These distributions may serve as a quantitative reference for clinical interpretation and population-based imaging research.
BackgroundIntratumoral Fusobacterium nucleatum (Fn) infection is closely associated with poor prognosis in esophageal cancer (EC) due to its impact on the tumor microenvironment (TME). The tumor cell-intrinsic cyclic GMP-AMP synthase (cGAS)-stimulator of interferon genes (STING) pathway is critical for regulating immune cell activation in the TME. However, the link between intratumoral Fn infection and the activation of the cGAS-STING pathway in tumor cells, as well as its effects on EC progression, remains largely unknown.MethodsIn the present study, we investigated the impact of intratumoral Fn infection on the activation of the tumor cell-intrinsic cGAS-STING pathway and EC progression by analyzing our own EC cohort and performing in vitro experiments using co-cultures of EC-cell lines and Fn.ResultsThe expression of tumor cell-intrinsic STING was significantly associated with worse prognosis in Fn-high EC patients. Exposure to Fn significantly activated the STING pathway in EC cells. RNA-seq analysis revealed that exposure to Fn markedly activated cytokine-chemokine-related signaling pathways and induced the expression of several cytokines and chemokines in STING-expressing EC cells. Among the differentially expressed cytokine and chemokine genes in EC cells co-cultured with Fn, analysis of TCGA datasets demonstrated that the expression of CCL20, CXCL10, and CSF2 may be associated with poor prognosis in EC patients.ConclusionWe revealed that the activation of the STING signaling pathway and the subsequent expression of cytokines and chemokines in EC cells induced by Fn infection may be closely associated with poor prognosis in EC patients.
Background: Anti-programmed death 1 receptor (PD-1) therapy is a promising treatment strategy for patients with unresectable advanced or recurrent gastric/gastroesophageal junction (G/GEJ) cancer. However, its response rate and survival benefits are still limited; an immunological analysis of the residual tumor after anti-PD-1 therapy would be important. Methods: We evaluated the clinical efficacy of tumor resection (TR) after chemotherapy or anti-PD-1 therapy in patients with unresectable advanced or recurrent G/GEJ cancer and analyzed the immune status of tumor microenvironment (TME) by immunohistochemistry using their surgically resected specimens. Results: Patients treated with TR after anti-PD-1 therapy had significantly longer survival compared to those treated with chemotherapy and anti-PD-1 therapy alone. Expression of human leukocyte antigen (HLA) class I and major histocompatibility complex (MHC) class II on tumor cells was markedly downregulated after anti-PD-1 therapy compared to chemotherapy. Furthermore, the downregulation of HLA class I may be associated with the activation of transforming growth factor-β signaling pathway in the TME. Conclusions: Immune escape from cytotoxic T lymphocytes may be induced in the TME in patients with unresectable advanced or recurrent G/GEJ cancer after anti-PD-1 therapy due to the downregulation of HLA class I and MHC class II expression on tumor cells. TR may be a promising treatment strategy for these patients when TR is feasible after anti-PD-1 therapy.
BACKGROUND/AIM:Esophageal squamous cell carcinoma (ESCC) significantly affects nutritional status. While neoadjuvant chemotherapy (NAC) is the standard treatment for clinical stage II/III ESCC, its impact on nutritional status and postoperative outcomes remains unclear. This study investigated the relationship between nutritional deterioration during NAC and outcomes following esophagectomy. PATIENTS AND METHODS:This single-center retrospective study included 85 patients with thoracic ESCC who received NAC followed by esophagectomy between January 2019 and December 2023. The NAC regimens included cisplatin plus 5-fluorouracil (CF), administered with or without radiation, and docetaxel, cisplatin, and fluorouracil (DCF). Nutritional status and postoperative outcomes were evaluated. RESULTS:Of the 85 patients, 20 received DCF, 36 received CF alone, and 29 received CF with radiation. Nutritional deterioration was noted during NAC with significant decreases in Prognostic Nutritional Index (PNI), Geriatric Nutritional Risk Index (GNRI), and hemoglobin-albumin-lymphocyte-platelet (HALP) score. Grade 3-4 hematologic toxicities correlated with reductions in GNRI and increases in neutrophil-lymphocyte ratio (NLR) (p=0.004 and 0.020, respectively). An increase in NLR post-NAC and decrease in PNI post-NAC were associated with prolonged hospital stays (p=0.024 and 0.042, respectively). Post-NAC NLR of 2.3 or over and HALP score below 20 were significantly associated with poorer overall survival (p=0.012 and <0.001, respectively). CONCLUSION:Although NAC reduces tumor burden and eliminates micrometastases, it potentially worsens nutritional status. Chemotherapy-induced hematologic toxicities are a risk factor for this decline. Therefore, comprehensive nutritional assessment and timely intervention during NAC are essential to optimize patient outcomes.
Large-scale medical visual question answering (MedVQA) datasets are critical for training and deploying vision–language models (VLMs) in radiology. Ideally, such datasets should be automatically constructed from routine radiology reports and their corresponding images. However, no existing method directly links free-text findings to the most relevant 2D slices in volumetric computed tomography (CT) scans. To address this gap, a contrastive language–image pre-training (CLIP)-based key slice selection framework is proposed, which matches each sentence to its most informative CT slice via text–image similarity. This experiment demonstrates that models pre-trained in the medical domain already achieve competitive slice retrieval accuracy and that fine-tuning them on a small dual-supervised dataset that imparts both lesion- and organ-level awareness yields further gains. In particular, the best-performing model (fine-tuned BiomedCLIP) achieved a Top-1 accuracy of 51.7% for lesion-aware slice retrieval, representing a 20-point improvement over baseline CLIP, and was accepted by radiologists in 56.3% of cases. By automating the report-to-slice alignment, the proposed method facilitates scalable, clinically realistic construction of MedVQA resources.