Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary prediction target decoupled from diagnosis. We propose GazeWorld, a medical imaging world model that treats the image as the world and the radiologist's fixation sequence as a trajectory through it. GazeWorld autoregressively predicts the latent representation of the next fixated patch from all previously visited ones, while a spatial-completion branch covers unvisited regions. At inference, GazeWorld generates a sequence of patch representations from the image alone without requiring real gaze data. Frozen GazeWorld features achieve state-of-the-art diagnostic accuracy across all nine supervised settings on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, as well as the highest zero-shot accuracy on all three benchmarks. On the GazeSearch benchmark, a generic decoder trained on the same frozen features outperforms the purpose-built LogitGaze-Med by over 16% in ScanMatch and 22% in SED, despite not being explicitly trained to predict gaze. GazeWorld demonstrates that modeling how experts read, not just what they conclude, offers a promising pretraining paradigm for medical imaging AI.
Alzheimer's disease (AD) and Lewy body dementia (LBD) present overlapping clinical features yet require distinct diagnostic strategies. While neuroimaging-based brain network analysis is promising, atlas-based representations may obscure individualized anatomy. Gyral folding-based networks using three-hinge gyri provide a biologically grounded alternative, but inter-individual variability in cortical folding results in inconsistent landmark correspondence and highly irregular network sizes, violating the fixed-topology and node-alignment assumptions of most existing graph learning methods, particularly in clinical datasets where pathological changes further amplify anatomical heterogeneity. We therefore propose a probability-invariant random-walk-based framework that classifies individualized gyral folding networks without explicit node alignment. Cortical similarity networks are built from local morphometric features and represented by distributions of anonymized random walks, with an anatomy-aware encoding that preserves permutation invariance. Experiments on a large clinical cohort of AD and LBD subjects show consistent improvements over existing gyral folding and atlas-based models, demonstrating robustness and potential for dementia diagnosis.
Quantum Artificial Intelligence (QAI) has emerged at the nexus of quantum computing and AI, promising to redefine computational frontiers. This survey critically synthesizes the state-of-the-art through 2024, elucidating the profound bidirectional synergy between these fields. We analyze how classical machine learning is accelerating quantum hardware control, circuit optimization, and error correction. Conversely, we assess the potential quantum advantage of algorithms, including variational and kernel-based methods, across domains such as drug discovery, financial modeling, and cybersecurity. Our analysis reveals a critical trade-of between the utility of near-term Noisy Intermediate-Scale Quantum (NISQ) devices and the long-term promise of fault-tolerant architectures. We identify fundamental obstacles to QAI's advancement, including hardware decoherence, algorithmic barren plateaus, and data-encoding bottlenecks. While QAI's potential is transformative, achieving practical quantum advantage requires a concerted effort to overcome these core challenges at the hardware-software interface. This work provides a roadmap for navigating the current landscape and prioritizing future research in this rapidly evolving discipline.
Automated chest X-ray report generation requires not only clinical accuracy but also transparent and interpretable diagnostic reasoning. In this work, we propose FiR-Rad, a two-stage framework that combines explicit structured reasoning with targeted fine-grained optimization. In the first stage, a supervised chain-of-thought approach guides the model to sequentially analyze and describe a comprehensive range of clinically significant thoracic abnormalities, ensuring clinically meaningful coverage. In the second stage, we introduce a segment-level reinforcement learning strategy based on Group Relative Policy Optimization (GRPO), which assigns precise rewards to each disease-specific reasoning step by evaluating the accuracy of corresponding findings in the synthesized report. This design provides direct feedback for intermediate reasoning and encourages consistency between detailed abnormality analysis and final diagnostic conclusions. Experimental results on the MIMIC-CXR and IU-Xray datasets demonstrate that our framework achieves state-of-the-art performance across clinical and linguistic metrics, with strong zero-shot generalization on IU-Xray. The proposed method significantly enhances interpretability and clinical accuracy, effectively addressing key limitations in automated radiology report generation.
Medical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work like SBM-Transformer attempts to incorporate such structures through stochastic binary masking, they suffer from non-differentiability, training instability, and the inability to model complex community structure. We present DCMM-Transformer, a novel ViT architecture for medical image analysis that incorporates a Degree-Corrected Mixed-Membership (DCMM) model as an additive bias in self-attention. Unlike prior approaches that rely on multiplicative masking and binary sampling, our method introduces community structure and degree heterogeneity in a fully differentiable and interpretable manner. Comprehensive experiments across diverse medical imaging datasets, including brain, chest, breast, and ocular modalities, demonstrate the superior performance and generalizability of the proposed approach. Furthermore, the learned group structure and structured attention modulation substantially enhance interpretability by yielding attention maps that are anatomically meaningful and semantically coherent.
Digital twins, as precise digital representations of physical systems, have evolved from passive simulation tools into intelligent and autonomous entities through the integration of artificial intelligence technologies. This paper presents a unified four-stage framework that systematically characterizes AI integration across the digital twin lifecycle, spanning modeling, mirroring, intervention, and autonomous management. By synthesizing existing technologies and practices, we distill a unified four-stage framework that systematically characterizes how AI methodologies are embedded across the digital twin lifecycle: (1) modeling the physical twin through physics-based and physics-informed AI approaches, (2) mirroring the physical system into a digital twin with real-time synchronization, (3) intervening in the physical twin through predictive modeling, anomaly detection, and optimization strategies, and (4) achieving autonomous management through large language models, foundation models, and intelligent agents. We analyze the synergy between physics-based modeling and data-driven learning, highlighting the shift from traditional numerical solvers to physics-informed and foundation models for physical systems. Furthermore, we examine how generative AI technologies, including large language models and generative world models, transform digital twins into proactive and self-improving cognitive systems capable of reasoning, communication, and creative scenario generation. Through a cross-domain review spanning eleven application domains, including healthcare, aerospace, smart manufacturing, robotics, and smart cities, we identify common challenges related to scalability, explainability, and trustworthiness, and outline directions for responsible AI-driven digital twin systems.
Background:The accuracy and safety of generating medication orders by large language models (LLMs) must be demonstrated. Without standardization, performance evaluation is limited to time and resource-intensive clinician grading. This evaluation aimed to develop a standardized medication format that supports automated performance evaluation (MedMatch). Methods:First, a survey of 40 medication prompts was given to clinicians to assess agreement in medication order communication. Second, a clinician panel developed a standardized medication format (MedMatch) for oral and intravenous medications. Third, a clinician-annotated dataset of medication prompts and standardized answers in the MedMatch format was developed for LLM testing. Finally, LLMs were retested with the same dataset, adjusted to exclude route information, to evaluate the appropriate categorization of medication route. Results:The formal medication orders consistently showed low omission rates and high overlap for all entities, compared to the verbal and brief written communication types. Lexical overlap results demonstrated pattern norms amongst clinicians with entities appearing most commonly in positions 1-5 in the order of drug name, dose, unit, route, and frequency. In the second survey, the formal written group performed the highest with 78.3% of prompts considered appropriate as a computer-generated response. LLM accuracy on MedMatch order standardization was highest in oral solid (64.2-72.5%), intravenous intermittent (72.5-84.3%), and intravenous push (62.7-74.5%) categories. LLMs performed the worst at categorizing medication orders accurately into intravenous push (18-61%) and intravenous intermittent (51-100%) routes. Conclusions:Standardized format for computer-based outputs may support automated performance analysis and enhance the clarity of medication communication.
Quantum and hybrid quantum-classical workflows repeatedly need to predict how ideal circuit distributions appear as finite-shot noisy hardware histograms, but refitting high-dimensional noise simulators at each calibration point can be costly. We cast this task as finite-shot distribution learning for structured quantum outputs: given ideal distributions and noisy counts, learn a low-dimensional Markov surrogate and evaluate held-out distributional risk. Existing fast closed-form kernels lack identifiability guarantees, while richer physics-motivated simulators fit approximately 20 BO-tuned noise parameters without parameter-identifiability guarantees. We instantiate this middle ground with SA-IBF, a two-parameter closed-form kernel mixing independent bit-flip noise with a Hamming-sector-preserving component. For the support-face regimes induced by symmetry-preserving circuits, we show a support-face bridge and a rank-conditioned local guarantee at fitted operating points; open-interior genericity establishes non-degeneracy of the full-simplex model. We also show a small-q Fisher upper bound in which the Fisher information for alpha conditional on q scales on the order of N times q squared for N shots and per-qubit flip rate q, certified for 2 to 8 qubits, yielding Cramer-Rao divergence as q approaches zero; analytic-Jacobian diagnostics verify the fitted-point rank premise at all seven representative operating points and all 230 panel fits. On an 890-execution IBM Heron hardware dataset spanning four backends, SA-IBF gives a diagnosable accuracy-cost Pareto point: in matched fitted-physics comparisons, it meets the sector-aware 21-parameter Ji-BO Hellinger accuracy budget on three of four molecular qubit settings while reducing refit cost by 231-447 times under matched implementations. The remaining setting flags a uniform-flip-rate boundary where the tested Ji-BO comparator's heterogeneous parameters lower risk. Supporting calibration evidence shows lower held-out Hellinger distance than plain bit-flip noise on all seven primary configurations and throughout a 400-execution cross-backend molecular panel. These results make SA-IBF a lightweight, identifiability-guided learned distribution-level surrogate for repeated quantum noise simulation in symmetry-structured quantum and hybrid quantum-classical workloads, including symmetry-preserving quantum-machine-learning settings.
Interactions among building blocks in physical, chemical, and biological systems follow structured patterns interpretable as a learnable language: just as language models learn which words tend to follow others, one can learn which physical phenomena follow others and under what conditions. We introduce an Interaction Language Model (ILM): a framework that treats interaction as a statistically structured, learnable language for design-to-function reasoning across fields. By learning statistical dependencies and ordering among components, ILMs infer interaction pathways, identify missing steps, and predict the next likely interaction. We demonstrate it through two complementary components, PhenoLink and PhenoSeq, applied to molecular diffusion, a ubiquitous mechanism for energy and particle transport. PhenoLink is a directed graph of interactions among diffusion-based events extracted from publications, where each edge aggregates paragraph-level evidence and carries a transition probability interpreted as an information cost. PhenoSeq complements this with a sequence model that proposes missing mechanistic steps under endpoint constraints. Together they form a generate-then-verify pipeline returning every explanation with a reproducible, per-step, corpus-grounded audit trail. Across four held-out suites of 172 queries, the pipeline matches a zero-shot Claude Opus 4.7 baseline on in-distribution accuracy, yet unlike the baseline, which fabricates a chain for every impossible query, it refuses up to 93.5
In recent years, the advances of large language models and autonomous agents have revolutionized the healthcare field, facilitating diagnosis and improving treatment results. However, most existing AI systems rely on pre-trained knowledge and predefined pipelines, which struggle to learn dynamically from the interactive chat session history that contains patient outcomes and past failures. To address this limitation, we propose VIBEMed, a multi-agent framework with a built-in self-evolution mechanism and architecture-level safety sandbox for robust clinical decision support. The system integrates three specialized agents, including a Clinical Diagnostic Agent (CDA) for hypothesis generation, a Therapeutic Execution Agent (TEA) for treatment planning, and a Clinical Evolution Manager Agent (CEMA) that distills longitudinal clinical feedback into reusable knowledge, transforming multimodal patient information into personalized medical decisions. Through self-evolution mechanism, the framework enables iterative updates across memory, model behavior, and decision strategies, allowing the system to improve over time. Experimental results show that VIBEMed demonstrates superior performance through its evolving mechanism in complex clinical cases, particularly in tasks that require integrated decision-making and longitudinal planning. The framework also supports reliable end-to-end decisions in challenging scenarios such as oncology treatment planning, highlighting its feasibility in real-world clinical contexts. Overall, VIBEMed provides a practical path beyond static AI systems toward adaptive, experience-driven clinical decision support, demonstrating the value of combining multi-agent collaboration with continuous evolution for advancing precision medicine.
Large language models (LLMs) have transformative potential in radiology, including textual summaries, diagnostic decision support, proofreading, and image analysis. However, the rapid increase in studies investigating these models, along with the lack of standardized LLM-specific reporting practices, affects reproducibility, reliability, and clinical applicability. To address this, reporting guidelines for LLM studies in radiology were developed using a two-step process. First, a systematic review of LLM studies in radiology was conducted across PubMed, IEEE Xplore, and the ACM Digital Library, covering publications between May 2023 and March 2024. Of 511 screened studies, 57 were included to identify relevant aspects for the guidelines. Then, in a Delphi process, 20 international experts developed the final list of items for inclusion. Items consented as relevant were summarized into a structured checklist containing 32 items across six key categories: general information and data input; prompting and fine-tuning; performance metrics; ethics and data transparency; implementation, risks, and limitations; and further/optional aspects. The final FLAIR (Framework for LLM Assessment in Radiology) checklist aims to standardize reporting of LLM studies in radiology, fostering transparency, reproducibility, comparability, and clinical applicability to enhance clinical translation and patient care. © The Author(s) 2026. Published by the Radiological Society of North America under a CC BY 4.0 license. Supplemental material is available for this article.
Distinguishing causal relationships from statistical correlations remains a fundamental challenge in clinical research, limiting the translation of observational findings into interventional treatment guidelines. Here we investigate whether causal machine learning can be used to explore estimated causal effects of radiation dose parameters on mandibular osteoradionecrosis (ORN), using a case study of 931 head and neck cancer patients treated with volumetric-modulated arc therapy. Using generalized random forests, all examined dosimetric factors showed positive estimated causal effects on ORN development (average treatment effects: 0.092–0.141). Integration with explainable machine learning suggested substantial treatment effect heterogeneity, with the largest estimated conditional average treatment effects observed in patients aged 50-60 years and smaller estimates in patients over 70 years. These results suggest that causal machine learning can help quantify dose-related effects and explore heterogeneity across patient characteristics. More broadly, this work provides a methodological framework for toxicity studies in oncology and other clinical settings where complex dose–response relationships warrant further prospective validation.
Alzheimer’s disease (AD) risk prediction relies on accurately characterizing pathological mechanisms underlying AD progression. However, existing methods struggle with heterogeneous multi-omics data and often fail to capture the spatiotemporal dynamics of the disease, limiting their predictive performance. In this paper, an integrated framework fusing spatial and temporal information is proposed to improve prediction capability. First, brain region-gene directed networks are constructed based on large foundation model-enhanced features. Second, a context perception attention model is designed to characterize topological changes of directed networks during AD progression. Based on this model, we develop a Context Perception Attention Generative Adversarial Network (CPA-GAN) that leverages adversarial training to mine AD evolutionary patterns, thereby supporting risk prediction and pathogeny extraction. Finally, the superiority, effectiveness, and robustness of CPA-GAN are validated by extensive experiments. Overall, this work provides a robust and effective modeling framework tailored for early-stage AD risk prediction.
Alzheimer's disease (AD) and Lewy body dementia (LBD) are common neurodegenerative dementias with overlapping clinical presentations, making differential diagnosis challenging. While structural magnetic resonance imaging (MRI) has revealed characteristic regional atrophy patterns, regional morphometric measures alone may not fully capture distributed cortical alterations. Morphometric similarity networks (MSNs) offer a systems-level framework to characterize coordinated structural organization, but existing approaches typically rely on atlas-based parcellations that may obscure individual-specific cortical folding geometry. Here, we propose a fine-scale, folding-informed cortical similarity network framework based on automatically detected three-hinge gyral (3HG) landmarks. Using a thickness-constrained arealization strategy in native surface space, we define individualized cortical regions and construct subject-specific MSNs without cross-subject registration. We then investigate how network topology relates to landmark-defined node count and how these properties differ between AD and LBD. We find that several graph theoretical metrics, particularly global efficiency and characteristic path length, exhibit clear associations with the number of detected landmarks, indicating that topology in individualized networks is partly shaped by node availability. When accounting for landmark count, several apparent group differences in global topology are attenuated, whereas multiple heterogeneity-related metrics remain significant, indicating that node-count scaling substantially influences the interpretation of individualized network topology. Nevertheless, multivariate topological patterns remain informative for AD/LBD classification after residualizing for node count, and landmark count itself provides modest diagnostic information. These findings highlight node-count scaling as a key methodological consideration in individualized structural networks and suggest that folding-based MSNs capture disease-related variation in cortical network organization between AD and LBD.
This narrative review with structured literature screening combines comprehensive research on the rapid adoption of object detection computer vision models, particularly “You Only Look Once” (YOLO), used alone or in conjunction with other machine learning models, to advance Precision Poultry Farming (PPF), which refers to the application of data-driven and automated technologies to monitor, manage, and optimize poultry health, welfare, and production efficiency. A literature search across search engines, such as Google Scholar, was used because of its broad interdisciplinary coverage, allowing retrieval of literature spanning animal science, computer vision, and agricultural engineering, which are often indexed across different publications venues, on October 15 2024, which revealed 408 results when searching with search expression “YOLO + broilers + layers” and publications dated from 2015 to October 15, 2024. We removed 200 articles during screening, and 126 articles were excluded after eligibility evaluation, resulting in 82 eligible research papers to be included for this review. The YOLO object detection models have evolved from YOLOv1 to YOLO11 by 2024, progressively improving in model performance, speed, accuracy, and robustness through the refinement of key architectural components, including backbone networks, detection heads, and loss functions. This review highlights how YOLO models have been applied to broiler chickens and laying hens across diverse housing systems to support key tasks such as identification, behavior detection, counting, tracking, health and disease monitoring, flock distribution pattern, and calculating activity index, often in combination with other machine vision models. The analysis shows that it took 4 years to apply YOLO models for the object detection task in poultry since the release of the first version of the YOLO model in 2015. The application of YOLO models in poultry from 2019 to 2021 was very slow and sporadic while it took rapid growth in publications since 2021, led primarily by research groups in China and the USA, and mainly concentrated in journals such as Computers and Electronics in Agriculture (10), Institute of Electrical and Electronics Engineers (IEEE) Conference (10), Poultry Science (9), Animals (6), and AgriEngineering (5). Major opportunities and challenges are identified around deploying these models for reliable, real-time decision support on commercial farms, particularly for animal welfare assessment, disease and wild bird detection, and integration with complementary sensing and analytics frameworks.
While gradient-based optimizers that incorporate randomization often demonstrate superior performance on complex optimizations, the theoretical foundations of this advantage remain underexplored. A central question arises: What role does randomization play in dimension-free, nonsmooth, nonconvex optimization? To address this gap, we examine both the theoretical and empirical impact of permutation randomization within gradient-based optimization frameworks, using it as a representative case to investigate broader implications. From a theoretical perspective, our analysis reveals that permutation randomization disrupts the shrinkage behavior characteristic of gradient-based optimizers, allowing for continued progress toward the global optimum with sufficient iterations. Moreover, we prove that permutation randomization preserves the convergence rate of the underlying optimizer. Empirically, we conduct extensive numerical experiments comparing permutation-randomized optimizers with three baseline methods. These experiments span tasks such as training deep neural networks with stacked architectures and optimizing noisy objective functions. The results not only support our theoretical findings but also demonstrate the practical benefits of permutation randomization. In summary, this work provides both rigorous theoretical justification and compelling empirical evidence for the effectiveness of permutation randomization, establishing a foundation for extending such analyses to broader implications of randomized strategies.
The 3-hinge gyrus (3HG), where three gyri converge, acts as a structural and functional hub in the brain. Connecting these hubs forms the GyralNet, but traditionally extraction requires complex geometric steps, limiting scalability. To address this challenge, we propose Deep-GyralNet, a novel deep learning framework for efficient 3HG and GyralNet extraction. Deep-GyralNet leverages a Spherical U-Net to learn gyral crest representations from cortical morphological features, coupled with a Connectivity-Aware Path Enforcement loss that enforces topological continuity of predicted ridges. A minimum spanning tree-based refinement is then applied to ensure structural connectivity and yield a topologically valid GyralNet. Experiments on the Human Connectome Project dataset demonstrate that Deep-GyralNet achieves an average completeness of 98%, while reducing processing time by 94% compared to the conventional pipeline. These establish DeepGyralNet the first deep learning solution for fast and accurate 3HG identification, enabling large-scale studies of gyral folding networks in brain development, aging and disorders.
BACKGROUND AND AIMS:Osteoradionecrosis (ORN) of the mandible is one of the most severe adverse events (AEs) for head and neck (H&N) cancer radiotherapy. Previous retrospective investigations on real-world data relied heavily on conventional statistical models that primarily elucidate correlation rather than establishing causal relationships. Through the novel causal machine learning method, we aim to obtain empirical relative biological effectiveness (RBE) for mandible ORN in head and neck (H&N) cancer patients treated with pencil-beam-scanning proton therapy (PBSPT). METHODS:1,266 H&N cancer patients were included: 335 patients treated by PBSPT and 931 patients treated by volumetric-modulated arc therapy (VMAT). We used 1:1 propensity-score case matching to minimize imbalance in clinical factors between patients treated with PBSPT and VMAT. Standardized mean differences (SMDs) were used to assess residual clinical-factor imbalance within the case-matched cohorts. Causal forest (CF) was adopted to investigate the causal effects between dosimetric factors and the incidence of ORN. For each modality and each prespecified DVH index, candidate dose-volume thresholds were evaluated systematically, and the volume threshold yielding the largest CF-estimated average treatment effect (ATE) was selected as the DVC volume threshold. Empirical RBE values were derived from equal-volume intersections on modality-specific volume-tolerance curves after converting PBSPT Gy[RBE] values to physical dose. RESULTS:335 VMAT patients were case-matched to 335 PBSPT patients; however, standardized mean bias analysis revealed persistent covariate imbalances within each group, indicating residual confounding influence. Using CF modeling, we identified candidate DVC volume thresholds for mandibular ORN and found that PBSPT had lower selected DVC volume thresholds than VMAT. The threshold-stability analyses supported the robustness of the DVC thresholds emphasized in the empirical RBE analysis. The resulting empirical RBE exceeded 1.1 in the moderate dose range (1.61 at 40 Gy[RBE], 1.30 at 50 Gy[RBE], and 1.13 at 60 Gy[RBE]). CONCLUSION:This study presents a novel application of causal machine learning to evaluate mandibular ORN in radiotherapy, identifying candidate DVC volume thresholds linked to the strongest threshold-defined causal effects and deriving empirical RBEs from equal-volume equivalent constraint dose analysis based on volume-tolerance curves. The results indicate that proton RBE may exceed 1.1 in the moderate dose range (40-60 Gy[RBE]), underscoring the importance of considering endpoint-specific variable RBE in PBSPT treatment planning. These CF-identified DVC volume thresholds should be interpreted as hypothesis-generating risk regions rather than definitive clinical cutoffs, pending independent prospective validation.
Manually auditing gait scores of broiler chickens is labor-intensive and subjective, necessitating the development of an automated and objective alternative. This study presents a novel three-dimensional (3D) deep learning pipeline to assess the walking ability of broilers by predicting gait scores ranging from 0 (optimal mobility) to 2 (severely impaired mobility). A total of 540 broiler chicken videos, sampled from 6 to 7 weeks of age, were recorded as the chickens traversed a 1.75-meter wooden platform. An Intel RealSense L515 LiDAR camera, mounted at a 2.5-meter height, was used to capture synchronized RGB-Depth data. The data were recorded using the Robot Operating System (ROS) Noetic, ensuring efficient and structured data acquisition. The proposed pipeline consisted of multiple sequential steps: (1) RGB-Depth frame extraction and synchronization, (2) pose estimation using a custom-trained YOLOv11-based chicken pose detection model, (3) back-projection of 2D keypoints into 3D space using camera intrinsics, (4) frame validation using a convolutional neural network to filter out occlusions and artifacts, (5) platform orientation detection via Hough line transformation, (6) segmentation of the chicken’s body using the Segment Anything Model (SAM) to extract 3D point clouds, and (7) kinematic feature extraction for gait analysis. Key 3D features, including velocity, acceleration, and head turn frequency, were fed into a multi-layer perceptron classifier to predict gait scores. The classifier predicted broiler gait scores with 93.34
The Perturbed Utility Model framework offers a powerful generalization of discrete choice analysis, unifying models like Multinomial Logit and Sparsemax through convex optimization. However, standard Maximum Likelihood Estimation (MLE) faces severe theoretical and numerical challenges when applied to this broader class, particularly regarding non-convexity and instability in sparse regimes. To resolve these issues, this paper introduces a unified estimation framework based on the Fenchel-Young loss. By leveraging the intrinsic convex conjugate structure of PUMs, we demonstrate that the Fenchel-Young estimator guarantees global convexity and bounded gradients, providing a mathematically natural alternative to MLE. Addressing the critical challenge of data scarcity, we further extend this framework via Wasserstein Distributionally Robust Optimization. We first derive an exact finite-dimensional reformulation of the infinite-dimensional primal problem, establishing its theoretical convexity. However, recognizing that the resulting worst-case constraints involve computationally intractable inner maximizations, we subsequently construct a tractable safe approximation by exploiting the global Lipschitz continuity of the Fenchel-Young loss. Through this tractable formulation, we uncover a rigorous geometric unification: two canonical regularization techniques, standard L2-regularization and the margin-enforcing Hinge loss, emerge mathematically as specific limiting cases of our distributionally robust estimator. Extensive experiments on synthetic data and the Swissmetro benchmark validate that the proposed framework significantly outperforms traditional methods, recovering stable preferences even under severe data limitations.