Evaluating the quality of open-domain chatbots has become increasingly reliant on LLMs acting as automatic judges. However, existing meta-evaluation benchmarks are static, outdated, and lacking in multilingual coverage, limiting their ability to fully capture subtle weaknesses in evaluation. We introduce MEDAL, an automated multi-agent framework for curating more representative and diverse open-domain dialogue evaluation benchmarks. Our approach leverages several state-of-the-art LLMs to generate user-chatbot multilingual dialogues, conditioned on varied seed contexts. Then, a strong LLM (GPT-4.1) is used for a multidimensional analysis of the performance of the chatbots, uncovering noticeable cross-lingual performance differences. Guided by this large-scale evaluation, we curate a new meta-evaluation multilingual benchmark and human-annotate samples with nuanced quality judgments. This benchmark is then used to assess the ability of several reasoning and non-reasoning LLMs to act as evaluators of open-domain dialogues. Using MEDAL, we uncover that state-of-the-art judges fail to reliably detect nuanced issues such as lack of empathy, commonsense, or relevance.
OBJECTIVE:Forensic dental age assessment is required when documentary evidence is absent or unreliable, and judicial decisions often depend on whether an individual falls below or above legally defined age thresholds. This pilot study evaluates ImageNet pretrained convolutional neural network architectures for binary age threshold classification. DESIGN:Orthopantomograms from individuals aged 0-25 years (1887 males and 1664 females) was selected using predefined inclusion criteria and labelled by chronological age. For each legal threshold (10, 12, 14, 16, 18 and 21 years), images were stratified into training and validation sets using an 80 20 split while preserving class distributions. Preprocessing included contrast enhancement with contrast limited adaptive histogram equalization, automated cropping, resizing to model specific input dimensions and ImageNet normalisation. Data augmentation was applied only to training images. Seven convolutional neural network architectures (ResNet 50, ResNet 152, VGG19, DenseNet 121, DenseNet 169, EfficientNetV2 M and Xception) were fine tuned in PyTorch using binary cross entropy loss with early stopping. Hyperparameters were optimised through Bayesian search targeting macro F1 score. Model interpretability was assessed using Grad CAM heatmaps reviewed by dental experts. RESULTS:EfficientNetV2 M showed the best performance for the 10-year threshold (accuracy: 0.935), DenseNet 169 for the 12-year threshold (accuracy: 0.945) and Xception for the remaining thresholds (accuracy: 14-year: 0.924; 16-year: 0.947; 18-year: 0.906; 21-year: 0.890). Errors clustered near age cut offs and increased at higher thresholds. Grad CAM highlighted posterior dento alveolar regions associated with root development. CONCLUSION:These results support convolutional neural network based orthopantomogram analysis as a judicial decision support approach and guided model selection for large scale evaluation.
Large language models (LLMs) have demonstrated remarkable capabilities in analyzing textual data to support disease diagnosis. Extending these advances to audio, Multimodal Large Language Models (MLLMs) open new opportunities for addressing speech-impairing conditions such as Parkinson’s Disease (PD). Using the NeuroVoz corpus, which includes speech recordings and expert evaluations across 14 perceptual dimensions of voice quality, phonation, and prosody, we validated the clinical utility of these annotations and assessed the ability of MLLMs to replicate them. The models generated acoustic macro-descriptors that showed good agreement with expert ratings and enabled effective PD classification. Feeding these descriptors into machine learning classifiers for PD detection achieved up to 80.47% UAR. Overall, the findings highlight the potential of MLLMs to provide reliable and interpretable features directly from audio, thereby supporting scalable and cross-domain speech-based diagnosis of neurodegenerative conditions.
In this work, we contribute the first approach to solve infinite-horizon discounted general-utility Markov decision processes (GUMDPs) in the single-trial regime, i.e., when the agent's performance is evaluated based on a single trajectory. First, we provide some fundamental results regarding policy optimization in the single-trial regime, investigating which class of policies suffices for optimality, casting our problem as a particular MDP that is equivalent to our original problem, as well as studying the computational hardness of policy optimization in the single-trial regime. Second, we show how we can leverage online planning techniques, in particular a Monte-Carlo tree search algorithm, to solve GUMDPs in the single-trial regime. Third, we provide experimental results showcasing the superior performance of our approach in comparison to relevant baselines.
We propose a provably correct Monte Carlo tree search (MCTS) algorithm for solving \textit{risk-aware} Markov decision processes (MDPs) with \textit{entropic risk measure} (ERM) objectives. We provide a \textit{non-asymptotic} analysis of our proposed algorithm, showing that the algorithm: (i) is \textit{correct} in the sense that the empirical ERM obtained at the root node converges to the optimal ERM; and (ii) enjoys \textit{polynomial regret concentration}. Our algorithm successfully exploits the dynamic programming formulations for solving risk-aware MDPs with ERM objectives introduced by previous works in the context of an upper confidence bound-based tree search algorithm. Finally, we provide a set of illustrative experiments comparing our risk-aware MCTS method against relevant baselines.