• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    I

    Instituto de Engenharia de Sistemas e Computadores Investigação e Desenvolvimento,Institute for Systems Engineering and Computers

    EST. 2000
    928论文总数
    1.7万引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Isabel Trancoso
    Isabel Trancoso
    Instituto Superior Tecnico, University of Lisbon
    论文:38引用:0H-index:0
    José Borbinha
    José Borbinha
    Department of Computer Science and Engineering, Instituto Superior Técnico, Universidade de Lisboa;INESC-ID
    论文:34引用:0H-index:0
    Fernando Batista
    Fernando Batista
    Spoken Language Systems Laboratory
    论文:24引用:0H-index:0
    Nuno Freire
    Nuno Freire
    Instituto Superior Técnico, Technical University of Lisbon
    论文:20引用:0H-index:0
    José Fernando Alves da Silva
    José Fernando Alves da Silva
    Department of Electrical and Computer Engineering, Instituto Superior Técnico, University of Lisbon
    论文:18引用:0H-index:0
    Joao Miranda Lemos
    Joao Miranda Lemos
    INESC-ID, Instituto Superior Tecnico, Universidade de Lisboa
    论文:17引用:0H-index:0
    Leonel Sousa
    Leonel Sousa
    Electrical and Computer Engineering Department, Instituto Superior Técnico, Universidade de Lisboa;INESC-ID
    论文:14引用:0H-index:0
    João Paulo Carvalho
    João Paulo Carvalho
    Department of Electrical and Computer Engineering, Instituto Superior Técnico, Universidade de Lisboa
    论文:13引用:0H-index:0
    David Martins de Matos
    David Martins de Matos
    Department of Computer Science and Engineering, Universidade de Lisboa Instituto Superior Técnico;Spoken Language Systems Lab, Instituto de Engenharia de Sistemas e Computadores Investigação e Desenvolvimento em Lisboa;Human Language Technology Lab, INESC-ID Lisboa
    论文:12引用:0H-index:0

    论文(928)

    年份
    起
    –
    止
    排序
    1MEDAL: A Framework for Benchmarking LLMs As Multilingual Open-Domain Dialogue Evaluators
    John Mendonça,Alon Lavie,Isabel Trancoso

    Evaluating the quality of open-domain chatbots has become increasingly reliant on LLMs acting as automatic judges. However, existing meta-evaluation benchmarks are static, outdated, and lacking in multilingual coverage, limiting their ability to fully capture subtle weaknesses in evaluation. We introduce MEDAL, an automated multi-agent framework for curating more representative and diverse open-domain dialogue evaluation benchmarks. Our approach leverages several state-of-the-art LLMs to generate user-chatbot multilingual dialogues, conditioned on varied seed contexts. Then, a strong LLM (GPT-4.1) is used for a multidimensional analysis of the performance of the chatbots, uncovering noticeable cross-lingual performance differences. Guided by this large-scale evaluation, we curate a new meta-evaluation multilingual benchmark and human-annotate samples with nuanced quality judgments. This benchmark is then used to assess the ability of several reasoning and non-reasoning LLMs to act as evaluators of open-domain dialogues. Using MEDAL, we uncover that state-of-the-art judges fail to reliably detect nuanced issues such as lack of empathy, commonsense, or relevance.

    2026Conference of the European Chapter of the Association for Computational Linguistics(2026)引用:1
    引用
    AI阅读
    加入学术空间
    2Assessing Convolutional Neural Networks Architectures for Orthopantomography Image Analysis: Judicial Dental Age Threshold Classification
    Cristiana Palmela Pereira, António Figueiras, Valon Nushi, Mohamed Elbawab, Alexandre Francisco, Arlindo Oliveira, Cátia Vaz,Rui Santos

    OBJECTIVE:Forensic dental age assessment is required when documentary evidence is absent or unreliable, and judicial decisions often depend on whether an individual falls below or above legally defined age thresholds. This pilot study evaluates ImageNet pretrained convolutional neural network architectures for binary age threshold classification. DESIGN:Orthopantomograms from individuals aged 0-25 years (1887 males and 1664 females) was selected using predefined inclusion criteria and labelled by chronological age. For each legal threshold (10, 12, 14, 16, 18 and 21 years), images were stratified into training and validation sets using an 80 20 split while preserving class distributions. Preprocessing included contrast enhancement with contrast limited adaptive histogram equalization, automated cropping, resizing to model specific input dimensions and ImageNet normalisation. Data augmentation was applied only to training images. Seven convolutional neural network architectures (ResNet 50, ResNet 152, VGG19, DenseNet 121, DenseNet 169, EfficientNetV2 M and Xception) were fine tuned in PyTorch using binary cross entropy loss with early stopping. Hyperparameters were optimised through Bayesian search targeting macro F1 score. Model interpretability was assessed using Grad CAM heatmaps reviewed by dental experts. RESULTS:EfficientNetV2 M showed the best performance for the 10-year threshold (accuracy: 0.935), DenseNet 169 for the 12-year threshold (accuracy: 0.945) and Xception for the remaining thresholds (accuracy: 14-year: 0.924; 16-year: 0.947; 18-year: 0.906; 21-year: 0.890). Errors clustered near age cut offs and increased at higher thresholds. Grad CAM highlighted posterior dento alveolar regions associated with root development. CONCLUSION:These results support convolutional neural network based orthopantomogram analysis as a judicial decision support approach and guided model selection for large scale evaluation.

    2026Archives of oral biology(2026)引用:1
    引用
    AI阅读
    加入学术空间
    3Multimodal LLMs As Expert Speech Annotators: Acoustic Macro-Descriptors for Parkinson's Detection
    David Ortiz-Perez,Catarina Botelho,Anna Pompili,Alberto Abad,Jose Garcia-Rodriguez

    Large language models (LLMs) have demonstrated remarkable capabilities in analyzing textual data to support disease diagnosis. Extending these advances to audio, Multimodal Large Language Models (MLLMs) open new opportunities for addressing speech-impairing conditions such as Parkinson’s Disease (PD). Using the NeuroVoz corpus, which includes speech recordings and expert evaluations across 14 perceptual dimensions of voice quality, phonation, and prosody, we validated the clinical utility of these annotations and assessed the ability of MLLMs to replicate them. The models generated acoustic macro-descriptors that showed good agreement with expert ratings and enabled effective PD classification. Feeding these descriptors into machine learning classifiers for PD detection achieved up to 80.47% UAR. Overall, the findings highlight the potential of MLLMs to provide reliable and interpretable features directly from audio, thereby supporting scalable and cross-domain speech-based diagnosis of neurodegenerative conditions.

    2026ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)(2026)引用:1
    引用
    AI阅读
    加入学术空间
    4Solving General-Utility Markov Decision Processes in the Single-Trial Regime with Online Planning
    Pedro Pinto Santos,Alberto Sardinha,Francisco S. Melo

    In this work, we contribute the first approach to solve infinite-horizon discounted general-utility Markov decision processes (GUMDPs) in the single-trial regime, i.e., when the agent's performance is evaluated based on a single trajectory. First, we provide some fundamental results regarding policy optimization in the single-trial regime, investigating which class of policies suffices for optimality, casting our problem as a particular MDP that is equivalent to our original problem, as well as studying the computational hardness of policy optimization in the single-trial regime. Second, we show how we can leverage online planning techniques, in particular a Monte-Carlo tree search algorithm, to solve GUMDPs in the single-trial regime. Third, we provide experimental results showcasing the superior performance of our approach in comparison to relevant baselines.

    ICLR 2026引用:1
    引用
    AI阅读
    加入学术空间
    5Entropic Risk-Aware Monte Carlo Tree Search
    Pedro P. Santos, Jacopo Silvestrin,Alberto Sardinha,Francisco S. Melo

    We propose a provably correct Monte Carlo tree search (MCTS) algorithm for solving \textit{risk-aware} Markov decision processes (MDPs) with \textit{entropic risk measure} (ERM) objectives. We provide a \textit{non-asymptotic} analysis of our proposed algorithm, showing that the algorithm: (i) is \textit{correct} in the sense that the empirical ERM obtained at the root node converges to the optimal ERM; and (ii) enjoys \textit{polynomial regret concentration}. Our algorithm successfully exploits the dynamic programming formulations for solving risk-aware MDPs with ERM objectives introduced by previous works in the context of an upper confidence bound-based tree search algorithm. Finally, we provide a set of illustrative experiments comparing our risk-aware MCTS method against relevant baselines.

    2026CoRR(2026)引用:1
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 928 篇论文

    合作机构(100)

    里斯本大学合作论文 118
    波尔图大学合作论文 32
    Instituto Superior Técnico合作论文 28
    卡内基梅隆大学合作论文 21
    阿尔加夫大学合作论文 18
    Universidade Nova de Lisboa合作论文 17
    葡萄牙里斯本理工大学合作论文 13
    Instituto Politecnico de Setubal合作论文 11
    Instituto Politécnico de Lisboa合作论文 11
    阿威罗大学合作论文 10

    机构统计