• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    法

    法尔茅斯大学

    Falmouth University
    院校EST. 1902falmouth.ac.uk
    723论文总数
    7,978引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Rupert Loydell
    Rupert Loydell
    Falmouth University
    论文:50引用:0H-index:0
    Minhua Ma
    Minhua Ma
    Falmouth University
    论文:15引用:0H-index:0
    Michael James Scott
    Michael James Scott
    Falmouth University
    论文:12引用:0H-index:0
    Mubarak Shah
    Mubarak Shah
    Center for Research in Computer Vision, University of Central Florida;Department of Computer Science, College of Engineering and Computer Science, University of Central Florida
    论文:10引用:0H-index:0
    Hayes Mawindi Mabweazara
    Hayes Mawindi Mabweazara
    School of Media and Performance, University College Falmouth
    论文:9引用:0H-index:0
    Richard J George
    Richard J George
    Department of Agriculture
    论文:8引用:0H-index:0
    Fiona Hackney
    Fiona Hackney
    Manchester Metropolitan University
    论文:8引用:0H-index:0
    Kamaldeep S Bhui
    Kamaldeep S Bhui
    Department of Psychiatry & Nuffield Department of Primary Care Health Sciences Medical Sciences Division, University of Oxford
    论文:8引用:0H-index:0
    Jason Whittaker
    Jason Whittaker
    Sch English & Journalism, Univ Lincoln
    论文:7引用:0H-index:0

    论文(723)

    年份
    起
    –
    止
    排序
    1AdaTooler-V: Adaptive Tool-Use for Images and Videos
    Chaoyang Wang,Kaituo Feng, Dongyang Chen, Zhongyu Wang, Zhixun Li, Sicheng Gao, Meng, Xu Zhou, Manyuan Zhang,Yuzhang Shang,Xiangyu Yue

    Recent advances have shown that multimodal large language models (MLLMs) benefit from multimodal interleaved chain-of-thought (CoT) with vision tool interactions. However, existing open-source models often exhibit blind tool-use reasoning patterns, invoking vision tools even when they are unnecessary, which significantly increases inference overhead and degrades model performance. To this end, we propose AdaTooler-V, an MLLM that performs adaptive tool-use by determining whether a visual problem truly requires tools. First, we introduce AT-GRPO, a reinforcement learning algorithm that adaptively adjusts reward scales based on the Tool Benefit Score of each sample, encouraging the model to invoke tools only when they provide genuine improvements. Moreover, we construct two datasets to support training: AdaTooler-V-CoT-100k for SFT cold start and AdaTooler-V-300k for RL with verifiable rewards across single-image, multi-image, and video data. Experiments across twelve benchmarks demonstrate the strong reasoning capability of AdaTooler-V, outperforming existing methods in diverse visual reasoning tasks. Notably, AdaTooler-V-7B achieves an accuracy of 89.8\% on the high-resolution benchmark V*, surpassing the commercial proprietary model GPT-4o and Gemini 1.5 Pro. All code, models, and data are released.

    2026Annual Meeting of the Association for Computational Linguistics(2026)引用:18
    引用
    AI阅读
    加入学术空间
    2Blurring the Frame: Youth-defined Photographic Expression in Participatory Research on Adverse Childhood Experiences
    Syeda Sana Batool,Anna Mankee-Williams, Kamaldeep Bhui, Syed Ali Jafar Naqvi, Jack Hanrahan, Grace Bennett

    Adverse childhood experiences (ACEs) contribute to 75% of mental illness cases in the UK before age 24, yet their emotional impacts are rarely explored through young people's perspectives. This study investigates how youth with ACEs use blurriness in photography as a form of emotional expression and narrative control. Using a participatory methodology, young people acted as co-researchers through photography tasks, Jamboard discussions and blog reflections. Blurred images - emerging spontaneously - became key artefacts for reflection and meaning-making. Thematic analysis, informed by Constructivist and Chaos Theories, found that blurriness symbolised confusion, fragmentation, vulnerability and distance. It also offered a way to express difficult emotions while avoiding overexposure. Participants associated blurred photography with youth visual culture, especially social media aesthetics that value imperfection and authenticity. This research demonstrates the potential of arts-based, co-produced methods to amplify marginalised youth voices and proposes a participatory visual framework for exploring emotional expression among young people affected by ACEs.

    2026MEDIA INTERNATIONAL AUSTRALIA(2026)引用:17
    引用
    AI阅读
    加入学术空间
    3Test-Time Scaling in Diffusion LLMs Via Hidden Semi-Autoregressive Experts
    Jihoon Lee, Hoyeon Moon, Kevin Zhai, Arun Kumar Chithanar,Anit Kumar Sahu,Soummya Kar, Chul Lee,Souradip Chakraborty,Amrit Singh Bedi

    Diffusion-based large language models (dLLMs) are trained to model extreme flexibility/dependence in the data-distribution; however, how to best utilize this at inference time remains an open problem. In this work, we uncover an interesting property of these models: dLLMs {trained on textual data} implicitly learn a mixture of semi-autoregressive experts, where different generation orders reveal different specialized behaviors. We show that committing to any single, fixed inference time schedule, a common practice, collapses performance by failing to leverage this latent ensemble. To address this, we introduce HEX (Hidden semi-autoregressive EXperts for test-time scaling), a training-free inference method that ensembles across heterogeneous block schedules. By doing a majority vote over diverse block-sized generation paths, HEX robustly avoids failure modes associated with any single fixed schedule. On reasoning benchmarks such as GSM8K, it boosts accuracy by up to 3.56× (from 24.72\% to 88.10\%), outperforming top-K margin inference and specialized fine-tuned methods like GRPO, without additional training. HEX even yields significant gains on MATH benchmark from 16.40\% to 40.00\%, scientific reasoning on ARC-C from 54.18\% to 87.80\%, and TruthfulQA from 28.36\% to 57.46\%. Our results establish test-time scaling as a powerful principle for dLLMs, showing that the sequence in which masking is done can play a significant role in test-time scaling/inferencing of dLLMs.

    ICLR 2026引用:6
    引用
    AI阅读
    加入学术空间
    4VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment
    Shaina Raza, Ashmal Vayani, Aditya Jain, Aravind Narayanan,Vahid Reza Khazaie,Syed Raza Bashir,Elham Dolatabadi, Gias Uddin, Christos Emmanouilidis,Rizwan Qureshi,Mubarak Shah

    Detecting disinformation that blends manipulated text and images has become increasingly challenging, as AI tools make synthetic content easy to generate and disseminate. While most existing AI-safety benchmarks focus on single-modality misinformation (i.e., false content shared without intent to deceive), intentional multimodal disinformation, such as propaganda or conspiracy theories that imitate credible news; remains largely unaddressed. In this work, we introduce the Vision-Language Disinformation Detection Benchmark (VLDBench), the first large-scale resource supporting both unimodal (text-only) and multimodal (text + image) disinformation detection. VLDBench comprises approximately 62,000 labeled text-image pairs across 13 categories, curated from 58 news outlets. Using a semi-automated pipeline followed by expert review, 22 domain experts invested over 500 hours to produce high-quality annotations with substantial inter-annotator agreement. Evaluation of state-of-the-art LLMs and VLMs on VLDBench shows that adding visual cues improves detection accuracy, with gains ranging from 5 points for strong baselines (e.g., LLaMA-3.2-11B-Vision 74.82% vs. LLaMA-3.2-1B-Instruct 70.29%) to 25-30 points for smaller families (e.g., LLaVA-v1.5-Vicuna7B 72.32% vs. Vicuna-7B-v1.5 55.21%), reflecting complementary evidence from images (e.g., meme-like visuals, image-text consistency) that text alone cannot capture. We provide data and code for evaluation, fine-tuning and robustness tests to support disinformation analysis. Developed in alignment with the AI Goverance frameworks (MIT AI Risk Repository), VLDBench offers a principled foundation for advancing trustworthy disinformation detection in multimodal media.

    2026INFORMATION FUSION(2026)引用:4
    引用
    AI阅读
    加入学术空间
    5Monkey Jump : MoE-Style PEFT for Efficient Multi-Task Learning
    Nusrat Jahan Prottasha, Md Kowsher, Chun-Nam Yu, Chen Chen, Ozlem Garibay

    Mixture-of-experts variants of parameter-efficient fine-tuning enable per-token specialization, but they introduce additional trainable routers and expert parameters, increasing memory usage and training cost. This undermines the core goal of parameter-efficient fine-tuning. We propose Monkey Jump, a method that brings mixture-of-experts-style specialization to parameter-efficient fine-tuning without introducing extra trainable parameters for experts or routers. Instead of adding new adapters as experts, Monkey Jump treats the adapters already present in each Transformer block (such as query, key, value, up, and down projections) as implicit experts and routes tokens among them. Routing is performed using k-means clustering with exponentially moving averaged cluster centers, requiring no gradients and no learned parameters. We theoretically show that token-wise routing increases expressivity and can outperform shared adapters by avoiding cancellation effects. Across multi-task experiments covering 14 text, 14 image, and 19 video benchmarks, Monkey Jump achieves competitive performance with mixture-of-experts-based parameter-efficient fine-tuning methods while using 7 to 29 times fewer trainable parameters, up to 48 percent lower memory consumption, and 1.5 to 2 times faster training. Monkey Jump is architecture-agnostic and can be applied to any adapter-based parameter-efficient fine-tuning method.

    2026CoRR(2026)引用:2
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 723 篇论文

    合作机构(100)

    佛罗里达中央大学合作论文 33
    牛津大学合作论文 15
    约克大学合作论文 15
    利兹大学合作论文 13
    伦敦大学合作论文 10
    肯特大学合作论文 9
    爱丁堡大学合作论文 8
    莫纳什大学合作论文 7
    埃克塞特大学合作论文 7
    布鲁内尔大学合作论文 6

    机构统计