• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    维

    维萨卡

    Visa Inc.
    企业
    183论文总数
    5,667引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Huiyuan Chen
    Huiyuan Chen
    Amazon
    论文:9引用:0H-index:0
    Michael Chin-Chia Yeh
    Michael Chin-Chia Yeh
    Visa Research
    论文:8引用:0H-index:0
    Junpeng Wang
    Junpeng Wang
    Visa Research
    论文:8引用:0H-index:0
    Colin Baptie
    Colin Baptie
    Visa international
    论文:6引用:0H-index:0
    Hao Yang
    Hao Yang
    Splunk
    论文:6引用:0H-index:0
    Yu (Jason) Gu
    Yu (Jason) Gu
    Visa
    论文:5引用:0H-index:0
    Zhang Wei
    Zhang Wei
    Department of Civil Engineering, Fujian University of Technology
    论文:5引用:0H-index:0
    Hanghang Tong
    Hanghang Tong
    Department of Computer Science, University of Illinois at Urbana-Champaign;Siebel School of Computing and Data Science, The Grainger College of Engineering, University of Illinois at Urbana-Champaign
    论文:4引用:0H-index:0
    Yuzhong Chen
    Yuzhong Chen
    College of Computer and Data Science, Fuzhou University
    论文:4引用:0H-index:0

    论文(184)

    年份
    起
    –
    止
    排序
    1ALERT: Zero-shot LLM Jailbreak Detection Via Internal Discrepancy Amplification
    Xiao Lin, Philip Li,Zhichen Zeng, Tingwei Li,Tianxin Wei,Xuying Ning,Gaotang Li, Yuzhong Chen,Hanghang Tong

    Despite rich safety alignment strategies, large language models (LLMs) remain highly susceptible to jailbreak attacks, which compromise safety guardrails and pose serious security risks. Existing detection methods mainly detect jailbreak status relying on jailbreak templates present in the training data. However, few studies address the more realistic and challenging zero-shot jailbreak detection setting, where no jailbreak templates are available during training. This setting better reflects real-world scenarios where new attacks continually emerge and evolve. To address this challenge, we propose a layer-wise, module-wise, and token-wise amplification framework that progressively magnifies internal feature discrepancies between benign and jailbreak prompts. We uncover safety-relevant layers, identify specific modules that inherently encode zero-shot discriminative signals, and localize informative safety tokens. Building upon these insights, we introduce ALERT (Amplification-based Jailbreak Detector), an efficient and effective zero-shot jailbreak detector that introduces two independent yet complementary classifiers on amplified representations. Extensive experiments on three safety benchmarks demonstrate that ALERT achieves consistently strong zero-shot detection performance. Specifically, (i) across all datasets and attack strategies, ALERT reliably ranks among the top two methods, and (ii) it outperforms the second-best baseline by at least 10

    2026CoRR(2026)引用:5
    引用
    AI阅读
    加入学术空间
    2The Organization of Innovation: Property Rights and the Outsourcing Decision
    Thomas Jungbauer,Sean Nicholson,June Pan,Michael Waldman,Lucy Xiaolu Wang

    Why do firms outsource research and development (R&D) for some products while conducting R&D in-house for similar ones?An innovating firm risks cannibalizing its existing products.The more profitable these products, the more the firm wants to limit cannibalization.We apply this logic to the organization of R&D by introducing a novel theoretical model in which developing in-house provides the firm more control over the new product's location in product space.An empirical analysis of our testable predictions using pharmaceutical data concerning patents, patent expiration, and outsourcing at various stages of the R&D process supports our theoretical findings.

    2026AMERICAN ECONOMIC JOURNAL-MICROECONOMICS(2026)引用:2
    引用
    AI阅读
    加入学术空间
    3Prompt Overflow: What the Guardrail Inspects is Not What the Model Infers
    Yuanbo Zhou, Changjia Zhu, Junyu Wang, Xu He, Yan Zhai,Kun Sun, Mingkui Wei, Junjie Xiong

    Guardrail models (a.k.a. safety checkers) are widely deployed to screen user inputs before they reach large language models (LLMs), serving as a primary defense against prompt injection attacks. Due to strict context constraints, these models handle overlength prompts through truncation or segmentation-based inspection. While prior work has focused on semantic adversarial inputs, the security implications of these long-input processing mechanisms remain largely unexplored. In this paper, we identify a critical blind spot arising from the mismatch between the limited inspection windows of guardrail models and the substantially larger context inference windows of downstream LLMs. We introduce a novel Prompt Overflow Attack, which exploits this mismatch by fragmenting malicious instructions and interleaving them with benign filler content across an overlong prompt, such that no individual inspected segment appears malicious while the full context remains actionable to the LLM. Through a systematic evaluation against state-of-the-art guardrail models, including Meta Llama Prompt Guard, IBM Granite Guardian, and DeBERTa-based detectors, we demonstrate that prompts reliably detected in short-context settings can evade guardrail models once adversarially manipulated into over-length inputs, yet remain fully actionable by downstream LLMs. We further propose potential defense strategies and outline mitigation directions to strengthen guardrail models.

    2026引用:2
    引用
    AI阅读
    加入学术空间
    4Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce
    Shimaa Ahmed, Yiwei Cai, Mohsen Minaei, Rahul Rachuri

    Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, where customer preferences are hidden, competitors adapt in real time, and demand can shift without warning, has not been systematically tested. We introduce Bazaar, a dynamic sealed-bid benchmark for multi-attribute auction under these conditions. Despite its dynamics, the benchmark is grounded in closed-form customer utilities, enabling exact evaluation. Across 11 frontier LLMs from four providers, the leading agents on customer acquisition (e.g. Gemini 3.1 Pro) are often not the leading agents on profit (e.g. Opus 4.6). The ranking shifts again under demand shocks: agents that learned fastest pre-shock are typically the slowest to revise their beliefs afterwards, while Gemini 3.1 Pro recovers fastest despite not leading on profit. However, even the strongest agent captures less than a third of hindsight-optimal profit, suggesting current LLMs are progressing in agentic commerce but leave substantial headroom.

    2026引用:1
    引用
    AI阅读
    加入学术空间
    5Fairness-Aware Test-Time Prompt Tuning
    Yoann Launay, Parameswaran Kamalaruban, Tom Kempton, Stuart Burrell, David Sutton

    Vision-language models have displayed remarkable capabilities in multi-modal understanding and are increasingly used in critical applications where economic and practical deployment constraints prohibit re-training or fine-tuning. However, these models can also exhibit systematic biases that disproportionately affect protected demographic groups and existing approaches to addressing these biases require extensive model retraining and access to demographic attributes. There is a clear need to develop test-time adaptation (TTA) approaches that improve the fairness characteristics of pretrained models under distributional shift. In this paper, we evaluate how episodic TTA affects fairness in CLIP classification under subpopulation shifts and develop FairTPT, a novel fairness-aware episodic TTA method that jointly minimizes target marginal entropy while maximizing spurious marginal entropy through soft-prompt tuning. We find that standard episodic TTA generally exacerbates disparities between majority and minority groups, that blinding a model to spurious attributes without degrading target performance is inherently challenging, and that excessive blinding can lead to catastrophic forgetting. This model collapse can be prevented by monitoring test-time changes in target loss within the linear regime, while still achieving fairness improvements on reactive data and preserving overall performance. Thus refined, FairTPT outperforms all state-of-the-art episodic test-time debiasing methods and establishes a foundation for robust TTA—essential for achieving fairness in practice.

    ICLR 2026引用:1
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 184 篇论文

    合作机构(100)

    伊利诺伊大学香槟分校合作论文 7
    明尼苏达大学合作论文 5
    乔治梅森大学合作论文 4
    国际商业机器公司合作论文 4
    德克萨斯 A&M 大学合作论文 4
    曼彻斯特大学合作论文 4
    谷歌合作论文 4
    莱斯大学合作论文 3
    亚马逊合作论文 3
    万事达卡合作论文 3

    机构统计