• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    国

    国立国语研究所

    National Institute for Japanese Language and Linguistics
    EST. 1948ninjal.ac.jp
    235论文总数
    1,634引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Masayuki Asahara
    Masayuki Asahara
    Nara Institute of Science and Technology
    论文:37引用:0H-index:0
    Kikuo Maekawa
    Kikuo Maekawa
    Dept . Language Research;National;Dept . Language Research, National
    论文:24引用:0H-index:0
    Haruo Kubozono
    Haruo Kubozono
    Dept . of Linguistics;Kobe University;Dept . of Linguistics, Kobe University
    论文:13引用:0H-index:0
    Yasuharu Den
    Yasuharu Den
    Graduate School of Humanities, Chiba University
    论文:11引用:0H-index:0
    Hanae Koiso
    Hanae Koiso
    National Institute for Japanese Language and Linguistics
    论文:11引用:0H-index:0
    Yuji Matsumoto
    Yuji Matsumoto
    Graduate School of Information Science, Nara Institute of Science and Technology;RIKEN Center for Advanced Intelligence Project
    论文:11引用:0H-index:0
    Prashant Pardeshi
    Prashant Pardeshi
    Promotion of Science;Kobe University;Promotion of Science, Kobe University
    论文:10引用:0H-index:0
    Yuichi Ishimoto
    Yuichi Ishimoto
    Institute of Technologists
    论文:10引用:0H-index:0
    Toshinobu Ogiso
    Toshinobu Ogiso
    Department of Corpus Studies, National Institute for Japanese Language and Linguistics (NINJAL),
    论文:9引用:0H-index:0

    论文(235)

    年份
    起
    –
    止
    排序
    1Linear Representations of Hierarchical Concepts in Language Models
    Masaki Sakata,Benjamin Heinzerling, Takumi Ito,Sho Yokoi,Kentaro Inui

    We investigate how and to what extent hierarchical relations (e.g., Japan ⊂ Eastern Asia ⊂ Asia) are encoded in the internal representations of language models. Building on Linear Relational Concepts, we train linear transformations specific to each hierarchical depth and semantic domain, and characterize representational differences associated with hierarchical relations by comparing these transformations. Going beyond prior work on the representational geometry of hierarchies in LMs, our analysis covers multi-token entities and cross-layer representations. Across multiple domains we learn such transformations and evaluate in-domain generalization to unseen data and cross-domain transfer. Experiments show that, within a domain, hierarchical relations can be linearly recovered from model representations. We then analyze how hierarchical information is encoded in representation space. We find that it is encoded in a relatively low-dimensional subspace and that this subspace tends to be domain-specific. Our main result is that hierarchy representation is highly similar across these domain-specific subspaces. Overall, we find that all models considered in our experiments encode concept hierarchies in the form of highly interpretable linear representations.

    2026引用:4
    引用
    AI阅读
    加入学术空间
    2Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
    Tomomasa Hara, Hiroto Kurita,Masaaki Imaizumi,Kentaro Inui,Sho Yokoi

    For constructing text embeddings, mean pooling, which averages token embeddings, is the standard approach. This paper examines whether mean pooling actually works well in real models. First, we note that mean pooling can collapse information beyond the first-order statistics of the token embeddings, such as second-order statistics that capture their spatial structure, potentially mapping distinct token embedding distributions to similar text embeddings. Motivated by this concern, we propose a simple metric to quantify such a collapse induced by mean pooling. Then, using this metric, we empirically measure how often this collapse occurs in actual models and texts, and find that modern text encoders are robust to this collapse. In particular, contrastive fine-tuned text encoders tend to be less prone to the collapse than their pretrained backbone models. We also find that the robustness of these text encoders lies in the concentration of token embeddings within each text. In addition, we find that robustness to the collapse, as quantified by our proposed metric, correlates with downstream task performance. Overall, our findings offer a new perspective on why modern text encoders remain effective despite relying on seemingly coarse mean pooling.

    2026ACL 2026(2026)引用:2
    引用
    AI阅读
    加入学术空间
    3An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects Via Surprisal
    Ryo Yoshida, Shinnosuke Isono, Taiga Someya,Yohei Oseki,Tatsuki Kuribayashi

    Surprisal theory hypothesizes that the difficulty of human sentence processing increases linearly with surprisal, the negative log-probability of a word given its context. Computational psycholinguistics has tested this hypothesis using language models (LMs) as proxies for human prediction. While surprisal derived from recent neural LMs generally captures human processing difficulty on naturalistic corpora that predominantly consist of simple sentences, it severely underestimates processing difficulty on sentences that require syntactic disambiguation (garden-path effects). This leads to the claim that the processing difficulty of such sentences cannot be reduced to surprisal, although it remains possible that neural LMs simply differ from humans in next-word prediction. In this paper, we investigate whether it is truly impossible to construct a neural LM that can explain garden-path effects via surprisal. Specifically, instead of evaluating off-the-shelf neural LMs, we fine-tune these LMs on garden-path sentences so as to better align surprisal-based reading-time estimates with actual human reading times. Our results show that fine-tuned LMs do not overfit and successfully capture human reading slowdowns on held-out garden-path items; they even improve predictive power for human reading times on naturalistic corpora and preserve their general LM capabilities. These results provide an existence proof for a neural LM that can explain both garden-path effects and naturalistic reading times via surprisal, but also raise a theoretical question: what kind of evidence can truly falsify surprisal theory?

    2026ACL 2026(2026)引用:1
    引用
    AI阅读
    加入学术空间
    4Timesteps of Mamba Align with Human Reading Times
    Yuji Yamamoto, Shinnosuke Isono, Yoshinobu Kawahara,Sho Yokoi

    This study demonstrates an alignment of per-word processing time in a popular state-space language model Mamba and human readers. In Mamba, the recurrent state transition at each layer conceptually takes some duration of time, the discretization timestep Δ_t, determined dynamically in response to the input. Using a naturalistic reading dataset, we show that the per-word timestep from Mamba is a significant predictor of human reading times, and remains significant even when known predictors such as GPT-2 surprisal are controlled for. We further suggest, through formal analysis of Mamba's architecture and internal dynamics, that Mamba can serve as a new, valuable lens to look at human real-time language processing with ever-updated memory, because it allows us to look at how each module (layer) weighs short- and long-term information retention, and how noise may interact with dynamic, continuous memory representation. Code is available online.

    2026Annual Meeting of the Association for Computational Linguistics(2026)
    引用
    AI阅读
    加入学术空间
    5LLM-Based Dependency Parsing with Step-by-Step Instructions and a Simple Tabular Output
    Hiroshi Matsuda, Chunpeng Ma,Masayuki Asahara
    2026Journal of Natural Language Processing(2026)
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 235 篇论文

    合作机构(100)

    京都大学合作论文 16
    筑波大学合作论文 15
    东京大学合作论文 13
    名古屋大学合作论文 12
    东北大学(日本)合作论文 12
    千叶大学合作论文 10
    奈良科学技术学院合作论文 9
    庆应义塾大学合作论文 7
    国立情报学研究所合作论文 6
    东京工科大学合作论文 5

    机构统计