• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    马

    马卡罗夫海军上将国立造船大学

    Admiral Makarov National University of Shipbuilding
    院校EST. 1920
    1,414论文总数
    1.2万引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Inna Irtyshcheva
    Inna Irtyshcheva
    Admiral Makarov Natl Univ Shipbldg
    论文:40引用:0H-index:0
    Andrii Radchenko
    Andrii Radchenko
    Department of Air Conditioning and Refrigeration, Admiral Makarov National University of Shipbuilding
    论文:40引用:0H-index:0
    Dmytro Konovalov
    Dmytro Konovalov
    Kherson Educ Sci Inst, Admiral Makarov Natl Univ Shipbuilding
    论文:25引用:0H-index:0
    Mykola Radchenko
    Mykola Radchenko
    Admiral Makarov National University of Shipbuilding
    论文:22引用:0H-index:0
    Roman Radchenko
    Roman Radchenko
    Conditioning and Refrigeration Department, Admiral Makarov National University of Shipbuilding
    论文:21引用:0H-index:0
    Iryna Kramarenko
    Iryna Kramarenko
    Admiral Makarov National University of Shipbuilding
    论文:18引用:0H-index:0
    Богдан Сергійович Портной
    Богдан Сергійович Портной
    Admiral Makarov Natl Univ Shipbldg, 9 Heroes Ukraine Ave, UA-54025 Mykolayiv, Ukraine
    论文:15引用:0H-index:0
    Сергій Анатолійович Кантор
    Сергій Анатолійович Кантор
    ПАТ "Завод "Екватор",
    论文:14引用:0H-index:0
    Nataliya Hryshyna
    Nataliya Hryshyna
    Admiral Makarov National University of Shipbuilding
    论文:13引用:0H-index:0

    论文(1414)

    年份
    起
    –
    止
    排序
    1Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
    Yue Liao, Pengfei Zhou,Siyuan Huang, Donglin Yang, Shengcong Chen, Yuxin Jiang,Yue Hu,Si Liu,Jianlan Luo, Liliang Chen,Shuicheng YAN, Maoqing Yao,

    We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that jointly learns visual representations and action policies within a single video-generative framework. At its core, GE-Base is a large-scale instruction-conditioned video diffusion model that captures the spatial, temporal, and semantic dynamics of real-world robotic interactions in a structured latent space. Building on this foundation, GE-Act employs a lightweight flow-matching decoder to map latent representations into executable action trajectories, enabling precise and generalizable policy inference across diverse embodiments with minimal supervision. Trained on over 1 million manipulation episodes, GE supports both short- and long-horizon tasks, and generalizes across embodiments. All code, models, and benchmarks will be released publicly.

    ICLR 2026引用:118
    引用
    AI阅读
    加入学术空间
    2AgenTracer: Who is Inducing Failure in the LLM Agentic Systems?
    Guibin Zhang, Junhao Wang, Junjie Chen,Wangchunshu Zhou,Kun Wang, Shuicheng YAN

    Large Language Model (LLM)-based agentic systems, often comprising multiple models, complex tool invocations, and orchestration protocols, substantially outperform monolithic agents. Yet this very sophistication amplifies their fragility, making them more prone to system failure. Pinpointing the specific agent or step responsible for an error within long execution traces defines the task of \textbf{agentic system failure attribution}. Current state-of-the-art reasoning LLMs, however, remain strikingly inadequate for this challenge, with accuracy generally below $10\\%$. To address this gap, we propose AgenTracer, the first automated framework for annotating failed multi-agent trajectories via counterfactual replay and programmed fault injection, producing the curated dataset TracerTraj. Leveraging this resource, we develop AgenTracer-8B, a lightweight failure tracer trained with multi-granular reinforcement learning, capable of efficiently diagnosing errors in verbose multi-agent interactions. On {Who\&When} benchmark, AgenTracer-8B outperforms giant proprietary LLMs like Gemini-2.5-Pro and Claude-4-Sonnet by up $18.18\\%$, setting a new standard in LLM agentic failure attribution. More importantly, AgenTracer-8B delivers actionable feedback to off-the-shelf multi-agent systems like MetaGPT and MaAS with $4.8\sim14.2\\%$ performance gains, empowering self-correcting and self-evolving agentic AI.

    ICLR 2026引用:75
    引用
    AI阅读
    加入学术空间
    3CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery
    Ao Qu, Han Zheng, Zijian Zhou, Yihao Yan,Yihong Tang, Shao Yong Ong, Fenglu Hong,Kaichen Zhou, Chonghe Jiang, Minwei Kong, Jiacheng Zhu,Sirui Li,

    Large language model (LLM)-based evolution is a promising approach for open-ended discovery, where progress requires sustained search and knowledge accumulation. Existing methods still rely heavily on fixed heuristics and hard-coded exploration rules, which limit the autonomy of LLM agents. We present CORAL, the first framework for autonomous multi-agent evolution on open-ended problems. CORAL replaces rigid control with long-running agents that explore, reflect, and collaborate through shared persistent memory, asynchronous multi-agent execution, and heartbeat-based interventions. It also provides practical safeguards, including isolated workspaces, evaluator separation, resource management, and agent session and health management. Evaluated on diverse mathematical, algorithmic, and systems optimization tasks, CORAL sets new state-of-the-art results on 10 tasks, achieving 3--10X higher improvement rates with far fewer evaluations than fixed evolutionary search baselines across tasks. On Anthropic's kernel engineering task, four co-evolving agents improve the best known score from 1363 to 1103 cycles. Mechanistic analyses further show how these gains arise from knowledge reuse and multi-agent exploration and communication. Together, these results suggest that greater agent autonomy and multi-agent evolution can substantially improve open-ended discovery. Code is available at \url{https://anonymous.4open.science/r/coral_anon-DA07/}.

    COLM 2026引用:34
    引用
    AI阅读
    加入学术空间
    4Kimi-Dev: Agentless Training As Skill Prior for SWE-Agents
    Zonghan Yang,Shengjie Wang, Kelin Fu, Wenyang He,Weimin Xiong,Yibo Liu,Yibo Miao,Bofei Gao,Yejie Wang,YINGWEI MA, Yanhao Li,Yue Liu,

    Large Language Models (LLMs) are increasingly applied to software engineering (SWE), with SWE-bench as a key benchmark. Solutions are split into SWE-Agent frameworks with multi-turn interactions and workflow-based Agentless methods with single-turn verifiable steps. We argue these paradigms are not mutually exclusive: reasoning-intensive Agentless training induces skill priors, including localization, code edit, and self-reflection that enable efficient and effective SWE-Agent adaptation. In this work, we first curate the Agentless training recipe and present Kimi-Dev, an open-source SWE LLM achieving 60.4\% on SWE-bench Verified, the best among workflow approaches. With additional SFT adaptation on 5k publicly-available trajectories, Kimi-Dev powers SWE-Agents to 48.6\% pass@1, on par with that of Claude 3.5 Sonnet (241022 version). These results show that structured skill priors from Agentless training can bridge workflow and agentic frameworks for transferable coding agents.

    ICLR 2026引用:31
    引用
    AI阅读
    加入学术空间
    5RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
    Yang Shi,Yuhao Dong, Yue Ding, Yuran Wang, Xuanyu Zhu,Sheng Zhou, Wenting Liu, Haochen Tian, Rundong Wang, Huanqian Wang,Zuyan Liu, Bohan Zeng,

    The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question remains unanswered by existing benchmarks: does this architectural unification actually enable synergetic interaction between the constituent capabilities? Existing evaluation paradigms, which primarily assess understanding and generation in isolation, are insufficient for determining whether a unified model can leverage its understanding to enhance its generation, or use generative simulation to facilitate deeper comprehension. To address this critical gap, we introduce RealUnify, a benchmark specifically designed to evaluate bidirectional capability synergy. RealUnify comprises 1,000 meticulously human-annotated instances spanning 10 categories and 32 subtasks. It is structured around two core axes: 1) Understanding Enhances Generation, which requires reasoning (e.g., commonsense, logic) to guide image generation, and 2) Generation Enhances Understanding, which necessitates mental simulation or reconstruction (e.g., of transformed or disordered visual inputs) to solve reasoning tasks. A key contribution is our dual-evaluation protocol, which combines direct end-to-end assessment with a diagnostic stepwise evaluation that decomposes tasks into distinct understanding and generation phases. This protocol allows us to precisely discern whether performance bottlenecks stem from deficiencies in core abilities or from a failure to integrate them. Through large-scale evaluations of 12 leading unified models and 6 specialized baselines, we find that current unified models still struggle to achieve effective synergy, indicating that architectural unification alone is insufficient. These results highlight the need for new training strategies and inductive biases to fully unlock the potential of unified modeling.

    ICLR 2026引用:28
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 1414 篇论文

    合作机构(99)

    新加坡国立大学合作论文 69
    诺丁汉特伦特大学合作论文 38
    Petro Mohyla Black Sea National University合作论文 28
    香港科技大学合作论文 26
    Mykolayiv National Agrarian University合作论文 23
    乌克兰国家科学院合作论文 22
    香港中文大学合作论文 20
    江苏科技大学合作论文 19
    中国科学技术大学合作论文 19
    北京大学合作论文 18

    机构统计