• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    M

    Moscow Technical University of Communication and Informatics

    院校EST. 1921
    1,305论文总数
    3,004引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Mikhail Gorodnichev
    Mikhail Gorodnichev
    Dept Math Cybernet & Informat Technol, Moscow Tech Univ Commun & Informat
    论文:41引用:0H-index:0
    V.B. Kreyndelin
    V.B. Kreyndelin
    Moscow Technical University of Communications and Informatics
    论文:24引用:0H-index:0
    O. V. Varlamov
    O. V. Varlamov
    Science and Research Department, Moscow Technical University of Communications and Informatics
    论文:24引用:0H-index:0
    Alexander Gavrilovich Kyurkchan
    Alexander Gavrilovich Kyurkchan
    Moscow Technical University of Communications and Informatics
    论文:21引用:0H-index:0
    S.A. Manenkov
    S.A. Manenkov
    Moscow Technical University of Communications and Informatics
    论文:19引用:0H-index:0
    Artem Sergeevich Adzhemov
    Artem Sergeevich Adzhemov
    Moscow Technical University of Communications and Informatics
    论文:19引用:0H-index:0
    A. M. Raitsin
    A. M. Raitsin
    Moscow Technical University of Communications and Informatics
    论文:18引用:0H-index:0
    S Yu Kazantsev
    S Yu Kazantsev
    Moscow Technical University of Communications and Informatics
    论文:17引用:0H-index:0
    Sergey Yablochnikov
    Sergey Yablochnikov
    Moscow Technical University of Communications and Informatics
    论文:17引用:0H-index:0

    论文(1305)

    年份
    起
    –
    止
    排序
    1Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance
    Ivan Viakhirev, Kirill Borodin,Mikhail Gorodnichev, Grach Mkrtchian

    Multi-branch deep neural networks like AASIST3 achieve state-of-the-art comparable performance in audio anti-spoofing, yet their internal decision dynamics remain opaque compared to traditional input-level saliency methods. While existing interpretability efforts largely focus on visualizing input artifacts, the way individual architectural branches cooperate or compete under different spoofing attacks is not well characterized. This paper develops a framework for interpreting AASIST3 at the component level. Intermediate activations from fourteen branches and global attention modules are modeled with covariance operators whose leading eigenvalues form low-dimensional spectral signatures. These signatures train a CatBoost meta-classifier to generate TreeSHAP-based branch attributions, which we convert into normalized contribution shares and confidence scores (Cb) to quantify the model’s operational strategy. By analyzing 13 spoofing attacks from the ASVspoof 2019 benchmark, we identify four operational archetypes—ranging from “Effective Specialization” (e.g., A09, Equal Error Rate (EER) 0.04%, C=1.56) to “Ineffective Consensus” (e.g., A08, EER 3.14%, C=0.33). Crucially, our analysis exposes a “Flawed Specialization” mode where the model places high confidence in an incorrect branch, leading to severe performance degradation for attacks A17 and A18 (EER 14.26% and 28.63%, respectively). These quantitative findings link internal architectural strategy directly to empirical reliability, highlighting specific structural dependencies that standard performance metrics overlook.

    2026MATHEMATICS(2026)引用:1
    引用
    AI阅读
    加入学术空间
    2From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales
    Ivan Viakhirev, Kirill Borodin, Grach Mkrtchian

    Hallucinations in large ASR models present a critical safety risk. In this work, we propose the Spectral Sensitivity Theorem, which predicts a phase transition in deep networks from a dispersive regime (signal decay) to an attractor regime (rank-1 collapse) governed by layer-wise gain and alignment. We validate this theory by analyzing the eigenspectra of activation graphs in Whisper models (Tiny to Large-v3-Turbo) under adversarial stress. Our results confirm the theoretical prediction: intermediate models exhibit Structural Disintegration (Regime I), characterized by a 13.4% collapse in Cross-Attention rank. Conversely, large models enter a Compression-Seeking Attractor state (Regime II), where Self-Attention actively compresses rank (-2.34%) and hardens the spectral slope, decoupling the model from acoustic evidence.

    2026引用:1
    引用
    AI阅读
    加入学术空间
    3Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
    Dmitrii Korzh, Dmitrii Tarasov, Artyom Iudin, Elvir Karimov, Matvey Skripkin, Nikita Kuzmin,Andrey Kuznetsov,Oleg Rogov,Ivan Oseledets

    Conversion of spoken mathematical expressions is a challenging task that involves transcribing speech into a strictly structured symbolic representation while addressing the ambiguity inherent in the pronunciation of equations. Although significant progress has been achieved in automatic speech recognition (ASR) and language models (LM), the problem of converting spoken mathematics into LaTeX remains underexplored. This task directly applies to educational and research domains, such as lecture transcription or note creation. Based on ASR post-correction, prior work requires 2 transcriptions, focuses only on isolated equations, has a limited test set, and provides neither training data nor multilingual coverage. To address these issues, we present the first fully open-source large-scale dataset, comprising over 66,000 human-annotated audio samples of mathematical equations and sentences in English and Russian, drawn from diverse scientific domains. In addition to the ASR post-correction models and few-shot prompting, we apply audio language models, demonstrating comparable character error rate (CER) results on the MathSpeech benchmark (28\% vs. 30\%) for the equations conversion. In contrast, on the proposed S2L-equations benchmark, our models outperform the MathSpeech model by a substantial margin of more than 36 percentage points, even after accounting for LaTeX formatting artifacts (27\% vs. 64\%). We establish the first benchmark for mathematical sentence recognition (S2L-sentences) and achieve an equation CER of 40\%. This work lays the groundwork for future advances in multimodal AI, with a particular focus on mathematical content recognition.

    ICLR 2026引用:1
    引用
    AI阅读
    加入学术空间
    4Towards Robust Speech Deepfake Detection Via Human-Inspired Reasoning
    Artem Dvirniak, Evgeny Kushnir, Dmitrii Tarasov, Artem Iudin, Oleg Kiriukhin,Mikhail Pautov, Dmitrii Korzh,Oleg Y. Rogov

    The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to evolve. Unfortunately, current SDD methods generally suffer from the lack of generalization to new audio domains and generators. More than that, they lack interpretability, especially human-like reasoning that would naturally explain the attribution of a given audio to the bona fide or spoof class and provide human-perceptible cues. In this paper, we propose HIR-SDD, a novel SDD framework that combines the strengths of Large Audio Language Models (LALMs) with the chain-of-thought reasoning derived from the novel proposed human-annotated dataset. Experimental evaluation demonstrates both the effectiveness of the proposed method and its ability to provide reasonable justifications for predictions.

    2026引用:1
    引用
    AI阅读
    加入学术空间
    5LLM-Guided Prompt Evolution for Password Guessing
    Vladimir A. Mazin, Mikhail A. Zorin, Dmitrii S. Korzh, Elvir Z. Karimov, Dmitrii A. Bolokhov, Oleg Y. Rogov

    Passwords still remain a dominant authentication method, yet their security is routinely subverted by predictable user choices and large-scale credential leaks. Automated password guessing is a key tool for stress-testing password policies and modeling attacker behavior. This paper applies LLM-driven evolutionary computation to automatically optimize prompts for the LLM password guessing framework. Using OpenEvolve, an open-source system combining MAP-Elites quality-diversity search with an island population model we evolve prompts that maximize cracking rate on a RockYou-derived test set. We evaluate three configurations: a local setup with Qwen3 8B, a single compact cloud model Gemini-2.5 Flash, and a two-model ensemble of frontier LLMs. The approach raises the cracking rates from 2.02% to 8.48%. Character distribution analysis further confirms how evolved prompts produce statistically more realistic passwords. Automated prompt evolution is a low-barrier yet effective way to strengthen LLM-based password auditing and underlining how attack pipelines show tendency via automated improvements.

    2026引用:1
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 1305 篇论文

    合作机构(100)

    俄罗斯科学院合作论文 35
    莫斯科动力工程研究所合作论文 22
    Financial University合作论文 15
    MIREA - Russian Technological University合作论文 14
    莫斯科国立大学合作论文 13
    All-Russian Research Institute for Optical and Physical Measurements合作论文 13
    莫斯科工程物理研究所合作论文 11
    Don State Technical University合作论文 11
    莫斯科罗蒙诺索夫国立大学合作论文 10
    North Ossetian State University合作论文 10

    机构统计