• 学术搜索
  • 科研智能体
    • Research Labs
    • AI 阅读
    • AI 文库
    • 深度研究
    • 学者亮点
  • 学术资源
    • AI2000
    • 期刊/会议
    • 学者库
    • 学术API
    • 溯源树
    • 数据集
  • 知识沉淀
    • 学术空间
订阅小程序
旧版功能
aminer vip
开通会员低至0.73元/天
一次搞定AI科研
立即登录
  • English
  • 联系方式
    L

    Laboratoire d'Informatique, du Traitement de l'Information et des Systèmes

    EST. 2006
    662论文总数
    6,546引用总数

    论文量&引用量时间轴

    机构学者

    排序
    Emmanuel Trouvé
    Emmanuel Trouvé
    LISTIC Laboratory, Université de Savoie
    论文:31引用:0H-index:0
    Sylvie Galichet
    Sylvie Galichet
    LISTIC - Polytech Annecy-Chambéry, University of Savoie Mont-Blanc
    论文:23引用:0H-index:0
    Stéfan Darmoni
    Stéfan Darmoni
    Département de Médecine, Université de Rouen Normandie
    论文:20引用:0H-index:0
    Laurent Foulloy
    Laurent Foulloy
    Laboratoire d’InformatiqueSystèmesTraitement de l’Information et de la Connaissance, Université de Savoie
    论文:19引用:0H-index:0
    Lamia Berrah
    Lamia Berrah
    Laboratoire d'Informatique, Université de Savoie
    论文:16引用:0H-index:0
    H. Verjus
    H. Verjus
    Computer Science Department
    论文:16引用:0H-index:0
    Thierry Paquet
    Thierry Paquet
    Universite de Rouen
    论文:14引用:0H-index:0
    Philippe Bolon
    Philippe Bolon
    LISTIC, Université Savoie Mont-Blanc;Polytech School of Engineering, University Savoie Mont-Blanc
    论文:13引用:0H-index:0
    Cyrille Bertelle
    Cyrille Bertelle
    LITIS, Univ Le Havre
    论文:13引用:0H-index:0

    论文(662)

    年份
    起
    –
    止
    排序
    1Global and Local Mamba Network for Multi-Modality Medical Image Super-Resolution
    Zexin Ji,Beiji Zou,Xiaoyan Kui,Sebastien Thureau,Su Ruan

    Convolutional neural networks and Transformer have made significant progress in multi-modality medical image super-resolution. However, these methods either have a fixed receptive field for local learning or significant computational burdens for global learning, limiting the super-resolution performance. To solve this problem, State Space Models, notably Mamba, is introduced to efficiently model long-range dependencies in images with linear computational complexity. Relying on the Mamba and the fact that low-resolution images rely on global information to compensate for missing details, while high-resolution reference images need to provide more local details for accurate super-resolution, we propose a global and local Mamba network (GLMamba) for multi-modality medical image super-resolution. To be specific, our GLMamba is a two-branch network equipped with a global Mamba branch and a local Mamba branch. The global Mamba branch captures long-range relationships in low-resolution inputs, and the local Mamba branch focuses more on short-range details in high-resolution reference images. We also use the deform block to adaptively extract features of both branches to enhance the representation ability. A modulator is designed to further enhance deformable features in both global and local Mamba blocks. To effectively incorporate reference guidance into low-resolution image super-resolution(SR), we further develop a multi-modality feature fusion block to adaptively fuse features by considering similarities, differences, and complementary aspects between modalities. In addition, a contrastive edge loss (CELoss) is developed for sufficient enhancement of edge textures and contrast in medical images. Quantitative and qualitative experimental results show that our GLMamba achieves superior super-resolution performance on BraTS2021, IXI and fastMRI datasets. We also validate the effectiveness of our approach on the downstream tumor segmentation task.

    2026PATTERN RECOGNITION(2026)引用:4
    引用
    AI阅读
    加入学术空间
    2PILOT: A Promptable Interleaved Layout-aware OCR Transformer
    Laziz Hamdi, Amine Tamasna, Pascal Boisson,Thierry Paquet

    Classical OCR pipelines decompose document reading into detection, segmentation, and recognition stages, which makes them sensitive to localization errors and difficult to extend to interactive querying. This work investigates whether a single compact model can jointly perform text recognition and spatial grounding on both handwritten and printed documents. We introduce PILOT, a 155M-parameter prompt-conditioned generative model that formulates document OCR as unified sequence generation. A lightweight depthwise-separable CNN encodes the page, and a Transformer decoder autoregressively emits a single stream of subword and quantized absolute-coordinate tokens on a 10 px grid, enabling full-page OCR, region-conditioned reading, and query-by-string spotting within the same architecture. A three-stage curriculum, progressing from plain transcription to joint text-and-box generation and finally to prompt-controlled extraction, stabilizes training and improves spatial grounding. Experiments on IAM, RIMES 2009, SROIE 2019, and the heterogeneous MAURDOR benchmark show that PILOT achieves competitive or superior performance in text recognition and line-level detection compared with traditional OCR systems, recent end-to-end HTR models, and compact vision–language models, while remaining substantially smaller than billion-scale multimodal models. Additional evaluations on fine-grained OCR and query-by-string spotting further confirm that a unified text–layout decoder can provide accurate and efficient promptable OCR in a compact setting. To support reproducibility, we will release the synthetic SROIE generator, the 500k annotated IDL/PDFA pages, and the harmonized line-level annotations for IAM, RIMES 2009, and MAURDOR, with code made publicly available upon acceptance.

    2026International Journal on Document Analysis and Recognition (IJDAR)(2026)引用:3
    引用
    AI阅读
    加入学术空间
    3DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates
    Laziz Hamdi, Amine Tamasna, Thierry Paquet

    Tables condense key transactional and administrative information into compact layouts, but practical extraction requires more than text recognition: systems must also recover structure (rows, columns, merged cells, headers) and interpret roles such as line items, subtotals, and totals under common capture artifacts. Many existing resources for table structure recognition and TableVQA are built from clean digital-born sources or rendered tables, and therefore only partially reflect noisy administrative conditions. We introduce DenTab, a dataset of 2,000 cropped table images from dental estimates with high-quality HTML annotations, enabling evaluation of table recognition (TR) and table visual question answering (TableVQA) on the same inputs. Its TableVQA test split contains 2,208 questions across eleven categories spanning retrieval, aggregation, and logic/consistency checks. We benchmark 16 systems spanning general vision–language models, document parsers, and OCR-oriented models. Across models, strong structure recovery does not consistently translate into reliable performance on multi-step arithmetic and consistency questions, and these reasoning failures persist even when using ground-truth HTML table inputs. To improve arithmetic reliability without training, we propose the Table Router Pipeline, which routes arithmetic questions to deterministic execution. The pipeline combines direct VQA, a TSR-derived structured table representation, a VLM-generated constrained table program, and a rule-based executor that performs exact computation over the parsed table. The source code and dataset will be made publicly available at https://github.com/hamdilaziz/DenTab .

    2026Pattern Recognition(2026)
    引用
    AI阅读
    加入学术空间
    4TableSeq: Unified Generation of Structure, Content, and Layout
    Laziz Hamdi, Amine Tamasna, Pascal Boisson, Thierry Paquet

    We present TableSeq, an image-only, end-to-end framework for joint table structure recognition, content recognition, and cell localization. The model formulates these tasks as a single sequence-generation problem: one decoder produces an interleaved stream of HTML tags, cell text, and discretized coordinate tokens, thereby aligning logical structure, textual content, and cell geometry within a unified autoregressive sequence. This design avoids external OCR, auxiliary decoders, and complex multi-stage post-processing. TableSeq combines a lightweight high-resolution FCN-H16 encoder with a minimal structure-prior head and a single-layer transformer encoder, yielding a compact architecture that remains effective on challenging layouts. Across standard benchmarks, TableSeq achieves competitive or state-of-the-art results while preserving architectural simplicity. It reaches 95.23 TEDS / 96.83 S-TEDS on PubTabNet, 97.45 TEDS / 98.69 S-TEDS on FinTabNet, and 99.79 / 99.54 / 99.66 precision / recall / F1 on SciTSR under the CAR protocol, while remaining competitive on PubTables-1 M under GriTS. Beyond TSR/TCR, the same sequence interface generalizes to index-based table querying without task-specific heads, achieving the best IRDR score and competitive ICDR/ICR performance. We also study multi-token prediction for faster blockwise decoding and show that it reduces inference latency with only limited accuracy degradation. Overall, TableSeq provides a practical and reproducible single-stream baseline for unified table recognition.

    2026International Journal on Document Analysis and Recognition (IJDAR)(2026)
    引用
    AI阅读
    加入学术空间
    5E-Alliance: a Software Infrastructure for Concurrent Inter-Organisational Alliances
    Ilham Alloui,Jean Marc Andreoli,Olivier Boissier,Mihnea Bratu,Stefania Castellani,Karim Megzari

    In this paper we introduce e-Alliance, a software infrastructure we are defining for supporting negotiation activities in concurrent inter-organisational alliances. e-Alliance main intent is to preserve autonomy of organisations within an alliance while enabling concurrency of their activities, flexibility of their negotiations and dynamism/evolution of their environment. The IT infrastructure we propose combines different technologies, such as software engineering techniques, middleware-level coordination facilities and multi-agent systems support. We present our approach in the context offered by a sample scenario where business-to-business interactions hold among printshops grouped into an alliance to better answer customers' demands.

    2026Advances in Concurrent Engineering(2026)
    引用
    AI阅读
    加入学术空间
    立即登录,查看全部 662 篇论文

    合作机构(100)

    Institut des Sciences de l''Ingénierie et des Systèmes,French National Centre for Scientific Research合作论文 12
    法国国家科学研究中心合作论文 7
    诺曼底大学合作论文 6
    鲁昂大学合作论文 6
    Government of France合作论文 6
    Groupe de Recherche en Informatique, Image, Automatique et Instrumentation de Caen合作论文 5
    法国国立计算机科学及自动化研究院合作论文 5
    巴黎萨克雷大学合作论文 5
    Centre Hospitalier Universitaire de Rouen合作论文 5
    Département Mathématiques et Informatique Appliquées,National Research Institute for Agriculture, Food and Environment合作论文 4

    机构统计