TENCON 2025 - 2025 IEEE Region 10 Conference (TENCON)(2025)
Department of Electronics and Communication Engineering
被引用0|浏览1
摘要
In today's digital world, processing multilingual documents is critical for business, legal tasks, and information retrieval. This study describes a Multilingual Document Processing System that uses Optical Character Recognition (OCR) and Retrieval-Augmented Generation (RAG) to extract, query and summarize text in multiple languages. The system employs advanced OCR models to correctly recognize text from scanned documents, images, and handwriting in various scripts. By incorporating RAG, it improves comprehension and response generation, allowing users to retrieve and summarize information in English even when the original language is different. This approach takes advantage of recent advances in natural language processing, large language models (LLM), and multimodal AI to address challenges in multilingual data accessibility, knowledge synthesis, and real-time communication. The system provides a scalable AI-driven solution to improve document processing, eliminate language barriers, and increase global user engagement. AWS services support scalable document processing but cold starts in AWS Lambda hinder real time tasks.
更多
查看译文
关键词
Amazon AI Services,Optical Character Recognition,Retrieval Augmented Generation,