The classification of Intangible Cultural Heritage (ICH) imagery remains challenging due to subtle semantics, fine-grained inter-class boundaries, and culturally diverse representations. Although Kolmogorov–Arnold Networks (KANs) exhibit strong representational power, their practical use is limited by the computational overhead of B-spline activations and training instability from high variance. To overcome these issues, we propose HeritKAN, a computationally efficient KAN variant employing compactly supported Wendland C^2 kernels for sparse and accelerated inference. We further introduce BaggingHeritKAN, an ensemble model using bootstrap aggregation to reduce variance and improve generalization. In addition, we curate the Vietnamese ICH dataset with 10,143 images across 11 categories and develop a web-based recognition system for real-time heritage identification. Experiments show that HeritKAN improves efficiency while BaggingHeritKAN achieves the highest accuracy ( 96.25% ), surpassing CNN ( 93.36% ) and standard KAN ( 94.94% ), providing a principled framework for cultural heritage analysis and digital preservation.
Image-text retrieval has become a fundamental component in intelligent multimedia systems; however, most existing vision-language models are optimized for highresource languages and remain suboptimal for low-resource settings such as Vietnamese. This work introduces ViCLIP-OT, a foundation vision-language model specifically designed for Vietnamese image-text retrieval. The proposed framework integrates CLIP-style contrastive learning with a Similarity-Graph Regularized Optimal Transport (SIGROT) loss to enhance global cross-modal consistency and mitigate modality gap issues. Extensive experiments on three Vietnamese benchmarks (UITOpenViIC, KTVIC, and Crossmodal-3600) demonstrate that ViCLIP-OT consistently outperforms CLIP and SigLIP baselines in both in-domain and zero-shot settings. On UIT-OpenViIC, the model achieves an average Recall@K of 67.34
Transparent educational question answering asks for answers that are not only correct but explainable, and doing so with small models rules out the reasoning power of the largest proprietary systems. The EXACT 2026 competition poses this problem concretely: open-weight language models of at most 8B parameters, self-hosted, with a natural-language explanation for every answer. It pairs two tasks: logical reasoning over university regulations, and multi-step physics problem solving. We describe the system that team developed to address both, a neuro-symbolic Program-of-Thought pipeline in which a 4B backbone writes a program rather than stating an answer directly: for regulation queries it emits a Z3 encoding whose entailment verdict grounds the deduction, and for physics it emits numerical Python, both wrapped in a shared self-correction loop and a unified explained-JSON output. Answer-type routing, distillation-based task fine-tuning, and a latency-aware serving stack – SGLang with speculative decoding – keep the system within the 60-second per-query limit. The system achieved a perfect score on the physics task in both automated selection rounds and obtained the highest final-round technical score of any team – 13.44/15, combining automated answer evaluation with expert-judged reasoning depth – with the equally weighted presentation score included, placed 3rd overall. Grounding answers in a symbolic solver yields correct, verifiable deductions at the 4B scale, and the residual difficulty lies in premise selection rather than the deduction itself.
The classification of Intangible Cultural Heritage (ICH) images in the Mekong Delta poses unique challenges due to limited annotated data, high visual similarity among classes, and domain heterogeneity. In such low-resource settings, conventional deep learning models often suffer from high variance or overfit to spurious correlations, leading to poor generalization. To address these limitations, we propose a robust framework that integrates the hybrid CoAtNet architecture with model soups, a lightweight weight-space ensembling technique that averages checkpoints from a single training trajectory without increasing inference cost. CoAtNet captures both local and global patterns through stage-wise fusion of convolution and self-attention. We apply two ensembling strategies - greedy and uniform soup - to selectively combine diverse checkpoints into a final model. Beyond performance improvements, we analyze the ensembling effect through the lens of bias-variance decomposition. Our findings show that model soups reduces variance by stabilizing predictions across diverse model snapshots, while introducing minimal additional bias. Furthermore, using cross-entropy-based distance metrics and Multidimensional Scaling (MDS), we show that model soups selects geometrically diverse checkpoints, unlike Soft Voting, which blends redundant models centered in output space. Evaluated on the ICH-17 dataset (7,406 images across 17 classes), our approach achieves state-of-the-art results with 72.36
Chest X-ray classification is limited by scarce annotations and the heavy cost of transformer models. We propose a three-stage framework: (1) self-supervised pretraining of a Vision Transformer (ViT) with DINOv2 on 880k unlabeled radiographs, (2) fine-tuning on ChestX-ray14, and (3) knowledge transfer into MobileViT using a combination of Binary Cross-Entropy (BCE) and Multi-Label Distillation (MLD) loss. The distilled MobileViT achieves a mean AUROC of 0.8404, surpassing its supervised counterpart by 1.9
Document clustering remains a fundamental task in information retrieval, yet accurately capturing semantic structure in long and context-rich texts poses persistent challenges. In this paper, we propose SMoC-LC (Segment-based Mixture of Clusters with Late Chunking), a novel clustering framework that addresses two key limitations of prior methods: fixed-length segmentation and hard cluster assignments. Our approach introduces Late Chunking to produce flexible, variable-length text segments using long-context embeddings, and employs Gaussian Mixture Models (GMM) to enable soft-probabilistic clustering. We benchmark SMoC-LC and its variants (SBoC, SBoC-LC, SMoC) on seven datasets spanning different domains and structural complexity, including AGNews, 20News-10K, BBCNews, Reuters-21578, and DBpedia (L1-L3). Results show that SMoC-LC consistently improves clustering quality across accuracy (ACC), normalized mutual information (NMI), and adjusted Rand index (ARI), with statistically significant gains observed in complex, hierarchical datasets. Our analysis reveals that Late Chunking is especially beneficial for short, structured documents, while soft clustering excels in ambiguous or multi-topic contexts. These findings underscore the need for adaptable clustering strategies aligned with textual granularity and semantic ambiguity.
This study presents an innovative AI-driven system designed to generate interactive mind-maps tailored for history education. At its core, the system constructs a hierarchical ontology using a novel clustering algorithm that groups semantically related historical concepts. This structured ontology is then visualized as dynamic, interactive mind maps, enabling learners to comprehend and explore historical knowledge more intuitively and effectively. In addition, a RAG-based chatbot is integrated into the system, allowing users to ask questions and receive context-aware responses derived from the ontology. To support this functionality, the study introduces a framework called CAHIM, which automatically extracts and organizes knowledge from raw documents (e.g., PDFs) to construct the ontology used for both visualization and question answering. Experimental results demonstrate that this approach significantly improves both learning efficiency and user engagement when compared with conventional methods.
Nghiên cứu này được thực hiện nhằm đề xuất một mô hình giáo dục tích hợp các phương pháp giảng dạy hiện đại như WebQuest, Học qua Giảng dạy (Learning by Teaching - LbT), Xây dựng lớp học tư duy (Building Thinking Classrooms - BTC), và Truy vấn dựa trên khái niệm (Concept-Based Inquiry - CBI). Mô hình kết hợp hiệu quả các công cụ giáo dục (OKMindmap, YouTube), công nghệ giáo dục (ChatGPT, Khai phá dữ liệu học tập - EDM) và truyền thông giáo dục (VideoTeach), hướng tới phát triển toàn diện năng lực 6Cs: phẩm chất, công dân toàn cầu, hợp tác, giao tiếp, sáng tạo và tư duy phản biện. Thông qua tổng quan tài liệu và thiết kế mô hình, kết quả nghiên cứu giúp chứng minh tính khả thi và tiềm năng của mô hình trong việc thúc đẩy đổi mới sư phạm và phát triển năng lực người học trong bối cảnh giáo dục hiện đại.
Multi-label classification of chest radiographs presents unique challenges due to noisy labels, severe class imbalance, and co-occurring pathologies. In this study, MS-CXR is proposed, a soft-voting ensemble that combines convolutional and transformer-based architectures to improve robustness and generalization. The ensemble integrates four diverse models: a DenseNet-121 trained with MixUp and CutMix, a Swin Transformer, a CoAtNet, and a Vision Transformer pretrained via self-supervised learning on radiographic data (ViT-DINOv2). Evaluation on the ChestX-ray14 dataset under a patient-level split shows that MS-CXR achieves a new state-of-the-art macro-average AUROC of 86.37
The quest for accurate traffic density estimation is gaining momentum globally, with Vietnam distinguished by its ranking among the top ten nations for private vehicle usage. Rapid advancements in computer vision, particularly through the development of convolutional neural network (CNN) methodologies, underscore the pressing need to incorporate these techniques into traffic density estimation efforts. In this study, three convolutional neural network (CNN) models—W-Net, UASD-Net (a fusion of U-Net with Adaptive Scenario Discovery), and CSR-Net (Congested Scene Recognition Network)—are employed to quantify and assess traffic density based on images captured in Vietnam. Furthermore, a novel approach for reallocating label points to generate more accurate density maps is proposed. Experimental results on a composite dataset, integrating the TRANCOS, TayDo, and KienGiang datasets, demonstrate promising mean absolute error rates of 3.67, 4.42, and 3.82 for W-Net, UASD-Net, and CSR-Net, respectively.
Document clustering plays a crucial role in various information retrieval tasks. Existing approaches often struggle with capturing the semantic relationships between documents, especially when dealing with long and complex texts. To address this issue, we propose SBoC, a novel Segment-based Bag-of-Clusters approach. SBoC first divides documents into segments, capturing local semantic information. It then applies clustering algorithms to these segments, forming clusters that represent distinct semantic concepts. Finally, a Bag-of-Clusters representation is constructed for each document, encoding its semantic content based on the assigned segment clusters. SBoC shows promising results, particularly in terms of capturing semantic relationships in document clustering. While not surpassing all existing methods, SBoC demonstrates competitive performance on benchmark datasets, particularly when handling long and complex texts. This approach provides a potential solution for enhancing document clustering for various information retrieval tasks.
Việc phát hiện kịp thời khối u hỗ trợ các bác sĩ trong quá trình chẩn đoán và điều trị cho bệnh nhân được thực hiện hiệu quả trong tình trạng các bệnh viện luôn quá tải là rất cần thiết. Ứng dụng Slicer cho phép dựng hình ảnh 2D vùng tổn thương thành dữ liệu khối 3D giúp các bác sĩ có cái nhìn trực quan hơn trong việc chẩn đoán và điều trị. Tuy nhiên, ứng dụng Slicer chưa cho phép phát hiện tự động vùng bất thường và yêu cầu máy tính đủ mạnh để thực thi các mô hình này. Trong nghiên cứu này, tiện ích mở rộng Billow AISA cho Slicer được đề xuất nhằm xây dựng một cổng dịch vụ phân tích, dự đoán từ dữ liệu ảnh do người dùng cung cấp. Chức năng phân tích, dự đoán được thử nghiệm trong nghiên cứu này là phát hiện vùng bất thường trên ảnh MRI não với mô hình Swin-Unet. Kết quả thực nghiệm trên tập dữ liệu thu thập từ Bệnh viện Trường Đại học Y Dược Cần Thơ cho thấy tính khả thi và hiệu quả của mô hình Billow AISA.
Để xác định vùng bất thường trên ảnh MRI sọ não, bác sĩ chẩn đoán hình ảnh cần khảo sát nhiều lát cắt từ bộ ảnh. Nghiên cứu này giúp tự động phát hiện vùng bất thường của não trên ảnh MRI. Các mô hình Unet, ResNet, Swin-Unet được huấn luyện trên bộ dữ liệu của Bệnh viện Trường Đại học Y Dược Cần Thơ kết hợp bộ dữ liệu LGG để phân đoạn ảnh có hoặc không có vùng bất thường. Sau đó mô hình sẽ đề xuất vùng bất thường thông qua đường biên được vẽ xung quanh. Kết quả thực nghiệm cho thấy, khi chia dữ liệu ngẫu nhiên theo ảnh, mô hình Swin-Unet đạt được độ chính xác cao nhất là 0,88, cùng với Recall, Precision và F1 Score lần lượt là 0,96, 0,71, và 0,82. Đối với việc xác định vị trí và hình dạng của vùng bất thường, Swin-Unet cũng thể hiện hiệu suất cao với mIoU đạt 0,89 và mDSC là 0,91. Khi chia dữ liệu theo bệnh nhân, mô hình Swin-Unet lại một lần nữa thể hiện hiệu suất tốt với độ chính xác (Accuracy) đạt 0,86, cùng với Recall là 0,88, Precision là 0,79, F1 Score là 0,83, còn đối với mIoU đạt 0,84 và mDSC đạt 0,89. Kết quả nghiên cứu cho thấy mô hình Swin-Unet có kết quả tốt trong bài toán phát hiện vùng bất thường trên ảnh MRI não.
Các mô hình phát hiện đối tượng dựa trên mạng nơ-ron tích chập đang phát triển liên tục và được áp dụng rộng rãi trong nhiều lĩnh vực, đặc biệt là trong hệ thống giao thông thông minh. Trong nghiên cứu này, các kỹ thuật học sâu đã được áp dụng, đặc biệt là các mô hình phát hiện phương tiện giao thông trong thời gian thực: dựa trên “anchor” (điển hình như mô hình You Only Look Once - YOLO), dựa trên “keypoint”(điển hình như mô hình CenterNet), và dựa trên “transformer”(điển hình như mô hình Detection Transformers - DETR). Các mô hình đã được tinh chỉnh và huấn luyện thông qua kỹ thuật học chuyển tiếp để cải thiện khả năng phát hiện phương tiện giao thông. Kết quả của các thử nghiệm đã chỉ ra rằng mô hình YOLO đạt được độ chính xác cao nhất (98,3%) với thời gian thực thi là 11,7 ms. Trong khi đó, mô hình DETR thực hiện thời gian thực thi nhanh nhất (2,3 ms), nhưng độ chính xác thấp nhất (62,4%). Mô hình CenterNet là lựa chọn tốt nhất (94,11% - 8 ms) vì cân đối được giữa độ chính xác và thời gian thực thi, có thể được sử dụng trong các ứng dụng thời gian thực.
Gene expression classification plays a crucial role in diagnosing diseases. In response to this critical challenge, the research community has developed a variety of methods. Among these, machine learning approaches, particularly those based on Support Vector Machine (SVM) algorithms, stand out for their effectiveness. However, these algorithms encounter major challenges due to the nature of gene expression datasets, which are characterized by extremely high dimensionality and a relatively small number of samples. This situation significantly challenges machine learning algorithms, as it increases the risk of overfitting and complicates the task of extracting meaningful patterns from a high-dimensional space with limited samples. To address these challenges, we propose an advanced ensemble framework based on SVM techniques. This framework begins with an extension of the Newton SVM, named NSVMX. Building on this foundation, we introduce an ensemble of NSVMX models, called E-NSVMX. We detail our methods through mathematical formulations and algorithmic procedures. Our comprehensive experiments across various gene expression datasets reveal that our proposed methods significantly outperform the LibSVM benchmark in terms of training speed. Moreover, they deliver competitive, and in certain instances, superior classification accuracy. These results make our methods particularly useful for applications that necessitate quick model updates or fast model retraining with new or augmented data. Beyond advancing theoretical knowledge, our research underscores the practical benefits, leading to more efficient and effective machine learning solutions for urgent real-world challenges.
Lưu lượng giao thông là một lĩnh vực quan trọng trong sự phát triển của kinh tế, xã hội và môi trường. Để đánh giá lưu lượng giao thông, việc ước lượng tốc độ luồng giao thông là quan trọng. Trong bài nghiên cứu này, chúng tôi đề xuất một mô hình ước lượng tốc độ luồng giao thông dựa trên dữ liệu thu thập từ các camera giám sát giao thông. Mục tiêu chính là đếm và theo dõi các phương tiện để ước lượng lưu lượng giao thông bằng cách kết hợp mô hình Yolov8 và ByteTrack, sau đó tính toán tốc độ trung bình của các phương tiện. Để huấn luyện và đánh giá hiệu suất của mô hình, dữ liệu thu thập từ Công an phường Vĩnh Thanh Vân – Thành phố Rạch Giá, bao gồm 10 092 ảnh và hơn 96 024 đối tượng được gán nhãn trong nhiều điều kiện khác nhau được sử dụng. Trong nghiên cứu đã thử nghiệm và so sánh hiệu suất của mô hình của mình với các mô hình kết hợp Yolov8 và DeepSort. Kết quả thực nghiệm cho thấy rằng mô hình đề xuất có thời gian thực thi thấp nhất và có khả năng ước lượng lưu lượng giao thông gần với thực tế, với độ chính xác là 91,39%. Bộ dữ liệu sử dụng trong nghiên cứu này có thể được nghiên cứu sử dụng như một tập kiểm thử đối với các bài toán tương tự.
Based on the images of dishes in the Mekong Delta along with questions about the dishes such as: What is the name of this dish? Where is it famous? What are the main ingredients? How is it made? An application chatbot will be built to promote the speciality dishes of the Mekong Delta. This report outlines a method for training a Visual Question Answering (VQA) model for classification tasks using Transformer-based models, such as ViT for image data, BERT/PhoBERT for text data, or ViLT for simultaneous processing of image and text data. After that, a Visual Encoder-Decoder model for the task of generating sentences will be built using the VQA model as a Visual Encoder and a GPT-2 as a Decoder. The experimental dataset, which includes 7,694 photos of dishes from the Mekong Delta, is a subset of the datasets 30VNFoods and VinaFood21. The accuracy metric was used to evaluate the VQA models, and the results were relatively good. For Model 1: ViT and BERT, the accuracy scores for English and Vietnamese are 94
Annie Morin合作论文数IFSIC( Institut de Formation en Informatique et Communication)7