Contrastive learning and Siamese embedding models have become the foundation of modern verification systems, where decisions are governed not by discrete classification boundaries, but by relational geometry in embedding space. However, existing adversarial attacks remain fundamentally classification-centric, overlooking the vulnerability of relational geometry. In this paper, we introduce a geometry-aware adversarial attack framework that reformulates attacks on contrastive systems as manifold-level relational corruption. Instead of targeting individual predictions, the proposed framework systematically distorts similarity organization within the embedding manifold by pushing positive pairs apart while simultaneously pulling negative pairs closer, ultimately collapsing and inverting pairwise similarity structure. To enable scalable deployment, we shift iterative online optimization into an offline adversarial geometry deformation prior learning stage and train a lightweight feed-forward generator that learns generalized geometry deformation patterns from the victim model. Once trained, the generator produces adversarial perturbations through a single forward pass without requiring online gradient computation, enabling real-time online attacks against similarity-based verification systems. Experimental results across multiple verification architectures demonstrate substantial degradation of verification performance together with severe manifold-level relational corruption. On the Markmatch verification system, the proposed attack reduces accuracy from 95.4
Hallucinations in vision-language models (VLMs) hinder reliability and real-world applicability, usually stemming from distribution shifts between pretraining data and test samples. Existing solutions, such as retraining or fine-tuning on additional data, demand significant computational resources and labor-intensive data collection, while ensemble-based methods incur additional costs by introducing auxiliary VLMs. To address these challenges, we propose a novel test-time adaptation framework using reinforcement learning to mitigate hallucinations during inference without retraining or any auxiliary VLMs. By updating only the learnable parameters in the layer normalization of the language model (approximately 0.003 reduces distribution shifts between test samples and pretraining samples. A CLIP-based hallucination evaluation model is proposed to provide dual rewards to VLMs. Experimental results demonstrate a 15.4 hallucination rates on LLaVA and InstructBLIP, respectively. Our approach outperforms state-of-the-art baselines with a 68.3 hallucination mitigation, demonstrating its effectiveness.
Background:Approximately 70% of adults with chronic obstructive pulmonary disease (COPD) remain undiagnosed. Opportunistic screening using chest computed tomography (CT) scans, commonly acquired in clinical practice, may be used to improve COPD detection through simple, clinically applicable deep-learning models. We developed a lightweight, convolutional neural network (COPDxNet) that utilizes minimally processed chest CT scans to detect COPD. Methods:We analyzed 13,043 inspiratory chest CT scans from the COPDGene participants, (9,675 standard-dose and 3,368 low-dose scans), which we randomly split into training (70%) and test (30%) sets at the participant level to no individual contributed to both sets. COPD was defined by postbronchodilator FEV /FVC < 0.70. We constructed a simple, four-block convolutional model that was trained on pooled data and validated on the held-out standard- and low-dose test sets. External validation was performed using standard-dose CT scans from 2,890 SPIROMICS participants and low-dose CT scans from 7,893 participants in the National Lung Screening Trial (NLST). We evaluated performance using the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, Brier scores, and calibration curves. Findings:On COPDGene standard-dose CT scans, COPDxNet achieved an AUC of 0.92 (95% CI: 0.91 to 0.93), sensitivity of 80.2%, and specificity of 89.4%. On low-dose scans, AUC was 0.88 (95% CI: 0.86 to 0.90). When the COPDxNet model was applied to external validation datasets, it showed an AUC of 0.92 (95% CI: 0.91 to 0.93) in SPIROMICS and 0.82 (95% CI: 0.81 to 0.83) on NLST. The model was well-calibrated, with Brier scores of 0.11 for standard-dose and 0.13 for low-dose CT scans in COPDGene, 0.12 in SPIROMICS, and 0.17 in NLST. Interpretation:COPDxNet demonstrates high discriminative accuracy and generalizability for detecting COPD on standard- and low-dose chest CT scans, supporting its potential for clinical and screening applications across diverse populations.
Rationale: Emphysema progression is heterogeneous. Predicting temporal changes in lung density and detecting rapid progressors may facilitate the selection of individuals for targeted therapies. Objectives: To test whether computed tomography (CT) radiomics can be used to predict changes in lung density and detect rapid progressors. Methods: We extracted radiomics features from inspiratory chest CT in 4,575 subjects with and without airflow obstruction at enrollment, who completed a follow-up visit at approximately 5 years. We quantified emphysema using adjusted lung density (ALD) and estimated emphysema progression as the annualized change in ALD (ΔALD/yr) between visits. We categorized participants into rapid progressors (>1% ΔALD/yr) and stable disease (≤1% ΔALD/yr). A gradient boosting model was used 1) to predict ALD at 5 years and 2) to identify rapid progressors. Four models using demographics (base clinical model), CT density, radiomics, and combined features (clinical, radiomics, and CT density) were evaluated and tested. Results: There were 1,773 (38.8%) rapid progressors. For predicting ALD at 5 years in the 20% held-out data, the base model explained 31% of the variance (adjusted R2 = 0.31), whereas R2 was 0.74 for the CT density model, 0.66 for the radiomics-only model, and 0.77 for the combined-features model. For detecting rapid progressors, the base model (area under the receiver operating characteristic curve [AUC], 0.57 [95% confidence interval (CI), 0.53-0.61]) was outperformed by the radiomics-only model (AUC, 0.73 [95% CI, 0.69-0.76]; Δ = 0.15; P < 0.001) and the combined model (AUC, 0.74 [95% CI, 0.71-0.77]; Δ = 0.17; P < 0.001). Conclusions: Parenchymal and airway radiomics features derived from inspiratory scans can be used to predict temporal changes in lung density and help identify rapid progressors.
We present MarkMatch, a retrieval system for detecting whether two paper ballot marks were filled by the same hand. Unlike the previous SOTA method BubbleSig, which used binary classification on isolated mark pairs, MarkMatch ranks stylistic similarity between a query mark and a mark in the database using contrastive learning. Our model is trained with a dense batch similarity matrix and a dual loss objective. Each sample is contrasted against many negatives within each batch, enabling the model to learn subtle handwriting difference and improve generalization under handwriting variation and visual noise, while diagonal supervision reinforces high confidence on true matches. The model achieves an F1 score of 0.943, surpassing BubbleSig's best performance. MarkMatch also integrates Segment Anything Model for flexible mark extraction via box- or point-based prompts. The system offers election auditors a practical tool for visual, non-biometric investigation of suspicious ballots.
RATIONALE:Spirometry is only 50 % accurate for the detection of true ventilatory restriction, necessitating additional lung volume tests. OBJECTIVE:To develop a detection tool for true lung restriction using spirometry and patient demographics. METHODS:We analyzed spirometry and lung volume data from 21,062 participants. Restrictive spirometric pattern (RSP) was defined by FEV1/FVC ≥0.70 and FVC %predicted <80. Lung volumes were acquired using multi-breath nitrogen washout. True ventilatory restriction (TVR) was defined by total lung capacity <80 % predicted. We developed a LightGBM machine-learning model incorporating five spirometry (FEV1, FVC, FEV1/FVC, FEV1 % predicted and FVC % predicted), and three demographic (age, sex, and BMI) features. The model was trained on 80 % of the cohort (n = 16,849) and evaluated on 20 % (n = 4213) held-out set. The performance of the model was assessed using receiver operating characteristic (ROC) analyses. RESULTS:Of 21,062 participants, 12,643 (60 %) had TVR, of whom 5,255 (41.6 %) had RSP. The accuracy of RSP alone in detection of TVR was 0.61 (95% CI 0.60-0.63) with sensitivity of 0.42 (95% CI 0.40-0.43) and specificity of 0.91 (95% CI 0.90-0.92). The LightGBM model outperformed RSP alone, with an accuracy of 0.78 (95% CI 0.77-0.80), area under the ROC curve (AUC) of 0.89 (95% CI 0.88-0.90), sensitivity of 0.74 (95% CI 0.72-0.75), and specificity of 0.86 (95% CI 0.84-0.87). CONCLUSIONS:A machine learning model using demographics and spirometry can accurately detect true ventilatory restriction and lower the need for additional lung volume testing.
In response to the urgent need for rapid and precise post-disaster damage evaluation, this study introduces the Visual Prompt Damage Evaluation (ViPDE) framework, a novel contrastive learning-based approach that leverages the embedded knowledge within the Segment Anything Model (SAM) and pairs of remote sensing images to enhance building damage assessment. In this framework, we propose a learnable cascaded Visual Prompt Generator (VPG) that provides semantic visual prompts, guiding SAM to effectively analyze pre- and post-disaster image pairs and construct a nuanced representation of the affected areas at different stages. By keeping the foundation model’s parameters frozen, ViPDE significantly enhances training efficiency compared with traditional full-model fine-tuning methods. This parameter-efficient approach reduces computational costs and accelerates deployment in emergency scenarios. Moreover, our model demonstrates robustness across diverse disaster types and geographic locations. Beyond mere binary assessments, our model distinguishes damage levels with a finer granularity, categorizing them on a scale from 1 (no damage) to 4 (destroyed). Extensive experiments validate the effectiveness of ViPDE, showcasing its superior performance over existing methods. Comparative evaluations demonstrate that ViPDE achieves an F1 score of 0.7014. This foundation model-based approach sets a new benchmark in disaster management. It also pioneers a new practical architectural paradigm for foundation model-based contrastive learning focused on specific objects of interest.
Crack segmentation is vital for structural health monitoring, ensuring infrastructure maintenance such as roads, bridges, and buildings. In these scenarios, every improvement in accuracy and every second saved in processing time can make a significant difference in real-time monitoring, potentially preventing catastrophic failures and saving lives. However, cracks' irregularity, thinness, and low contrast make segmentation challenging, particularly as minority-class features. We present dCrack++, a deep learning architecture incorporating the proposed Edge-Guided Attention (EGA), specifically designed to enhance the detection of thin and irregular crack features across multiple scales, thus significantly improving segmentation accuracy while maintaining computational efficiency. To support model development, we curated dCrack61k, a comprehensive dataset comprising over 61 K images to address annotation inconsistencies, dataset overlap, and minority-class under-representation. Our experimental results demonstrate that dCrack++ consistently outperforms SOTA models in segmenting minority-class features across various surface types and conditions. Moreover, cross-domain evaluations on the CHASE_DB1 dataset for retinal vessel segmentation confirm dCrack++'s adaptability to similar tasks involving elongated structures, underscoring its robustness and generalizability to real-world applications.
Translation-based Video Synthesis (TVS) has emerged as a transformative technology that enables sophisticated manipulation and generation of dynamic visual content. This comprehensive survey systematically examines the evolution of TVS methodologies, encompassing both image-to-video (I2V) and video-to-video (V2V) translation approaches. We analyze the progression from domain-specific facial animation techniques to generalizable diffusion-based frameworks, investigating architectural innovations that address fundamental challenges in temporal consistency and cross-domain adaptation. Our investigation categorizes V2V methods into paired approaches, including conditional GAN-based frameworks and world-consistent synthesis, and unpaired approaches organized into five distinct paradigms: 3D GAN-based processing, temporal constraint mechanisms, optical flow integration, content-motion disentanglement learning, and extended image-to-image frameworks. Through comprehensive evaluation across diverse datasets, we analyze the performance using spatial quality metrics, temporal consistency measures, and semantic preservation indicators. We present a qualitative analysis comparing methods evaluated on identical benchmarks, revealing critical trade-offs between visual quality, temporal coherence, and computational efficiency. Current challenges persist in long-term temporal coherence, with future research directions identified in long-range video generation, audio-visual synthesis for enhanced realism, and development of comprehensive evaluation metrics that better capture human perceptual quality. This survey provides a structured understanding of methodological foundations, evaluation frameworks, and future research opportunities in TVS. We identify pathways for advancing cross-domain generalization, improving computational efficiency, and developing enhanced evaluation metrics for practical deployment, contributing to the broader understanding of temporal video synthesis technologies.
Rising incidences of paper check fraud, particularly with checks illicitly sold on platforms such as Telegram, pose significant challenges in financial security. Despite investigators' capability to gain access to these platforms, manually pinpointing checks in images and extracting necessary details to alert banks are inefficient and unscalable. Traditional optical character recognition-based (OCR) systems for extracting textual details from checks specifically struggle with handwritten content and are constrained by their dependency on predefined check layouts, limiting their effectiveness across varied and evolving check designs. To address these challenges, we introduce GenCheck, a generative AI-based framework that automates both the check detection and accurate extraction of check information, ensuring robust performance across various check layouts or styles. GenCheck operates through a two-stage pipeline: the preliminary stage encompasses multiple sub-tasks including check image classification, single check segmentation, image rectification, and check element detection, while the main stage focuses on the key task of check information extraction. Central to our pipeline is the strategic enhancement of a state-of-the-art (SOTA) multimodal large language model (LLaVA-NeXT) using Low-Rank Adaptation (LoRA). This fine-tuning leverages the model's pre-trained knowledge, applying a targeted, parameter-efficient approach that significantly enhances its ability to accurately extract key details such as dates, amounts, and payee information from paper check images. Our framework achieves exceptional accuracy rates in extracting date information with 92.07 % for year, 85.16% for month, and 82.72% for day. It also obtains an accuracy of 80.61 % in extracting monetary amounts and a normalized edit distance of 0.2583 for payee information, demonstrating substantial improvements over pure OCR-based methods. As the first framework of its kind, GenCheck estab-lishes a methodological base that supports continuous innovation and enhancement, allowing for independent updates of each component model. This also sets a new standard in automated check analysis, reducing the need for labor-intensive, rule-based processes and significantly advancing fraud prevention initiatives.
The prevalence of check fraud, particularly with stolen checks sold on platforms such as Telegram, creates significant challenges for both individuals and financial institutions. This underscores the urgent need for innovative solutions to detecting and preventing such fraud on social media platforms. While deep learning techniques show great promise in detecting objects and extracting information from images, their effectiveness in addressing check fraud is hindered by the lack of comprehensive, open-source, large training datasets specifically for check information extraction. To bridge this gap, this paper introduces "CheckGuard," a large labeled image-to-text cross-modal dataset designed for check information extraction. CheckGuard comprises over 7,000 real-world stolen check image segments from more than 15 financial institutions, featuring a variety of check styles and layouts. These segments have been manually labeled, resulting in over 50,000 samples across seven key elements: Drawer, Payee, Amount, Date, Drawee, Routing Number, and Check Number. This dataset supports various tasks such as visual question answering (VQA) on checks and check image captioning. Our paper details the rigorous data collecting, cleaning, and annotation processes that make CheckGuard a valuable resource for researchers in check fraud detection, machine learning, and multimodal large language models (MLLMs). We not only benchmark state-of-the-art (SOTA) methods on this dataset to assess their performance but also explore potential enhancements. Our application of parameter-efficient fine-tuning (PEFT) techniques on the SOTA MLLMs demonstrates significant performance improvements, providing valuable insights and practical approaches for enhancing model efficacy on this task. As an evolving project, CheckGuard will continue to be updated with new data, enhancing its utility and driving further advancements in the field. Our PEFT-based MLLM code is available at: https://github.com/feizhao19/CheckGuard. For data access, researchers are required to contact the authors directly.
The accurate and timely assessment of building damage is critical for effective post-disaster response efforts. However, traditional methods, reliant on manual inspection, are time-consuming and impractical in the face of large affected ar-eas. This work introduces a novel vision foundation model-based framework (SAM- RS) that leverages the knowledge embedded in pre-trained Segment Anything Model (SAM) with the remote sensing imagery to enhance building damage assessment performance. The proposed SAM-RS exploits pairs of high-resolution satellite images captured pre- and post-disaster, processing them through the SAM to boost the downstream damaged building segmentation task. Meanwhile, the parameter-efficient fine-tuning paradigm (PEFT) is adopted to accelerate the training process. We propose a multi-stage fusion adapter module (MSFA), in-jected at the end of each Transformer block to merge visual information from different temporal states of the observed area, thereby enhancing the model's ability to discern subtle differences indicative of damage levels. Comparative evaluations demonstrate that SAM-RS achieves a F1 score of 0.68 surpassing state-of-the-art (SOTA) models by 18.63 % on building damage assessment, marking this approach as a pioneering application of foundation models in the domain of disaster management technology. This advancement not only sets a new benchmark for rapid and precise damage evaluation but also significantly reduces the computational and storage demands typically associated with such tasks, paving the way for its adoption in operational disaster response workflows.
This article studies the stabilization problem for a category of nonlinear systems. It introduces a novel hybrid control strategy specifically for nonlinear high-order fully actuated (HOFA) systems. The stabilization problem for these nonlinear HOFA systems is transformed into a stabilization problem for impulsive switched systems. This issue is addressed by applying impulsive control in discrete-time and switching control based on the HOFA system approach in continuous-time. Using switched Lyapunov functions, we obtain the criteria for the exponential and asymptotic stabilization of these systems. The effectiveness of this newly presented hybrid control strategy is exemplified by simulations of a robotic manipulator system and a reduced model of the lac operon.
Translation-based Video Synthesis (TVS) has emerged as a vital research area in computer vision, aiming to facilitate the transformation of videos between distinct domains while preserving both temporal continuity and underlying content features. This technique has found wide-ranging applications, encompassing video super-resolution, colorization, segmentation, and more, by extending the capabilities of traditional image-to-image translation to the temporal domain. One of the principal challenges faced in TVS is the inherent risk of introducing flickering artifacts and inconsistencies between frames during the synthesis process. This is particularly challenging due to the necessity of ensuring smooth and coherent transitions between video frames. Efforts to tackle this challenge have induced the creation of diverse strategies and algorithms aimed at mitigating these unwanted consequences. This comprehensive review extensively examines the latest progress in the realm of TVS. It thoroughly investigates emerging methodologies, shedding light on the fundamental concepts and mechanisms utilized for proficient video synthesis. This survey also illuminates their inherent strengths, limitations, appropriate applications, and potential avenues for future development.
Abstract Current clinical tools to assess neonatal pain, including various pain scales such as Neonatal Infant Pain Scale (NIPS) and Neonatal Pain, Agitation, and Sedation Scale (N-PASS), are overly reliant on nurses’ subjective observation and analysis. Emerging deep learning approaches seek to fully automate this, but face chal- lenges including massive training data and computational resources, and potential public mistrust. Our study prioritizes facial information for pain detection, as facial muscles exhibit distinct patterns during pain events. This approach, using a single camera, avoids challenges associated with multimodal methods, such as data synchronization, larger training datasets, deployment issues, and high computational costs. We propose a deep learning-based neonatal pain detection framework that can alert a neonate pain management team when a pain event occurs, consisting of two main components: a transfer learning-based end-to-end pain detection neural network, and a manual assessment branch. The proposed neural network requires much less data to train and can evaluate whether a neonate is in a pain state based on facial information only. Additionally, the man- ual assessment branch can specifically handle the borderline/hard cases where the pain detection network is less confident. The integration of both machine detection and manual evaluation can increase the recall rate of true pain events, reduce the manual evaluation effort, and increase public trust in such applications. Experimental results show our neural network sur- passes state-of-the-art algorithms by at least 25% in accuracy on the MNPAD dataset, with overall framework accuracy reaching 82.35% with integration of manual assessment branch.
To enhance the integrity of mail-in voting, we introduce BubbleSig, an AI-assisted framework that detects samehand ballot stuffing by analyzing "signature-like" patterns on ballot marks. This method is voter-independent that relies solely on discrepancies in marking styles across different ballots, thus preserving voter anonymity and avoiding the use of biometric or historical voter data, e.g., fingerprints or signatures. Capable of handling diverse ballot formats and layouts with a single model, BubbleSig eliminates the need for retraining across election cycles. Its efficacy is demonstrated through real election data, achieving an F1 score of 0.925 on Same-Hand Ballot Stuffing Detection dataset, 100% accuracy in both mark and ballot level stuffing detection for a small set of real ballots known filled by the same person, and notable mean Average Precision (mAP) and Hit Rate (HR) scores in retrieving ballots suspected of stuffing, using expanded test sets containing both real and synthetic ballots. Our experimental results demonstrate the model's promise to handle diverse data collection processes, including variations in scanner types and scanning resolutions, and generalizability to various real-world ballot formats and layouts, underscoring its practical applicability. While the AI tool significantly aids in flagging potential ballot stuffing activities, final adjudications on ballot legitimacy remain with election officials, who examine suspicious ballots returned by the AI tool using the physical evidence on paper ballots.
Multimodal Artificial Intelligence (Multimodal AI), in general, involves various types of data (e.g., images, texts, or data collected from different sensors), feature engineering (e.g., extraction, combination/fusion), and decision-making (e.g., majority vote). As architectures become more and more sophisticated, multimodal neural networks can integrate feature extraction, feature fusion, and decision-making processes into one single model. The boundaries between those processes are increasingly blurred. The conventional multimodal data fusion taxonomy (e.g., early/late fusion), based on which the fusion occurs in, is no longer suitable for the modern deep learning era. Therefore, based on the main-stream techniques used, we propose a new fine-grained taxonomy grouping the state-of-the-art (SOTA) models into five classes: Encoder-Decoder methods, Attention Mechanism methods, Graph Neural Network methods, Generative Neural Network methods, and other Constraint-based methods. Most existing surveys on multimodal data fusion are only focused on one specific task with a combination of two specific modalities. Unlike those, this survey covers a broader combination of modalities, including Vision + Language (e.g., videos, texts), Vision + Sensors (e.g., images, LiDAR), and so on, and their corresponding tasks (e.g., video captioning, object detection). Moreover, a comparison among these methods is provided, as well as challenges and future directions in this area.
Traditional computer vision-based methods for estimating body fat percentage (BFP) rely on RGB images of the whole body, posing a risk of inadvertently leaking personal and private user information. Moreover, these methods often depend on hand-crafted features that are unreliable, highly customized, and computationally expensive. To address these challenges, this paper introduces a novel two-stage framework. In the first stage, we adopt a VGG19-UNet model to segment and extract only the body contours from RGB images, effectively removing irrelevant environmental content and ensuring that no re-identifiable details are disclosed. In the second stage, we implement spatial attention mechanisms to focus the model on crucial areas of body shapes. Additionally, the visual and demographic features are fused in an automated manner, enhancing the robustness of the model. Our experiments achieved a Root Mean Square Error (RMSE) of 4.13, representing a 29.5% enhancement from the state-of-the-art (SOTA) benchmark of 5.86.
Harmful Algal Blooms (HABs) present significant environmental and public health threats. Recent machine learning-based HABs monitoring methods often rely solely on unimodal data, e.g., satellite imagery, overlooking crucial environmental factors such as temperature. Moreover, existing multi-modal approaches grapple with real-time applicability and generalizability challenges due to the use of ensemble methodologies and hard-coded geolocation clusters. Addressing these gaps, this paper presents a novel deep learning model using a single-model-based multi-task framework. This framework is designed to segment water bodies and predict HABs severity levels concurrently, enabling the model to focus on areas of interest, thereby enhancing prediction accuracy. Our model integrates multimodal inputs, i.e., satellite imagery, elevation data, temperature readings, and geolocation details, via a dual-branch architecture: the Satellite-Elevation (SE) branch and the Temperature-Geolocation (TG) branch. Satellite and elevation data in the SE branch, being spatially coherent, assist in water area detection and feature extraction. Meanwhile, the TG branch, using sequential temperature data and geolocation information, captures temporal algal growth patterns and adjusts for temperature variations influenced by regional climatic differences, ensuring the model's adaptability across different geographic regions. Additionally, we propose a geometric multimodal focal loss to further enhance representation learning. On the Tick-Tick Bloom (TTB) dataset, our approach outperforms the SOTA methods by 15.65%.
John K. Johnstone合作论文数University of Alabama;Department of Computer and Information Sciences7
Rangasami L. Kashyap合作论文数School of Electrical Engineering, Purdue University, West Lafayette, IN 47907, USA3