
Infrared imagery is crucial for surveillance, object detection, and autonomous navigation applications. However, the inherent low resolution of IR sensors limits the extraction of fine-grained details necessary for downstream tasks. Previous super-resolution approaches have predominantly optimized performance for IR single-image super-resolution; however, existing methods overlook spatial error analysis and real-world validation. This article presents an Edge Enhanced Deep Super-Resolution Network specifically designed for IR images by integrating a three-component loss function; pixel-wise L1 loss, edge-preserving loss, and perceptual loss. EEDSR is specifically tailored to ensure color fidelity, sharp edge reconstruction, and perceptually realistic outputs, unlike prior methods. Experiments performed on benchmark datasets and real-world captured IR images demonstrate significant improvements over existing methods, achieving a PSNR of 30.47 and an SSIM of 0.8810. Furthermore, comprehensive error map analysis highlights superior structural preservation and reduced spatial artifacts. The proposed approach bridges the gap between high-fidelity IR reconstruction and practical applicability, paving the way for more vigorous IR-based vision systems. Suárez, Patricia L., Dario Carpio, and Angel D. Sappa. “Enhancement of Guided Thermal Image Super-Resolution Approaches.” Neurocomputing 573 (March 2024): 127197. https://doi.org/10.1016/j.neucom.2023.127197.
We propose SepScene, a proxy-guided framework implementing inherently separable 3D scenes with native per-component manipulation. Using semantically annotated proxy scenes, our approach separates scene elements based on movability attributes, i.e., distinguishing immovables such as walls and terrain from user-relocatable movables including furniture and props. Environment-anchored immovables undergo geometry-constrained depth prediction to ensure structural permanence, while user-relocatable movables are processed through occlusion-aware diffusion inpainting followed by single-view mesh reconstruction, preserving their inherent reconfigurability. Critical challenges including occlusion handling and semantic consistency are resolved through proxy-guided raster masks and ID-mapped text embeddings, with object details refined via ratio-adaptive hallucination. The SepScene system establishes substantial post-generation control by enabling direct transformation, replacement, or removal of scene components without full regeneration. Our method surpasses existing approaches by achieving superior scene-text alignment and layout plausibility while also offering enhanced perceptual quality and faster generation speed.
Accurate camera calibration is crucial for vision-based 3D reconstruction tasks, especially in outdoor or uncontrolled environments. This paper introduces a novel calibration approach for planar grid patterns, including both printed calibration chessboards and planar grids formed by natural tiles. The method employs Triplet Geometry Constraints (TGC) to enforce second-order geometric consistency among feature points explicitly. In contrast to classical regularization methods, TGC adaptively assigns weights to local feature correspondences, enhancing robustness against measurement noise and scenarios with limited calibration data. Our calibration pipeline incorporates a TGC-based homography initialization stage, followed by nonlinear optimization and bundle adjustment, yielding globally consistent estimates of both intrinsic and extrinsic camera parameters. We validate the proposed method through extensive experiments on synthetic data and real-world outdoor scenes with natural tile grids. The results demonstrate substantial improvements in calibration accuracy and stability, directly contributing to more reliable 3D reconstructions of architectural structures from natural outdoor environments. Overall, the method provides a practical calibration solution that enables rapid deployment in the field without dedicated calibration targets, while maintaining high precision and robustness.
Both video and motion capture have been used to capture and reconstruct human movements. Combining them can improve the accuracy of human pose estimation. Here, we move beyond using these signals to refine the motion of a single individual to contribute to real world issues relevant to human-computer interaction. We describe an information fusion approach, specifically a post-process software and algorithm based pipeline to synchronize different streams of data. We then show its potential application in three case studies to demonstrate how the video-mocap information fusion approach can be beneficial to human-centered interaction research.
VR exergames combine fitness benefits with engaging gameplay, and adaptive AI offers opportunities to dynamically personalize intensity and feedback. We present iRow, a generative AI–powered VR rowing exergame that adapts scenes and AI-generated music in real time based on users’ physiological signals. In a mixed-methods study (N = 14), we examined how participants understood and responded to the system’s adaptive logic and how this related to trust, engagement, and performance. Results showed that participants generally trusted and adjusted to the system, even with partial understanding. Understanding was weakly associated with engagement but unrelated to exercise performance. Interviews revealed diverse preferences for explanation necessity and style. These findings suggest that providing intuitive, embodied feedback aligned with adaptive AI-generated content in VR exergames is more desirable for players than offering detailed explanations for fully understanding the adaptation mechanism.