Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, they often suffer from geometric distortions and limited camera controllability. To overcome these challenges, we introduce GeoNVS, a geometry-grounded novel-view synthesizer that enhances both geometric fidelity and camera controllability through explicit 3D geometric guidance. Our key innovation is the Gaussian Splat Feature Adapter (GS-Adapter), which lifts input-view diffusion features into 3D Gaussian representations, renders geometry-constrained novel-view features, and adaptively fuses them with diffusion features to correct geometrically inconsistent representations. Unlike prior methods that inject geometry at the input level, GS-Adapter operates in feature space, avoiding view-dependent color noise that degrades structural consistency. Its plug-and-play design enables zero-shot compatibility with diverse feed-forward geometry models without additional training, and can be adapted to other video diffusion backbones. Experiments across 9 scenes and 18 settings demonstrate state-of-the-art performance, achieving 11.3
We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an agent to plan, question, and eliminate candidates under information asymmetry. An ARGOS agent receives a vague witness statement and must decide what to ask, when to invoke spatial or temporal tools, and how to interpret ambiguous responses, all within a limited turn budget. Reasoning is grounded in a Spatio-Temporal Topology Graph (STTG) encoding camera connectivity and empirically validated transition times. The benchmark comprises 2,691 tasks across 14 real-world scenarios in three progressive tracks: semantic perception (Who), spatial reasoning (Where), and temporal reasoning (When). Experiments with four LLM backbones show the benchmark is far from solved (best TWS: 0.383 on Track 2, 0.590 on Track 3), and ablations confirm that removing domain-specific tools drops accuracy by up to 49.6 percentage points.
Multi-target multi-camera tracking is a crucial task that involves identifying and tracking individuals over time using video streams from multiple cameras. This task has practical applications in various fields, such as visual surveillance, crowd behavior analysis, and anomaly detection. However, due to the difficulty and cost of collecting and labeling data, existing datasets for this task are either synthetically generated or artificially constructed within a controlled camera network setting, which limits their ability to model real-world dynamics and generalize to diverse camera configurations. To address this issue, we present MTMMC, a real-world, large-scale dataset that includes long video sequences captured by 16 multi-modal cameras in two different environments - campus and factory - across various time, weather, and season conditions. This dataset provides a challenging test-bed for studying multi-camera tracking under diverse real-world complexities and includes an additional input modality of spatially aligned and temporally synchronized RGB and thermal cameras, which enhances the accuracy of multi-camera tracking. MTMMC is a super-set of existing datasets, benefiting independent fields such as person detection, re-identification, and multiple object tracking. We provide baselines and new learning setups on this dataset and set the reference scores for future studies. The datasets, models, and test server will be made publicly available.
Building identification refers to the recognition of the identity of a building when the building is sighted. With proper identification, additional useful information can be gained, such as the geographical information of the building and facility information within the building based on the recognized identity. However, existing studies related to building identification require additional information such as building images in advance, or have constraints on the types of possible input images. Therefore, we propose an approach that undertakes building identification using a building boundary map and smartphone sensors. Our approach measures the position of the user and the orientation of the user's view using sensors embedded in a smartphone. We find the building sighted by the user by reducing the area of buildings that can exist in the user's orientation in a step-by-step manner. During the validation of our approach, it identified the buildings with accuracy of up to 83.3% despite the inaccuracy of the building boundary map used and the error inherent in the user's position and orientation based on the smartphone.
Video streaming services strive to support high-quality videos at higher resolutions and frame rates to improve the quality of experience (QoE). However, high-quality videos consume considerable amounts of energy on mobile devices. This paper proposes NeuSaver, which reduces the power consumption of mobile devices when streaming videos by applying an adaptive frame rate to each video chunk without compromising the user experience. NeuSaver generates a policy that can determine the appropriate frame rate for each video chunk using reinforcement learning (RL). The RL model automatically learns the policy that optimizes the QoE goals based on previous observations. NeuSaver also uses an asynchronous advantage actor-critic algorithm to reinforce the RL model quickly and robustly. Streaming servers that support NeuSaver preprocess videos into segments with various frame rates, which is similar to the process of creating videos with multiple bit rates in dynamic adaptive streaming over HTTP. NeuSaver utilizes the commonly used H.264 video codec. We evaluated NeuSaver in various experiments and a user study through four video categories along with the previously proposed model. Our experiments showed that NeuSaver effectively reduces the power consumption of mobile devices when streaming video by an average of 16.14% and up to 23.12% while maintaining high QoE.
The goal of this work is to develop self-sufficient framework for Continuous Sign Language Recognition (CSLR) that addresses key issues of sign language recognition. These include the need for complex multi-scale features such as hands, face, and mouth for understanding, and absence of frame-level annotations. To this end, we propose (1) Divide and Focus Convolution (DFConv) which extracts both manual and non-manual features without the need for additional networks or annotations, and (2) Dense Pseudo-Label Refinement (DPLR) which propagates non-spiky frame-level pseudo-labels by combining the ground truth gloss sequence labels with the predicted sequence. We demonstrate that our model achieves state-of-the-art performance among RGB-based methods on large-scale CSLR benchmarks, PHOENIX-2014 and PHOENIX-2014-T, while showing comparable results with better efficiency when compared to other approaches that use multi-modality or extra annotations.
The existing methods of ransomware detection have limitations. To be specific, static analysis is not effective to obfuscated binaries, while dynamic analysis is usually restricted to a certain platform and often takes tens of minutes. In this paper, we propose a block-level monitoring system to detect potentially malicious cryptographic operations. We carry out statistical analysis to find heuristic rules to distinguish between normal and encrypted blocks. In order to apply the heuristic rule to the filesystem without kernel modification, we adopt Filesystem in Userspace (FUSE) and define our filesystem Rcryptect for real-time detection of cryptographic function. We demonstrate the protection of well-known ransomware and show that various cryptographic functions can be detected with about 13% overhead.
Selective video encryption schemes for battery-powered video devices have been proposed to prevent illegal acquisition of recorded video during transmission. However, the transmission security of conventional schemes is lower than the fully encrypted transmission schemes because a substantial part of the recorded video remains unencrypted including significant metadata for decoding recorded video. In order to address this limitation, we propose a secure video transmission framework that selectively encrypts video metadata and frames on a per packet basis with a consideration of currently available computing power and transmission distance. Evaluation on a testbed demonstrates that the proposed framework provides enhanced video transmission security with less than half of the battery power consumption required by the fully encrypted transmission scheme.
With the Internet of Things becoming mainstream, connectivity among cars has become mandatory. Although connectivity in a Vehicular Ad-hoc NETwork can improve users’ safety in transit, it creates an attack surface for cyber-crime such as impersonation and identity revealing attacks. Thus, anonymity and traceability are necessary when authenticating with other network users. A Fast ID Tracking (FIT) scheme with Static Random Access Memory Physically Unclonable Function (SRAM PUF)-based authentication, implemented as a System on Chip, is proposed. An anonymous ID is generated as a challenge and response pair of the SRAM PUF. Vehicle tracking is possible with simple eXclusive OR (XOR) with the values stored in a Road Side Unit and Trust Authority without additional cryptographic operation. The PUF generates unique IDs depending on internal devices with non-replicable features as a security primitive. The proposed FIT requires less than 1% of the tracking time of conventional schemes.
Due to the 4th industrial revolution and the strength of the 5 th Generation (5G) era, the Internet of Things (IoT) industry is growing significantly. As a result, the number of IoT devices in various industries, such as smart cars, smart homes, and smart healthcare, and the importance of security for these devices are increasing. This study proposes a design method for a secure cryptographic system on a chip (SecSoC) that can be used in the IoT industry and presents the results of the performance and security evaluations of the implemented chipset. The experimental results demonstrated that the SecSoC is a low-power high-performance cryptographic chip that is safe from external attacks. Compared to conventional smart card integrated circuits, the proposed design includes intrusion detection circuits that can respond to external attacks. At the same time, it supports a physical unclonable function for hiding secret data and cryptographic logic for maintaining integrity and confidentiality. The SecSoC ensured a fast transfer rate up to 110 Mbps and consumed only 95.8 mW when operating at maximum frequency.
Pursuing a more coherent scene understanding towards real-time vision applications, single-stage instance segmentation has recently gained popularity, achieving a simpler and more efficient design than its two-stage counterparts. Besides, its global mask representation often leads to superior accuracy to the two-stage Mask R-CNN which has been dominant thus far. Despite the promising advances in single-stage methods, finer delineation of instance boundaries still remains unexcavated. Indeed, boundary information provides a strong shape representation that can operate in synergy with the fully-convolutional mask features of the single-stage segmenter. In this work, we propose Boundary Basis based Instance Segmentation(B2Inst) to learn a global boundary representation that can complement existing global-mask-based methods that are often lacking high-frequency details. Besides, we devise a unified quality measure of both mask and boundary and introduce a network block that learns to score the per-instance predictions of itself. When applied to the strongest baselines in single-stage instance segmentation, our B2Inst leads to consistent improvements and accurately parse out the instance boundaries in a scene. Regardless of being single-stage or two-stage frameworks, we outperform the existing state-of-the-art methods on the COCO dataset with the same ResNet-50 and ResNet-101 backbones.
Fault attacks (FA) intentionally inject some fault into the encryption process for analyzing a secret key based on faulty intermediate values or faulty ciphertexts. One of the easy ways for software-based countermeasures is to use time redundancy. However, existing methods can be broken by skipping comparison operations or by using non-uniform distributions of faulty intermediate values. In this paper, we propose a secure software-based redundancy, aptly named table redundancy, applying different linear and nonlinear transformations to redundant computations of table-based block cipher structures. To reduce the table size and the number of lookups, some outer tables that are not subjected to FA are shared, while the inner tables are protected by table redundancy. The basic idea is that different transformations protecting redundant computations are correctly decoded if the redundant outcomes are combined without faulty values. In addition, this recombination provides infective computations because a faulty byte is likely to propagate its error to adjacent bytes due to the use of 32-bit linear transformations. Our method also presents a stateful feature in the connection with detected faults and subsequent plaintexts for preventing iterative fault injection. We demonstrate the protection of AES-128 against FA and show a negligible advantage of FA.
Advances in deep learning recognition have led to accurate object detection with 2D images. However, these 2D perception methods are insufficient for complete 3D world information. Concurrently, advanced 3D shape estimation approaches focus on the shape itself, without considering metric scale. These methods cannot determine the accurate location and orientation of objects. To tackle this problem, we propose a framework that jointly estimates a metric scale shape and pose from a single RGB image. Our framework has two branches: the Metric Scale Object Shape branch (MSOS) and the Normalized Object Coordinate Space branch (NOCS). The MSOS branch estimates the metric scale shape observed in the camera coordinates. The NOCS branch predicts the normalized object coordinate space (NOCS) map and performs similarity transformation with the rendered depth map from a predicted metric scale mesh to obtain 6D pose and size. Additionally, we introduce the Normalized Object Center Estimation (NOCE) to estimate the geometrically aligned distance from the camera to the object center. We validated our method on both synthetic and real-world datasets to evaluate category-level object pose and shape.
A linear transformation is applied to the white-box cryptographic implementation for the diffusion effect to prevent key-dependent intermediate values from being analyzed. However, it has been shown that there still exists a correlation before and after the linear transformation, and thus this is not enough to protect the key against statistical analysis. So far, the Hamming weight of rows in the invertible matrix has been considered the main cause of the key leakage from the linear transformation. In this study, we present an in-depth analysis of the distribution of intermediate values and the characteristics of block invertible binary matrices. Our mathematical analysis and experimental results show that the balanced distribution of the key-dependent intermediate value is the main cause of the key leakage.
Connectivity is one of the most challenging issues in Wireless Sensor Network (WSN). Connectivity problems in WSN seek to guarantee a satisfactory communication capability where all mobile sensors can connect to a base station via relay nodes in all data gathering events. In this paper, we focus on minimizing the number of relay nodes while ensuring connectivity in Mobile Wireless Sensor Networks. We propose an improved heuristic algorithm named Clustered Steiner Tree Heuristic (CSTH) to solve this problem in two phases. The first phase is Node Anchoring, which utilizes a greedy approach to find anchor points among clusters of mobile sensors. The second phase is called Steiner Relay Placement, in which a Steiner tree-based heuristic is used to minimize the number of relay nodes while maintaining connectivity in each cluster. Experiments were performed to compare CSTH with previous state-of-the-art heuristics for the problem. Results show that our algorithm can significantly improve the number of required relay nodes as well as computation time.
In general, Internet of Things (IoT) devices collect status information or operate according to control commands from other devices. If the safety and reliability of externally accessed devices are compromised, the risk of exposure of internally collected privacy information or abnormal operation of internal devices increases. This paper proposes a method of building a safe smart home environment by pre-blocking devices that may cause a risk by performing mutual safety verification between devices prior to data transmission and reception through the Session Initiation Protocol (SIP) of the home network. Using a Samsung’s commercial smartphone, not a development board to implement the device’s own verification function, and using an open source application and a SIP server providing free service, we established a test environment that is practically applicable and proved the feasibility of the attestation operation of the device. As a result of an operation test involving the capturing of packet data on a communication channel between two devices, it was confirmed that the transmission of parameter data for the actual attestation in SIP/Session Description Protocol packets succeeded without any problems. It was also confirmed that the final verification result of the target device was correctly derived. With the proposed method, it is possible to establish a safe trust relationship between smart home devices and external smart devices or between various IoT devices while also securing the smart home environment by blocking communications with devices that intentionally seek to do harm.
White-box cryptography is a software technique to protect secret keys of cryptographic algorithms from attackers who have access to memory. By adapting techniques of differential power analysis to computation traces consisting of runtime information, Differential Computation Analysis (DCA) has recovered the secret keys from white-box cryptographic implementations. In order to thwart DCA, a masked white-box implementation was suggested. It was a customized masking technique that randomizes all the values in the lookup tables with different masks. However, the round output was only permuted by byte encodings, not protected by masking. This is the main reason behind the success of DCA variants on the masked white-box implementation. In this paper, we improve the masked white-box cryptography in such a way to protect against DCA variants by obfuscating the round output with random masks. Specifically, we introduce a white-box AES (WB-AES) implementation applying the masking technique to the key-dependent intermediate value and the several outer-round outputs computed by partial bits of the key. Our analysis and experimental results show that the proposed WB-AES can protect against DCA variants including DCA with a 2-byte key guess, collision, and bucketing attacks. This work requires approximately 3.7 times the table size and 0.7 times the number of lookups compared to the previous masked WB-AES.
In this paper, a method to generate the downward image of current vehicle location using a commercial around view monitor (AVM) is proposed. The proposed system consists of three stages, namely feature tracking, obstacle filtering, and downward view generation. In the feature tracking stage, the Shi-Tomasi corner detection is used and feature tracking is performed with a Kanade-Lucas-Tomasi (KLT) tracker to determine the transformation between AVM images. In the obstacle filtering stage, features on obstacles are filtered by using the difference in a tracking distance between features detected on the ground and features detected on the obstacle. This is performed by using a histogram based on the tracking distance of features. Finally, in the downward view generation stage, transformation between the current AVM image and the previous AVM image is obtained by using the tracked feature pair refined via the aforementioned steps. Specifically, the random sample consensus (RANSAC) method is used to obtain transformations from which the influence of outliers is removed. The downward image of the current vehicle is generated by the obtained transformation and is synthesized to the existing AVM image. The results indicate that the proposed system synthesizes the existing AVM image and the generated downward image of a vehicle in a seamless way by determining the exact pair of matching points between current and the previous AVM images. We believe the proposed system can be utilized in efficient wireless charging system and safe driving system.
Linear transformations are often applied to the table-based cryptographic implementation including white-box cryptography in order to prevent key-dependent intermediate values from being analyzed. However, it has been shown that there still exists a correlation before and after the linear transformations, and thus this is not enough to protect the key against gray-box attacks such as power analysis. So far, the Hamming weight of rows in the invertible matrix has been considered the main cause of the key leakage from the linear transformation. In this study, we present an in-depth analysis of the distribution of intermediate values and the characteristics of block invertible binary matrices. Our mathematical analysis and experimental results show that the balanced distribution of the key-dependent intermediate value is the main cause of the key leakage.
Streaming services gradually support high-quality videos for the better user experience. However, streaming high-quality video on mobile devices consumes a considerable amount of energy. This paper presents the design and prototype of EVSO, which achieves power saving by applying adaptive frame rates to parts of videos with a little degradation of the user experience. EVSO utilizes a novel perceptual similarity measurement method based on human visual perception specialized for a video encoder. We also extend the media presentation description, in which the video content is selected based only on the network bandwidth, to allow for additional consideration of the user's battery status. EVSO's streaming server preprocesses the video into several processed videos according to the similarity intensity of each part of the video and then provides the client with the processed video suitable for the network bandwidth and the battery status of the client's mobile device. The EVSO system was implemented on the commonly used H.264/AVC encoder. We conduct various experiments and a user study with nine videos. Our experimental results show that EVSO effectively reduces the energy consumption when mobile devices uses streaming services by 22% on average and up to 27% while maintaining the quality of the user experience.
Sungwon Kang合作论文数Computer Science Department, KAIST14
Charles J. Alpert合作论文数IBM Austin Research Laboratory;IBM Research Division6
Samuel T. Chanson合作论文数Computer Science;Department of HKUST5
Shih-Hsu Huang合作论文数Dept. of Electronic Engineering
Chung-Yuan Christian University3
Rungbin Lin合作论文数Ta-Tung University3