Animal re-identification (ReID) faces critical challenges due to viewpoint variations, particularly in Aerial-Ground (AG-ReID) settings where models must match individuals across drastic elevation changes. However, existing datasets lack the precise angular annotations required to systematically analyze these geometric variations. To address this, we introduce the Multi-view Oriented Observation (MOO) dataset, a large-scale synthetic AG-ReID dataset of $1,000$ cattle individuals captured from $128$ uniformly sampled viewpoints ($128,000$ annotated images). Using this controlled dataset, we quantify the influence of elevation and identify a critical elevation threshold, above which models generalize significantly better to unseen views. Finally, we validate the transferability to real-world applications in both zero-shot and supervised settings, demonstrating performance gains across four real-world cattle datasets and confirming that synthetic geometric priors effectively bridge the domain gap. Collectively, this dataset and analysis lay the foundation for future model development in cross-view animal ReID. MOO is publicly available at https://github.com/TurtleSmoke/MOO.
This paper presents iMatcher, a fully differentiable framework for feature matching in point cloud registration. The proposed method leverages learned features to predict a geometrically consistent confidence matrix that incorporates both local and global consistency. First, a local graph embedding module initializes the score matrix. A subsequent repositioning step refines this matrix by considering bilateral source-to-target and target-to-source matching via nearest neighbor search in 3D space. The paired point features are then stacked and refined through global geometric consistency reasoning to predict a point-wise matching probability. Extensive experiments on real-world outdoor (KITTI, KITTI-360) and indoor (3DMatch) datasets, as well as on 6-DoF pose estimation (TUD-L) and partial-to-partial matching (MVP-RG), demonstrate that iMatcher significantly improves rigid registration performance. The method achieves state-of-the-art inlier ratios, scoring 95%-97% on KITTI, 94%-97% on KITTI-360, and up to 81.1% on 3DMatch, highlighting its robustness across diverse settings. The source code will be publicly available after publication.
Aerial-Ground Re-Identification (AG-ReID) is constrained by the viewpoint-domain gap, as drastic viewpoint disparities occlude or distort discriminative features, making cross-viewpoint image retrieval challenging. While existing methods rely on paired cross-view annotations, real-world deployments, such as wilderness search-and-rescue (SAR), often lack target-domain data, requiring retrieval from ground-level references alone. To our knowledge, we are the first to address this challenge by formalizing the Single-View AG-ReID (SV AG-ReID) setting, where models trained on a single real viewpoint must generalize to an unseen viewpoint. We propose 3D Lifting-based Elevated Novel-view Synthesis (3D-LENS), a unified framework combining geometrically-consistent novel view synthesis that leverages large-scale 3D mesh reconstruction, with a robust representation learning scheme to mitigate synthetic-to-real bias. Unlike 2D generative baselines that suffer from geometric inconsistencies or prior 3D methods that are restricted to class-specific templates, our approach ensures view-consistent synthesis across diverse categories without predefined templates that fail to capture fine-grained details, such as carried objects. Extensive experiments demonstrate that our method achieves state-of-the-art performance on SV AG-ReID scenarios. Code and data will be released at https://github.com/TurtleSmoke/3D-LENS.
Myocardial strain plays a crucial role in diagnosing heart failure and myocardial infarction. Its computation relies on assessing heart muscle motion throughout the cardiac cycle. This assessment can be performed by following key points on each frame of a cine Magnetic Resonance Imaging (MRI) sequence. The use of segmentation labels yields more accurate motion estimation near heart muscle boundaries. However, since few frames in a cardiac sequence usually have segmentation labels, most methods either rely on annotated pairs of frames/volumes, greatly reducing available data, or use all frames of the cardiac cycle without segmentation supervision. Moreover, these techniques rarely utilize more than two phases during training. In this work, a new semi-supervised motion estimation algorithm using all frames of the cardiac sequence is presented. The distance map generated from the end-diastolic segmentation label is used to weight loss functions. The method is tested on an in-house dataset containing 271 patients. Several deep learning image registration and tracking algorithms were retrained on our dataset and compared to our approach. The proposed approach achieves an average End Point Error (EPE) of 1.02mm, against 1.19mm for RAFT (Recurrent All-Pairs Field Transforms). Using the end-diastolic distance map further improves this metric to 0.95mm compared to 0.91 for the fully supervised version. Correlations in systolic peak were 0.83 and 0.90 for the left ventricular global radial and circumferential strain respectively, and 0.91 for the right ventricular circumferential strain.
This paper introduces a new hybrid descriptor for 3D point matching and point cloud registration, combining local geometrical properties and learning-based feature propagation for each point's neighborhood structure description. The proposed architecture first extracts prior geometrical information by computing each point's planarity, anisotropy, and omnivariance using a Principal Components Analysis (PCA). This prior information is completed by a descriptor based on the normal vectors estimated thanks to constructing a neighborhood based on triangles. The final geometrical descriptor is propagated between the points using local graph convolutions and attention mechanisms. The new feature extractor is evaluated on ModelNet40, Bunny Stanford dataset, KITTI, and MVP (Multi-View Partial)-RG for point cloud registration and shows interesting results, particularly on noisy and low overlapping point clouds.The code will be released after publication.
This paper introduces RoCNet++, a point cloud registration method with two main contributions, one concerning the design of a robust descriptor and another concerning the estimation of the rigid transformation. First, to robustly capture the local geometric properties of the surface, i.e., each point is characterized by all the triangles formed by itself and its nearest neighbours in the 3D point cloud. The idea is to assist the learning of the descriptor by introducing a priori information about interesting geometric properties such as the invariance of triangle angles under rigid transformations. This local triangle-based descriptor is integrated into the recently developed RoCNet architecture for estimating the correspondences between source and target point clouds. We then introduce the Farthest Sampling-guided Registration (FSR), which relies on successive farthest point samplings to estimate the global rigid transformation between 3D point clouds. The new proposed architecture RoCNet++ has been evaluated in different configurations: clean, noisy and partial data on both synthetic and real databases such as ModelNet40, KITTI, and 3DMatch. RoCNet++ shows improved performances on these benchmark datasets in favourable and unfavourable conditions. Furthermore, both the local triangle-based descriptor and the Farthest Sampling-guided Registration (FSR) can be used in other registration algorithms.
Context: Deep learning algorithms have been widely used for cardiac image segmentation. However, most of these architectures rely on convolutions that hardly model long-range dependencies, limiting their ability to extract contextual information. Moreover, the traditional U-net architecture suffers from the difference of semantic information between feature maps of the encoder and decoder (also known as the semantic gap). Material and method: To address this issue, a new network architecture relying on attention mechanism was introduced. Swin Filtering Blocks (SFB), that use Swin Transformer blocks in a cross-attention manner, were added between the encoder and the decoder to filter information coming from the encoder based on the feature map from the decoder. Attention was also employed at the lowest resolution in the form of a transformer layer to increase the receptive field of the network. We conducted experiments to assess both generalization capability and to evaluate how training on all frames of the cardiac cycle rather than only the end-diastole and end-systole impacts strain and segmentation performances. Results and conclusion: Visual inspection of feature maps suggested that Swin Filtering Blocks contribute to the reduction of the semantic gap. Performing attention between all patches using a transformer layer brought higher performance than convolutions. Training the model with all phases of the cardiac cycle resulted in slightly more accurate segmentations while leading to a more noticeable improvement for strain estimation. A limited decrease in performance was observed when testing on out-of-distribution data, but the gap widens for the most apical slices. (c) 2024 AGBM. Published by Elsevier Masson SAS. This is an open access article under the CC BY-NC-ND license (http://creativecommons .org /licenses /by-nc-nd/4.0/).
When we converse, we adapt our behaviors to our interlocutors. The adaptation can serve to indicate our engagement which can also elicit enhancement of the involvement of others. Virtual agents (or socially interactive virtual agents) that play the role of interaction partners can improve the human users’ interaction experience by displaying continuous and adaptive behaviors in real time. Virtual agents have been used in multiple domains to improve user interaction and performance. The promising results of the endowment of adaptation to agents in increasing the agents’ perception and user experience were shown in previous studies. In this paper, we develop an adaptive virtual agent that renders real-time adaptive behaviors based on the behaviors shown by its human interlocutor. The ASAP model rendering reciprocally adaptive agent behavior was employed to realize the system. The system consists of four main parts: perception of social signals, agent adaptive behavior generation, agent visualization (i.e. rendering of the agent’s verbal and nonverbal behavior), and communication of signals. To showcase the usefulness of our adaptive agent, as a proof-of-concept we choose the e-health application of cognitive behavior therapy (CBT), which identifies and rectifies biased and irrational thoughts (or automatic thoughts). Through this study, we show the importance of giving the agent reciprocal adaptation capability notably in enhancing the user experience and the effectiveness of the CBT session. We validate the importance of endowing such adaptation capability by studying the difference between agents that are reciprocally adaptive, solely expressive (with mismatched behavior), and inexpressive (in a still posture) via questionnaires and measures related to the agent perception (naturalness, human-likeliness, synchrony, and engagement) for user experience and the CBT effectiveness (mood, anxiety, stress, and cognitive change). These results highlight the value of making virtual agents adapt in real time. This could lead to agents being capable of providing more personalized and interactive experiences for a wide range of applications. Also, we have collected a new human-agent interaction (HAI) database, HAI-CBT database, which is publicly available to the research community.
Information about the motion of pixels between images is crucial for many computer vision tasks. When dealing with cardiac sequences, information about the heart’s motion can help physicians diagnose pathologies. Most methods that try to estimate this motion rely on pair of frames. This can lead to suboptimal performance when the amount of motion between them is important as it is the case when considering distant frames in a video sequence. Moreover, performing registration image by image leads to the integration of registration errors and is also suboptimal. In this work, a new registration method that uses all the frames in a video sequence is presented and applied to cardiac cine-MRI in short-axis views. A first neural network is used to compute motion flows between adjacent frames. Then, a second one processes the output of the first network to merge motion flows according to the time dimension throughout the sequence. Estimated flows are used to propagate segmentation masks across the sequence. The method is tested on an in-house dataset containing 271 patients. Segmentation, similarity and motion flow regularization metrics are computed to assess the model performance. The proposed approach achieves an average registration Dice score and SSIM between the end-diastole and end-systole frame of 95.26 ± 0.01 and 86 ± 0.05 respectively against 93.24 ± 0.02 and 82.75 ± 0.06 for the best Voxelmorph version.
This paper introduces a new method for 3D point cloud registration based on deep learning. The architecture is composed of three distinct blocs: (i) an encoder composed of a convolutional graph-based descriptor that encodes the immediate neighbourhood of each point and an attention mechanism that encodes the variations of the surface normals. Such descriptors are refined by highlighting attention between the points of the same set and then between the points of the two sets. (ii) a matching process that estimates a matrix of correspondences using the Sinkhorn algorithm. (iii) Finally, the rigid transformation between the two point clouds is calculated by RANSAC using the Kc best scores from the correspondence matrix. We conduct experiments on the ModelNet40 dataset, and our proposed architecture shows very promising results, outperforming state-of-the-art methods in most of the simulated configurations, including partial overlap and data augmentation with Gaussian noise.
When conversing, people adapt their behaviors to one another to show their engagement. Virtual agents, acting as interaction partners, should also adapt to their interlocutors in real time. In this paper, we introduce a virtual agent delivering Cognitive Behavioral Therapy (CBT) and adapting its behaviors in real time. The system focuses on the real-time generation of adaptive behavior and management of natural CBT dialogue.
Interruptions are an important aspect of human-human communication. They help to adjust the conversation flow. Our aim is to equip virtual agents with the ability to handle interruptions, that is to decide when and how to interrupt their human interlocutor. In this paper, we focus on predicting when interruptions may occur during the conversation using multimodal features only from the speaker and propose a model trained on a corpus of dyadic interactions. To assess the model's accuracy, we conduct a perceptual study where we compare different timings (ground truth, randomly chosen or predicted by our model).
Interlocutors adapt their verbal and nonverbal behaviors as signs of engagement during face-to-face interaction. We aim to build engaging Socially Interactive Agents, SIAs, that can adapt their behaviors during interaction. With an adaptive behavior generation model, we drive SIAs’ upper face and head movements in real-time. We evaluate this platform through a scenario for Social Skills Training, SST.
During a conversation, individuals take turns speaking and engage in exchanges, which can occur smoothly or involve interruptions. Listeners have various ways of participating, such as displaying backchannels, signalling the aim to take a turn, waiting for the speaker to yield the floor, or even interrupting and taking over the conversation. These exchanges are commonplace in natural interactions. To create realistic and engaging interactions between human participants and embodied conversational agents (ECAs), it is crucial to equip virtual agents with the ability to manage these exchanges. This includes being able to initiate or respond to signals from the human user. In order to achieve this, we annotate, analyze and characterize these exchanges in human-human conversations. In this paper, we present an analysis of multimodal features, with a focus on prosodic features such as pitch (F0) and loudness, as well as facial expressions, to describe different types of exchanges.
During an interaction, interlocutors emit multimodal social signals to communicate their intent by exchanging speaking turns smoothly or through interruptions, and adapting to their interacting partners which is referred to as interpersonal synchrony. We are interested in understanding whether the synchrony of multimodal signals could help to distinguish different types of turn-shifts. We consider three types of turn-shifts: smooth turn exchange, interruption and backchannel in this paper. We segmented each turn-shift into three phases: before, during and after, we calculated the synchrony measures of the three phases for multimodal signals (facial expression, head pose, and low-level acoustic features). In this paper, a brief analysis of synchronization during turn-shifts is presented, we also study the evolution of interpersonal synchrony before, during and after the turn-shifts. We proposed computational models for the turn-shift classification task only using synchrony measures. The best performance was obtained with an FNN model using the three phases’ synchrony score of all features (accuracy of 0.75).
During an interaction, people exchange speaking turns by coordinating with their partners.Exchanges can be done smoothly, with pauses between turns or through interruptions.Previous studies have analyzed various modalities to investigate turn shifts and their types (smooth turn exchange, overlap, and interruption).Modality analyses were also done to study the interpersonal synchronization which is observed throughout the whole interaction.Likewise, we intend to analyze different modalities to find a relationship between the different turn switch types and interpersonal synchrony.In this study, we provide an analysis of multimodal features, focusing on prosodic features (F0 and loudness), head activity, and facial action units, to characterize different switch types.
Socially Interactive Agents (SIAs) offer users with interactive face-to-face conversations. They can take the role of a speaker and communicate verbally and nonverbally their intentions and emotional states; but they should also act as active listener and be an interactive partner. In human-human interaction, interlocutors adapt their behaviors reciprocally and dynamically. The endowment of such adaptation capability can allow SIAs to show social and engaging behaviors. In this paper, we focus on modelizing the reciprocal adaptation to generate SIA behaviors for both conversational roles of speaker and listener. We propose the Augmented Self-Attention Pruning (ASAP) neural network model. ASAP incorporates recurrent neural network, attention mechanism of transformers, and pruning technique to learn the reciprocal adaptation via multimodal social signals. We evaluate our work objectively, via several metrics, and subjectively, through a user perception study where the SIA behaviors generated by ASAP is compared with those of other state-of-the-art models. Our results demonstrate that ASAP significantly outperforms the state-of-the-art models and thus shows the importance of reciprocal adaptation modeling.
: Recent works focus on creating socially interactive agents (SIAs) that are social, engaging, and human-like. SIA development is mainly on endowing the agent with human capacities such as communication and behavior adaptation skills. Nevertheless, the task of evaluating the agent’s quality remains as a challenge. Especially, the way of objectively evaluating human-agent interactions is not evident. To address this problem, we pro-pose new measures to evaluate the agent’s interaction quality. This paper focuses on interlocutors’ continuous, dynamic, and reciprocal behavior adaptation during an interaction, which we refer to as reciprocal adaptation. Our reciprocal adaptation measures capture this adaptation by measuring the synchrony of behaviors including their absence of response and by assessing the behavior entrainment loop. We investigate the nonverbal adaptation, notably for smile, in dyads. Statistical analyses are conducted to improve the understanding of the adaptation phenomenon. We also studied how the presence of reciprocal adaptation may be related to different aspects of the interaction dynamics and conversational engagement. We investigate how the influence of the social dimensions of warmth and competence along with the engagement is related to reciprocal adaptation.