The topological semimetal FeSn antiferromagnet, characterized by its kagome lattice, two-dimensional flat bands, and Dirac-like surface states, holds immense promise for spintronic applications. In this work, for the first time, we investigate the spin pumping behavior in epitaxial-FeSn/Py (Ni_80Fe_20) heterostructures. We report a giant effective spin mixing conductance (g^↑↓_eff) of (116± 7) nm^-2, which is nearly one order of magnitude higher than that of standard Pt/Py heterostructures. The insertion of a 3 nm Al spacer layer results in a two-fold reduction in the effective damping, confirming the interfacial origin of the large g^↑↓_eff. Consistently, we observe an order-of-magnitude higher inverse spin Hall effect voltage in the FeSn/Py system compared to a reference Pt/Py film stack. We attribute the giant g^↑↓_eff to the direct interfacing of the Py layer with the topologically active [001]-kagome surface of epitaxial-FeSn. These findings establish the critical role of topologically active interfaces for advanced quantum-material-based spintronic devices.
Networks of coupled oscillators underpin fundamental studies of collective dynamics and emerging paradigms in physical computing. Spin Hall nano-oscillators (SHNOs) are particularly attractive due to their scalability and fast spin-wave-mediated interactions, yet mutual synchronization has so far been limited to small arrays and predominantly steady-state characterization. Here we demonstrate nanosecond phase ordering in lattices of up to N = 105,000 constriction-type SHNOs with widths of 10-20 nm. Microwave spectra reveal full mutual synchronization, a quality factor exceeding 106, power scaling as N and linewidth scaling as N-1. Time-resolved Brillouin light scattering shows a weak, approximately logarithmic increase in the synchronization time with array size. The synchronization time varies from 10 ns in arrays of 100 SHNOs to 45 ns for the largest arrays and is consistent with Kuramoto-type collective phase-ordering dynamics in a large two-dimensional oscillator lattice. These results establish spin-wave-mediated SHNO lattices as an experimentally accessible platform for exploring collective oscillator physics and for developing embedded-Ising and reservoir-computing architectures operating at tens of gigahertz.
Appearance-based gait recognition have achieved strong performance on controlled datasets, yet systematic evaluation of its robustness to real-world corruptions and silhouette variability remains lacking. We present RobustGait, a framework for fine-grained robustness evaluation of appearance-based gait recognition systems. RobustGait evaluation spans four dimensions: the type of perturbation (digital, environmental, temporal, occlusion), the silhouette extraction method (segmentation and parsing networks), the architectural capacities of gait recognition models, and various deployment scenarios. The benchmark introduces 15 corruption types at 5 severity levels across CASIA-B, CCPG, and SUSTech1K, with in-the-wild validation on MEVID, and evaluates six state-of-the-art gait systems. We came across several exciting insights. First, applying noise at the RGB level better reflects real-world degradation, and reveal how distortions propagate through silhouette extraction to the downstream gait recognition systems. Second, gait accuracy is highly sensitive to silhouette extractor biases, revealing an overlooked source of benchmark bias. Third, robustness is dependent on both the type of perturbation and the architectural design. Finally, we explore robustness-enhancing strategies, showing that noise-aware training and knowledge distillation improve performance and move toward deployment-ready systems.
Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual decision-making relatively underexplored. Addressing this gap, we introduce Corrective Sequence Planning (CoSPlan) benchmark, where VLMs must plan a sequence of visual actions from an initial scene to a target scene. CoSPlan evaluates models on their ability to imagine and execute a coherent set of visual steps required to reach the goal (Step Completion). To prevent any shortcuts that simply describe the final scene, we introduce an erroneous action in decision-making, which must be detected (Error Detection) and corrected to reach the goal, enabling a deeper understanding of the task. CoSPlan spans across 4 tasks: maze navigation, block re-arrangement, image reconstruction, and object re-organization. Despite using advanced reasoning strategies such as Chain-of-Thought and Scene Graphs, VLMs struggle on CoSPlan, while still showing promising performance in the text domain. Addressing this, we propose Scene Graph Incremental updates (SGI), a novel training-free method to transform images into `textual' scene graphs, enabling step-by-step reasoning through iterative scene graph refinement. SGI yields an average of 4.4
Constriction-based spin Hall nano-oscillators (SHNOs) show great promise for application as highly tunable microwave sources with straightforward scalability toward large coupled networks. However, details of the magnetization dynamics within SHNOs have thus far not been addressed experimentally, due to the minute time and length scales involved. In this work, we present direct imaging of the magnetization dynamics within a single CoFeB-based SHNO using time-resolved scanning transmission x-ray microscopy (TR-STXM). Our measurements reveal that the magnon amplitude is strongest at the two constriction edges, with a pronounced asymmetry favoring one edge, and that the emitted spin waves (SWs) exhibit strongly anisotropic propagation. Micromagnetic simulations suggest that grain boundaries and the Dzyaloshinskii-Moriya interaction (DMI) play a key role in both effects. Furthermore, the magnetodynamics changed during measurement, indicating that the CoFeB/MgO interface may be more susceptible to x-ray-induced modifications than previously recognized, challenging its presumed radiation hardness.
In this work, we focus on semi-supervised learning for video action detection. Video action detection requires spatio-temporal localization in addition to classification, and a limited amount of labels makes the model prone to unreliable predictions. We present Stable Mean Teacher, a simple end-to-end student-teacher-based framework that benefits from improved and temporally consistent pseudo labels. It relies on a novel ErrOr Recovery (EoR) module, which learns from students' mistakes on labeled samples and transfers this to the teacher to improve pseudo labels for unlabeled samples. Moreover, existing spatio-temporal losses do not take temporal coherency into account and are prone to temporal inconsistencies. To overcome this, we present Difference of Pixels (DoP), a simple and novel constraint focused on temporal consistency, which leads to coherent temporal detections. We evaluate our approach on four different spatio-temporal detection benchmarks: UCF101-24, JHMDB21, AVA, and Youtube-VOS. Our approach outperforms the supervised baselines for action detection by an average margin of 23.5% on UCF101-24, 16% on JHMDB21, and 3.3% on AVA. Using merely 10% and 20% of data, it provides a competitive performance compared to the supervised baseline trained on 100% annotations on UCF101-24 and JHMDB21 respectively. We further evaluate its effectiveness on AVA for scaling to large-scale datasets and Youtube-VOS for video object segmentation, demonstrating its generalization capability to other tasks in the video domain.
Spin Hall nano-oscillators (SHNOs) are emerging spintronic oscillators with significant potential for technological applications, including microwave signal generation, and unconventional computing. Despite their promising applications, SHNOs face various challenges, such as high energy consumption and difficulties in growing high-quality thin film heterostructures with clean interfaces. Here, single-layer topological magnetic Weyl semimetals open a possible solution as they possess both intrinsic ferromagnetism and a large spin-orbit coupling due to their topological properties. However, producing such high-quality thin films of magnetic Weyl semimetals that retain their topological properties and Berry curvature remains a challenge. We address these issues with high-quality single-layer epitaxial ferromagnetic Co2MnGa Weyl semimetal thin film-based SHNOs. We observe a giant spin Hall conductivity, σSHC = (6.08 ± 0.02) × 105 (ℏ/2e) Ω-1 m-1, which is an order of magnitude higher than previous reports. Theoretical calculations corroborate the experimental results with a large intrinsic spin Hall conductivity due to presence of a strong Berry curvature. Further, self spin-orbit torque driven magnetization auto-oscillations are demonstrated for the first time, at an ultralow threshold current density of Jth = 6.2 × 1011 A m-2. These findings indicate that magnetic Weyl semimetals have tremendous application potential for developing energy-efficient spintronic devices.
Self-supervised learning has emerged as a powerful paradigm for label-free model pretraining, particularly in the video domain, where manual annotation is costly and time-intensive. However, existing self-supervised approaches employ diverse experimental setups, making direct comparisons challenging due to the absence of a standardized benchmark. In this work, we establish a unified benchmark that enables fair comparisons across different methods. Additionally, we systematically investigate five critical aspects of self-supervised learning in videos: (1) dataset size, (2) model complexity, (3) data distribution, (4) data noise, and (5) feature representations. To facilitate this study, we evaluate six self-supervised learning methods across six network architectures, conducting extensive experiments on five benchmark datasets and assessing performance on two distinct downstream tasks. Our analysis reveals key insights into the interplay between pretraining strategies, dataset characteristics, pretext tasks, and model architectures. Furthermore, we extend these findings to Video Foundation Models (ViFMs), demonstrating their relevance in large-scale video representation learning. Finally, leveraging these insights, we propose a novel approach that significantly reduces training data requirements while surpassing state-of-the-art methods that rely on 10 pretraining data. We believe this work will guide future research toward a deeper understanding of self-supervised video representation learning and its broader implications.
In this work, we focus on Weakly Supervised Spatio-Temporal Video Grounding (WSTVG). It is a multimodal task aimed at localizing specific subjects spatio-temporally based on textual queries without bounding box supervision. Motivated by recent advancements in multi-modal foundation models for grounding tasks, we first explore the potential of state-of-the-art object detection models for WSTVG. Despite their robust zero-shot capabilities, our adaptation reveals significant limitations, including inconsistent temporal predictions, inadequate understanding of complex queries, and challenges in adapting to difficult scenarios.We propose CoSPaL (Contextual Self-Paced Learning), a novel approach which is designed to overcome these limitations. CoSPaL integrates three core components: (1) Tubelet Phrase Grounding (TPG), which introduces spatio-temporal prediction by linking textual queries to tubelets; (2) Contextual Referral Grounding (CRG), which improves comprehension of complex queries by extracting contextual information to refine object identification over time; and (3) Self-Paced Scene Understanding (SPS), a training paradigm that progressively increases task difficulty, enabling the model to adapt to complex scenarios by transitioning from coarse to fine-grained understanding.
Video action detection requires dense spatio-temporal annotations which are challenging as well as expensive to obtain. However, real-world videos often have varying level of difficulty and may not require equal level of annotations. In this paper we analyze the types of annotation appropriate for each sample and how it affects spatio-temporal video action detection. We focus on two different aspects affecting video action detection; 1) how to obtain varying level of annotations for videos, and 2) how to learn video action detection with different types of annotations. We study several annotation types including i) video level tags, ii) points iii) scribbles, iv) bounding box, and v) pixel level masks. First, we propose a simple active learning strategy which estimates appropriate types of annotations required for each video sample. Next, we propose a novel learning based spatio-temporal 3D-superpixel approach which generates pseudo-labels from different types of annotations and enables learning of video action detection from such annotations. We validate our approach on two different datasets, UCF101-24 and JHMDB-21, for video action detection, significantly reducing the annotation cost without significant drop in performance.
While mutually interacting spin Hall nano-oscillators (SHNOs) hold great promise for wireless communication, neural networks, neuromorphic computing, and Ising machines, the highest number of synchronized SHNOs remains limited to N = 64. Using ultra-narrow 10 and 20-nm nano-constrictions in W-Ta/CoFeB/MgO trilayers, we demonstrate mutually synchronized SHNO networks of up to N = 105,000. The microwave power and quality factor scale as N with new record values of 9 nW and 1.04 × 10^6, respectively. An unexpectedly strong array size dependence of the frequency-current tunability is explained by magnon exchange between nano-constrictions and magnon losses at the array edges, further corroborated by micromagnetic simulations and Brillouin light scattering microscopy. Our results represent a significant step towards viable SHNO network applications in wireless communication and unconventional computing.
Abstract Fe $$_3$$ 3 Sn $$_{2}$$ 2 is a topological kagome ferromagnet that possesses numerous Weyl points close to the Fermi energy, which can manifest various unique transport phenomena such as chiral anomaly, anomalous Hall effect, and giant magnetoresistance. However, the magnetodynamic properties of Fe $$_3$$ 3 Sn $$_{2}$$ 2 have not yet been explored. Here, we report, for the first time, the measurements of the intrinsic Gilbert damping constant ( $$\alpha _\text{int}$$ α int ), and the effective spin mixing conductance (g $$_\text{eff}^{\uparrow \downarrow }$$ eff ↑ ↓ ) of Pt/Fe $$_3$$ 3 Sn $$_2$$ 2 bilayers for Fe $$_3$$ 3 Sn $$_{2}$$ 2 thicknesses down to 2 nm, for which $$\alpha _\text{int}$$ α int is $$(3.8 \pm 0.2) \times 10^{-2}$$ ( 3.8 ± 0.2 ) × 10 - 2 , and g $$_\text{eff}^{\uparrow \downarrow }$$ eff ↑ ↓ is $$(11.7 \pm 0.6)~\text{nm}^{-2}$$ ( 11.7 ± 0.6 ) nm - 2 . The films have a high saturation magnetization, $$M_\text{S}=620~\mathrm{emu~cm^{-3}}$$ M S = 620 emu cm - 3 , and large anomalous Hall coefficient, $$R_\text{S}=4.6\times 10^{-10}~{\Omega ~\rm cm~G^{-1}}$$ R S = 4.6 × 10 - 10 Ω cm G - 1 . The large values of g $$_\text{eff}^{\uparrow \downarrow }$$ eff ↑ ↓ , together with the topological properties of Fe $$_3$$ 3 Sn $$_2$$ 2 , make Fe $$_3$$ 3 Sn $$_2$$ 2 /Pt bilayers useful heterostructures for the study of topological spintronic devices.
Ultra-fast spectrum analysis concept based on rapidly tuned spintronic nano-oscillators has been under development for the last few years and has already demonstrated promising results. Here, we demonstrate an ultra-fast microwave spectrum analyzer based on a chain of five mutually synchronized nano-constriction spin Hall nano-oscillators (SHNOs). As mutual synchronization affords the chain a much improved signal quality, with linewidths well below 1 MHz at close to a 10 GHz operating frequency, we observe an order of magnitude better frequency resolution bandwidth compared to previously reported spectral analysis based on single magnetic tunnel junction based spin torque nano-oscillators. The high-frequency operation and ability to synchronize long SHNO chains and large arrays make SHNOs ideal candidates for ultra-fast microwave spectral analysis.
This chapter reviews the state of the art in mutually synchronized spin-torque and spin Hall nano-oscillator (STNO and SHNO) arrays. After briefly introducing the underlying physics, we discuss different nano-oscillator implementations and their functional properties with respect to frequency range, output power, phase noise, and modulation rates. We then introduce the concepts and the theory of mutual synchronization and discuss the possible coupling mechanisms in spintronic nano-oscillators, such as dipolar, electrical, and spin-wave coupling. We review the experimental literature on mutually synchronized STNOs and SHNOs in one- and two-dimensional arrays and discuss ways to increase the number of mutually synchronized nano-oscillators. Finally, the potential for applications ranging from microwave signal sources/detectors and ultrafast spectrum analyzers to neuromorphic computing elements and Ising machines is discussed together with the specific electronic circuitry that has been designed so far to harness this potential.
Spin-torque nano-oscillators (STNOs) have emerged as an intriguing category of spintronic devices based on spin transfer torque to excite magnetic moment dynamics. The ultra-wide frequency tuning range, nanoscale size, and rich nonlinear dynamics have positioned STNOs at the forefront of advanced technologies, holding substantial promise in wireless communication, and neuromorphic computing. This review surveys recent advances in STNOs, including architectures, experimental methodologies, magnetodynamics, and device properties. Significantly, we focus on the exciting applications of STNOs, in fields ranging from signal processing to energy-efficient computing. Finally, we summarize the recent advancements and prospects for STNOs. This review aims to serve as a valuable resource for readers from diverse backgrounds, offering a concise yet comprehensive introduction to STNOs. It is designed to benefit newcomers seeking an entry point into the field and established members of the STNOs community, providing them with insightful perspectives on future developments.
Fe $$_3$$ Sn $$_{2}$$ is a topological kagome ferromagnet that possesses numerous Weyl points close to the Fermi energy, which can manifest various unique transport phenomena such as chiral anomaly, anomalous Hall effect, and giant magnetoresistance. However, the magnetodynamic properties of Fe $$_3$$ Sn $$_{2}$$ have not yet been explored. Here, we report, for the first time, the measurements of the intrinsic Gilbert damping constant ( $$\alpha _\text{int}$$ ), and the effective spin mixing conductance (g $$_\text{eff}^{\uparrow \downarrow }$$ ) of Pt/Fe $$_3$$ Sn $$_2$$ bilayers for Fe $$_3$$ Sn $$_{2}$$ thicknesses down to 2 nm, for which $$\alpha _\text{int}$$ is $$(3.8 \pm 0.2) \times 10^{-2}$$ , and g $$_\text{eff}^{\uparrow \downarrow }$$ is $$(11.7 \pm 0.6)~\text{nm}^{-2}$$ . The films have a high saturation magnetization, $$M_\text{S}=620~\mathrm{emu~cm^{-3}}$$ , and large anomalous Hall coefficient, $$R_\text{S}=4.6\times 10^{-10}~{\Omega ~\rm cm~G^{-1}}$$ . The large values of g $$_\text{eff}^{\uparrow \downarrow }$$ , together with the topological properties of Fe $$_3$$ Sn $$_2$$ , make Fe $$_3$$ Sn $$_2$$ /Pt bilayers useful heterostructures for the study of topological spintronic devices.
Self-supervised learning is an effective way for label-free model pre-training, especially in the video domain where labeling is expensive. Existing self-supervised works in the video domain use varying experimental setups to demonstrate their effectiveness and comparison across approaches becomes challenging with no standard benchmark. In this work, we first provide a benchmark that enables a comparison of existing approaches on the same ground. Next, we study five different aspects of self-supervised learning important for videos; 1) dataset size, 2) complexity, 3) data distribution, 4) data noise, and, 5)feature analysis. To facilitate this study, we focus on seven different methods along with seven different network architectures and perform an extensive set of experiments on 5 different datasets with an evaluation of two different downstream tasks. We present several interesting insights from this study which span across different properties of pretraining and target datasets, pretext-tasks, and model architectures among others. We further put some of these insights to the real test and propose an approach that requires a limited amount of training data and outperforms existing state-of-the-art approaches which use 10x pretraining data. We believe this work will pave the way for researchers to a better understanding of self-supervised pretext tasks in video representation learning.
Nano-constriction based spin Hall nano-oscillators (SHNOs) are at the forefront of spintronics research for emerging technological applications, such as oscillator-based neuromorphic computing and Ising Machines. However, their miniaturization to the sub-50 nm width regime results in poor scaling of the threshold current. Here, it shows that current shunting through the Si substrate is the origin of this problem and studies how different seed layers can mitigate it. It finds that an ultra-thin Al2 O3 seed layer and SiN (200 nm) coated p-Si substrates provide the best improvement, enabling us to scale down the SHNO width to a truly nanoscopic dimension of 10 nm, operating at threshold currents below 30 μ $\umu$ A. In addition, the combination of electrical insulation and high thermal conductivity of the Al2 O3 seed will offer the best conditions for large SHNO arrays, avoiding any significant temperature gradients within the array. The state-of-the-art ultra-low operational current SHNOs hence pave an energy-efficient route to scale oscillator-based computing to large dynamical neural networks of linear chains or 2D arrays.
In this work, we focus on semi-supervised learning for video action detection. We present Enhanced Mean Teacher, a simple end-to-end student-teacher based framework which rely on pseudo-labels to learn from unlabeled samples. Limited amount of data make the teacher prone to unreliable boundaries while detecting the spatio-temporal actions. We propose a novel auxiliary module, which learns from student’s mistakes on labeled samples and improve the spatio-temporal pseudo-labels generated by the teacher on unlabeled set. The proposed framework utilize spatial and temporal augmentations to generate pseudo-labels where both classification as well as spatio-temporal consistencies are used to train the model. We evaluate our approach on two action detection benchmark datasets, UCF101-24, and JHMDB-21. On UCF101-24, our approach outperforms the supervised baseline by an approximate margin of 19% on f-mAP@0.5 and 25% on v-mAP@0.5. Using merely 10-15% of the annotations in UCF-101-24, the proposed approach provides a competitive performance compared to the supervised baseline trained on 100% annotations. We also evaluate the effectiveness of Enhanced Mean Teacher for video object segmentation demonstrating its generalization capability to other tasks in video domain.