
Top-view fisheye cameras are widely used in personnel surveillance for their broad field of view, but their unique imaging characteristics pose challenges like distortion, complex scenes,scale variations, and small objects near image edges. To tackle these, we proposed peripheral focus you only look once(PF-YOLO), an enhanced YOLOv8n-based method. Firstly, we introduced a cutting-patch data augmentation strategy to mitigate the problem of insufficient small-object samples in various scenes. Secondly, to enhance the model's focus on small objects near the edges, we designed the peripheral focus loss, which uses dynamic focus coefficients to provide greater gradient gains for these objects, improving their regression accuracy. Finally, we designed the three dimensional(3D) spatial-channel coordinate attention C2f module, enhancing spatial and channel perception, suppressing noise, and improving personnel detection. Experimental results demonstrate that PF-YOLO achieves strong performance on the challenging events for person detection from overhead fisheye images(CEPDTOF) and in-the-wild events for people detection and tracking from overhead fisheye cameras(WEPDTOF) datasets. Compared to the original YOLOv8n model, PFYOLO achieves improvements on CEPDTOF with increases of 2.1%, 1.7% and 2.9% in mean average precision 50(mAP 50), mAP 50-95, and tively. On WEPDTOF, PF-YOLO achieves substantial improvements with increases of 31.4%,14.9%, 61.1% and 21.0% in 91.2% and 57.2%, respectively.
With the rapid growth of connected devices, traditional edge-cloud systems are under overload pressure. Using mobile edge computing(MEC) to assist unmanned aerial vehicles(UAVs)as low altitude platform stations(LAPS) for communication and computation to build air-ground integrated networks(AGINs) offers a promising solution for seamless network coverage of remote internet of things(IoT) devices in the future. To address the performance demands of future mobile devices(MDs), we proposed an MEC-assisted AGIN system. The goal is to minimize the long-term computational overhead of MDs by jointly optimizing transmission power, flight trajectories, resource allocation, and offloading ratios, while utilizing non-orthogonal multiple access(NOMA) to improve device connectivity of large-scale MDs and spectral efficiency. We first designed an adaptive clustering scheme based on K-Means to cluster MDs and established communication links, improving efficiency and load balancing. Then, considering system dynamics, we introduced a partial computation offloading algorithm based on multi-agent deep deterministic policy gradient(MADDPG), modeling the multi-UAV computation offloading problem as a Markov decision process(MDP). This algorithm optimizes resource allocation through centralized training and distributed execution, reducing computational overhead. Simulation results show that the proposed algorithm not only converges stably but also outperforms other benchmark algorithms in handling complex scenarios with multiple devices.
Poly(m-phenylene isophthalamide)(PMIA), a key aromatic polyamide, is widely used for its outstanding mechanical strength, high thermal stability, and excellent insulation properties.However, different applications demand varying dielectric properties, so tailoring its dielectric performance is essential. PMIA was first synthesized in this study, followed by introducing pores and developing porous PMIA films and PMIA-based composites with reduced dielectric constants.Porous PMIA films were fabricated using the wet phase inversion process with N, N-dimethylacetamide(DMAC) solvent and water as the non-solvent. The impact of casting solution composition and coagulation bath temperature on pore structures was analyzed. A film produced with 18%PMIA and 5% LiCl in a 35 ℃ coagulation bath achieved the lowest dielectric constant of 1.76 at 1 Hz, 48% lower than the standard PMIA film, which had a tensile strength of 18.5 MPa and an initial degradation temperature of 320 ℃.
With the progression of photolithography processes, the present technology nodes have attained 3 nm and even 2 nm, necessitating a transition in the precision standards for displacement measurement and alignment methodologies from the nanometer scale to the sub-nanometer scale.Metasurfaces, owing to their superior light field manipulation capabilities, exhibit significant promise in the domains of displacement measurement and positioning, and are anticipated to be applied in the advanced alignment systems of lithography machines. This paper primarily provides an overview of the contemporary alignment and precise displacement measurement technologies employed in photolithography stages, alongside the operational principles of metasurfaces in the context of precise displacement measurement and alignment. Furthermore, it explores the evolution of metasurface systems capable of achieving nano/sub-nano precision, and identifies the critical issues associated with sub-nanometer measurements using metasurfaces, as well as the principal obstacles encountered in their implementation within photolithography stages. The objective is to provide initial guidance for the advancement of photolithography technology.
The integrated waveguide polarizer is essential for photonic integrated circuits, and various designs of waveguide polarizers have been developed. As the demand for dense photonic integration increases rapidly, new strategies to minimize the device size are needed. In this paper, we have inversely designed an integrated transverse electric pass(TE-pass) polarizer with a footprint of 2.88 μm× 2.88 μm, which is the smallest footprint ever achieved. A direct binary search algorithm is used to inversely design the device for maximizing the transverse electric(TE) transmission while minimizing transverse magnetic(TM) transmission. Finally, the inverse-designed device provides an average insertion loss of 0.99 dB and an average extinction ratio of 33 dB over a wavelength range of 100 nm.
Due to the limitations of existing imaging hardware, obtaining high-resolution hyperspectral images is challenging. Hyperspectral image super-resolution(HSI SR) has been a very attractive research topic in computer vision, attracting the attention of many researchers. However, most HSI SR methods focus on the tradeoff between spatial resolution and spectral information, and cannot guarantee the efficient extraction of image information. In this paper, a multidimensional features network(MFNet) for HSI SR is proposed, which simultaneously learns and fuses the spatial,spectral, and frequency multidimensional features of HSI. Spatial features contain rich local details,spectral features contain the information and correlation between spectral bands, and frequency feature can reflect the global information of the image and can be used to obtain the global context of HSI. The fusion of the three features can better guide image super-resolution, to obtain higher-quality high-resolution hyperspectral images. In MFNet, we use the frequency feature extraction module(FFEM) to extract the frequency feature. On this basis, a multidimensional features extraction module(MFEM) is designed to learn and fuse multidimensional features. In addition, experimental results on two public datasets demonstrate that MFNet achieves state-of-the-art performance.
Single-photon sensors are novel devices with extremely high single-photon sensitivity and temporal resolution. However, these advantages also make them highly susceptible to noise. Moreover, single-photon cameras face severe quantization as low as 1 bit/frame. These factors make it a daunting task to recover high-quality scene information from noisy single-photon data. Most current image reconstruction methods for single-photon data are mathematical approaches, which limits information utilization and algorithm performance. In this work, we propose a hybrid information enhancement model which can significantly enhance the efficiency of information utilization by leveraging attention mechanisms from both spatial and channel branches. Furthermore, we introduce a structural feature enhance module for the FFN of the transformer, which explicitly improves the model's ability to extract and enhance high-frequency structural information through two symmetric convolution branches. Additionally, we propose a single-photon data simulation pipeline based on RAW images to address the challenge of the lack of single-photon datasets. Experimental results show that the proposed method outperforms state-of-the-art methods in various noise levels and exhibits a more efficient capability for recovering high-frequency structures and extracting information.
Optical microscopes are essential tools for scientific research, but traditional microscopes are restricted to capturing only two-dimensional(2D) texture information, lacking comprehensive three-dimensional(3D) morphology capabilities. Additionally, traditional microscopes are inherently constrained by the limited space-bandwidth product of optical systems, resulting in restricted depth of field(DOF) and field of view(FOV). Attempts to expand DOF and FOV typically come at the cost of diminished resolution. In this paper, we propose a texture-driven FOV stitching algorithm specifically designed for extended depth-of-field(EDOF) microscopy, allowing for the integration of 2D texture and 3D depth data to achieve high-resolution, high-throughput multimodal imaging. Experimental results demonstrate an 11-fold enhancement in DOF and an 8-fold expansion in FOV compared to traditional microscopes, while maintaining axial resolution after FOV extension.
Among hyperspectral imaging technologies, interferometric spectral imaging is widely used in remote sening due to advantages of large luminous flux and high resolution. However, with complicated mechanism, interferometric imaging faces the impact of multi-stage degradation. Most exsiting interferometric spectrum reconstruction methods are based on tradition model-based framework with multiple steps, showing poor efficiency and restricted performance. Thus, we propose an interferometric spectrum reconstruction method based on degradation synthesis and deep learning.Firstly, based on imaging mechanism, we proposed an mathematical model of interferometric imaging to analyse the degradation components as noises and trends during imaging. The model consists of three stages, namely instrument degradation, sensing degradation, and signal-independent degradation process. Then, we designed calibration-based method to estimate parameters in the model, of which the results are used for synthesizing realistic dataset for learning-based algorithms.In addition, we proposed a dual-stage interferogram spectrum reconstruction framework, which supports pre-training and integration of denoising DNNs. Experiments exhibits the reliability of our degradation model and synthesized data, and the effectiveness of the proposed reconstruction method.
Panoramic images, offering a 360-degree view, are essential in virtual reality(VR) and augmented reality(AR), enhancing realism with high-quality textures. However, acquiring complete and high-quality panoramic textures is challenging. This paper introduces a method using generative adversarial networks(GANs) and the contrastive language-image pretraining(CLIP) model to restore and control texture in panoramic images. The GAN model captures complex structures and maintains consistency, while CLIP enables fine-grained texture control via semantic text-image associations. GAN inversion optimizes latent codes for precise texture details. The resulting low dynamic range(LDR) images are converted to high dynamic range(HDR) using the Blender engine for seamless texture blending. Experimental results demonstrate the effectiveness and flexibility of this method in panoramic texture restoration and generation.
In response to challenges posed by complex backgrounds, diverse target angles, and numerous small targets in remote sensing images, alongside the issue of high resource consumption hindering model deployment, we propose an enhanced, lightweight you only look once version 8small(YOLOv8s) detection algorithm. Regarding network improvements, we first replace traditional horizontal boxes with rotated boxes for target detection, effectively addressing difficulties in feature extraction caused by varying target angles. Second, we design a module integrating convolutional neural networks(CNN) and Transformer components to replace specific C2f modules in the backbone network, thereby expanding the model's receptive field and enhancing feature extraction in complex backgrounds. Finally, we introduce a feature calibration structure to mitigate potential feature mismatches during feature fusion. For model compression, we employ a lightweight channel pruning technique based on localized mean average precision(LMAP) to eliminate redundancies in the enhanced model. Although this approach results in some loss of detection accuracy, it effectively reduces the number of parameters, computational load, and model size. Additionally, we employ channel-level knowledge distillation to recover accuracy in the pruned model, further enhancing detection performance. Experimental results indicate that the enhanced algorithm achieves a 6.1% increase in mAP50 compared to YOLOv8s, while simultaneously reducing parameters, computational load, and model size by 57.7%, 28.8%, and 52.3%, respectively.
Second harmonic generation(SHG), a fundamental and widely-studied phenomenon in nonlinear optics, has attracted significant attention for its ability to convert fundamental frequencies into their second harmonics. While the dominant SHG research has been focused on the optical and infrared regimes, its investigation in the microwave range presents challenges due to the requirements of materials with higher nonlinear coefficients and high-power microwave sources.Here, we provide an overview of methods together with underlying mechanisms for SHG in microwave frequencies, and discuss prospects and insights into the future developments of SHGbased technologies. The discussions on both numerical analyses and experimental studies will offer guidance for further SHG research and communication advancements in microwave regime.
When tracking a unmanned aerial vehicle(UAV) in complex backgrounds, environmental noise and clutter often obscure it. Traditional radar target tracking algorithms face multiple limitations when tracking a UAV, including high vulnerability to target occlusion and shape variations,as well as pronounced false alarms and missed detections in low signal-to-noise ratio(SNR) environments. To address these issues, this paper proposes a UAV detection and tracking algorithm based on a low-frequency communication network. The accuracy and effectiveness of the algorithm are validated through simulation experiments using field-measured point cloud data. Additionally,the key parameters of the algorithm are optimized through a process of selection and comparison,thereby improving the algorithm's precision. The experimental results show that the improved algorithm can significantly enhance the detection and tracking performance of the UAV under high clutter density conditions, effectively reduce the false alarm rate and markedly improve overall tracking performance metrics.
Depth maps play a crucial role in various practical applications such as computer vision,augmented reality, and autonomous driving. How to obtain clear and accurate depth information in video depth estimation is a significant challenge faced in the field of computer vision. However,existing monocular video depth estimation models tend to produce blurred or inaccurate depth information in regions with object edges and low texture. To address this issue, we propose a monocular depth estimation model architecture guided by semantic segmentation masks, which introduces semantic information into the model to correct the ambiguous depth regions. We have evaluated the proposed method, and experimental results show that our method improves the accuracy of edge depth, demonstrating the effectiveness of our approach.
Topological insulators represent a new phase of matter, characterized by conductive surfaces, while their bulk remains insulating. When the dimension of the system exceeds that of the topological state by at least two, the insulators are classified as higher-order topological insulators(HOTI). The appearance of higher-order topological states, such as corner states, can be explained by the filling anomaly, which leads to the fractional spectral charges in the unit cell. Previously reported fractional charges have been quite limited in number and size. In this work, based on the two-dimensional(2D) Su-Schrieffer-Heeger model lattice, we demonstrated a new class of HOTIs with adjustable fractional charges that can take any value ranging from 0 to 1, achieved by utilizing the Lorentz transformation. Furthermore, this transformation generates novel bound-state-incontinuum-like corner states, even when the lattice is in a topological trivial phase, offering a new approach to light beam localization. This work paves the way for fabricating HOTIs with diverse corner states that offer promising applicative potential.
Passive bistatic radar(PBR) frequently experiences interference from direct signal waves when detecting maritime targets, which can completely mask target echoes, particularly for distant targets or weak targets with low radar cross-section(RCS). To mitigate this, the paper proposes a direct signal interference(DSI) suppression method. The approach involves dual-channel reception of digital video broadcast satellites(DVB-S) signals from the China Sat-9, followed by signal preprocessing. The reference and surveillance channel signals are then segmented. After segmentation,the signals undergo fast Fourier transformation(FFT), and an adaptive filtering clutter suppression method is applied at each frequency point. Finally, an inverse fast Fourier transform(IFFT) is performed on the suppressed signals to obtain the DSI-suppressed output. Compared to traditional clutter suppression techniques, this method is not only faster but also achieves more effective suppression. Simulation experiments involving both single and multiple targets validate the superiority of the proposed algorithm.
Video snapshot compressive imaging(Video SCI) modulates scenes using various encoding masks and captures compressed measurements with a low-speed camera during a single exposure. Subsequently, reconstruction algorithms restore image sequences of dynamic scenes, offering advantages such as reduced bandwidth and storage space requirements. The temporal correlation in video data is crucial for Video SCI, as it leverages the temporal relationships among frames to enhance the efficiency and quality of reconstruction algorithms, particularly for fast-moving objects.This paper discretizes video frames to create image datasets with the same data volume but differing temporal correlations. We utilized the state-of-the-art(SOTA) reconstruction framework, EfficientSCI++, to train various compressed reconstruction models with these differing temporal correlations. Evaluating the reconstruction results from these models, our simulation experiments confirm that a reduction in temporal correlation leads to decreased reconstruction accuracy. Additionally, we simulated the reconstruction outcomes of datasets devoid of temporal correlation, illustrating that models trained on non-temporal data affect the temporal feature extraction capabilities of transformers, resulting in negligible impacts on the evaluation of reconstruction results for non-temporal correlation test datasets.
Terahertz(THz)metamaterials,with their exceptional ability to precisely manipulate the phase,amplitude,polarization and orbital angular momentum(OAM)of electromagnetic waves,have demonstrated significant application potential across a wide range of fields.However,traditional design methodologies often rely on extensive parameter sweeps,making it challenging to address the increasingly complex and diverse application requirements.Recently,the integration of artificial intelligence(AI)techniques,particularly deep learning and optimization algorithms,has introduced new approaches for the design of THz metamaterials.This paper reviews the fundamen-tal principles of THz metamaterials and their intelligent design methodologies,with a particular focus on the advancements in AI-driven inverse design of THz metamaterials.The AI-driven inverse design process allows for the creation of THz metamaterials with desired properties by working backward from the unit structures and array configurations of THz metamaterials,thereby acceler-ating the design process and reducing both computational resources and time.It examines the criti-cal role of AI in improving both the functionality and design efficiency of THz metamaterials.Finally,we outline future research directions and technological challenges,with the goal of provid-ing valuable insights and guidance for ongoing and future investigations.
This paper introduces a lightweight remote sensing image dehazing network called multi-dimensional weight regulation network(MDWR-Net),which addresses the high computational cost of existing methods.Previous works,often based on the encoder-decoder structure and utilizing multiple upsampling and downsampling layers,are computationally expensive.To improve effi-ciency,the paper proposes two modules:the efficient spatial resolution recovery module(ESRR)for upsampling and the efficient depth information augmentation module(EDIA)for downsampling.These modules not only reduce model complexity but also enhance performance.Additionally,the partial feature weight learning module(PFWL)is introduced to reduce the computational burden by applying weight learning across partial dimensions,rather than using full-channel convolution.To overcome the limitations of convolutional neural networks(CNN)-based networks,the haze dis-tribution index transformer(HDIT)is integrated into the decoder.We also propose the physical-based non-adjacent feature fusion module(PNFF),which leverages the atmospheric scattering model to improve generalization of our MDWR-Net.The MDWR-Net achieves superior dehazing performance with a computational cost of just 2.98×109 multiply-accumulate operations(MACs),which is less than one-tenth of previous methods.Experimental results validate its effectiveness in balancing performance and computational efficiency.
A polarization converter with broadband polarization characteristics and capable of dynamic reconfiguration is proposed. By introducing out-of-plane degrees of freedom, dynamically tunable broadband and high-efficiency linear polarization conversion within the wavelength range of2 000 –2 800 nm is achieved. Research results indicate that when a two-dimensional(2D) split-ring resonator(SRR) is irradiated by a low-dose focused ion beam, it will deform upward and transform into a three-dimensional(3D) SRR, achieving a linear polarization conversion efficiency of over 90%. The 3D SRR can be driven by electrostatic force to return to the 2D SRR state, thereby realizing the dynamic reconfiguration of this polarization converter. By changing the applied voltage and adjusting the structural parameters, a tailored polarization converter that exhibits broadband performance and high polarization conversion efficiency is also achieved. The results may provide novel ideas and technical methodologies for various applications such as polarized optical imaging, emerging display technologies, polarized optical communication, and optical sensing.