
Superdirective beamforming can achieve a narrower mainlobe width compared to conventional methods. However, its high sensitivity to errors and interferences limits its practical applications. This paper proposes an efficient implementation method for robust superdirective beamforming in linear arrays. First, a tridiagonal matrix is constructed based on the noise covariance matrix, and the optimal superdirective beam for the linear array is decomposed by exploiting the structural properties of Toeplitz matrices. The resulting eigenbeams exhibit different directivity and robustness. Subsequently, robust superdirective beampatterns of arbitrary orders are synthesized by selecting specific eigenbeams. Simulation results demonstrate that the proposed method achieves robust superdirectivity with a favorable trade-off between robustness and directivity.
This paper introduces an automated method for Japanese pitch-accent transcription that combines Whisper-derived hidden-state embeddings with neural network classifiers. Three architectures—multilayer perceptron (MLP), convolutional neural network (CNN), and bidirectional long short-term memory (BiLSTM)—are evaluated on corpora of Tokyo and Osaka dialects under identical preprocessing and modeling protocols. In within-dialect experiments, the MLP attains the best mora-level accuracy for Tokyo, whereas the CNN performs best for Osaka. A cross-dialect transfer condition is also tested in which all Tokyo data are used for training and all Osaka data for evaluation with an unchanged pipeline. Under this setting, the best configuration reaches 69.13% mora accuracy, 0.678 macro-F1, and 22.82% word-level accuracy, with the other configurations slightly lower. These findings indicate that cross-dialect transfer remains challenging without explicit adaptation, motivating future work on simple, reproducible strategies for target-dialect robustness.
Electroencephalography (EEG) signals contain rich person-specific information that can be used for biometric applications. This raises both opportunities and concerns in areas such as personalized health care and data privacy. This paper presents an approach for EEG based subject verification, where a transformer-based model learns to generate subject specific embeddings to later determine whether two EEG recordings belong to the same person or not. The proposed method generates subject-specific embeddings using pre-trained models and performs subject verification using a distance-based matching strategy with threshold tuning. To assess the models generalization, cross-session and cross-dataset experiments are conducted using two EEG datasets with different recording montages with one being made up of subjects with medical pathologies. Results indicate high performance even when trained and evaluated across different subject groups, and recording configurations, reaching accuracies above 80%. These findings show that identity-related information can be extracted from EEG signals independent of health status or recording protocols, while also highlighting the need for privacy preservation of EEG data to protect sensitive subject-specific information.
Intracytoplasmic Sperm Injection (ICSI) is a widely used assisted reproduction technique involving the precise injection of a single spermatozoon into a mature oocyte. One of the most critical and challenging steps in this process is the manual manipulation and positioning of the sperm inside the injection pipette, which is prone to human error, variability, and potential damage to biological material. This work proposes a novel automated system based on Artificial Intelligence (AI) techniques to optimize and control sperm movement and positioning inside the ICSI needle. The system employs real-time AI-driven control algorithms, including reinforcement learning strategies, to enable adaptive and precise manipulation of sperm cells. This approach reduces operator dependency, enhances reproducibility, and improves the accuracy of sperm aspiration and delivery. By integrating AI and automation, the system aims to minimize procedural variability, reduce the risk of cellular damage, and ultimately improve fertilization success rates. This development represents a significant advancement in the automation of micromanipulation tasks within assisted reproductive technologies, opening new avenues for improving clinical outcomes in infertility treatments.
This paper addresses the problem of centralized fusion estimation for signals in the quaternion domain, based on observations from multiple sensors. It is assumed that the observations from each sensor may be affected by possible Denial-of-Service (DoS) attacks, which can block the transmission of sensor data. In such cases, to mitigate these threats, a compensation strategy is employed, where the missing observation is replaced by its predictor, based on the information available up to the previous time instant. Additionally, the correlation between the additive noise of the state and the observations at the same time instant is considered. To combine information from multiple sensors, a centralized fusion method is used. Under C-properness conditions, a recursive centralized fusion filtering algorithm is proposed to obtain the optimal estimator and corresponding error variance, reducing the problem dimension by half and thus significantly decreasing the computational load. Numerical simulation examples are presented to demonstrate the effectiveness and advantages of the proposed algorithm.
Instance-Level image retrieval (ILIR) requires fine-grained discrimination to identify all images depicting the same object or scene as a query, regardless of variations in viewpoint, scale, or background. While recent methods based on convolutional and transformer architectures have achieved high accuracy, most rely on supervised training and dataset-specific tuning, limiting their generalization to real-world, open-set scenarios. Pretrained Vision Transformers such as CLIP [1] offer a promising zero-shot alternative, but current approaches for image-to-image retrieval yet far from the supervised method’s performances for benchmark datasets, as well as they focus only to the ranking task, not offering a solution for the amount of the images to retrieve, which is an important point for real-world retrieval task to prevent false retrievals and obtaining unnecessary amount of retrieved data at the end.In this paper, we propose Chain Retrieve, a fully zero-shot retrieval framework that iteratively refines the query embedding by integrating information from top-k retrieved images. Unlike existing methods, our approach does not assume a fixed result count or rely on handcrafted thresholds. Instead, it expands the query representation in a retrieval loop and stops automatically when similarity scores fall below a fixed, interpretable threshold. Built entirely on pre-trained CLIP [1] ViTs and requiring no training or fine-tuning, Chain Retrieve achieves superior performance over existing zero-shot image-to-image methods on both custom and benchmark datasets, narrowing the gap with supervised models and offering a scalable, plug-and-play solution for ILIR.
Medical information is very important and delicate due to its relevance in clinical diagnosis, treatments, research and other commercial and non-commercial applications of governmental and private organizations. Recent advances in information and communication technologies have transformed the paradigm of medical data management, but also introduced new risks related to unauthorized access, manipulation, and distribution. To solve part of these problems, this paper proposes a method that employs two information security levels based on robust watermarking and binary data hiding, respectively. Firstly, clinical data is embedded into a logo using binary data hiding. Then, this logo is inserted into the spatial domain of a medical image through block segmentation, statistical information, and convolutional encoding for ownership authentication. Experimental results demonstrate logo robustness against JPEG DICOM compression in conjunction with aggressive attacks like cropping and copy-move, using bit error rate as a metric. The imperceptibility has been measured via peak signal-to-noise ratio and structural similarity index metrics.
We present Grid-Shift, a lightweight image pre-processing approach to counteract overfitting when training Convolutional Neural Networks. Grid-Shift solves the problem that tiling large images for training disrupts coherent features (i.e. an object may be split at the edge of a sub-image) and thus leads to information loss. Existing augmentation methods that reduce overfitting do not solve this problem explicitly. In our case study of Land Use and Land Cover Classification, Grid-Shift outperforms all other approaches tested (a raw UNet, a UNet with Batch Normalization, and various augmentation methods). Grid-Shift achieves a Categorical Accuracy of 95%, which is almost 20% better than a raw UNet and still 4% better than the best augmentation approach tested.
Indoor sound source localization enables innovative and cost-effective applications, such as needs-based automation of building functions. This paper presents a measurement system for 3D positional sound source localization in the near-field using a single, distributed microphone array with non-invasive geometry. Two localization techniques are evaluated, both relying on Time-Difference-of-Arrival (TDOA) between microphone pairs, which are estimated using the well-known beta-GCC-PHAT frequency weighting. The first algorithm is a closed-form analytical 3D Sound Source Localization (3D-SSL) solution, while the second employs a 3D grid-based Steered Response Power (SRP) method. Experimental results using both human speech and clapping signals in a highly reverberant environment show high localization accuracy, indicating strong potential for practical deployment. The 3D-SSL algorithm, when combined with stability and spatial filtering, achieves a localization error standard deviation of less than 30cm in both the x- and y-directions. The SRP algorithm exhibits lower localization accuracy, with errors below 70cm in the x- and y-directions, but offers greater robustness and does not require filtering. Both methods are particularly well-suited for 3D indoor applications that require high planar accuracy.
Novel view synthesis is a technique for generating free-view omni-directional images using multiple cameras. Traditionally, it has primarily been applied to synthesize novel views from inward-facing cameras placed around the center. This paper discusses whether novel view synthesis can also be achieved using outward-facing cameras placed around the center. In outward-facing camera configurations, the density of rays that can be captured is significantly lower compared to that in inward-facing cameras. As a result, achieving comparable reconstruction accuracy requires high-resolution cameras. The simplest approach to address this issue is to combine novel view synthesis with super-resolution techniques. In this paper, we present OF-NeRF (Outward-Facing cameras-based Neural Radiance Field), which is a novel method for synthesizing views from sparsely-placed outward-facing cameras. OF-NeRF integrates novel view synthesis and super-resolution techniques to achieve high-accuracy view generation.
With the growing demand for secure and efficient digital image processing, the integration of data hiding and encryption techniques has become essential to ensure both confidentiality and integrity of visual information. This paper proposes a hybrid approach that combines a low complexity steganographic scheme for binary and color images with a computationally efficient encryption technique inspired by quantum operators. In the data hiding component, an RGB channel of a color image is selected, and its least significant bit plane is extracted. This bit plane is then divided into suitable blocks, and HVS-based encoding tables are used to embed the information within these blocks, achieving high embedding capacity while preserving visual quality. In parallel, the encryption stage operates on color images and is based on a novel operator model that simulates quantum behaviors such as creation, exchange, and annihilation, enabling effective protection without requiring intensive mathematical processing. Experimental validation confirms the robustness of the combined approach through quality metrics such as PSNR, SSIM and entropy.
This paper aims to establish an evaluation index system for the innovation capacity of the high-end equipment manufacturing industry in Sichuan Province, providing a scientific basis for the evaluation and cultivation of innovation capacity for related enterprises. Through a comprehensive analysis of the characteristics of the high-end equipment manufacturing industry in Sichuan Province and drawing on existing research results, an initial evaluation index system was constructed, which includes four primary indicators, ten secondary indicators, and twenty-five tertiary indicators. After expert-interviewing and indicator screening, a simplified index system was finally formed, including three primary indicators, six secondary indicators, and nine tertiary indicators. Then ordering relation analysis method was applied to determine the weights of each indicator, among which the weight of innovation input capacity is the highest, indicating its crucial and fundamental role in innovation capacity. To provide high-quality development references for Sichuan’s high-end equipment manufacturing industry, this study also proposed countermeasures and suggestions such as increasing innovation input, elevating innovation management level, optimizing innovation output and value realization, promoting the construction of innovation cooperation mechanisms, and strengthening policy support and guidance.
Accurate parameter estimation from noisy, nonlinear time-series is a fundamental challenge in signal processing. Time-delay estimation, a critical step in techniques like time-delay embedding, is particularly susceptible to noise when using standard methods like mutual information (MI). This paper introduces a novel signal processing framework to overcome this limitation. We propose a multi-stage feature extraction front-end, termed GFMIEME, that transforms the raw signal into a more informative and noise-resilient feature space. This is achieved through a fine-grained multiscale analysis combined with the extraction of local statistical features (mean, standard deviation, root-mean-square). By computing MI in this enhanced feature domain, our method significantly boosts estimation accuracy. We benchmark our approach using chaotic signals, a canonical example of complex nonlinear data. Results show that our method successfully estimates the optimal time-delay at SNRs as low as 25 dB, a regime where traditional MI fails, demonstrating its superior robustness for practical signal analysis.
Hybrid accelerator systems have the potential to increase the useful lifespan of hardware designs. These systems allow parts of a hardware accelerator to be replaced with software running on a general purpose CPU. This allows the design to be extended or fixed in the field, after the fact, compared to traditional fixed function hardware accelerators.However, when implemented naively, such systems have significant overheads due to cache and interrupt management, limited memory bandwidth, and single threaded operation. This paper details three mechanisms that can reduce these overheads, and evaluates their impact using full-system simulation of an ARM-based embedded multicore architecture, extended with a cycle-accurate hardware accelerator. We show that efficiency can be increased by 5-20%, leading to decreased power and energy consumption, extending the potential applications of such systems.
This article provides an overview of the evolution of FPGA layout and routing algorithms, covering traditional methods (such as K-L algorithm), open-source frameworks ( Yosys+NextPNR), and 3D integration technologies. Research has shown that traditional K-L algorithms suffer from local convergence defects and computational efficiency bottlenecks; The open-source Yosys+NextPNR framework significantly reduces the design threshold and promotes transparency innovation; The 3D layered hybrid algorithm (partitioning+simulated annealing) outperforms traditional 2D solutions in terms of online length (down arrow 21% and latency (down arrow 24%), but requires addressing thermal management and vertical through-hole optimization issues. In the future, it is necessary to integrate reinforcement learning, multi-objective optimization, and heterogeneous acceleration technologies to unleash the potential of next-generation reconfigurable computing.
The Industrial Internet of Things (IIoT) has been extensively studied and applied across various industries. The integration of IoT, cloud computing, sensor technologies, and data processing has significantly advanced the digitization and automation of manufacturing processes. In automated manufacturing, the smooth operation of production along the line is essential for the overall efficiency and effectiveness of the system. This research proposes a smart factory monitoring system based on IIoT. We focus on the roles and functions of each component within the automated manufacturing environment, referred to as smart manufacturing. The proposed system integrates cloud services, smart sensors, data processing, and actuators/intelligent agents to work collaboratively, ensuring seamless operation of the automated production line. Simulation results demonstrate the effective and efficient design of the system and offer practical insights for implementing such a system across a wide range of industrial applications.
Electroencephalogram (EEG) based emotion recognition faces significant challenges in real-world deployment due to non-stationary data distributions across sessions and subjects, leading to catastrophic forgetting in static deep learning models. This study proposes a unified continual learning (CL) framework combining a temporal context encoder (LSTM-Attention) with two CL strategies: Fisher-Guided Synaptic Stability (EWC) for parameter regularization and Episodic Experience Integration (Replay) for memory retention. Evaluated on the SEED-IV dataset under session-based (SeCL) and subject-based (SuCL) incremental learning scenarios, our framework outperforms Sequential Adaptation (SA) and Naive Fine-tuning, with Replay excelling in sessions (92% accuracy, 0.02 forgetting) and EWC dominating subject adaptation (0.955 stability ratio, 0.045 forgetting). Both strategies significantly reduce forgetting (>50%) versus baselines, demonstrating robust cross- session/subject generalization for real-world EEG applications.
This article discusses the issue of constructing a new and straightforward multiple-input multiple- output (MIMO) system in which only one of several transmitting antennas is active (MIMO-1 system). A regular method for constructing such systems is described, based on the formation of so-called partner signals of different types by partitioning (from some base signal), but of the same size and distance characteristics for each antenna. For example, using the presented method, new MIMO-1 systems with spectral efficiency SE. {3, 4, 5, 6} bpcu ( bits per channel use) are constructed; their error rate characteristics, which were obtained by computer simulation, are presented in the form of simulation curves. The Gaussian channel is represented by Nakagami-m fading. Four-dimensional hybrid signals with frequency-phase modulation were used as base signals. Comparison of the obtained results with other known results showed the high efficiency of the new MIMO-1 system at different fading depths m. {0.5, 0.6, 0.75, 1, 1.3, 2.5} with different spectral efficiencies.
Recent advances in vision-language models (VLMs) like CLIP have revolutionized zero-shot image understanding by aligning visual and textual semantics. However, extending this success to zero-shot video action recognition remains challenging due to the inherent temporal complexity of videos. Critical actions are often diluted by redundant frames, while subtle interframe variations are overlooked by global pooling-based methods. In this paper, we propose Cross-Frame Semantic Alignment Network (CFSAN), a novel framework that jointly models local fine-grained action dynamics and global semantic consistency. We develop a Local Adjacent Frame Semantic Alignment Module to capture inter-frame semantic changes by aggregating multi-scale semantic differences between adjacent frames. In addition, the Global Semantic Aware Spectral Cluster Module is designed to divide frames into semantic clusters, and a gate mechanism is employed to filter noise, which retains only action key segments, to reduce the dilution effect of irrelevant frames on the global pool. We conduct a zero-shot video action recognition evaluation on HMDB51 and UCF101 datasets and experimental results show that the CFSAN model achieves competitive results.
The Intelligent Reflecting Surface (IRS) is an enabling technology for beyond 5G and 6G. IRSs are passive reflective surfaces that can modify signals in a controlled manner. They are utilized to enhance coverage and spectrum with high energy efficiency and, therefore, are referred to as passive relays. This energy efficiency makes them a promising technology. The reflective nature of these surfaces allows for a line-of-sight path through the IRS between a transmitter and receiver when the direct path is obstructed with low-energy consumption. However, this requires careful positioning. Proper IRS positioning is crucial for optimizing network performance. Most studies consider a fixed placement of the IRS, but the IRS can also be mobile to better adapt to the dynamics of certain networks. In this paper, we propose to compare the fixed and mobile placements of several IRSs. The problem is modeled using an IRS activation strategy for each period, employing different comparison criteria.