Deep neural networks (DNN) are commonly used for electroencephalogram (EEG) decoding tasks. However, since these EEG data are collected using well-designed paradigms, there is growing concern that the reported decoding performance may be overestimated due to the inherent temporal autocorrelations (TAs) in EEG signals. The impact of TAs on various EEG decoding tasks has not been thoroughly explained, and the coupling between task-driven neural responses and TAs complicates understanding the limitations of using DNN. In this study, a unified framework is presented to formulate the impact of TAs on various EEG decoding tasks and decouple TAs features from task-driven features by shuffling EEG signals across EEG datasets collected for different purposes. The results show that similar performances were achieved with the shuffled datasets, highlighting the pervasive role of TAs in DNN-based decoding. Strategies are also suggested to mitigate the TAs effects, tailored to the specific type of EEG decoding tasks.
Decoding speech from neural recordings has critical importance in application and scientific research. However, this task is still challenging with non-invasive recordings. Previous research has shown significant improvement in speech perception decoding task by leveraging wav2vec vectors and gives the potential for applications. To further explore this problem, we proposed a novel multimodal method by using functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG). In our method, separate encoders for fMRI and MEG are considered, then features extracted from both modalities are integrated and aligned with wav2vec vectors that were extracted from the speech. The multimodal method reaches averaged performance of 72.6% in top-10 accuracy with a negative sample size of 128. Performance evaluated with various metrics achieves steady improvement across subjects, demonstrating the effectiveness of the proposed data fusion method. Interpretation of the performance increment was also investigated by testing the correlation between encoder hidden outputs and different level of features extracted from the speech. Results demonstrate that MEG encoder learns more low-level information and fMRI encoder learns more high-level information, which indicates both complementary characteristics lead to the improvement. The result of this work shows the potential of multimodal methods for speech decoding.
We achieved a microwave photonic RF receiver with pre-amplification on Er-doped lithium niobate platform for the first time. This scheme exhibits improved signal recovery quality compared to off-chip gain. © 2025 The Author(s)
We demonstrate the first self-amplified high-speed integrated photonic transmitter based on Erbium-doped lithium niobate platform. The proposed system integrates optical gain and electro-optic dynamics monolithically, supporting up to 170 GHz ultra-high bandwidth electro-optic operation.
With the rapid advancement of integrated photonic systems,rare-earth-doped waveguide amplifiers have emerged as a critical research focus due to their low-noise characteristics,extended luminescence lifetimes,and superior thermal stability.This article provides a comprehensive review of recent progress in rare-earth-doped waveguide amplifiers,exploring their fundamental photonic emission mechanisms,pivotal technological advancements,and system-level applications.We elucidate the impact of pumping schemes,material systems,and waveguide geometries on gain performance,establishing design principles for performance optimization.Furthermore,we demonstrate the versatility of rare-earth-doped waveguide amplifiers in enabling high-speed coherent communications,on-chip femtosecond pulse amplification,and high-energy Q-switched lasing,underscoring their transformative potential in broadband optical networks and high-power photonic-integrated systems.Future research will prioritize multi-ion co-doping strategies,dynamically tunable gain spectra,and high-efficiency on-chip pumping techniques to drive breakthroughs in integrated photonic devices and expand their application frontiers across next-generation communication and sensing platforms.
Auditory Attention Decoding (AAD) identifies a listener's focus in complex auditory scenes based on cortical neural responses. High decoding performance using DNN-based methods has been achieved with public EEG datasets. However, performance may be overestimated as models might learn temporal-autocorrelation features rather than auditory attention-related features. While data splitting risks have been discussed, experimental design risks have not. In this work, we collected a non-block design (NBD) scalp-EEG and ear-EEG joint dataset and compared it to previous block design (BD) datasets using DNN-based models. Results show a significant accuracy drop from BD to NBD dataset, while a linear stimulus reconstruction model remains robust. Inter-trial phase coherence analysis confirms stronger neural phase-locking to attended speech in BD dataset. These findings suggest BD enhances coherence of neural response but risks overestimating AAD accuracy. Code and data are released.
The brain is dynamic, associative and efficient. It reconfigures by associating the inputs with past experiences, with fused memory and processing. In contrast, AI models are static, unable to associate inputs with past experiences, and run on digital computers with physically separated memory and processing. We propose a hardware-software co-design, a semantic memory-based dynamic neural network (DNN) using memristor. The network associates incoming data with the past experience stored as semantic vectors. The network and the semantic memory are physically implemented on noise-robust ternary memristor-based Computing-In-Memory (CIM) and Content-Addressable Memory (CAM) circuits, respectively. We validate our co-designs, using a 40nm memristor macro, on ResNet and PointNet++ for classifying images and 3D points from the MNIST and ModelNet datasets, which not only achieves accuracy on par with software but also a 48.1 Moreover, it delivers a 77.6
Relating speech to EEG holds considerable importance but is challenging. In this study, a deep convolutional network was employed to extract spatiotemporal features from EEG data. Self-supervised speech representation and contextual text embedding were used as speech features. Contrastive learning was used to relate EEG features to speech features. The experimental results demonstrate the benefits of using self-supervised speech representation and contextual text embedding. Through feature fusion and model ensemble, an accuracy of 60.29% was achieved, and the performance was ranked as No.2 in Task 1 of the Auditory EEG Challenge (ICASSP 2024). The code to implement our work is available on Github: https://github.com/bobwangPKU/EEG-Stimulus-Match-Mismatch.
Auditory spatial attention detection (ASAD) aims to decode the attended spatial location with EEG in a multiple-speaker setting. ASAD methods are inspired by the brain lateralization of cortical neural responses during the processing of auditory spatial attention, and show promising performance for the task of auditory attention decoding (AAD) with neural recordings. In the previous ASAD methods, the spatial distribution of EEG electrodes is not fully exploited, which may limit the performance of these methods. In the present work, by transforming the original EEG channels into a two-dimensional (2D) spatial topological map, the EEG data is transformed into a three-dimensional (3D) arrangement containing spatial-temporal information. And then a 3D deep convolutional neural network (DenseNet-3D) is used to extract temporal and spatial features of the neural representation for the attended locations. The results show that the proposed method achieves higher decoding accuracy than the state-of-the-art (SOTA) method (94.3% compared to XANet's 90.6%) with 1-second decision window for the widely used KULeuven (KUL) dataset, and the code to implement our work is available on Github: https://github.com/xuxiran/ASAD_DenseNet
Decoding language from neural signals holds considerable theoretical and practical importance. Previous research has indicated the feasibility of decoding text or speech from invasive neural signals. However, when using non-invasive neural signals, significant challenges are encountered due to their low quality. In this study, we proposed a data-driven approach for decoding semantic of language from Magnetoencephalography (MEG) signals recorded while subjects were listening to continuous speech. First, a multi-subject decoding model was trained using contrastive learning to reconstruct continuous word embeddings from MEG data. Subsequently, a beam search algorithm was adopted to generate text sequences based on the reconstructed word embeddings. Given a candidate sentence in the beam, a language model was used to predict the subsequent words. The word embeddings of the subsequent words were correlated with the reconstructed word embedding. These correlations were then used as a measure of the probability for the next word. The results showed that the proposed continuous word embedding model can effectively leverage both subject-specific and subject-shared information. Additionally, the decoded text exhibited significant similarity to the target text, with an average BERTScore of 0.816.
The comprehensive management of light polarization states has significantly advanced various fields into a new era. With the advent of photonic integration, there has been a persistent desire to replace the bulky optical components with compact chip‐scale circuits. Nonetheless, the complete integration of polarization‐dependent systems has not yet been accomplished due to the absence of a mature polarization management scheme that possesses a tiny form factor and high foundry process compatibility meanwhile maintaining low operation complexity. Here, to overcome these limitations a novel concept called polarization phase mapping, which encodes the information between the light polarization in one waveguide and the relative light phase shift in another two waveguides, is proposed. With this bi‐directional mapping approach, the fundamental basis of polarization management has shifted from polarization adjustment to phase regulation. All essential polarization‐related functions including synthesizing, stabilizing, measuring, rotating, splitting, and mixing are demonstrated with the standard process in foundries. The size of the polarization rotating unit is pushed down to a few light wavelengths while keeping a competitive performance. Moreover, the proposed concept can be readily applied to other integrated photonics platforms. It is expected to unlock new opportunities for complex polarization‐related applications.
We demonstrate an erbium-doped lithium niobate on insulator waveguide amplifier which achieved the highest internal net gain of 38 dB with a 9.16 cm waveguide at 1531.7 nm.
To investigate the processing of speech in the brain, simple linear models are commonly used to establish a relationship between brain signals and speech features. However, these linear models are illequipped to model a highly dynamic and complex non-linear system like the brain. Although non-linear methods with neural networks have been developed recently, reconstructing unseen stimuli from unseen subjects’ EEG is still a highly challenging task. This work presents a novel method, ConvConcatNet, to reconstruct mel-spectrograms from EEG, in which the deep convolution neural network and extensive concatenation operation were combined. With our ConvConcatNet model, the Pearson correlation between the reconstructed and the target mel-spectrogram can achieve 0.0420, which was ranked as No.1 in the Task 2 of the Auditory EEG Challenge. The codes and models to implement our work will be available on Github: https://github.com/xuxiran/ConvConcatNet
In recent years, with the continuous development of silicon photonics technology, more and more optoelectronic devices have integrated into silicon-base platform. However, the inter-device transmission and coupling loss is serious, which results in an urgent need for on-chip waveguide amplifiers to compensate for the loss. Erbium silicate is an ideal gain material because of its extremely high Er 3+ concentration (10 22 cm −3 ). Nevertheless, Erbium silicate must be annealed above 1000 °C to activate the Er 3+ , which would damage other on-chip optoelectronic components. To solve this problem, here, we report a low-fabrication-temperature erbium-based waveguide amplifier. Erbium-ytterbium silicate and Bi 2 O 3 mixed film is used as gain material, which can activate the Er 3+ at 600 °C annealing condition. The waveguide amplifier was fabricated by lift-off process due to that the mixed film is hard to etch. Finally, 2 dB signal enhancement has been observed at 1550nm. This work shows that the proposed material has great potential for on-chip waveguide amplifier.
A magneto‐mechano‐electric (MME) generator that can harvest ambient magnetic noise plays a significant role in powering Internet of Things (IoT) sensor networks. However, it is still a challenge to capture sufficient energy and continuously drive IoT nodes from extremely low‐intensity magnetic noise below 1 Oe. To circumvent the close dependence of the resonant frequency on the magnetic proof mass in conventional MME generators, a new clamped‐clamped (C‐C) MME generator is proposed, that allows a much heavier magnetic mass to be attached at the beam center. Under weak magnetic fields of 0.48 and 0.96 Oe at 50 Hz, optimized output powers of 370 and 970 μWRMS, respectively are achieved, which shows an enhancement of ≈120% over that of cantilevered MME generators. The underlying mechanics are theoretically revealed by comparing the lumped parameters with a cantilevered MME generator and by calculating their deflection gain. Finally, it is demonstrated that the harvested energy from the proposed C‐C MME generator from a 0.48 Oe magnetic field at 50 Hz is sufficient to continuously drive an IoT sensor without any additional intervals for recharging. It is believed that this work will open new possibilities for designing MME generators suitable for weak field energy harvesting.
In recent years, silicon photonics has achieved great success in optical communication area. More and more on-chip optoelectronic devices have been realized and commercialized on silicon photonics platform, such as silicon-based modulators, filters and detectors. However, on-chip light sources are still not achieved because that silicon is an indirect bandgap material. To solve this problem, the rare earth element erbium (Er) is considered, which emits light covering 1.5 μm to 1.6 μm and has been widely used in fiber amplifiers. Compared to Er-doped fiber amplifiers (EDFA), the Er ion concentration needs to be more than two orders higher for on-chip Er-based light sources due to the compact size integration requirements. Therefore, the choice of the host material is crucially important. In this paper, we review the recent progress in on-chip Er-based light sources and the advantages and disadvantages of different host materials are compared and analyzed. Finally, the existing challenges and development directions of the on-chip Er-based light sources are discussed.
The increasing prevalence of integrated on-chip optoelectronic devices has identified serious issues regarding inter-device transmission and coupling losses, highlighting an urgent need for on-chip waveguide amplifiers to compensate for these losses. Compared with other Er-based optical materials, erbium silicate is ideally suited to high-efficiency on-chip amplifiers and lasers because of its extremely high Er3+ concentration (1022 cm−3). Nevertheless, erbium silicate must be annealed above 1000°C to crystallize and activate the Er3+, which damages other on-chip optoelectronic components and is not conducive to device integration. Here, we report a low-fabrication-temperature, high-luminescence-efficiency gain material by adding Bi2O3 to an erbium-ytterbium silicate mixed film. Our experiments demonstrate that the proposed film crystallizes at 600° C while the activation of Er3+ is also achieved, which is the lowest activation temperature of on-chip waveguide amplifier to our knowledge. This material forms the basis for a new chip-scale waveguide amplifier design, with a theoretical multi-energy-level model of Bi-Er-Yb in the mixed thin films used to analyze its signal enhancement properties. We achieve a peak on-chip gain of 23 dB in a 3.3-mm-long waveguide under the pump and signal powers of 300 mW and 1 µW, respectively. These results highlight the potential of the proposed material for realizing on-chip amplifiers and lasers for large-scale nanophotonic integrated circuits.
Integrated waveguides with slot structures have attracted increasing attention due to their advantages of tight mode confinement and strong light-matter interaction. Although extensively studied, the issue of mode mismatch with other strip waveguide-based optical devices is a huge challenge that prevents integrated waveguides from being widely utilized in large-scale photonic-based circuits. In this paper, we demonstrate an ultra-compact low-loss slot-strip converter with polarization insensitivity based on the multimode interference (MMI) effect. Sleek sinusoidal profiles are adopted to allow for smooth connection between the slot and strip waveguide, resulting reflection reduction. By manipulating the MMI effect with structure optimization, the self-imaging positions of the TE0 and TM0 modes are aligned with minimized footprint, leading to low-loss transmission for both polarizations. The measurement results show that high coupling efficiencies of − 0.40 and − 0.64 dB are achieved for TE0 and TM0 polarizations, respectively. The device has dimensions as small as 1.1 μm × 1.2 μm and composed of factory-available structures. The above characteristics of our proposed compact slot-strip converter makes it a promising device for future deployment in multi-functional integrated photonics systems.
In this paper, a spike time reduction circuit is proposed for improving the transient response of low-dropout (LDO) voltage regulator. Test results show that the circuit can significantly reduce the overshoot recovery time of output voltage after the sudden load current decrease. The main part of presented circuit includes an asymmetrical comparator and a discharging transistor. By setting a lower comparator input reference voltage value or increasing the size of the discharging transistor, the overshoot recovery time can be reduced from $240.5\mu \mathrm{s}$ to $126.1\mu \mathrm{s}$ during load current change from 1mA to 200mA. The prototype of the circuit is fabricated by 0.18um CMOS processes. Its consumption is only 102.4nA. The circuit has been verified on a commercial LDO production.