Spectrum, as the 'fingerprint' of materials, reveals their composition and properties. However, traditional spectral imaging systems are hindered by their bulkiness, high cost, time-consuming and high complexity, limiting their widespread use. By combining quantum dot (QD) spectral sensing with multi-spectral filter array designs, QD mosaic snapshot spectral imaging provides a real-time, compact, low-cost, and low-complexity alternative to traditional spectral imaging systems. Yet, the high flexibility of QD response curves presents a significant challenge: how to achieve an efficient, principle-guided, and data-adaptive design of high reconstruction performance? In this work, we propose a two-stage design method. Firstly, a set of absorption spectra with distinct particle sizes is selected using QR decomposition. Then, by leveraging the similarity between the imaging model and the convolutional operation, we optimize the concentration of QDs via gradient descent during the training stage of the reconstruction network, which serves as the software decoder to recover the spectral images from the encoded measurements. The proposed method is validated on the CAVE and Harvard datasets across multiple cases, achieving up to a 6.82% improvement in PSNR over the baseline, along with consistent gains in SSIM and SAM metrics. These results confirm the effectiveness of the proposed principle-guided method in achieving high reconstruction performance for spectral imaging systems.
Information security is crucial in today's era of big data, particulary in applications involving property, privacy, and military operations. While cryptography secures data content, steganography conceals its existence. Recently, optical steganography has gained attention for hiding data within optical elements. This method offers enhanced security, as the hidden text cannot be retrieved directly like a digital file but requires specific light-field environments or spectrometers for recovery. However, the concepts of cryptography and steganography are often confused in optical systems, and existing optical steganography methods rarely achieve large capacity, robustness, and security simultaneously. In this paper, we establish a clear boundary for cryptographic and steganographic optical systems based on the different functions performed by optical components. We then construct a quantum-dot film array-based steganographic system that simultaneously achieves high capacity, robustness, and security. The system is founded on a nonlinear mathematical model of quantum-dot fluorescence superposition and a spatial-spectral-temporal triple key to strengthen security. We perform theoretical analysis, simulations, and experiments to demonstrate its superior behavior. Finally, we propose a system solution for practical steganographic applications that can adapt to different security and capacity requirements. Moreover, the advancement of thin film technology in the future will represent a new step forward in the performance of the system.
It has been known for 100 years that light is both a wave and a particle, yet our approach to measuring light spectrum is still rooted only in the wave nature of light. Spectrometers built in this way suffer from the compromise between spectral features and device sizes. Recently, a particle-based spectrum detection method, coupled with advanced algorithms, has emerged. These new spectrometers offer high spectral resolution, wide spectral range, as well as compact size. This review outlines the principles, implementation, characteristics, performance, and extension of both wave and particle-based spectral detection technologies.
Colorimetric sensing methods are extensively utilized for rapid and sensitive detection of various biomedical and environmental targets, with higher dimensional spectra resulting in more accurate results. Miniature reconstructive spectrometers, as portable colorimetric sensing devices, show promise in capturing high-dimension spectra signals, while facing the challenges of noise-sensitive spectrum reconstruction and complex pre-calibration. To address these issues, we present a virtual barcode method, which is directly based on the utilization of a high-dimension quantum dot (QD) spectrometer intensity vector. Any spectral changes of the analytes can be reflected in the corresponding barcode, without the redundant operation for spectral analysis. We demonstrate the QD barcode method in quantitatively detecting multiple biomarkers in the artificial human urine, including urinary calcium, glucose, nitrite, and creatinine, with lower limits of detection compared to the RGB sensing method (2.4-14.4-fold). To simplify the preparation of the QD spectrometer, we optimize both the number and the spectral distribution of QD filters. Furthermore, an artificial neural network model is also established to achieve 8-fold improved quantitative recognition performance. This QD barcode method can greatly broadens the biosensing application of miniaturized reconstructive spectrometers.
Optical computing accelerators, with high parallelism, large bandwidth, and low transmission loss, have the potential to enhance electronic computing in both computational power and energy efficiency. Photonic acceleration plays a crucial role in supporting computationally intensive operations, such as dynamic-static matrix multiplication, significantly improving overall efficiency. Existing photonic architectures for dynamic-static matrix multiplication depend on complex coherent optical systems or costly nano-optics fabrication, limiting scalability. This study introduces a novel quantum dot fluorescence-based dynamic-static matrix multiplication photonic acceleration architecture that eliminates the need for coherent light sources or intricate fabrication. By leveraging simple, cost-effective quantum dot preparation and printing techniques, this architecture has significant potential for large-scale, high-performance, low-cost photonic accelerators. We detail the mathematical and physical mechanisms of the proposed architecture, experimentally validate the key physical processes, and demonstrate its application in template matching for image recognition, achieving 95% accuracy.
Liver cancer is a serious threat to people all over the world. Surgical resection is the preferred treatment for long-term survival. The residual malignant tissue will cause frequent recurrence of the disease and lead to high mortality. Therefore, the accurate determination of tumor resection margin is very important for surgical treatment. On the one hand, the ultimate goal of surgical resection is to remove all liver cancer tissue as much as possible to avoid residual; On the other hand, it is necessary to preserve as much healthy tissue as possible and maintain the basic functions of the organs. Hyperspectral imaging technology can quickly and accurately capture the spatial and spectral information of liver tumor tissue, and realize the accurate division of liver tumor incisal margin during surgery. The two goals mentioned above essentially correspond to the sensitivity and specificity in performance evaluation. In this work, we used the idea of contrastive learning to improve the classical U-net backbone network and optimize the overall performance of the liver tumor classification network. Specifically, a dataset containing 36 specimens was collected from 19 patients and pathological results were used as ground truth. The mean value of the improved network classification results increased and the standard deviation decreased in terms of accuracy, sensitivity and specificity, especially the sensitivity and specificity were significantly improved, with the overall sensitivity reaching 96.61% and the specificity reaching 95.00%, which met the original intention of surgical resection, and was of great significance for the practical application.
Dense Retrieval (DR) is now considered as a promising tool to enhance the memorization capacity of Large Language Models (LLM) such as GPT3 and GPT-4 by incorporating external memories. However, due to the paradigm discrepancy between text generation of LLM and DR, it is still an open challenge to integrate the retrieval and generation tasks in a shared LLM. In this paper, we propose an efficient LLM-Oriented Retrieval Tuner, namely LMORT, which decouples DR capacity from base LLM and non-invasively coordinates the optimally aligned and uniform layers of the LLM towards a unified DR space, achieving an efficient and effective DR without tuning the LLM itself. The extensive experiments on six BEIR datasets show that our approach could achieve competitive zero-shot retrieval performance compared to a range of strong DR models while maintaining the generation ability of LLM.
Dense Retrieval (DR) is now considered as a promising tool to enhance the memorization capacity of Large Language Models (LLM) such as GPT3 and GPT-4 by incorporating external memories. However, due to the paradigm discrepancy between text generation of LLM and DR, it is still an open challenge to integrate the retrieval and generation tasks in a shared LLM. In this paper, we propose an efficient LLM-Oriented Retrieval Tuner, namely LMORT, which decouples DR capacity from base LLM and non-invasively coordinates the optimally aligned and uniform layers of the LLM towards a unified DR space, achieving an efficient and effective DR without tuning the LLM itself. The extensive experiments on six BEIR datasets show that our approach could achieve competitive zero-shot retrieval performance compared to a range of strong DR models while maintaining the generation ability of LLM.
In 2023, the Nobel Prize in Chemistry was awarded to Bawendi, Brus, and Ekimov, three scientists who have made great contributions to the discovery and synthesis of quantum dots (QDs), heralding a new era for these nanomaterials. The inception of QDs dates back more than 40 years, during which the theory of QDs has been continuously refined, the manufacturing techniques have significantly flourished, and the applications have largely expanded. Recently, QDs have become important optical devices, playing key roles in numerous fields such as display, energy, and biomedical applications. To celebrate the outstanding achievements of QDs over the years, we dedicate this paper to QDs. In the information field, QDs have been extensively utilized to design devices related to domains like transmission and storage, achieving many breakthroughs in performance. This paper proposes a comprehensive set of methodologies and paradigms for designing information systems using QDs. The proposed approach embodies two characteristics of QDs: 1) QDs play a central role in every aspect of the system and possess the capability to construct an all-quantum-dot (All-QD) information system. 2) QDs possess tunability and wavelength flexibility, which can significantly enhance the information density. Finally, we construct a prototype model of an All-QD information system and validate its feasibility through simulation. We believe that with the continued development of quantum dot (QD) technology, the realization of an All-QD information system is on the horizon.
Spectrometer miniaturization has become a significant trend driven by the demand for distributed and continuous spectral sensing. Broadband encoding spectrometers, which utilize broadband encoder arrays to extract spectral features and algorithms to reconstruct spectra, are among the most competitive candidates for high-performance miniaturized spectrometers. Enhancing the spectral feature extraction capability of these broadband encoders is essential for improving spectrometer performance. However, the strategies and approaches for optimizing these encoders are not yet well-defined. This study analyzes the effectiveness of improving the basis orthogonality of the encoders for their optimization and proposes a dual-layer broadband encoder structure to implement this optimization strategy. The designed dual-layer broadband encoders consist of vertically stacked quantum dot encoders and TiO2/SiO2 encoders, with the corresponding basis being mixed Gaussian. Simulation experiments demonstrate that the proposed dual-layer broadband encoder structure significantly improves the encoders' basis orthogonality, leading to enhanced spectral detection accuracy of the spectrometer constructed with these dual-layer encoders. Experimental fabrication of the dual-layer encoders confirms their physical feasibility and basis orthogonality enhancement.
We propose a design scheme of dual-layer broadband filter spectrometers, and analyze the advantages of this dual-layer filter structure in improving the ill-conditioned spectrum reconstruction processes of micro-spectrometers. Quantum dot filters and silicon film filters are stacked vertically to form the dual-layer filter array. The results show that in the case of different scales of filter arrays, the dual-layer filter structure can provide a less ill-conditioned spectral description matrix than the monolayer filter structure. This usually means that the spectrometers based on this dual-layer structure can have better spectrum reconstruction performance and noise resistance.
Few-shot dense retrieval (DR) aims to effectively generalize to novel search scenarios by learning a few samples. Despite its importance, there is little study on specialized datasets and standardized evaluation protocols. As a result, current methods often resort to random sampling from supervised datasets to create "few-data" setups and employ inconsistent training strategies during evaluations, which poses a challenge in accurately comparing recent progress. In this paper, we propose a customized FewDR dataset and a unified evaluation benchmark. Specifically, FewDR employs class-wise sampling to establish a standardized "few-shot" setting with finely-defined classes, reducing variability in multiple sampling rounds. Moreover, the dataset is disjointed into base and novel classes, allowing DR models to be continuously trained on ample data from base classes and a few samples in novel classes. This benchmark eliminates the risk of novel class leakage, providing a reliable estimation of the DR model's few-shot ability. Our extensive empirical results reveal that current state-of-the-art DR models still face challenges in the standard few-shot scene. Our code and data will be open-sourced at https://github.com/OpenMatch/ANCE-Tele.
Large-capacity information encryption has attracted significant interest in the information age. The diversity and controllability of spectra have positioned them to be widely applied for information encryption. Current spectra-based information encryption methods commonly rely on either spectral alteration induced by external stimuli or the utilization of narrowband channels within spectra. However, these methods encounter a common challenge in attaining both high security and large capacity simultaneously. To address these issues, we propose a multiple-channel information encryption system based on quantum dot (QD) absorption spectra. The diversity of QD absorption spectra and their broadband features ensure that the encrypted spectra can hardly be decrypted without knowing the correct channel matrix. Meanwhile, the large capacity is realized through the combination of multiple QD spectral channels with a theoretical maximum capacity of 24.0 bits in a single spectrum. In order to optimize the performance of our proposed system, the selection principle of the channel matrix is established to achieve the rapid identification of the optimal channel matrix in several milliseconds. The additivity of QD spectral channels and the consistency of QD spectra are also explored to minimize the impact of errors on information decryption. Furthermore, two spectral encryption scenarios of spatial pattern and spectral pattern are applied to demonstrate the feasibility, showcasing their ability to achieve both a high level of security and large capacity. Owing to the advantages offered by QD spectra, the QD spectra-based information system exhibits excellent potential for broader applications in information storage, authentication, and computing.
In this paper, we investigate the instability in the standard dense retrieval training, which iterates between model training and hard negative selection using the being-trained model. We show the catastrophic forgetting phenomena behind the training instability, where models learn and forget different negative groups during training iterations. We then propose ANCE-Tele, which accumulates momentum negatives from past iterations and approximates future iterations using lookahead negatives, as "teleportations" along the time axis to smooth the learning process. On web search and OpenQA, ANCE-Tele outperforms previous state-of-the-art systems of similar size, eliminates the dependency on sparse retrieval negatives, and is competitive among systems using significantly more (50x) parameters. Our analysis demonstrates that teleportation negatives reduce catastrophic forgetting and improve convergence speed for dense retrieval training. Our code is available at https://github.com/OpenMatch/ANCE-Tele.
The effectiveness of Neural Information Retrieval (Neu-IR) often depends on a large scale of in-domain relevance training signals, which are not always available in real-world ranking scenarios. To democratize the benefits of Neu-IR, this paper presents MetaAdaptRank, a domain adaptive learning method that generalizes Neu-IR models from label-rich source domains to few-shot target domains. Drawing on source-domain massive relevance supervision, MetaAdaptRank contrastively synthesizes a large number of weak supervision signals for target domains and meta-learns to reweight these synthetic "weak" data based on their benefits to the target-domain ranking accuracy of Neu-IR models. Experiments on three TREC benchmarks in the web, news, and biomedical domains show that MetaAdaptRank significantly improves the few-shot ranking accuracy of Neu-IR models. Further analyses indicate that MetaAdaptRank thrives from both its contrastive weak data synthesis and meta-reweighted data selection. The code and data of this paper can be obtained from https://github.com/thunlp/MetaAdaptRank.
The issue of information security is closely related to every aspect of daily life. For pursuing a higher level of security, much effort has been continuously invested in the development of information security technologies based on encryption and storage. Current approaches using single-dimension information can be easily cracked and imitated due to the lack of sufficient security. Multidimensional information encryption and storage are an effective way to increase the security level and can protect it from counterfeiting and illegal decryption. Since light has rich dimensions (wavelength, duration, phase, polarization, depth, and power) and synergy between different dimensions, light as the input is one of the promising candidates for improving the level of information security. In this review, based on six different dimensional features of the input light, we mainly summarize the implementation methods of multidimensional information encryption and storage including material preparation and response mechanisms. In addition, the challenges and future prospects of these information security systems are discussed.
Surgical removal is the primary treatment for liver cancer, but frequent recurrence caused by residual malignant tissue remains an important challenge, as recurrence leads to high mortality. It is unreliable to distinguish tumors from normal tissues merely under visual inspection. Hyperspectral imaging (HSI) has been proved to be a promising technology for intra-operative use by capturing the spatial and spectral information of tissue in a fast, non-contact and label-free manner. In this work, we investigated the feasibility of HSI for liver tumor delineation on surgical specimens using a multi-task U-Net framework. Measurements are performed on 19 patients and a dataset of 36 specimens was collected with corresponding pathological results serving as the ground truth. The developed framework can achieve an overall sensitivity of 94.48% and a specificity of 87.22%, outperforming the baseline SVM method by a large margin. In particular, we propose to add explanations on the well-trained model from the spatial and spectral dimensions to show the contribution of pixels and spectral channels explicitly. On that basis, a novel saliency-weighted channel selection method is further proposed to select a small subset of 5 spectral channels which provide essentially as much information as using all 224 channels. According to the dominant channels, the absorption difference of hemoglobin and bile content in the normal and malignant tissues seems to be promising markers that could be further exploited.
Open-domain KeyPhrase Extraction (KPE) aims to extract keyphrases from documents without domain or quality restrictions, e.g., web pages with variant domains and qualities. Recently, neural methods have shown promising results in many KPE tasks due to their powerful capacity for modeling contextual semantics of the given documents. However, we empirically show that most neural KPE methods prefer to extract keyphrases with good phraseness, such as short and entity-style n-grams, instead of globally informative keyphrases from open-domain documents. This paper presents JointKPE, an open-domain KPE architecture built on pre-trained language models, which can capture both local phraseness and global informativeness when extracting keyphrases. JointKPE learns to rank keyphrases by estimating their informativeness in the entire document and is jointly trained on the keyphrase chunking task to guarantee the phraseness of keyphrase candidates. Experiments on two large KPE datasets with diverse domains, OpenKP and KP20k, demonstrate the effectiveness of JointKPE on different pre-trained variants in open-domain scenarios. Further analyses reveal the significant advantages of JointKPE in predicting long and non-entity keyphrases, which are challenging for previous neural KPE methods. Our code is publicly available at https://github.com/thunlp/BERT-KPE.
Cross-reactive sensor arrays have powerful abilities in distinguishing multiple analytes. The larger the effective data volume is, the better the detection performance is. However, current collected data volume depends heavily on the amount of sensing materials involved. It is data-inefficient for each material and causes heavy workload to prepare materials. Herein, by introducing the dimension of excitation wavelength, we report a cross-reactive sensor array by using only one fluorescent material (beta-cyclodextrin-modified quantum dots, QD-beta-CD). The newly added dimension is demonstrated by collecting emission signals under four excitation wavelengths. Through machine learning algorithms, the cross-reactive sensor array can be used for the detection of single nitrophenol (NP) isomer (e.g. a superior detectability with a classification accuracy of 94.9 % for 0.01 mu M pnitrophenol), and concurrent quantitative analysis of complex binary/ternary NP mixtures in environment analysis. Our results indicate that the additional dimension of excitation wavelength can provide an effective way to increase the differences of multiple analytes in the design of cross-reactive sensor array. The idea of adding an extra dimension can be generally applicable for the fields of multidimensional information collection.
Liver cancer is the fourth leading cause of cancer death in the world, and it is even more serious in China, occupying over 50% of new cases and deaths worldwide. One of the major challenges of cancer surgery remains the complete removal of the tumor. Therefore, numerous studies have proposed to design automatic systems aimed at supporting physicians during the surgery. As an optical imaging modality, hyperspectral imaging (HSI) simultaneously captures spatial and spectral features to bring additional insight to surgeons in a label-free manner. However, to the best of our knowledge, previous effort on medical HSI lacks attention to the interpretability of decisions, which greatly hinders the application of HSI in safety-critical scenarios. Through the intrinsic correlation between tissue composition and its spectrum, hyperspectral images can be interpreted according to different spectral characteristics. Generally, interpretability of spectral images often employs a physical model of light interact with tissues. Such an approach has been long practiced in tissue optics, e.g., diffuse reflectance spectroscopy (DRS). Recent advances in deep learning showed state-of-the-art performance on both detection and segmentation of malignant tissues, breaking through the bottleneck of traditional methods in real-time applications. However, due to the lack of interpretability of the deep learning model, the reliability of the results becomes an issue. In this work, we develop and evaluate a diffusion model-based and a deep learning-based interpretable spectral imaging model, respectively, in the context of liver tumor delineation from clinical hyperspectral image dataset. Notably, both methods point to a key sign of abnormality is the difference in bile volume fraction, which could lay the foundation for the interpretability and system design of the spectral imaging system in practical use.