The rapid advancement of large language models (LLMs) has significantly enhanced their reasoning abilities, enabling increasingly complex tasks. However, these capabilities often diminish in smaller, more computationally efficient models like GPT-2. Recent research shows that reasoning distillation can help small models acquire reasoning capabilities, but most existing methods focus primarily on improving teacher-generated reasoning paths. Our observations reveal that small models can generate high-quality reasoning paths during sampling, even without chain-of-thought prompting, though these paths are often latent due to their low probability under standard decoding strategies. To address this, we propose Self-Enhanced Reasoning Training (SERT), which activates and leverages latent reasoning capabilities in small models through self-training on filtered, self-generated reasoning paths under zero-shot conditions. Experiments using OpenAI's GPT-3.5 as the teacher model and GPT-2 models as the student models demonstrate that SERT enhances the reasoning abilities of small models, improving their performance in reasoning distillation.
Sand filters (SFs) are common treatment processes for nitrogen pollutant removal in drinking water treatment plants (DWTPs). However, the mechanisms on the nitrogen-cycling role of SFs are still unclear. In this study, 16S rRNA gene amplicon sequencing was used to characterise the diversity and composition of the bacterial community in SFs from DWTPs. Additionally, metagenomics approach was used to determine the functional microorganisms involved in nitrogen cycle in SFs. Our results showed that Pseudomonadota, Acidobacteria, Nitrospirae and Chloroflexi dominated in SFs. Subsequently, 85 high-quality metagenome-assembled genomes (MAGs) were retrieved from metagenome datasets of selected SFs involving nitrification, assimilatory nitrogen reduction, denitrification and anaerobic ammonia oxidation (anammox) processes. Read mapping to reference genomes of Nitrospira and the phylogenetic tree of the ammonia monooxygenase subunit A gene, amoA, suggested that Nitrospira is abundantly found in SFs. Furthermore, according to their genetic content, a nitrogen metabolic model in SFs was proposed using representative MAGs and pure culture isolate. Quantitative real-time polymerase chain reaction (qPCR) showed that ammonia-oxidising bacteria (AOB) and archaea (AOA), and complete ammonia oxidisers (comammox) were ubiquitous in the SFs, with the abundance of comammox being higher than that of AOA and AOB. Moreover, we identified a bacterial strain with a high NO3-N removal rate as Pseudomonas sp. DW-5, which could be applied in the bioremediation of micro-polluted drinking water sources. Our study provides insights into functional nitrogen-metabolising microbes in SFs of DWTPs.
Inspired by the contrastive predictive coding (CPC), we propose a feature representation scheme for automatic speech recognition (ASR), which encodes sequential dependency information from raw audio signals. Following the original CPC, for a given frame, mutual information (MI) lower bound is maximized between historical context and future prediction. While computing the MI lower bound, based on original CPC, we develop the sequential CPC (SEQ-CPC), which takes the sequential information between frames into consideration. Since speech frames are not independent events, incorporating sequential information leads to better recognition performance. Experimental results on WSJ corpus show that SEQ-CPC achieves the best performance than CPC and NCE which is the contrastive objective used in wav2vec.
In this work, we address the Text-to-Speech (TTS) task by proposing a non-autoregressive architecture called EfficientTTS. Unlike the dominant non-autoregressive TTS models, which are trained with the need of external aligners, EfficientTTS optimizes all its parameters with a stable, end-to-end training procedure, while allowing for synthesizing high quality speech in a fast and efficient manner. EfficientTTS is motivated by a new monotonic alignment modeling approach (also introduced in this work), which specifies monotonic constraints to the sequence alignment with almost no increase of computation. By combining EfficientTTS with different feed-forward network structures, we develop a family of TTS models, including both text-to-melspectrogram and text-to-waveform networks. We experimentally show that the proposed models significantly outperform counterpart models such as Tacotron 2 and Glow-TTS in terms of speech quality, training efficiency and synthesis speed, while still producing the speeches of strong robustness and great diversity. In addition, we demonstrate that proposed approach can be easily extended to autoregressive models such as Tacotron 2.
BackgroundSBRT is a treatment modality for intracranial lesions of oligometastatic NSCLC. Combined with immune checkpoint inhibitors (ICIs), SBRT can improve the anti-tumor effect and the systemic response to immunotherapy via the abscopal effect in brain metastases (BMs) NSCLC patients. Meanwhile, the combination of anlotinib(a novel multi-target tyrosine kinase inhibitor)and ICIs can promote the infiltration of the innate immune cells and conferred potentially synergistic anti-tumor activity. Therefore, triple therapy of SBRT, anlotinib, and toripalimab (anti-PD-1 antibody) may be more effective in driver mutation-negative NSCLC patients with untreated oligometastatic BMs.Trial DesignThis is a prospective, single-center, open-label, phase Ib study to evaluate the safety and feasibility of SBRT and anlotinib with/without toripalimab treatment. Ten patients will be 1:1 randomized into 2 groups. SBRT (35 Gy/5 F for 1∼5 lesions) plus anlotinib (12 mg, po, d1∼14, q3w) with or without toripalimab (240 mg, iv, d1, q3w) will be given as an induction therapy. Afterward, patients will be treated with toripalimab (240 mg, iv, d1, q3w) plus anlotinib (12 mg, po, d1∼14, q3w) on day 22 for 1 year until disease progression or intolerable toxicity. Key inclusion criteria: patients aged 18-70 years, ECOG ≤ 1, driver mutation-negative (EGFR, ALK, ROS1), untreated NSCLC BMs. The primary endpoints are intracranial response rate (iORR) and the incidence of treatment-related adverse events. Secondary endpoints include intracranial progression-free survival (iPFS), local control rate (LCR), overall survival (OS), quality of life (QoL) and cognitive function.Clinical trial identificationNCT05021328.Legal entity responsible for the studyThe authors.FundingHas not received any funding.DisclosureAll authors have declared no conflicts of interest. BackgroundSBRT is a treatment modality for intracranial lesions of oligometastatic NSCLC. Combined with immune checkpoint inhibitors (ICIs), SBRT can improve the anti-tumor effect and the systemic response to immunotherapy via the abscopal effect in brain metastases (BMs) NSCLC patients. Meanwhile, the combination of anlotinib(a novel multi-target tyrosine kinase inhibitor)and ICIs can promote the infiltration of the innate immune cells and conferred potentially synergistic anti-tumor activity. Therefore, triple therapy of SBRT, anlotinib, and toripalimab (anti-PD-1 antibody) may be more effective in driver mutation-negative NSCLC patients with untreated oligometastatic BMs. SBRT is a treatment modality for intracranial lesions of oligometastatic NSCLC. Combined with immune checkpoint inhibitors (ICIs), SBRT can improve the anti-tumor effect and the systemic response to immunotherapy via the abscopal effect in brain metastases (BMs) NSCLC patients. Meanwhile, the combination of anlotinib(a novel multi-target tyrosine kinase inhibitor)and ICIs can promote the infiltration of the innate immune cells and conferred potentially synergistic anti-tumor activity. Therefore, triple therapy of SBRT, anlotinib, and toripalimab (anti-PD-1 antibody) may be more effective in driver mutation-negative NSCLC patients with untreated oligometastatic BMs. Trial DesignThis is a prospective, single-center, open-label, phase Ib study to evaluate the safety and feasibility of SBRT and anlotinib with/without toripalimab treatment. Ten patients will be 1:1 randomized into 2 groups. SBRT (35 Gy/5 F for 1∼5 lesions) plus anlotinib (12 mg, po, d1∼14, q3w) with or without toripalimab (240 mg, iv, d1, q3w) will be given as an induction therapy. Afterward, patients will be treated with toripalimab (240 mg, iv, d1, q3w) plus anlotinib (12 mg, po, d1∼14, q3w) on day 22 for 1 year until disease progression or intolerable toxicity. Key inclusion criteria: patients aged 18-70 years, ECOG ≤ 1, driver mutation-negative (EGFR, ALK, ROS1), untreated NSCLC BMs. The primary endpoints are intracranial response rate (iORR) and the incidence of treatment-related adverse events. Secondary endpoints include intracranial progression-free survival (iPFS), local control rate (LCR), overall survival (OS), quality of life (QoL) and cognitive function. This is a prospective, single-center, open-label, phase Ib study to evaluate the safety and feasibility of SBRT and anlotinib with/without toripalimab treatment. Ten patients will be 1:1 randomized into 2 groups. SBRT (35 Gy/5 F for 1∼5 lesions) plus anlotinib (12 mg, po, d1∼14, q3w) with or without toripalimab (240 mg, iv, d1, q3w) will be given as an induction therapy. Afterward, patients will be treated with toripalimab (240 mg, iv, d1, q3w) plus anlotinib (12 mg, po, d1∼14, q3w) on day 22 for 1 year until disease progression or intolerable toxicity. Key inclusion criteria: patients aged 18-70 years, ECOG ≤ 1, driver mutation-negative (EGFR, ALK, ROS1), untreated NSCLC BMs. The primary endpoints are intracranial response rate (iORR) and the incidence of treatment-related adverse events. Secondary endpoints include intracranial progression-free survival (iPFS), local control rate (LCR), overall survival (OS), quality of life (QoL) and cognitive function. Clinical trial identificationNCT05021328. NCT05021328. Legal entity responsible for the studyThe authors. The authors. FundingHas not received any funding. Has not received any funding.
Bacterial, archaeal, and eukaryota diversity in mountainous areas varies along elevational gradients, but details remain unclear. Here, we use a next-generation sequencing method based on 16S/18S rRNA to reveal the soil microbial diversity and community compositions of alpine meadow ecosystems along an elevation span of nearly 2,000 m (1,936–3,896 m) in China’s Qilian Mountains. Both bacterial and eukaryota diversity increased linearly with increasing elevation, whereas archaeal diversity increased, but not significantly. The diversity patterns of several phyla in the bacterial, archaeal, and eukaryota communities were consistent with the overall elevational trend, but some phyla did not follow this pattern. The soil microbial community compositions were shaped by the coupled effects of regional climate and local soil properties. Intradomain links were more important than interdomain links in the microbial network of the alpine meadows, and these links were mostly positive. The bacteria formed more connections than either archaea or eukaryota, but archaea may be more important than bacteria in building the soil microbial co-occurrence network in this region. Our results provide new visions on the formation and maintenance of soil microbial diversity along an elevational gradient and have implications for microbial responses to climate change in alpine ecosystems.
In this paper, we propose a novel system based on wordlevel features and window-based attention for polyphone disambiguation, which is a fundamental task for Grapheme-tophoneme (G2P) conversion of Mandarin Chinese. The framework aims to combine a pre-trained language model with explicit word-level information in order to get meaningful context extraction. Particularly, we employ a pre-trained bidirectional encoder from Transformers (BERT) model to extract character-level features, and an external Chinese word segmentation (CWS) tool is used to obtain the word units. We adopt a mixed pooling mechanism to convert character-level features into word-level features based on the segmentation results. A window-based attention module is utilized to incorporate contextual word-level features for the polyphonic characters. Experimental results show that our method achieves an accuracy of 99.06% on an open benchmark dataset for Mandarin Chinese polyphone disambiguation, which outperforms the baseline systems.
In order to obtain high-quality images in natural conditions, natural degradation image enhancement has been a research hotspot in recent years. In this paper, we present a transfer learning approach for multiple types of natural degradation image enhancement. We propose to create a common source domain for various natural degradations and perform the transfer learning for every specific natural degradation individually. By reusing the general enhancement model, we can circumvent the scarcity of training dataset and the computation-intensive training process for deep learning methods. In the experiment, we transfer the general model to three target tasks: raining image enhancement, snowing image enhancement and underwater image enhancement. With the finetuning of only 5 epochs, the enhancement models have been able to outperform several state-of-the-art methods that designed for specific task.
Existing multi-style speech synthesis methods require either style labels or large amounts of unlabeled training data, making data acquisition difficult. In this paper, we present an unsupervised multi-style speech synthesis method that can be trained with limited data. We leverage instance discriminator to guide a style encoder to learn meaningful style representations from a multi-style dataset. Furthermore, we employ information bottleneck to filter out style-irrelevant information in the representations, which can improve speech quality and style similarity. Our method is able to produce desirable speech using a fairly small dataset, where the baseline GST-Tacotron fails. ABX tests show that our model significantly outperforms GST-Tacotron in both emotional speech synthesis task and multi-speaker speech synthesis task. In addition, we demonstrate that our method is able to learn meaningful style features with only 50 training samples per style.
Text Normalization (TN) is an essential part in conversational systems like text-to-speech synthesis (TTS) and automatic speech recognition (ASR). It is a process of transforming non-standard words (NSW) into a representation of how the words are to be spoken. Existing approaches to TN are mainly rule-based or hybrid systems, which require abundant hand-crafted rules. In this paper, we treat TN as a neural machine translation problem and present a pure data-driven TN system using Transformer framework. Partial Parameter Generator (PPG) and Pointer-Generator Network (PGN) are combined in our model to improve accuracy of normalization and act as auxiliary modules to reduce the number of simple errors. The experiments demonstrate that our proposed model reaches remarkable performance on various semiotic classes.
In this paper, we present EfficientSing, a Chinese singing voice synthesis (SVS) system based on a non-autoregressive durationfree acoustic model and HiFi-GAN neural vocoder. Different from many existing SVS methods, no auxiliary duration prediction module is needed in this work, since a newly proposed monotonic alignment modeling mechanism is adopted. Moreover, we follow the non-autoregressive architecture of EfficientTTS with some singing-specific adaption, making training and inference fully parallel and efficient. HiFi-GAN vocoder is adopted to improve the voice quality of synthesized songs and inference efficiency. Both objective and subjective experimental results show that the proposed system can produce quite natural and high-fidelity songs and outperform the Tacotron-based baseline in terms of pronunciation, pitch and rhythm.
In this paper we describe a novel replay detection system for the ASVspoof 2019 challenge. The objective of this challenge is to distinguish arbitrarily audio files from bona fide or spoofing attacks, where spoofing attacking includes replay attacks, text-tospeech and voice conversions. Our replay detection system is a pipeline system with three aspects: feature engineering, DNN models, and score fusion. Firstly, logspec is extracted as input features according to previous research works where spectrum augmentation is applied during training stage to boost performance under limited training data. Secondly, DNN models part includes three major models: SEnet, DenseNet, and our proposed model, channel consistency DenseNeXt, where binary cross entropy loss and center loss are applied as training objectives. Finally, score fusion is applied to all three DNN models in order to obtain primary system results. The experiment results show that for our best single system, channel consistency DenseNeXt, t-DCF and EER are 0.0137 and 0.46% on physical access evaluation set respectively. The performance of primary system obtains 0.00785 and 0.282% in terms of t-DCF and EER respectively. This is a 96.8% improvement compared to the baseline system CQCC-GMM and it achieves state-ofthe-art performance in PA challenge.
BACKGROUND:Procalcitonin (PCT), C-reactive protein (CRP), and neutrophil-to-lymphocyte ratio (NLR) have emerged as important markers of inflammation, and these markers, especially PCT and CRP, have been studied in patients with neutropenia. This study was designed to evaluate their value in differentiating infectious fever from tumor fever (TF) and to investigate their role in assessing outcomes in nonneutropenic lung cancer patients (NNLCPs).METHODS:This retrospective clinical study included 588 febrile NNLCPs between January 2019 and December 2019. The levels of PCT, CRP, and conventional inflammatory markers, including white blood cells (WBC) and neutrophils (NEU), were measured. NLR was defined as the ratio of the absolute neutrophil count to the absolute lymphocyte count. Patients' clinical and bacteriological data were recorded.RESULTS:This study included 311 NNLCPs with bacterial infections and 277 with TF. Inflammatory markers such as PCT, CRP, WBC, and NEU levels and NLR were significantly higher in patients with bacterial infections than in those with TF (p < 0.0001). However, PCT level was the best predictor of bacterial infections, with an area under the curve (AUC) of 0.874, followed by CRP level (AUC = 0.855) and NLR (AUC = 0.792) (p < 0.0001). Additionally, PCT level was significantly elevated in patients with bacterial infections with progressive disease after radiotherapy and chemotherapy (p < 0.01).CONCLUSIONS:The present study demonstrated the superiority of PCT over CRP and NLR in the diagnosis of febrile patients with bacterial infections. Additionally, PCT can be used to assess the clinical outcomes and cancer progression in NNLCPs.
Colorectal cancer (CRC) remains one of the most commonly diagnosed malignancies worldwide. Circular RNAs (circRNAs) are being found to play crucial roles in human cancer, including CRC. The purpose of this study was to explore the function and mechanism of circ_0007031 on CRC progression and 5-fluorouracil (5-FU) resistance. The levels of circ_0007031, ATP-binding cassette subfamily C member 5 (ABCC5) and miR-133b were assessed by quantitative real-time polymerase chain reaction (qRT-PCR) or western blot. Cell survival and proliferation were detected by the 3-(4,5-dimethylthiazol-2yl)-5-(3-carboxymethoxyphenyl)-2-(4-sulfophenyl)-2H-tetrazolium (MTS) assay. Cell colony formation was evaluated using a standard colony formation assay. Transwell assays were performed to determine cell migration and invasion. Targeted correlations among circ_0007031, miR-133b and ABCC5 were verified by dual-luciferase reporter, RNA immunoprecipitation (RIP) and RNA pulldown assays. Animal experiments were performed to observe the role of circ_0007031 in vivo. Our data indicated that circ_0007031 up-regulation was associated with CRC resistance to 5-FU. Circ_0007031 knockdown repressed CRC cell proliferation, migration and invasion and enhanced 5-FU sensitivity. Circ_0007031 directly interacted with miR-133b. Moreover, circ_0007031 knockdown regulated CRC cell progression and 5-FU sensitivity by miR-133b. ABCC5 was a direct target of miR-133b, and circ_0007031 mediated ABCC5 expression via acting as a miR-133b sponge. Furthermore, miR-133b overexpression regulated CRC cell progression and sensitivity to 5-FU by down-regulating ABCC5. Additionally, circ_0007031 knockdown suppressed tumor growth in vivo. Our current work had led to the identification of circ_0007031 knockdown that repressed CRC cell malignant progression and enhanced 5-FU sensitivity via regulating ABCC5 expression by sponging miR-133b.
Background Colorectal cancer (CRC) is the most common cause of cancer-related mortality in the world. Long non-coding RNAs (lncRNAs) are involved in the development of many cancers. However, studies on the effect of lncRNA small nucleolar RNA host gene 16 (SNHG16) on the proliferation, metastasis and apoptosis of CRC are still few. Methods Quantitative real-time polymerase chain reaction (qRT-PCR) was performed to determine the expression levels of SNHG16, microRNA-132-3p (miR-132-3p) and ubiquitin specific peptidase 22 (USP22). The proliferation, apoptosis, migration and invasion of CRC cells were evaluated by the 3-(4,5-dimethyl-2 thiazolyl)-2,5-diphenyl-2-H-tetrazolium bromide (MTT) assay, flow cytometry and transwell assay, respectively. Dual-luciferase reporter assay was used to verify the interactions among SNHG16, miR-132-3p and USP22. Also, Western blot analysis was used to assess the protein levels of USP22 and metastasis-related markers. Moreover, mice xenograft models were used to determine the effect of SNHG16 on CRC tumor growth in vivo. Results SNHG16 was highly expressed in CRC tissues and cells. Knockdown of SNHG16 reduced the proliferation, migration, invasion, and promoted the apoptosis of CRC cells. MiR-132-3p could interact with SNHG16, and its inhibitor recovered the suppression effect of silenced SNHG16 on CRC cell progression. Besides, USP22 was a target of miR-132-3p, and its overexpression restored the inhibition effect of miR-132-3p mimic on CRC cell progression. In addition, interference of SNHG16 reduced CRC tumor growth in vivo. Conclusion LncRNA SNHG16 might act as an oncogene in CRC. The discovery of the SNHG16/miR-132-3p/USP22 pathway provided new thinking for the treatment of CRC.
Emotion recognition in conversation (ERC) is an important topic for developing empathetic machines in a variety of areas including social opinion mining, health-care and so on. In this paper, we propose a method to model ERC task as sequence tagging where a Conditional Random Field (CRF) layer is leveraged to learn the emotional consistency in the conversation. We employ LSTM-based encoders that capture self and inter-speaker dependency of interlocutors to generate contextualized utterance representations which are fed into the CRF layer. For capturing long-range global context, we use a multi-layer Transformer encoder to enhance the LSTM-based encoder. Experiments show that our method benefits from modeling the emotional consistency and outperforms the current state-of-the-art methods on multiple emotion classification datasets.
In order to standardize the supervision process of systemic risk of Internet finance in China, this paper puts forward a novel supervision method of systemic risk of Internet finance in China based on big data technology. This method is based on big data technology, combines computer technology and Internet of things technology, can real-time monitor the systematic data of China's Internet finance, and quickly feedback the information to the relevant staff. The results show that this method can effectively improve the standardization of China's Internet financial systemic risk supervision and reduce the harm caused by China's Internet financial systemic risk.
Recent studies have shown remarkable success in voice conversion (VC) based on generative adversarial networks (GANs) without parallel data. In this paper, based on the conditional generative adversarial networks (CGANs), we propose a selfand semi-supervised method combined with mixup and data augmentation that allows non-parallel many-to-many voice conversion with fewer labeled data. In this method, the discriminator of CGANs learns to not only distinguish real/fake samples, but also classify attribute domains. We augment the discriminator with an auxiliary task to improve representation learning and introduce a training task to predict labels for the unlabeled samples. The proposed approach reduces the appetite for labeled data in voice conversion, which enables single generative network to implement many-to-many mapping between different voice domains. Experiment results show that the proposed method is able to achieve comparable voice quality and speaker similarity with only 10% of the labeled data.
This paper proposes a nonparallel emotional speech conversion (ESC) method based on Variational AutoEncoder-Generative Adversarial Network (VAE-GAN). Emotional speech conversion aims at transforming speech from one source emotion to that of a target emotion without changing the speaker's identity and linguistic content. In this work, an encoder is trained to elicit the content-related representations from acoustic features. Emotion-related representations are extracted in a supervised manner. Then the transformation between emotion-related representations from different domains is learned using an improved cycle-consistent Generative Adversarial Network (CycleGAN). Finally, emotion conversion is performed by eliciting and recombining the content-related representations of the source speech and the emotion-related representations of the target emotion. Subjective evaluation experiments are conducted and the results show that the proposed method outperforms the baseline in terms of voice quality and emotion conversion ability.