Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core focus on a strong rhythmic sense that harmonizes with accompanying beats. In this paper, we propose Freestyler, the first system that generates rapping vocals directly from lyrics and accompaniment inputs. Freestyler utilizes language model-based token generation, followed by a conditional flow matching model to produce spectrograms and a neural vocoder to restore audio. It allows a 3-second prompt to enable zero-shot timbre control. Due to the scarcity of publicly available rap datasets, we also present RapBank, a rap song dataset collected from the internet, alongside a meticulously designed processing pipeline. Experimental results show that Freestyler produces high-quality rapping voice generation with enhanced naturalness and strong alignment with accompanying beats, both stylistically and rhythmically.
In this paper, we propose a neural Text-To-Speech (TTS) system SoftSpeech, which employs a novel soft length regulated duration attention based decoder. It learns the encoder output mapping to decoder output simultaneously from an unsupervised duration model (Soft-LengthRegulator) without the requirement of external duration information. The Soft-LengthRegulator consists of a Feed-Forward Transformer (FFT) block with Conditional Layer Normalization (CLN), following a learned upsampling layer with multi-head attention and guided multi-head attention constraint, and it is integrated in each decoder layer and achieves accelerated training convergence and better naturalness within FastSpeech 2 framework. Soft Dynamic Time Warping (Soft-DTW) is adopted to align the mismatch spectrogram loss. Moreover, a Fine-Grained style Variational AutoEncoder (VAE) is designed to further improve the naturalness of synthesized speech. The experiments show SoftSpeech outperforms FastSpeech 2 in subjective tests, and can be successfully applied to minority languages with low resources.
This paper proposes ProsodySpeech, a novel prosody model to enhance encoder-decoder neural Text-To-Speech (TTS), to generate high expressive and personalized speech even with very limited training data. First, we use a Prosody Extractor built from a large speech corpus with various speakers to generate a set of prosody exemplars from multiple reference speeches, in which Mutual Information based Style content separation (MIST) is adopted to alleviate "content leakage" problem. Second, we use a Prosody Distributor to make a soft selection of appropriate prosody exemplars in phone-level with the help of an attention mechanism. The resulting prosody feature is then aggregated into the output of text encoder, together with additional phone-level pitch feature to enrich the prosody. We apply this method into two tasks: highly expressive multi style/emotion TTS and few-shot personalized TTS. The experiments show the proposed model outperforms baseline FastSpeech 2 + GST with significant improvements in terms of similarity and style expression.
Denoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses the state-of-the-art generative models, which invariably results in slow inference speed. Previous approaches aim to optimize the choice of inference schedule over a few iterations to speed up inference. However, this results in reduced generation quality, mainly because the inference process is optimized separately, without jointly optimizing with the training process. In this paper, we propose InferGrad, a diffusion model for vocoder that incorporates inference process into training, to reduce the inference iterations while maintaining high generation quality. More specifically, during training, we generate data from random noise through a reverse process under inference schedules with a few iterations, and impose a loss to minimize the gap between the generated and ground-truth data samples. Then, unlike existing approaches, the training of InferGrad considers the inference process. The advantages of InferGrad are demonstrated through experiments on the LJSpeech dataset showing that InferGrad achieves better voice quality than the baseline WaveGrad under same conditions while maintaining the same voice quality as the baseline but with 3x speedup (2 iterations for InferGrad vs 6 iterations for WaveGrad).
Cross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expressing all styles and to collect corresponding recordings for model training. However, the performances of existing style transfer methods are still far behind real application needs. The root causes are mainly twofold. Firstly, the style embedding extracted from single reference speech can hardly provide fine-grained and appropriate prosody information for arbitrary text to synthesize. Secondly, in these models the content/text, prosody, and speaker timbre are usually highly entangled, it’s therefore not realistic to expect a satisfied result when freely combining these components, such as to transfer speaking style between speakers. In this paper, we propose a crossspeaker style transfer text-to-speech (TTS) model with explicit prosody bottleneck. The prosody bottleneck builds up the kernels accounting for speaking style robustly, and disentangles the prosody from content and speaker timbre, therefore guarantees high quality cross-speaker style transfer. Evaluation result shows the proposed method even achieves on-par performance with source speaker’s speaker-dependent (SD) model in objective measurement of prosody, and significantly outperforms the cycle consistency and GMVAEbased baselines in objective and subjective evaluations.
In this paper, we propose a cycle consistent network based end-to-end TTS for speaking style transfer, including intra-speaker, inter-speaker, and unseen speaker style transfer for both parallel and unparallel transfers. The proposed approach is built upon a multi-speaker Variational Autoencoder (VAE) TTS model. The model is usually trained in a paired manner, which means the reference speech is totally paired with the output including speaker identity, text, and style. To achieve a better quality for style transfer, which for most cases is in an unpaired manner, we augment the model with an unpaired path with a separated variational style encoder. The unpaired path takes as input an unpaired reference speech and yields an unpaired output. The unpaired output, which lacks direct ground-truth target, is then successfully constrained by a delicately designed cycle consistent network. Specifically, the unpaired output of the forward transfer is fed into the model again as an unpaired reference input, and after the backward transfer yields an output expected to be the same as the original unpaired reference speech. Ablation study shows the effectiveness of the unpaired path, separated style encoders and cycle consistent network in the proposed model. The final evaluation demonstrates the proposed approach significantly outperforms the Global Style Token (GST) and VAE based systems for all the six style transfer categories, in metrics of naturalness, speech quality, similarity of speaker identity, and similarity of speaking style.
Autoimmune deficiency and destruction in either β-cell mass or function can cause insufficient insulin levels and, as a result, hyperglycemia and diabetes. Thus, promoting β-cell proliferation could be one approach toward diabetes intervention. In this report we describe the discovery of a potent and selective DYRK1A inhibitor GNF2133, which was identified through optimization of a 6-azaindole screening hit. In vitro, GNF2133 is able to proliferate both rodent and human β-cells. In vivo, GNF2133 demonstrated significant dose-dependent glucose disposal capacity and insulin secretion in response to glucose-potentiated arginine-induced insulin secretion (GPAIS) challenge in rat insulin promoter and diphtheria toxin A (RIP-DTA) mice. The work described here provides new avenues to disease altering therapeutic interventions in the treatment of type 1 diabetes (T1D).
Loss of beta-cell mass and function can lead to insufficient insulin levels and ultimately to hyperglycemia and diabetes mellitus. The mainstream treatment approach involves regulation of insulin levels; however, approaches intended to increase beta-cell mass are less developed. Promoting beta-cell proliferation with low-molecular-weight inhibitors of dual-specificity tyrosine-regulated kinase 1A (DYRK1A) offers the potential to treat diabetes with oral therapies by restoring beta-cell mass, insulin content and glycemic control. GNF4877, a potent dual inhibitor of DYRK1A and glycogen synthase kinase 3 beta (GSK3 beta) was previously reported to induce primary human beta-cell proliferationin vitroandin vivo. Herein, we describe the lead optimization that lead to the identification of GNF4877 from an aminopyrazine hit identified in a phenotypic high-throughput screening campaign measuring beta-cell proliferation.
Checkpoint inhibition has transformed immunotherapy by alleviating T cell exhaustion in a subset of patients. However, an important component of effective immune targeting to expand the benefit of immune response requires engagement of both innate and adaptive responses. Our understanding of safe and effective engagement of the innate immune system is evolving, with multiple preclinical and clinical agents targeting pathways such the Toll-like Receptors. Here we disclose the structure and preclinical activity of LHC165, a benzonapthyridine TLR7 agonist that is adsorbed to aluminum hydroxide. The interaction between LHC165 and aluminum hydroxide allows for a slow release from the injection site resulting in improved efficacy in mouse models compared with free LHC165. This localization allows for immune activation at the site of the tumor and also results in lower systemic exposure and cytokine induction. Intratumoral studies in syngeneic preclinical studies show single agent activity and a benefit when dosed in combination with checkpoint blockade. LHC165 as a single agent and in combination with PDR001 is currently enrolling patients with advanced malignancies in CLHC165X2101. Citation Format: Jonathan A. Deane, German A. Cortez, Chun Li, Nora Eifler, Shailaja Kasibhatla, Nehal Parikh, Shifeng Pan, Steven Bender. Identification and characterization of LHC165, a TLR7 agonist designed for localized intratumoral therapies [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 4128.
In this paper, we introduce the Variational Autoencoder (VAE) to an end-to-end speech synthesis model, to learn the latent representation of speaking styles in an unsupervised manner. The style representation learned through VAE shows good properties such as disentangling, scaling, and combination, which makes it easy for style control. Style transfer can be achieved in this framework by first inferring style representation through the recognition network of VAE, then feeding it into TTS network to guide the style in synthesizing speech. To avoid Kullback-Leibler (KL) divergence collapse in training, several techniques are adopted. Finally, the proposed model shows good performance of style control and outperforms Global Style Token (GST) model in ABX preference tests on style transfer.
A versatile flow synthesis method for in situ formation of organozinc reagents and subsequent cross-coupling with aryl halides and activated carboxylic acids is reported. Formation of organozinc reagents is achieved by pumping organic halides, in the presence of ZnCl2 and LiCl, through an activated Mg-packed column under flow conditions. This method provides efficient in situ formation of aryl, primary, secondary, and tertiary alkyl organozinc reagents, which are subsequently telescoped downstream to a Negishi or decarboxylative Negishi cross-coupling reaction. The described method offers access to a variety of C-C bond formations with organozinc reagents that are otherwise commercially unavailable or difficult to prepare under traditional batch reaction conditions.
Objectives Wnt signalling has been implicated in activating a fibrogenic programme in fibroblasts in systemic sclerosis (SSc). Porcupine is an O-acyltransferase required for secretion of Wnt proteins in mammals. Here, we aimed to evaluate the antifibrotic effects of pharmacological inhibition of porcupine in preclinical models of SSc. Methods The porcupine inhibitor GNF6231 was evaluated in the mouse models of bleomycin-induced skin fibrosis, in tight-skin-1 mice, in murine sclerodermatous chronic-graft-versus-host disease (cGvHD) and in fibrosis induced by a constitutively active transforming growth factor-β-receptor I. Results Treatment with pharmacologically relevant and well-tolerated doses of GNF6231 inhibited the activation of Wnt signalling in fibrotic murine skin. GNF6231 ameliorated skin fibrosis in all four models. Treatment with GNF6231 also reduced pulmonary fibrosis associated with murine cGvHD. Most importantly, GNF6231 prevented progression of fibrosis and showed evidence of reversal of established fibrosis. Conclusions These data suggest that targeting the Wnt pathway through inhibition of porcupine provides a potential therapeutic approach to fibrosis in SSc. This is of particular interest, as a close analogue of GNF6231 has already demonstrated robust pathway inhibition in humans and could be available for clinical trials.
TGR5 is a Gs-coupled GPCR activated by both primary and secondary bile acids (i.e. taurolithocholic acid) encoded by the G protein-coupled bile acid receptor 1 (GPBAR1) gene. The biological role of TGR5 is a function of both its pattern of expression and its ability to increase cellular cAMP [1-3]. Activation of TGR5 in intestinal enteroendocrine cells leads to secretion of GLP-1 and other related incretins, through a mechanism similar to GPR119 [4]. In monocytes/macrophages and dendritic cells (Mo/M /DC), TGR5 activation leads to decreases in LPS-induced Th1 cytokines like TNF-α and IL12 [5]. Together, these functions make TGR5 an attractive target in the treatment of type 2 diabetes, where strategies to enhance GLP-1 [6] and mitigate chronic inflammation [7-9] are already demonstrating benefit in the clinic. While several pharmaceutical organizations have pursued TGR5 agonists, no TGR5 drug has been approved for diabetes or metabolic diseases indication.
Abstract Wnt signaling is tightly controlled during cellular proliferation, differentiation and embryonic morphogenesis. Aberrant activation of this pathway plays a critical role in a variety of cancers. Blockade of Wnt signaling is therefore an attractive therapeutic approach for anticancer therapy. In this presentation, we will discuss our approach to search for inhibitors of Wnt ligand secretion. We developed and performed a cellular high-throughput screen using a co-culture system. Lead structure (GNF-1331) was identified and further target elucidation revealed Porcupine, a membrane bound O-acyl transferase, as its molecular target. Further structure-activity relationship studies led to the discovery of WNT974, a potent and specific Porcupine inhibitor. Treatment of WNT974 leads to tumor regression in a Wnt dependent MMTV-Wnt1 mouse model at well tolerated doses. WNT974 is currently in Phase 1 clinical trials. Citation Format: Shifeng Pan, Jun Liu, Dai Cheng, Dong Han, Guobao Zhang, Mindy Hsieh, Nicholas Ng, Chun Li, Shailaja Kasibhatla, Peter McNamara, H. Martin Seidel, Jennifer Harris. Discovery of porcupine inhibitors targeting Wnt signaling in cancer. [abstract]. In: Proceedings of the Fourth AACR International Conference on Frontiers in Basic Cancer Research; 2015 Oct 23-26; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Res 2016;76(3 Suppl):Abstract nr IA35.
Blockade of aberrant Wnt signaling is an attractive therapeutic approach in multiple cancers. We developed and performed a cellular high-throughput screen for inhibitors of Wnt secretion and pathway activation. A lead structure (GNF-1331) was identified from the screen. Further studies identified the molecular target of GNF-1331 as Porcupine, a membrane bound O-acyl transferase. Structure-activity relationship studies led to the discovery of a novel series of potent and selective Porcupine inhibitors. Compound 19, GNF-6231, demonstrated excellent pathway inhibition and induced robust antitumor efficacy in a mouse MMTV-WNT1 xenograft tumor model.
Aberrant Wnt pathway activation due to inactivating mutations in the gene encoding RNF43 (an E3 ubiquitin ligase that promotes degradation of the Wnt receptors Frizzled and LRP6) may contribute to the unresponsiveness of BRAF V600E -mutant colorectal cancer (CRC) to BRAF inhibitors. Analysis of The Cancer Genome Atlas (TCGA) CRC data set reveals a striking co-occurrence of the BRAF V600E mutation and truncating mutations in RNF43. RNF43 mutations are likely to be functionally significant, as RNF43 mutations and mutations in the β-catenin destruction complex component APC are almost completely mutually exclusive. The vast majority of BRAF V600E ;RNF43-mutant CRCs are hypermutable [microsatellite instability (MSI)-high phenotype]. The mismatch repair deficiency in these tumors may directly contribute to RNF43 mutagenesis, as RNF43 mutations tend to be small insertions/deletions in homopolymeric tracts. To determine if RNF43 mutations confer Wnt dependency in BRAF V600E -mutant CRC, we treated three BRAF V600E ;RNF43-mutant CRC patient-derived xenograft (PDX) models with the porcupine inhibitor WNT974 (formerly LGK974), which blocks the palmitoylation and secretion of Wnt ligands. Single agent WNT974 anti-tumor activity was observed in 2/3 PDX models, and correlated with decreased tumor cell proliferation and mucinous differentiation. Single agent anti-tumor activity with the BRAF inhibitor LGX818 was also observed in 2/3 PDX models. No single agent anti-tumor activity was observed with the EGFR inhibitor cetuximab. The double combinations of WNT974+LGX818 and LGX818+cetuximab, and the triple combination of WNT974+LGX818+cetuximab were efficacious in all three BRAF V600E ;RNF43-mutant CRC PDX models. In summary, the Wnt pathway and the EGFR-MAPK pathway may jointly promote tumorigenesis of BRAF V600E -mutant CRC, providing a strong rationale to treat patients with BRAF V600E -mutant CRCs harboring upstream Wnt pathway mutations with combinations of WNT974, LGX818 and/or cetuximab. Citation Format: Youzhen Wang, Michael Palmer, Savina Jaeger, Linda Bagdasarian, Shumei Qiu, Steve Woolfenden, Ronald Meyer, Guizhi Yang, John Green, Shifeng Pan, Jun Liu, Hui Gao, Z. Alexander Cao, Andrea Myers, Margaret E. McLaughlin. Dual Wnt and EGFR-MAPK dependency of BRAF V600E -mutant colorectal cancer. [abstract]. In: Proceedings of the 106th Annual Meeting of the American Association for Cancer Research; 2015 Apr 18-22; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Res 2015;75(15 Suppl):Abstract nr 2140. doi:10.1158/1538-7445.AM2015-2140
As China's BeiDou Navigation Satellite System (BDS) has become operational in the Asia-Pacific region, it is important to demonstrate the capabilities that a combination of GPS, BDS and GLONASS to high-precision positioning. Multi-constellation combination increases the available satellites and thus improves the positioning reliability. However at the same time, it will bring some challenges to the high-dimension ambiguity resolution (AR). In this contribution, a GPS/BDS/GLONASS combined real time kinematic (RTK) positioning method for middle-long baseline is proposed. In order to reduce the influence of troposphere and ionosphere delays on AR, a two-step AR strategy is adopted, where wide-lane and ionosphere-free observation model are used respectively. In the integer ambiguity search process, a partial ambiguity resolution (PAR) method is proposed to improve the AR performance. In the PAR method, satellite cutoff elevation, satellite number, AR success rate and ratio are used together to determine the ambiguity subset, which can be fixed reliably. A set of baselines ranging from about 30 to 60 km, which all contain GPS/BDS/GLONASS observations, are used to test RTK positioning performance. Experiment results demonstrate that GPS/BDS/GLONASS combined RTK positioning with partial ambiguity resolution can get much improved performance for middle-long baseline both in positioning speed and accuracy, as within about 20 sand 5 cm, respectively.
Emerging approaches to treat immune disorders target positive regulatory kinases downstream of antigen receptors with small molecule inhibitors. Here we provide evidence for an alternative approach in which inhibition of the negative regulatory inositol kinase Itpkb in mature T lymphocytes results in enhanced intracellular calcium levels following antigen receptor activation leading to T cell death. Using Itpkb conditional knockout mice and LMW Itpkb inhibitors these studies reveal that Itpkb through its product IP4 inhibits the Orai1/Stim1 calcium channel on lymphocytes. Pharmacological inhibition or genetic deletion of Itpkb results in elevated intracellular Ca2+ and induction of FasL and Bim resulting in T cell apoptosis. Deletion of Itpkb or treatment with Itpkb inhibitors blocks T-cell dependent antibody responses in vivo and prevents T cell driven arthritis in rats. These data identify Itpkb as an essential mediator of T cell activation and suggest Itpkb inhibition as a novel approach to treat autoimmune disease.