We introduce Seed-Music, a suite of music generation systems capable of producing high-quality music with fine-grained style control. Our unified framework leverages both auto-regressive language modeling and diffusion approaches to support two key music creation workflows: controlled music generation and post-production editing. For controlled music generation, our system enables vocal music generation with performance controls from multi-modal inputs, including style descriptions, audio references, musical scores, and voice prompts. For post-production editing, it offers interactive tools for editing lyrics and vocal melodies directly in the generated audio. We encourage readers to listen to demo audio examples at https://team.doubao.com/seed-music .
Progress in the task of symbolic music generation may be lagging behind other tasks like audio and text generation, in part because of the scarcity of symbolic training data. In this paper, we leverage the greater scale of audio music data by applying pre-trained MIR models (for transcription, beat tracking, structure analysis, etc.) to extract symbolic events and encode them into token sequences. To the best of our knowledge, this work is the first to demonstrate the feasibility of training symbolic generation models solely from auto-transcribed audio data. Furthermore, to enhance the controllability of the trained model, we introduce SymPAC (Symbolic Music Language Model with Prompting And Constrained Generation), which is distinguished by using (a) prompt bars in encoding and (b) a technique called Constrained Generation via Finite State Machines (FSMs) during inference time. We show the flexibility and controllability of this approach, which may be critical in making music AI useful to creators and users.
Freezing of Gait (FOG) is an episodic lower extremity movement disorder that is highly susceptible to falls and carries a serious risk of disability. Monitoring of FOG can assist in the diagnosis and treatment of FOG. Providing appropriate gait guidance along with monitoring can help reduce the frequency and duration of freezing epi-sodes. This study aims to improve the robustness of the monitoring model using multimodal fusion methods. The gait signals from 32 FOG patients are collected by the inertial measurement unit (IMU) and force-sensitive insole (FSI) simultaneously. A multimodal fused FOG monitoring model was constructed by using deep neural networks to extract complementary features from IMU and FSI signals respectively, and feature-level fusion of the two modalities by an adaptive weighting method. Experimental results show that the proposed multimodal fusion approach improves the F1 value by 0.029 in the FOG detection task compared to the unimodal model. In addition, to construct the pre-FOG dataset more accurately, an automatic labeling method of pre-FOG events based on the FOG index ratio is also proposed in this paper. Compared to directly labeling the data 2.5 s before the freezing episode as the pre-FOG event, the proposed labeling method obtained more samples and improved the freezing prediction accuracy by 1.4 %.
In this paper, we introduce Jointist, an instrument-aware multi-instrument framework that is capable of transcribing, recognizing, and separating multiple musical instruments from an audio clip. Jointist consists of an instrument recognition module that conditions the other two modules: a transcription module that outputs instrument-specific piano rolls, and a source separation module that utilizes instrument information and transcription results. The joint training of the transcription and source separation modules serves to improve the performance of both tasks. The instrument module is optional and can be directly controlled by human users. This makes Jointist a flexible user-controllable framework. Our challenging problem formulation makes the model highly useful in the real world given that modern popular music typically consists of multiple instruments. Its novelty, however, necessitates a new perspective on how to evaluate such a model. In our experiments, we assess the proposed model from various aspects, providing a new evaluation perspective for multi-instrument transcription. Our subjective listening study shows that Jointist achieves state-of-the-art performance on popular music, outperforming existing multi-instrument transcription models such as MT3. We conducted experiments on several downstream tasks and found that the proposed method improved transcription by more than 1 percentage points (ppt.), source separation by 5 SDR, downbeat detection by 1.8 ppt., chord recognition by 1.4 ppt., and key estimation by 1.4 ppt., when utilizing transcription results obtained from Jointist. Demo available at \url{https://jointist.github.io/Demo}.
Uranium is an important nuclear element, and its efficient recovery is of great significance to nuclear industry and to nuclear safety. Herein, we rationally design robust antifouling fiber membranes (Anti-NH2-AO FMs) by facilely grafting two novel ligands on porous surfaces. The efficient ligands, i.e., NH2-AO and GSH, endow the membranes with rapid U-capture rate, good hydrophilicity and antifouling ability. Anti-NH2-AO FMs display porous fiber network structure, good flexibility, and high mechanical strength (9.0 MPa). The uranium capture rate reaches up to 167 mg g-1 h-1 in the first two hours, and they can reduce uranium to 11 ppb in 2 ppm uranium-contaminated water, which is below the drinking water limit. Furthermore, Anti-NH2-AO FMs still exhibit very high seawater permeability (6186 L m-2 h-1 bar -1) and flux recovery ratio (95.12 %) after three seawater fouling cycles. The imaging tests visually demonstrate their antiadhesion to the bacteria. In addition, they have good selectivity and long service life (10 cycles of adsorption-desorption). Kinetics models and XPS analyses illustrate the uranium adsorption mechanisms. Compared with the current membranes and fibers, the Anti-NH2-AO displays the advantages of exceptional properties and easy preparation, and we believe that it possesses great potential in the practical uranium recovery.
In this paper, we introduce Jointist, an instrument-aware multi-instrument framework that is capable of transcribing, recognizing, and separating multiple musical instruments from an audio clip. Jointist consists of the instrument recognition module that conditions the other modules: the transcription module that outputs instrument-specific piano rolls, and the source separation module that utilizes instrument information and transcription results. The instrument conditioning is designed for an explicit multi-instrument functionality while the connection between the transcription and source separation modules is for better transcription performance. Our challenging problem formulation makes the model highly useful in the real world given that modern popular music typically consists of multiple instruments. However, its novelty necessitates a new perspective on how to evaluate such a model. During the experiment, we assess the model from various aspects, providing a new evaluation perspective for multi-instrument transcription. We also argue that transcription models can be utilized as a preprocessing module for other music analysis tasks. In the experiment on several downstream tasks, the symbolic representation provided by our transcription model turned out to be helpful to spectrograms in solving downbeat detection, chord recognition, and key estimation.
Separating a song into vocal and accompaniment components is an active research topic, and recent years witnessed an increased performance from supervised training using deep learning techniques. We propose to apply the visual information corresponding to the singers' vocal activities to further improve the quality of the separated vocal signals. The video frontend model takes the input of mouth movement and fuses it into the feature embeddings of an audio-based separation framework. To facilitate the network to learn audiovisual correlation of singing activities, we add extra vocal signals irrelevant to the mouth movement to the audio mixture during training. We create two audiovisual singing performance datasets for training and evaluation, respectively, one curated from audition recordings on the Internet, and the other recorded in house. The proposed method outperforms audio-based methods in terms of separation quality on most test recordings. This advantage is especially pronounced when there are backing vocals in the accompaniment, which poses a great challenge for audio-only methods.
Plantar pressure has been put in use in clinical research for decades, such as in digital human modeling, biomechanics studies, and foot surgeries. Plantar pressure indicates the stress distribution of the foot to the ground, which also could reflect the state of the arch height. Abnormal arch height would incur lower limb imbalance problem, and would further cause joint disorders if was not properly treated as soon as possible, so a measurement of the severity of abnormal arch height is important. In this paper, we present a new framework of automatic annotation that using plantar pressure for index calculation of arch height, a new approach that could share knowledge from and to clinical studies. We addressed plantar pressure parsing problem as two separate tasks, landmark detection and semantic segmentation, and proposed Plantar Parsing U-Net like (PPU-Net) fully convolutional network for the tasks. Experiment results will show that we have achieved an excellent precision with proposed single model in landmark detection while keeping the comparable performance in palm area segmentation compared to baseline models.
Freezing of gait (FOG) is a common symptom in the late stage of Parkinson’s disease, and it is an important cause of falls in patients. In this study, we have designed a FOG monitoring and management system that can be used in the home environment based on wearable and telemedicine technology. The system uses force-sensitive insoles and inertial sensors to obtain the motion signals of patients, and segments gait cycle and detects FOG in real time in the elastic computing server. The monitoring results can be fed back to the patient to help them restore normal walking ability. Doctors can also know the frequency and degree of FOG attacks in each patient through the system, so as to provide more accurate medical evaluation and drug treatment.
Occlusion runs through the whole process of prevention, diagnosis, and treatment of oral diseases. However, occlusion can only be analyzed by current clinical methods qualitatively, and the dental occlusion cannot be recorded continuously. Firstly, a digital occlusal analysis system based on a flexible force-sensitive sensor is developed, including the sensor, hardware system, and software system. Then, a semi-automatic method was used to establish the dental arch model from the photographs of the occlusal surface. The dental arch model can be used to calculate the evaluating indicators of occlusal contact. Finally, the system is verified according to Bland&Altman. Fifteen adults attended this experiment. It shows that the 1.96-fold measurement error (accuracy) is 15.39%, and the 2.77-fold measurement error (variability) is 21.74%. Compared with the other commercial product, the accuracy and repeatability of the system meet the requirements.
Wind direction variation with height (wind veer) plays an essential role in the inflow wind field as the wind turbine enlarges. We explore the wind veer characteristics and their impact on turbine performance using a 5-year field dataset measured at the Eolos Wind Energy Research Station of the University of Minnesota. Wind veer exhibits an appreciable diurnal variation that veering and backing winds tend to occur during nighttime and daytime, respectively. We further propose to divide the wind veer conditions into four scenarios based on their changes in turbine upper and lower rotors that influence the loading on different rotor sections: VV (upper rotor: veering, lower rotor: veering), VB (upper rotor: veering, lower rotor: backing), BV (upper rotor: backing, lower rotor: veering), and BB (upper rotor: backing, lower rotor: backing). Such a division allows us to elucidate better the impact of wind veer on turbine power generation. The clockwise-rotating turbines tend to yield substantial power losses in scenarios VV and VB and small power gains in scenarios BV and BB. The counterclockwise-rotating turbines follow exactly opposite trends to the clockwise turbine. The derived findings are generalizable to other wind sites for power evaluation and provide insights into the turbine type selections targeting the maximum profits.
This article examines whether and how the firms’ mergers and acquisitions (M&A) policies are influenced by the risk preference of local community where the firms’ headquarters are located. By utilizing the different preferences towards risk-taking from the county-level religiosity-based measure, we document a significantly positive relation between the local risk preference and the likelihood of firms’ M&A. Firms whose headquarters are located in the counties with a higher degree of risk tolerance are more likely to engage in takeovers or acquire riskier targets. Local risk preference also results in wealth transfer during M&A that reduces acquirers’ equity value. When interacting with firm CEOs’ career concern and financial compensation, we find that managerial risk-taking incentive can be significantly affected by the local risk preference, suggesting an important economic interplay between the social norms and financial decision-making.
To withstand the complex natural environment and ecological system in the ocean, adsorbents with high strength, rapid rate, and antifouling ability are urgently needed for uranium extraction. Herein, instead of the traditional poly(amidoxime) structures, we design an anti-adhesive adsorbent (Anti-LS/SA) by conjugating functional ligands of salicylaldoxime (SA) and hyaluronic acid (HA) on the porous biomass via a facile one-step process. The interconnected microchannels in the porous biomass can accelerate seawater transport, along with the micromolecular uranium-adsorbing ligand, endowing the Anti-LS/SA with a very fast uranium adsorption ability; the average adsorption rates reach up to 131 +/- 1.8 mg g-1 hour- 1 (in uranium spiked seawater with concentration of 8 ppm) and 0.1176 mg g-1 day- 1 (in natural seawater) in the main adsorption stages, exceeding the currently used oxime-based adsorbents. Owing to the high hydration of HA, the adsorbents exhibit excellent hydrophilicity and anti-bacterial adhesion properties, as demonstrated by the bacteria adhesion testing and extraction experiments in bacteria-spiked seawater. Furthermore, the Anti-LS/SA displays a long service life and a very high tensile strength (62.63 MPa) to withstand strong ocean waves. Kinetics models and XPS data are used to analyze the extraction mechanisms. Considering its high efficiencies, structural advantages, as well as the facile and massive construction, we believe that the Anti-LS/SA will be a promising candidate for the large-scale uranium extraction from seawater.
The modified electroconvulsive therapy is an effective method for the treatment of severe depression and some other mental and nervous system disorders. In this paper, an electrical stimulation device for modified electroconvulsive therapy is designed. The device uses STM32F103ZET6 as the main control chip, and the constant current source circuit and H-bridge drive circuit are designed to output the stimulus pulse. The serial port screen is used as the human-computer interaction medium to set the stimulation parameters and control the output of the electrical stimulator. The operation is simple and the human-computer interaction is friendly. The experimental results show that the output pulse waveform of the electric stimulation device is good and the stimulation parameters reach the set target value.
Automatic music transcription (AMT) is the task of transcribing audio recordings into symbolic representations. Recently, neural network-based methods have been applied to AMT, and have achieved state-of-the-art results. However, many previous systems only detect the onset and offset of notes frame-wise, so the transcription resolution is limited to the frame hop size. There is a lack of research on using different strategies to encode onset and offset targets for training. In addition, previous AMT systems are sensitive to the misaligned onset and offset labels of audio recordings. Furthermore, there are limited researches on sustain pedal transcription on large-scale datasets. In this article, we propose a high-resolution AMT system trained by regressing precise onset and offset times of piano notes. At inference, we propose an algorithm to analytically calculate the precise onset and offset times of piano notes and pedal events. We show that our AMT system is robust to the misaligned onset and offset labels compared to previous systems. Our proposed system achieves an onset F1 of 96.72% on the MAESTRO dataset, outperforming previous onsets and frames system of 94.80%. Our system achieves a pedal onset F1 score of 91.86%, which is the first benchmark result on the MAESTRO dataset. We have released the source code and checkpoints of our work at https://github.com/bytedance/piano_transcription
When an AGV (Automated Guided Vehicle) performs navigation tasks, it needs to run the path planning algorithm to obtain an optimal path in a current environment. In this paper, Dijkstra algorithm and A*algorithm with different heuristic functions are applied to static environment modeling with various types of obstacles. To solve the problem that there are many redundant points and inflection points in the search process of the A*algorithm, an improved A*algorithm with Manhattan distance as a heuristic function is selected as the path planning algorithm. In addition, a calculation method of optimizing a past cost function is proposed, and the weight of heuristic function is optimized simultaneously. Simulation results show that the improved algorithm has a higher efficiency and less path inflection points than the traditional A*algorithm has.
•A facile, universal method is designed to construct robust anti-biofouling AO gels.•The gels display porosity, large surface area and excellent mechanical strength.•The gels exhibit very high uranium capture capacity and long service life.•Bactericidal assays and simulated seawater tests verify their antifouling ability.•They have good practical prospects also due to the massive and low-cost production.
For Parkinson's disease, a degenerative neurological disease that is difficult to detect and has a high rate of misdiagnosis, telemedicine technology can undoubtedly bring great improvement to the real limitations of delayed follow-up diagnosis and unsustainable control of the disease. This paper takes the symptom of Parkinson's Dyskinesia as the focus of attention and designs a complete set of full-process available system which is based on home telemedicine and wearable technology. Finally, a system consists of intelligent hardware facilities, user terminals, and management terminals will be presented to construct a monitoring service scenario where the Parkinson's Dyskinesia is continuously controllable.
The realization of mobile robots' autonomous positioning and map constructing in unknown environments is crucial for the robots' obstacle avoidance and path planning. In this paper, an improved ORB (Oriented fast and Rotated Brief)-SLAM2 (Simultaneous Localization And Mapping 2) algorithm is used to construct a 3D (Three Dimensional) point cloud map of the robot's own positioning and environment. The improved ORB-SLAM2 algorithm is schemed as follows: firstly, after the environment map constructions, it adds the function of saving maps to help implementing map type conversion and navigation obstacle avoidance. Then we employ a PCL (Point Cloud Library) to convert the saved 3D point cloud map into an octomap. A path planning algorithm for mobile robots is implemented on the basis of the octomaps. The robot's dynamical global path planning is implemented using a RRT (Rapidly-exploring Random Tree) algorithm. The experimental results of map constructing and path planning show that the scheme proposed in this paper can effectively realize the obstacle avoidance and path planning of the mobile robot. Thus, the algorithm provides a basis for the further realizing the mobile robot' autonomous movement.
Symbolic music datasets are important for music information retrieval and musical analysis. However, there is a lack of large-scale symbolic datasets for classical piano music. In this article, we create a GiantMIDI-Piano (GP) dataset containing 38,700,838 transcribed notes and 10,855 unique solo piano works composed by 2,786 composers. We extract the names of music works and the names of composers from the International Music Score Library Project (IMSLP). We search and download their corresponding audio recordings from the internet. We further create a curated subset containing 7,236 works composed by 1,787 composers by constraining the titles of downloaded audio recordings containing the surnames of composers. We apply a convolutional neural network to detect solo piano works. Then, we transcribe those solo piano recordings into Musical Instrument Digital Interface (MIDI) files using a high-resolution piano transcription system. Each transcribed MIDI file contains the onset, offset, pitch, and velocity attributes of piano notes and pedals. GiantMIDI-Piano includes 90% live performance MIDI files and 10\% sequence input MIDI files. We analyse the statistics of GiantMIDI-Piano and show pitch class, interval, trichord, and tetrachord frequencies of six composers from different eras to show that GiantMIDI-Piano can be used for musical analysis. We evaluate the quality of GiantMIDI-Piano in terms of solo piano detection F1 scores, metadata accuracy, and transcription error rates. We release the source code for acquiring the GiantMIDI-Piano dataset at https://github.com/bytedance/GiantMIDI-Piano