Clinical event extraction is crucial for structuring medical data, supporting clinical decision-making, and enabling other intelligent healthcare services. Traditional approaches for clinical event extraction often use pipeline-based methods to identify event triggers and elements. However, these methods commonly suffer from error propagation and information loss, leading to suboptimal performance. To address this challenge, this paper proposes an end-to-end clinical event extraction method based on the large language models (LLMs). Specifically, we transform the clinical event extraction task into an end-to-end text generation task and design a prompt learning method based on the LLMs called LMCEE. Experimental results demonstrate a significant improvement over traditional pipeline methods, with the F1 score increasing by 12%. Additionally, the proposed method outperforms the generative-based method named UIE, showcasing a 5.7% improvement in F1 score. However, the experimental results also disclose certain limitations of the proposed method, such as its sensitivity to prompt templates and its heavy dependence on the type of LLMs. These findings highlight the need for further investigation and optimization to enhance performance and robustness.
A variety of sludge originating from wastewater treatment is increasingly accumulating worldwide, especially in the developed regions. Fortunately, phosphorus (P) in sludge can potentially be recovered via thermal treatment, and P forms are crucial to the quality of the recovered products. This review provides an overview related to the transformation of P and its regulation strategies in the different thermal treatment processes, highlighting the current knowledge regarding the influence factors and mechanisms of transformation. First, the current re-searches with P recovery by thermal treatments (i.e., incineration, pyrolysis, and hydrothermal carbonization) were illustrated and critically discussed, which indicates that the selection of appropriate method significantly affects the recovery effects, such as the proportion of Apatite Phosphorus (AP) or Ca-P (P combing with Ca). Furthermore, the effects of various factors (e.g., treatment techniques, temperature, additives, the composition of sludge, pH and metals ions) on the transformation or regulation of P forms were critically discussed, suggesting that temperature and Ca/Cl-based additives are the two crucial factors, which have the considerable effects on the emergence of Ca-P. Especially the Ca-based additives, which have the highest combining ability with P and the uppermost impact on hydroxyapatite formation. In addition, an overall comparison of incineration, pyrolysis, and hydrothermal carbonization technologies in practical applications is provided, focusing on the recovery efficiency, advantages, and disadvantages, demonstrating that HTC will be more predominant for P recovery from sludge in the future. Finally, challenges and potential development directions regarding the transformation mechanisms of P forms and regulate strategies are identified. This review will become the foundation for future research devoted to improving the quality of the recovered products.
Electronic medical record named entity recognition can extract important clinical information from unstructured text, which is helpful for clinical diagnosis and medical decision-making. However, due to the particularity of the medical field, it is difficult for researchers to obtain sufficient labeled electronic medical records. Models trained using traditional supervised learning methods with insufficient data are not promising. To solve this problem, this paper proposes two weakly supervised learning methods, sampling-based active learning and parameter-based transfer learning, to achieve better performance. In sampling-based active learning, two uncertainty sampling strategies, least confidence sampling and entropy sampling, are used to select data from unlabeled dataset for retraining. In parameter-based transfer learning, the parameters of word representation layer and encoding layer in the source domain are initialized to the corresponding layer of the target domain, and the objective is to learn generalized linguistic knowledge from the source domain. Finally, we use a voting mechanism to ensemble these individual models to get better prediction results. Experiment on the CCKS2017 official test set shows that our system for MER achieves 0.8972 F1 score and gets better performance than the supervised methods, which obtains 0.8921 F1 score and proves the effectiveness of our approaches. The experimental results show that the weakly supervised learning methods proposed in this paper achieve the satisfactory performance as the supervised methods under comparable conditions.
Catalytic hydrodeoxygenation (HDO) of oxygenated compounds under mild conditions is essential for chemical refining, which is always relied on improving hydrogenation activity or/and acid strength of catalytic system. Here, with the synthesis of functional TiO 2 (B) nanosheet, we developed a novel Pt/TiO 2 (B) catalyst for HDO of phenol. Through the synergistic effect of dispersed Pt isolated single atoms and nanoclusters, phenol was efficiently converted into cyclohexane without extra acid under 50 o C and atmospheric hydrogen. The catalytic diversity of two Pt species for C=O/C=C bonds leads to the formation of ketene/enol intermediates, as detected in the in-situ FT-IR tests. Due to the low C-O bond energy of enol intermediates, the HDO of phenol could be carried out at mild conditions through the implementable C-O bond cleavage of enol intermediates. This study provides a new understanding of green HDO process and a novel approach for HDO catalysts developing.
Automatic extraction of clinical named entities, such as body parts, drugs and surgeries, has been of great significance to understand clinical texts. Deep neural networks approaches have achieved remarkable success in named entity recognition task recently. However, most of these approaches train models from large, high-quality and labor-consuming labeled data. In order to reduce the labeling costs, we propose a weakly supervised learning method for clinical named entity recognition (CNER) tasks. We use a small amount of labeled data as seed corpus, and propose a bootstrapping method integrating external knowledge to iteratively generate the labels for unlabeled data. The external knowledge consists of domain specific dictionaries as well as a bunch of handcraft rules. We conduct experiments on CCKS-2018 CNER task dataset and our approach achieves competitive results comparing to the supervised approach with fully labeled data.
As a shining pearl in traditional Tibetan culture, historical Tibetan documents have received extensive attention from historians, linguists and Buddhist scholars. These documents are converted into digital form using Tibetan document segmentation and recognition methods. The document digitization is of great significance for the research, protection and inheritance of Tibetan history. This paper proposes an overall segmentation and recognition framework for historical Tibetan document images. Firstly, the historical Tibetan document image is preprocessed to correct imbalanced illumination, tilt and noises, and is further transformed into the binarized image. Secondly, we propose a layout segmentation method based on block projection to segment Tibetan document images into texts, lines and frames. Thirdly, in order to solve the problems of touching strokes between text-lines and curvilinear text-lines, we present a text-line segmentation method based on graph model for historical Tibetan text-line segmentation. Lastly, we present a touching segmentation method to segment touching Tibetan character string, and then recognize Tibetan characters. Experimental results show our proposed methods on layout segmentation, text-line segmentation and touching character string segmentation, achieve the satisfactory performance. The proposed methods can also be applied to other fonts in Tibetan font family.
As a tool to express common semantics of objects, language can be used to describe the attributes and locations of objects within the scope of human vision. Searching for the location of an object in the field of vision through natural language is an important capability of the human. Proposing a mechanism to learn this ability of human is a major challenge for computer vision. Most existing object localization methods usually use strong supervised information of the training set to train the model. However, these models lack interpretability and require expensive labels which are difficult to obtain. Facing these challenges, we propose a new method for locating object by natural language descriptions for fine-grained image. Firstly, we propose a model that can learn the semantically relevant parts between fine-grained images and languages, and achieve ideal localization accuracy without using strong supervisory signal. In addition, we have improved the contrast loss function to make natural language descriptions better match target regions of fine-grained images.The multi-scale fusion techniques are utilized to improve the ability of capturing details on fine-grained images. Comprehensive experiments demonstrate that the proposed method achieves ideal localization results on the CUB200-2011 dataset. And the proposed model has strong zero-shot learning ability on untrained data.
Pyrolysis experiments of sawdust with KOH and K2CO3 catalysts were carried out under different heating rate in nitrogen atmosphere using thermogravimetric analyzer. The distributed activation energy model (DAEM) was used to analyze pyrolysis kinetics of sawdust. The results showed that both KOH and K2CO3 had strong catalytic effect on sawdust pyrolysis, which reduced the pyrolysis temperature of sawdust and increased the yield of char. There was only one main peak in DTG curve, which means that the pyrolysis behavior of cellulose and hemicellulose in sawdust was greatly changed. The catalytic performance of KOH was found to be more excellent in sawdust pyrolysis. Also KOH could catalyze the pyrolysis of sawdust at low temperature. The kinetic analysis results showed that the two kinds of catalysts could reduce the activation energy of sawdust pyrolysis and maintain a similar catalytic trend, but KOH had a more stable catalytic performance.
In order to investigate the effect of potassium carbonate on biomass pyrolysis properties, sawdust was used as raw material and different amounts of K2CO3 were added by impregnation method to carry out thermogravimetric and pyrolysis experiments. The effects of pyrolysis temperature and the amount of K2CO3 addition on the pyrolysis of sawdust were studied using a self-made fixed-bed pyrolysis furnace. Calculation of pyrolysis kinetics shows that the existence of K2CO3 catalyst changes the pyrolysis path of sawdust, so that the activation energy of pyrolysis sawdust decreases at low temperature and increases at high temperature. The pyrolysis experiments shows that the addition of K2CO3 and the increase of pyrolysis temperature both reduce the yield of the pyrolysis oil of sawdust and increase the yield of the pyrolysis syngas. However, K2CO3 catalyst promotes the yield of char, the increase of pyrolysis temperature decreases the yield of char. Analysis of the pyrolysis products finds that the addition of K2CO3 and the increase of pyrolysis temperature both improve quality of the pyrolysis oil, form more microporous surface of char, and increase the hydrogen content in the pyrolysis syngas. It is considered that the optimal process for producing pyrolysis syngas is 900 degrees C of pyrolysis temperature and 10% of K2CO3 addition. (C) 2018 Hydrogen Energy Publications LLC. Published by Elsevier Ltd. All rights reserved.
The digitalization of historical documents attract increasing research interests in recent years.Focusing on layout analysis,the essential step in digitizing historical documents,this paper proposes a convolutional denoising auto-encoder approach to historical Tibetan documents.Firstly,the document images are clustered into superpixel blocks.Then,we use the convolutional autoencoder to extract features from these blocks.Finally,the superpixel blocks are classified by the SVM classifier,thus the different parts of the Tibetan historical document are identified.Experiments on the dataset of historical Tibetan documents show that our method can effectively separate the different layout elements of Tibetan historical documents.
Text extraction is an important initial step in digitizing the historical documents. In this paper, we present a text extraction method for historical Tibetan document images based on block projections. The task of text extraction is considered as text area detection and location problem. The images are divided equally into blocks and the blocks are filtered by the information of the categories of connected components and corner point density. By analyzing the filtered blocks’ projections, the approximate text areas can be located, and the text regions are extracted. Experiments on the dataset of historical Tibetan documents demonstrate the effectiveness of the proposed method.
Text-line segmentation is an important task in the historical Tibetan document recognition. Historical Tibetan document images usually contain touching or overlapping characters between consecutive text-lines, making text-line segmentation a difficult task. In this paper, we present a text-line segmentation method based on baseline detection. The initial positions for the baseline of each line are obtained by template matching, pruning algorithms and closing operation. The baseline is estimated using dynamic tracing within pixel points of each line and the context information between pixel points. The overlapping or touching areas are cut by finding the minimum width stroke. Finally, text-lines are extracted based on the estimated baseline and the cut position of touching area. The proposed algorithm has been evaluated on the dataset of historical Tibetan document images. Experimental result shows the effectiveness of the proposed method.
Because of the conglutinated characteristic of Mongolian words, it's difficult to realize online handwritten Mongolian word recognition with high recognition accuracy based on segmentation-based strategy. Meanwhile, as the vocabulary of Mongolian words is large, using a segmentation-free method with deep bidirectional long short term memory(DBLSTM) network is more suitable. We design a 5 bidirectional hidden level DBLSTM network for online handwritten Mongolian word recognition. This paper mainly proposes a novel sliding window method which selects frames with different intervals to enhance recognition rate. The novel method can generate hundreds of sequence data for each sample, while only one sequence data is generated using ordinary sliding window method. More sequence data and more abundant sequence information are helpful to raise the recognition rate. We evaluated the recognition performance on our online handwritten Mongolian database with 925 classes. The proposed method achieves the word level recognition rate of 89.24% with PCA feature extractor and best path decoding, compared to that of 88.45% using ordinary sliding window method. Further, several well trained DBLSTM models based on the proposed method are combined to vote the output, finally, the word-level recognition raises to 90.35%.
In this paper, we present a text extraction method for historical Tibetan document images. The task of text extraction is considered as text area detection and location problem. Firstly, the historical Tibetan document image is preprocessed to correct imbalanced illumination, tilt and noises, then get the binary image. Secondly, the regions of interest in historical Tibetan documents are divided into three categories using connected components. The images are divided equally into grids and the grids are filtered by the information of the categories of CCs and corner point density. The remaining grids are used to compute vertical and horizontal grid projections. Thirdly, by analyzing the projections, the approximate location of the text area can be detected. Finally, the text area is extracted accurately by correcting the bounding box of the approximate text area. Experiments on the dataset of historical Tibetan document images demonstrate the effectiveness of the proposed method.
A novel hierarchical hollow structured bismuth oxychloride (BiOCl) microspheres assembled by thin BiOCl nanosheets with active (001) facets were fabricated via facile hydrothermal route. The composition, morphology, and physical properties of as-prepared products were characterized by XRD, XPS, SEM, TEM, BET, and UV-vis spectroscopy, etc. The results indicated that morphology and microstructure of the BiOCl could be controlled by the surfactant poly(vinyl pyrrolidone) (PVP). With the presence of PVP, the evolution from solid to hollow microspheres has been easily achieved; moreover, the thickness of assembled nanosheets became thinner from 18 to 11 nm. The specific surface area of the BiOCl hollow microspheres reached to be 27.50 m(2)/g. In addition, the photocatalytic ability of the BiOCl microspheres were investigated, and experimental results showed that the hollow microspheres exhibited excellent catalytic activity for the RhB degradation with high photostability and rapid degradation rate of 0.1072 min(-1) under visible light irradiation, which were much higher than that of the commercial P25, BiOCl solid microspheres, the BiOCl semiconductors and even the BiOCl hierarchical microspheres in the literatures. Initial mechanism for the degradation of RhB over the BiOCl hollow microspheres under visible light irradiation was mainly attributed to the dye self-sensitization process.
K E Y W O R D S online handwritten Tibetan character recognition, de-noising, pre-processing, three-stage classification This is a contribution from Himalayan Linguistics, Vol. 15(1): 31–40. ISSN 1544-7502 © 2016. All rights reserved. This Portable Document Format (PDF) file may not be altered in any way. Tables of contents, abstracts, and submission guidelines are available at escholarship.org/uc/himalayanlinguistics Himalayan Linguistics, Vol. 15(1). © Himalayan Linguistics 2016 ISSN 1544-7502 31 Online unconstrained handwritten Tibetan character recognition using statistical recognition Long-Long Ma Jian Wu Chinese Academy of Sciences
This paper describes a recognition system for online handwritten Tibetan characters using advanced techniques in character recognition. To eliminate noise points of handwriting trajectories, we introduce a de-noising approach by using dilation, erosion, thinning operators of mathematical morphology. Selecting appropriate structuring elements, we can clear up large amounts of noises in the glyphs of the character. To enhance the recognition performance, we adopt a three-stage classification strategy, where the top rank output classes by the baseline classifier are re-classified by similarcharacter discrimination classifier. Experiments have been carried out on two databases MRG-OHTC and IIP-OHTC. Test results show the used recognition algorithm is effective and can be applied to pen-based mobile devices.
A new online handwritten Mongolian word database, MRG-OHMW, is introduced in this paper. This database contains 946 Mongolian words produced by 300 persons from Mongolian ethnic minority. These Mongolian words are composed of one to fourteen Mongolian characters, and selected from large-scale Mongolian text corpus according to the frequencies of usage. The current version of this database is collected using Anoto pen on paper. The database is further annotated using Mongolian word-level string alignment strategy. We partition the samples into training and test sets, and evaluate the database using the CNN-based recognizer as a baseline. Experimental results reveal a big challenge to higher recognition performance. To our knowledge, MRG-OHMW is the first publicly available database for online handwritten Mongolian research. It provides a basic database to compare empirically different algorithms for online handwritten Mongolian word recognition.
The co-combustion of pulverised coal and biomass is increasingly being used for environmental reasons, and a number of computational fluid dynamic investigations are being undertaken to understand the details of the combustion process. These investigations assume that the particle flow entering the burner is uniformly distributed across the burner mouth or inlet. In this paper, this assumption is examined for an industrial burner by numerically simulating the fuel particle flows in the tube leading to the burner mouth. While there is evidence of maldistribution of the particles at the burner mouth, it is concluded from the flame data that this effect does not significantly influence the combustion flame in the furnace for the cases investigated.
This paper presents a new component-based recognition method using conditional random field (CRF) for on-line handwritten Tibetan characters. The character pattern is over-segmented into a sequence of sub-structure blocks. Integrated segmentation and recognition method based on the CRF model is used to determine the component segmentation points from these block sequences. The CRF model combines component shape likelihood with geometrical likelihood. The parameters are learned using an energy minimization method. We build a component-based spelling rule model to ensure the correct component appearing at a specific structural position. A character-component generation model is presented to reduce component recognition error rate and accelerate the recognition process. Experimental results on MRG-OHTC database show that the proposed method gives promising performance comparing with the holistic method and the component-based conventional path evaluation method.