Large-scale multi-label text classification (LMTC) aims to associate a document with its relevant labels from a large candidate set. Most existing LMTC approaches rely on massive human-annotated training data, which are often costly to obtain and suffer from a long-tailed label distribution (i.e., many labels occur only a few times in the training set). In this paper, we study LMTC under the zero-shot setting, which does not require any annotated documents with labels and only relies on label surface names and descriptions. To train a classifier that calculates the similarity score between a document and a label, we propose a novel metadata-induced contrastive learning (MICoL) method. Different from previous text-based contrastive learning techniques, MICoL exploits document metadata (e.g., authors, venues, and references of research papers), which are widely available on the Web, to derive similar document–document pairs. Experimental results on two large-scale datasets show that: (1) MICoL significantly outperforms strong zero-shot text classification and contrastive learning baselines; (2) MICoL is on par with the state-of-the-art supervised metadata-aware LMTC method trained on 10K–200K labeled documents; and (3) MICoL tends to predict more infrequent labels than supervised methods, thus alleviates the deteriorated performance on long-tailed labels.
Scientific knowledge is evolving at an unprecedented rate of speed, with new concepts constantly being introduced from millions of academic articles published every month. In this paper, we introduce a self-supervised end-to-end system, SciConceptMiner, for the automatic capture of emerging scientific concepts from both independent knowledge sources (semi-structured data) and academic publications (unstructured documents). First, we adopt a BERT-based sequence labeling model to predict candidate concept phrases with self-supervision data. Then, we incorporate rich Web content for synonym detection and concept selection via a web search API. This two-stage approach achieves highly accurate (94.7%) concept identification with more than 740K scientific concepts. These concepts are deployed in the Microsoft Academic production system and are the backbone for its semantic search capability.
We focus on a recently deployed system built for summarizing academic articles by concept tagging. The system has shown great coverage and high accuracy of concept identification which could be contributed by the knowledge acquired from millions of publications. Provided with the interpretable concepts and knowledge encoded in a pre-trained neural model, we investigate whether the tagged concepts can be applied to a broader class of applications. We propose transforming the tagged concepts into sparse vectors as representations of academic documents. The effectiveness of the representations is analyzed theoretically by a proposed framework. We also empirically show that the representations can have advantages on academic topic discovery and paper recommendation. On these applications, we reveal that the knowledge encoded in the tagging system can be effectively utilized and can help infer additional features from data with limited information.
The whole globe has cranked up for coping with the COVID-19 situation. The hands-on tutorial targets at providing a comprehensive and pragmatic end-to-end walk-through for building an academic research paper recommender for the use case of COVID-19 related study, with the help of knowledge graph technology. The code examples that demonstrate the theories are reproducible and can hopefully provide value for researchers to build tools that support conducting research to find a cure to COVID-19.
An ongoing project explores the extent to which artificial intelligence (AI), specifically in the areas of natural language processing and semantic reasoning, can be exploited to facilitate the studies of science by deploying software agents equipped with natural language understanding capabilities to read scholarly publications on the web. The knowledge extracted by these AI agents is organized into a heterogeneous graph, called Microsoft Academic Graph (MAG), where the nodes and the edges represent the entities engaging in scholarly communications and the relationships among them, respectively. The frequently updated data set and a few software tools central to the underlying AI components are distributed under an open data license for research and commercial applications. This paper describes the design, schema, and technical and business motivations behind MAG and elaborates how MAG can be used in analytics, search, and recommendation scenarios. How AI plays an important role in avoiding various biases and human induced errors in other data sets and how the technologies can be further improved in the future are also discussed.
On the behest of the Office of Science and Technology Policy in the White House, six institutions, including ours, have created an open research dataset called COVID-19 Research Dataset (CORD-19) to facilitate the development of question-answering systems that can assist researchers in finding relevant research on COVID-19. As of May 27, 2020, CORD-19 includes more than 100,000 open access publications from major publishers and PubMed as well as preprint articles deposited into medRxiv, bioRxiv, and arXiv. Recent years, however, have also seen question-answering and other machine learning systems exhibit harmful behaviors to humans due to biases in the training data. It is imperative and only ethical for modern scientists to be vigilant in inspecting and be prepared to mitigate the potential biases when working with any datasets. This article describes a framework to examine biases in scientific document collections like CORD-19 by comparing their properties with those derived from the citation behaviors of the entire scientific community. In total, three expanded sets are created for the analyses: 1) the enclosure set CORD-19E composed of CORD-19 articles and their references and citations, mirroring the methodology used in the renowned “A Century of Physics” analysis; 2) the full closure graph CORD-19C that recursively includes references starting with CORD-19; and 3) the inflection closure CORD-19I, that is, a much smaller subset of CORD-19C but already appropriate for statistical analysis based on the theory of the scale-free nature of the citation network. Taken together, all these expanded datasets show much smoother trends when used to analyze global COVID-19 research. The results suggest that while CORD-19 exhibits a strong tilt toward recent and topically focused articles, the knowledge being explored to attack the pandemic encompasses a much longer time span and is very interdisciplinary. A question-answering system with such expanded scope of knowledge may perform better in understanding the literature and answering related questions. However, while CORD-19 appears to have topical coverage biases compared to the expanded sets, the collaboration patterns, especially in terms of team sizes and geographical distributions, are captured very well already in CORD-19 as the raw statistics and trends agree with those from larger datasets.
Since the relaunch of Microsoft Academic Services (MAS) 4 years ago, scholarly communications have undergone dramatic changes: more ideas are being exchanged online, more authors are sharing their data, and more software tools used to make discoveries and reproduce the results are being distributed openly. The sheer amount of information available is overwhelming for individual humans to keep up and digest. In the meantime, artificial intelligence (AI) technologies have made great strides and the cost of computing has plummeted to the extent that it has become practical to employ intelligent agents to comprehensively collect and analyze scholarly communications. MAS is one such effort and this paper describes its recent progresses since the last disclosure. As there are plenty of independent studies affirming the effectiveness of MAS, this paper focuses on the use of three key AI technologies that underlies its prowess in capturing scholarly communications with adequate quality and broad coverage: (1) natural language understanding in extracting factoids from individual articles at the web scale, (2) knowledge assisted inference and reasoning in assembling the factoids into a knowledge graph, and (3) a reinforcement learning approach to assessing scholarly importance for entities participating in scholarly communications, called the saliency, that serves both as an analytic and a predictive metric in MAS. These elements enhance the capabilities of MAS in supporting the studies of science of science based on the GOTO principle, i.e., good and open data with transparent and objective methodologies. The current direction of development and how to access the regularly updated data and tools from MAS, including the knowledge graph, a REST API and a website, are also described.
In modern web-scale applications that collect data from different sources, entity conflation is a challenging task due to various data quality issues. In this paper, we propose a robust and distributed framework to perform conflation on noisy data in the Microsoft Academic Service dataset. Our framework contains two major components. In the offline component, we train a GBDT model to determine whether two papers from different sources should be conflated to the same paper entity. In the online component, we propose a scalable shingling algorithm that can apply our offline model to over 100 million instances. The result shows that our algorithm can conflate noisy data robustly and efficiently.
In this research, conducting poly(3,4-ethylenedioxythiophene)-poly(styrenesulfonic acid) (PEDOT:PSS) aqueous dispersion was synthesized at first via chemical oxidative polymerization and followed by mixing it with poly(styrene-r-butyl acrylate) P(St-BA) aqueous latex, creating a conductive material with outstanding stretchability. The elastic conductive composite were then film formed on the glass and poly(ethylene terephthalate) (PET) nonwoven fabric substrate by spin coating and dip coating, respectively. Composite films with various contents of PEDOT:PSS polymer (10-100 wt.%) had been prepared. From the conductivity measurements, the conductivity was still kept as high as 88 S cm(-1) even the PEDOT:PSS content was lowered to 10 wt.%. Furthermore, the elasticity of conductive films on the PET-onwoven fabric substrate was evaluated by the 180 degrees bending test repeating 100 times. With introducing soft P(St-BA) material in the PEDOT:PSS phase, the surface resistance increased merely 3-6 times after bending 100 times, while the surface resistance for pure PEDOT:PSS film could reach 18-20 times. (C) 2013 Elsevier B.V. All rights reserved.
In this study, a highly conductive poly(3,4-ethylenedioxythiophene)–poly(styrenesulfonic acid) (PEDOT:PSS) dispersion served as a stabilizer for producing conductive PEDOT:PSS–poly(styrene-co-butyl acrylate) (PEDOT:PSS–P(St–BA)) composite latexes by emulsion polymerization. Furthermore, soft latex particles of poly(styrene-co-butyl acrylate), P(St–BA), were synthesized via emulsion polymerization and then mixed with the conductive PEDOT:PSS–P(St–BA), followed by casting on the substrate. After drying, conductive composite films with flexibility and transparency could be obtained. The particle size and morphology of PEDOT:PSS–P(St–BA) were observed by a scanning electron microscope. The film surface resistance, thickness, conductivity and transmittance of the composite films were measured and studied. Bending tests were also conducted to detect the flexibility. According to the results, the composite films showed high transmittance while possessing good conductivity and superior flexibility. The method may serve as a practical approach to fabricate conductive, flexible and transparent films.
We present a Compressive Sensing (CS) based Client-Cloud system describing the future Cloud Computing structure, which takes 3D depth reconstruction as an instance. First, a sparse representation for continuous depth data is exploited. Second, we propose a dynamic measurement generation method adapted to the variation of sparsity to reduce bandwidth requirements. Third, a feedback correction scheme is developed to detect the incorrectly reconstructed signals and perform supplementary reconstruction. According to the experimental results, the proposed system can reduce about 40% to 70% bandwidth requirements and lower about 50% error rate while reconstruction.
In this research, poly(3,4-ethylenedioxythiophene) (PEDOT) nanoparticles less than 100 nm were synthesized first and applied as the solid stabilizer for producing PEDOT-polystyrene (PEDOT-PSt) composite latex by Pickering emulsion polymerization. The results showed that most PEDOT nanoparticles adhered to the PSt core particles having the size from 100 to 250 nm. By casting the latex, the obtained PEDOT-PSt film had a surface resistance of 4-5 k Omega/square, almost the same as the pure PEDOT film, though its PEDOT content was only 6.2 wt%. Those PEDOT nanoparticles in the outer layer could contact with one another, forming a continuous network as the conductive passageway. Furthermore, soft latex particles of poly(styrene-co-butyl acrylate), P(St-BA), were synthesized and mixed with the conducting rigid PEDOT-PSt latex for improving the toughness and transparency of the casting film. A critical point at about 2 wt% of PEDOT content in the PEDOT-PSt/P(St-BA) film was observed in the surface resistance measurement. (C) 2012 Elsevier Ltd. All rights reserved.
Innovative elastic and flexible conductive composite materials PEDOT:PSS/P(BA-St) and PEDOT:PSS-PBA were prepared by two approaches based on poly(3,4-ethylenedioxythiophene):poly(styrenesulfonate) (PEDOT:PSS). In the first part, PEDOT:PSS/P(BA-St) was prepared by blending various soft poly(n-butyl acrylate-styrene) (P(BA-St)) latexes into a PEDOT:PSS conductive dispersion. The PEDOT:PSS conductive dispersion was prepared via oxidative polymerization of EDOT, and the soft P(BA-St) latexes were synthesized by emulsion polymerization using various types of surfactants. In the second part, poly(styrenesulfonate) (PSS) served as a surfactant to synthesize poly(styrenesulfonate)-poly(butyl acrylate) (PSS-PBA) soft latex via emulsion polymerization. Then, the PEDOT:PSS-PBA dispersion was prepared by using the PSS-PBA soft latex as a polymeric template to polymerize EDOT. The glass transition temperature (T-g) and particle size of P(BA-St) and PSS-PBA were measured. The optoelectronic properties of the conductive PEDOT:PSS/P(BA-St) and PEDOT:PSS-PBA films such as transmittance, surface resistance, and UV-vis absorbance were investigated. According to the elongation measurement, the fabricated PEDOT:PSS/P(BA-St) elastic film containing a P(BA-St) content as high as 83 wt% showed a high value of 97% elongation, while possessing good film conductivity (63 S cm(-1)) and superior transmittance (93%). On the other hand, the fabricated PEDOT: PSS-PBA flexible film containing 13 wt% PBA showed a low surface resistance increment (R/R-0 < 1.2) after two steps of bending test, while maintaining superior conductivity (300 S cm(-1)) and high light transmittance (>80%).
In this research, poly(3,4-ethylenedioxythiophene) (PEDOT) latex nanoparticles with good colloidal stability were prepared by emulsion polymerization and the conversions of EDOT were determined. Two kinds of oxidants, Iron(III) p-toluenesulfonate Fe(OTs)(3) and hydrogen peroxide H2O2, were introduced to decrease the use of iron salt and therefore reduce the particle coagulation. The ferrous ions (Fe2+) produced during the polymerization would be re-oxidized back to the reactive ferric ions (Fe3+) with the help of H2O2. This cyclic oxidation reduction process resulted in the sustained regeneration of Fe3+ ions and led to a higher conversion. A dark blue PEDOT latex with long-term dispersion stability was obtained when Fe (OTs)(3) and H2O2 were added in sequence. The results obtained from dynamic light scattering and TEM measurements showed that the sizes of nanoparticles were around 100 nm. To determine the conversion of EDOT, two methods (gravimetric analysis and UV-visible method) were used and compared. For the first time, the UV-visible method was established to quantitatively determine the conversion of EDOT monomer. From the measurement, the conversion of EDOT in this system was determined as 74-75%. The PEDOT film prepared by drying the latex solution had conductivity up to 6.3 S/cm. (C) 2011 Elsevier Ltd. All rights reserved.
Electroluminescent (EL) phenomenon has been widely used in the illumination industry such as light emitting diode (LED) and organic light emitting diode (OLED) nowadays. In this research, flexible EL devices were fabricated by using the flexible substrates and conductive polymer poly(3,4-ethylenedioxythiophene):poly(styrenesulfonate) (PEDOT:PSS) dispersions as the anode materials. The PEDOT:PSS dispersion was firstly prepared by oxidative polymerization while PSS was served as a polymeric template. Via introducing poly(vinyl pyrrolidone) (PVP) ethanol solution into the PEDOT:PSS dispersion, viscosity and wetting ability of the mixed PEDOT:PSS/PVP solution were enhanced in the same time. The three formed PEDOT:PSS/PVP thin films on glass substrates presented good conductivity and high transmittance (transmittance >76%). The wetting abilities of these PEDOT:PSS/PVP dispersions on the PET films coated with the low dielectric resin were evaluated from the contact angle measurements. It showed that the contact angle was decreased with increasing the addition of PVP ethanol solution. The different PEDOT:PSS/PVP dispersions were further used as the anode material in the EL flat plates and EL wires by dip coating process. Uniform and smooth conductive thin films with high transmittance were formed on the low dielectric resin layer. The luminance of these EL devices was observed and photographed after connecting to 110V AC power.
Andreas Argyriou合作论文数Ecole Centrale Paris1