The development of computational data science techniques in natural language processing (NLP) and machine learning (ML) algorithms to analyze large and complex textual information opens new avenues to study intricate policy processes at a scale unimaginable even a few years ago. We apply these scalable NLP and ML techniques to analyze the United States Government's regulation of the banking and financial services sector. First, we employ NLP techniques to convert the text of financial regulation laws into feature vectors and infer representative "topics" across all the laws. Second, we apply ML algorithms to the feature vectors to predict various attributes of each law, focusing on the amount of authority delegated to regulators. Lastly, we compare the power of alternative models in predicting regulators' discretion to oversee financial markets. These methods allow us to efficiently process large amounts of documents and represent the text of the laws in feature vectors, taking into account words, phrases, syntax, and semantics. The vectors can be paired with predefined policy features, thereby enabling us to build better predictive measures of financial sector regulation. The analysis offers policymakers and the business community alike a tool to automatically score policy features of financial regulation laws to and measure their impact on market performance.
The development of computational data science techniques in natural language processing and machine learning algorithms to analyze large and complex textual information opens new avenues for studying the interaction between economics and politics. We apply these techniques to analyze the design of financial regulatory structure in the United States since 1950. The analysis focuses on the delegation of discretionary authority to regulatory agencies in promulgating, implementing, and enforcing financial sector laws and overseeing compliance with them. Combining traditional studies with the new machine learning approaches enables us to go beyond the limitations of both methods and offer a more precise interpretation of the determinants of financial regulatory structure.
The development of computational data science techniques in natural language processing (NLP) and machine learning (ML) algorithms to analyze large and complex textual information opens new avenues to study intricate processes, such as government regulation of financial markets, at a scale unimaginable even a few years ago. This paper develops scalable NLP and ML algorithms (classification, clustering and ranking methods) that automatically classify laws into various codes/labels, rank feature sets based on use case, and induce best structured representation of sentences for various types of computational analysis. The results provide standardized coding labels of policies to assist regulators to better understand how key policy features impact financial markets.
(54) GENERATING PROSODIC CONTOURS FOR 6,871,178 B2 3/2005 Case et al. SYNTHESIZED SPEECH 6,975,987 B1 12/2005 Tenpaku et a1. 6,990,449 B2 1/2006 Case . 6,990,450 B2 l/2006 Case et al. (75) Inventors: Martin Jansclhe, New York, NY (US); 7,035,791 B2 400% Chazan et a1‘ Mlchael DRlley, New York, NY (Us); 7,062,439 B2 6/2006 Brittan et al. Andrew M. Rosenberg, Brooklyn, NY 7,076,426 B1 7/2006 Beutnagel et al. (Us); Terry Tai’ Jersey City’ N] (US) 7,191,132 B2 3/2007 Brittan et al. 7,200,558 B2 * 4/2007 Kato et al. .................. .. 704/244
This paper describes our recent improvements to IBM TRANSTAC speech-to-speech translation systems that address various issues arising from dealing with resource-constrained tasks, which include both limited amounts of linguistic resources and training data, as well as limited computational power on mobile platforms such as smartphones. We show how the proposed algorithms and methodologies can improve the performance of automatic speech recognition, statistical machine translation, and text-to-speech synthesis, while achieving low-latency two-way speech-to-speech translation on mobiles.
In natural language question answering (QA) systems, questions often contain terms and phrases that are critically important for retrieving or finding answers from documents. We present a learnable system that can extract and rank these terms and phrases (dubbed mandatory matching phrasesor MMPs), and demonstrate their utility in a QA system on Internet discussion forum data sets. The system relies on deep syntactic and semantic analysis of questions only and is independent of relevant documents. Our proposed model can predict MMPs with high accuracy. When used in a QA system features derived from the MMP model improve performance significantly over a state-of-the-art baseline. The final QA system was the best performing system in the DARPA BOLT-IR evaluation.
We present a novel formalism for introducing deep belief features to Hierarchical Machine Translation Model. The deep features are generated by unsupervised training of a deep belief network built with stacked sets of Restricted Boltzmann Machines. We show that our new deep feature based hierarchical model is better than the baseline hierarchical model with gains for two different languages pairs in two different data size settings. We obtain absolute BLEU score improvement of +1.13 on Darito-English and +0.66 on English-to-Dari Transtac Evaluation task. We also observe gains on English-to-Chinese translation task.
We present Power Mean Pyramid Scores (PMP), an evaluation metric that extends the Pyramid evaluation scheme for summarization by combining Sentence Content Units (SCU) scores using Power Mean. The Pyramid method generates a summarization score by linearly combining component SCU scores. We find that by combining SCU scores using Power Mean, we can optimize a single parameter, α, leading to significantly improved correlation with human judgements. We demonstrate this result through an empirical study based on TAC-08 evaluation.
Sparse representations (SRs) are often used to characterize a test signal using few support training examples, and allow the number of supports to be adapted to the specific signal being categorized. Given the good performance of SRs compared to other classifiers for both image classification and phonetic classification, in this paper, we extended the use of SRs for text classification, a method which has thus far not been explored for this domain. Specifically, we demonstrate how sparse representations can be used for text classification and how their performance varies with the vocabulary size of the documents. In addition, we also show that this method offers promising results over the Naive Bayes (NB) classifier, a standard baseline classifier used for text categorization, thus introducing an alternative class of methods for text categorization.
rogress in both speech and language processing has spurred efforts to sup-port applications that rely on spoken—rather than written—language input.A key challenge in moving from text-based documents to such “spoken doc-uments” is that spoken language lacks explicit punctuation and formatting,which can be crucial for good performance. This article describes differentlevels of speech segmentation, approaches to automatically recovering segment bound-ary locations, and experimental results demonstrating impact on several language pro-cessing tasks. The results also show a need for optimizing segmentation for the endtask rather than independently.
Most existing techniques for combining multiple alignment tables can combine only two alignment tables at a time, and are based on heuristics (Och and Ney, 2003), (Koehn et al., 2003). In this paper, we propose a novel mathematical formulation for combining an arbitrary number of alignment tables using their power mean. The method frames the combination task as an optimization problem, and finds the optimal alignment lying between the intersection and union of multiple alignment tables by optimizing the parameter p: the affinely extended real number defining the order of the power mean function. The combination approach produces better alignment tables in terms of both F-measure and BLEU scores.
Julia Hirschberg合作论文数Department of Computer Science, Columbia University9
Elizabeth Shriberg合作论文数Speech Technology & Research Laboratory (Wednesdays)2