This chapter addresses the question whether opinion-mining techniques can successfully be used to automatically retrieve political viewpoints from parliamentary proceedings. Two specific preprocessing tasks were identified and systematically evaluated: automatically determining subjectivity in the publications and automatically determining the semantic orientation of the subjective parts. A corpus of recent parliamentary proceedings was collected and a gold standard annotation was created on both subjectivity and orientation. Following this, a number of models based on subjectivity lexicons and machine-learning algorithms were evaluated. Machine-learning algorithms perform best, but methods based on subjectivity lexicons also provide promising results. Based on these results we can conclude that opinion-mining techniques applied to political data score just as well as the state of the art in other more traditional domains of opinion mining like product reviews and blogs.
We collect evidence to answer the following question: Is the quality of the XML documents found on the Web sufficient to apply XML technology like XQuery, XPath and XSLT? XML collections from the Web have been previously studied statistically, but no detailed information about the quality of the XML documents on the Web is available to date. We address this shortcoming in this study. We gathered 180K XML documents from the Web. Their quality is surprisingly good; 85.4% are well-formed and 99.5% of all specified encodings is correct. Validity needs serious attention. Only 25% of all files contain a reference to a DTD or XSD, of which just one-third are actually valid. Well-formedness errors and validity errors are studied in detail. Our study is well-documented, easily repeatable and all data is publicly available [21], (Grijzenhout, 2010) [52]. This paves the way for a periodic quality assessment of the XML Web.
The question is addressed if opinion mining techniques can be successfully used to automatically retrieve political viewpoints in Dutch parliamentary publications. Two specific tasks are identified: automatically determining subjectivity in the publications and automatically determining the semantic orientation of the subjective parts. A collection of recent parliamentary publications has been collected and a golden standard annotation is created on both subjectivity and orientation. Following this, models based on subjectivity lexicons and machine learning algorithms are evaluated on automatic classification. Overall results tend to be dominated by machine learning algorithms, but methods based on subjectivity lexicons provide promising results. Based on the results it is concluded that opinion mining techniques can indeed be successfully used to automatically retrieve political viewpoints in Dutch parliamentary publications.