ObjectiveThough applied widely in the fields of medicine, finance, ecology, psychology, and computer science, machine learning algorithmic‐based methods are a relatively novel approach to social scientific analysis that have yet to be extensively applied. Yet as we argue in this article, a specific form of algorithmic analysis known as C4.5 classification trees has much to offer social analysis and, specifically, the study of social and political violence.MethodThis article describes four novel classification model comparison techniques for the C4.5 classification method and applies them to the study of terrorism.ResultsOur state‐level analysis suggests that there is something fundamentally different in the targeting choices of religious and secular terrorists.ConclusionThis analysis highlights the ability of classification trees to heighten our understanding of terrorism and even provide recommendations to policymakers for avoiding future attacks.
An analysis of eight computing model curricula verifies that there are significant differences between computing disciplines. While there are many courses in the models with the same or similar names, the courses may be completely different. By reverse engineering model course descriptions, the courses are compared to determine the inclusiveness of each course in each of the others. Although expected, these results are significant for colleges and universities establishing or revising computing programs.
Interesting classification rules can be determined by a number of measures. When searching a domain for a characterisation of unique, different, but important data an appropriate measurement is diversity. Diversity as a measure of a classification rule is based on the relative distinctness of the rule to the other rules in the rule-set. The diversity measure is the sum of the inverse of commonness of a rule's items. In this paper, diversity is derived from the simplest classification trees using techniques from statistics and information retrieval, and demonstrated using sample datasets.
In data analysis, when data are unattainable, it is common to select a closely related attribute as a proxy. But sometimes substitution of one attribute for another is not sufficient to satisfy the needs of the analysis. In these cases, a classification model based on one dataset can be investigated as a possible proxy for another closely related domain's dataset. If the model's structure is sufficient to classify data from the related domain, the model can be used as a proxy tree. Such a proxy tree also provides an alternative characterization of the related domain. Just as important, if the original model does not successfully classify the related domain data the domains are not as closely related as believed. This paper presents a methodology for evaluating datasets as proxies along with three cases that demonstrate the methodology and the three types of results.
What is the relationship between religious liberty and faith-based terrorism? The wider literature on freedom and terrorism has failed to reach a conclusive verdict: some hold that restricting civil liberties is necessary to prevent acts of terrorism; others find that respecting such rights undermines support for terrorist groups, thus making terrorism less likely. This article moves the debate on liberty and terrorism forward by looking specifically at terrorism motivated by a religious imperative and a country's level of religious libertysomething not attempted in previous studies. Using classification data mining, we test a unique dataset on religious terrorism in order to discover the characteristics that contribute to a country experiencing religiously motivated terrorism. The analysis finds that religious terrorism is indeed a product of a dearth of religious liberty. The study concludes by discussing the implications of these findings for policy-makers.
AbstractThis essay introduces data mining as an analytical technique for novice to professional social and behavioral scientists. It presents data mining, which is also known as, among other things,data analyticsandpredictive analytics, as an effective tool for researchers who are interested in the analysis of “big data” as well as small, unique data sets. It addresses foundational elements of data mining such as how to avoid “data dredging” and the importance of theory as embodied in researcher domain expertise. It also briefly defines and describes classification analysis, association rules, and clustering, which are the major methodologies among a large number of methodologies that constitute data mining. This essay identifies analytical problems and data for which the techniques are best suited. It goes on to highlight a number of cutting‐edge studies that relied on data mining techniques in disciplines such as criminal justice, education, health sciences, linguistics, political science, and sociology. This essay concludes with a review of key considerations for future research to include discussions of the burgeoning of new analytical techniques and new data sets and sources, the importance and protection of data‐source privacy, and the ethical obligation researchers have to exploit to their fullest extent the costly data on social and behavioral issues collected by scientists and society.
ObjectiveTo understand what kind of individuals lead particular regimes, this study examines the most influential people in politics, the executives, to uncover the relationship between their characteristics and the type of regime they govern.MethodsThis article employs data mining with characteristics of executives worldwide against the state's Freedom House ranking.ResultsThrough data mining, the results indicate that while there are still many important factors that coincide with democracy, the length of time in office and to a lesser extent the religious beliefs of executives and the likelihood of being classified as a democracy are heavily related.ConclusionThis article concludes with a recommendation for supporting specific types of executives to increase the likelihood for successful democratization to minimize authoritarian rule.
Data mining involves the analysis of data to find interesting patterns and previously unknown relationships in data. Data mining not only predicts the results of a future event, but it also can provide knowledge about the structure and interrelationships among the data. These predictions and relationships are expressed as decision trees, classification rules, association rules, or clusters. But, data mining occurs in a domain. The data mining algorithms operate on data that was collected, or is now being used for, a specific purpose. The data is being used to study a domain question. Both the question and the data selected to answer the question need to be identified by a domain expert. This expert must drive the data cleaning process and act as a participant in the data mining process.
This chapter data mines the usage patterns of the ANGEL Learning Management System (LMS) at a comprehensive college. The data includes counts of all the features ANGEL offers its users for the Fall and Spring semesters of the academic years beginning in 2007 and 2008. Data mining techniques are applied to evaluate which LMS features are used most commonly and most effectively by instructors and students. Classification produces a decision tree which predicts the courses that will use the ANGEL system based on course specific attributes. The dataset undergoes association mining to discover the usage of one feature’s effect on the usage of another set of features. Finally, clustering the data identifies messages and files as the features most commonly used. These results can be used by this institution, as well as similar institutions, for decision making concerning feature selection and overall usefulness of LMS design, selection and implementation.
Social scientists address some of the most pressing issues of society such as health and wellness, government processes and citizen reactions, individual and collective knowledge, working conditions and socio-economic processes, and societal peace and violence. In an effort to understand these and many other consequential issues, social scientists invest substantial resources to collect large quantities of data, much of which are not fully explored. This chapter proffers the argument that privacy protection and responsible use are not the only ethical considerations related to data mining social data. Given (1) the substantial resources allocated and (2) the leverage these “big data” give on such weighty issues, this chapter suggests social scientists are ethically obligated to conduct comprehensive analysis of their data. Data mining techniques provide pertinent tools that are valuable for identifying attributes in large data sets that may be useful for addressing important issues in the social sciences. By using these comprehensive analytical processes, a researcher may discover a set of attributes that is useful for making behavioral predictions, validating social science theories, and creating rules for understanding behavior in social domains. Taken together, these attributes and values often present previously unknown knowledge that may have important applied and theoretical consequences for a domain, social scientific or otherwise. This chapter concludes with examples of important social problems studied using various data mining methodologies including ethical concerns.
Data mining is a collection of algorithms for finding interesting and unknown patterns or rules in data. However, different algorithms can result in different rules from the same data. The process presented here exploits these differences to find particularly robust, consistent, and noteworthy rules among much larger potential rule sets. More specifically, this research focuses on using association rules and classification mining to select the persistently strong association rules. Persistently strong association rules are association rules that are verifiable by classification mining the same data set. The process for finding persistent strong rules was executed against two data sets obtained from the American National Election Studies. Analysis of the first data set resulted in one persistent strong rule and one persistent rule, while analysis of the second data set resulted in 11 persistent strong rules and 10 persistent rules. The persistent strong rule discovery process suggests these rules are the most robust, consistent, and noteworthy among the much larger potential rule sets.
Firms need to deliver their products. In the Net Economy, delivery often has to leave the Net and be provided through traditional means. The firm’s delivery mechanism influences the design of the firm’s Net presence. This chapter examines the pursuit of e-entrepreneurial ventures by existing businesses with specific attention on the architecture of Web portals and the delivery mechanisms of products. Additionally, outlined are features and facets of Web portals necessary to sell and deliver mark-up based and production based products and services in the B2C sector of the Net Economy. Specifically, three case studies are examined: a catalog sales/brick-and-mortar business, a financial service institution, and a travel provider.
This research demonstrates the application of multiple data mining techniques to test theories of the macro-level causes of terrorism. The unique dataset is comprised of terrorist events and measures of social, political and economic contexts in 185 countries worldwide between the years 1970 and 2004. The theories are assessed using the iterative expert data mining (IEDM) methodology with classification mining and then association mining. The resulting 100 rules suggest that the level of democracy in a country is an integral part of the explanation for terrorism. This research shows that a multi-method data mining approach can be used to test competing theories in a discipline by analysing large, comprehensive datasets that capture multiple theories and include large numbers of records.
Data mining is a collection of algorithms for finding interesting and unknown patterns or rules in data. However, different algorithms can result in different rules from the same data. The process presented here exploits these differences to find particularly robust, consistent, and noteworthy rules among much larger potential rule sets. More specifically, this research focuses on using association rules and classification mining to select the persistently strong association rules. Persistently strong association rules are association rules that are verifiable by classification mining the same data set. The process for finding persistent strong rules was executed against two data sets obtained from the American National Election Studies. Analysis of the first data set resulted in one persistent strong rule and one persistent rule, while analysis of the second data set resulted in 11 persistent strong rules and 10 persistent rules. The persistent strong rule discovery process suggests these rules are the most robust, consistent, and noteworthy among the much larger potential rule sets.
Business marketers widely use data mining for segmenting and targeting markets. To assess data mining for use by political marketers, we mined the 1948 to 2004 American National Elections Studies data file to identify a small number of variables and rules that can be used to predict individual voting behavior, including abstention, with the intent of segmenting the electorate in useful and meaningful ways. The resulting decision tree correctly predicts vote choice with 66 percent accuracy, a success rate that compares favorably with other predictive methods. More importantly, the process provides rules that identify segments of voters based on their predicted vote choice, with the vote choice of some segments predictable with up to 87 percent success. These results suggest that the data mining methodology may increase efficiency for political campaigns, but they also suggest that, from a democratic theory perspective, overall participation may be improved by communicating more effective messages that better inform intended voters and that motivate individuals to vote who otherwise may abstain.
One often-noted difficulty in pre-election polling is the identification of likely voters. Our objective is to build a likely voter model for presidential elections that efficiently balances accuracy and number of questions used. We employ the Iterative Expert Data Mining technique and data from the American National Election Studies to identify a small number of survey questions that can be used to classify likely voters while maintaining or surpassing the accuracy rates of other models. Specifically, we propose two survey items that together correctly classify 78 percent of respondents as voters or nonvoters over a multielection, multidecade period. We argue that our proposed model compares favorably to competing models by capturing the successful elements of those models while ignoring other elements that constrain identification. We end by suggesting that our model offers a new approach to identifying and evaluating likely voters that may maintain or increase accuracy without also increasing cost.
The diversity of IS programs and research has been of interest to various professions. It has been argued that IS has developed to the extent where it does not have to rely on other reference disciplines, but should rather serve as a reference discipline for other disciplines. While IS may have developed its own discipline, its location in different academic units may influence the venue of faculty publications. The understanding of the relationships between venue of publication and location of IS programs will influence curriculum development especially at the doctoral level and inform faculty placement decisions. In this paper, we examine IS research that falls into the professional categories of business, engineering, education, and library science for faculty from information systems programs. We examined the research publications of the faculty from the twenty-four IS programs accredited by ABET Inc. The data shows that irrespective of the location of the IS program, over 50% of the faculty publications are in the Engineering venue. Further, the results indicate that the location of the IS program influences the publication venue. We also suggest that the tenure and promotion requirements also influence the venue of the publications of IS faculty. Our research contributes to both professional practice and scholarly research. In academia, we suggest that the interest of the faculty may influence their employment locations and research venues.
Information technology (IT) is an umbrella term that encompasses disciplines dealing with the computer and its functions. These disciplines originated from interests in using the computer to solve problems, the theory of computation, and the development of the computer and its components. Professionals from around the world with similar interests in IT came together and formed international professional organizations. The professional organizations span the disciplines of computer engineering (CE), computer science (CS), software engineering (SE), computer information systems (CIS), management information systems (MIS), and information technology (IT) (Freeman & Aspray, 1999). Note that information technology is both an umbrella term and a specific discipline under that umbrella. These organizations exist to promote their profession and one method of promotion is through education. So, these professional organizations defined bodies of knowledge around the computer, which have been formalized and shaped as model curriculums. The organizations hope that colleges and universities will educate students in the IT disciplines to become knowledgeable professionals. Because of the common interest in computing, there is a basic theory and a common technical core that exists among the model curricula (Denning, 1999; Tucker et al., 1991). Nevertheless each of the model curricula emphasizes a different perspective of IT. Each fills a different role in providing IT professionals. It falls upon the colleges and universities to select and modify the corresponding curriculum model to fit their needs.