%% This BibTeX bibliography file was created using BibDesk. %% https://bibdesk. sourceforge.io/ %% Created for jiaqi bao at 2020-02-06 20:34:06 -0800 %% Saved with string encoding Unicode (UTF-8) @url{optics, Author = {Chire}, Date-Added = {2020-03-06 15:10:01 -0800}, Date-Modified = {2020-03-06 15:11:31 -0800}, Lastchecked = {20 October 2011}, Urldate = {https://commons.wikimedia.org/wiki/File:DBSCAN- Illustration.svg}} @url{ae, Author = {Michela Massi}, Date-Added = {2020-03-06 15:07:04 -0800}, Date-Modified = {2020-03-06 15:11:37 -0800}, Lastchecked = {2019}, Urldate = {https://commons.wikimedia.org/wiki/File:Autoencoder_schema.png}} @url{featureset, Date-Added = {2020-02-06 20:32:52 -0800}, Date-Modified …
We discuss the most important database research advances, industry developments, role of relational and NoSQL databases, Computing Reality, Data Curation, Cloud Computing, Tamr and Jisto startups, what he learned as a chief Scientist of Verizon, Knowledge Discovery, Privacy Issues, and more.
The point isnot that thedata is necessarilybig, eventhough there are now some gargantuan data sets that did not exist in the past. Instead, it is big in a relative sense, not in an absolute sense—it is often big in relation to the phenomenon that we are trying to record and understand. So, if we are only looking at 64,000 data points, but that represents the totality or the universe of observations, where before we might have used a sampling technique, now we do not have to sample; we can use all the observations. That is what qualifies as big data. You do not have to have a hypothesis in advance before you collect your data. You have collected all there is—all the data there is about a phenomenon.
The primary focus of SIGKDD is to provide the premier forum for advancement and adoption of the "science" of knowledge discovery and data mining. SIGKDD main activity is to organize KDD, the leading conference on data mining and knowledge discovery , held since 1995. KDD conference is top-ranked in Data Mining, according to Microsoft Research Asia. KDD-2011 was held in San Diego, CA, USA was the largest data-mining meeting in the world, with over 1,100 participants from around the world.
Knowledge discovery and data mining provide a key technology for Web intelligence and have been applied in many newly developed intelligent Web information systems. New requirements arise from an increasing number of Web-based real-world applications (e-applications, for short), such as e-business (including e-commerce and e-finance), e-service, e-science, e-learning, e-government, and e-community and social networks. The main objective of this special issue is to present a showcase of the theory and applications of knowledge discovery and data mining for Web intelligence. This Special Issue includes 3 articles selected by 2 rounds of peer-reviews. The article by Jie Tang, Limin Yao, Duo Zhang, and Jing Zhang is entitled “A Combination Approach to Web User Profiling”. In the proposed combination approach, the authors use a novel method called TCRF (Tree-structured Conditional Random Fields) for profile extraction, a probabilistic model to solve the ambiguity problem for profile integration, and a topic model for user-interest discovery. Mohamed Bouguessa, Shengrui Wang and Benoit Dumoulin in their article, “Discovering Knowledge-Sharing Communities in Question-Answering Forums”, address the topic of identifying community structures in Web-based question-answering forums. By applying the methods presented in this article, a knowledge-sharing community can be discovered by identifying a set of questioners and authoritative users. Anon Plangprasopchok and Kristina Lerman in the article, “Modeling Social Annotation: A Bayesian Approach”, present a new approach for modeling social annotation on the social Web, which takes into account all essential entities, such as users, resources, and tags. The proposed approach has been applied to infer topics of Web resources from social annotations for resource discovery.
I survey the transformation of the data mining and knowledge discovery field over the last 10 years from the unique vantage point of KDnuggets as a leading chronicler of the field. Analysis of the most frequent words in KDnuggets News leads to revealing observations.
Interview with Jon Kleinberg, a pioneer in web mining, social network analysis, and other fields and a winner of many awards, including 2 KDD Best Papers and a MacArthur 'genius' award.
Interview with Simon Funk -- a Netflix prize leader, an outstanding hacker, and an original thinker.
We discuss what makes exciting and motivating Grand Challenge problems for Data Mining, and propose criteria for a good Grand Challenge. We then consider possible GC problems from multimedia mining, link mining, large-scale modeling, text mining, and proteomics. This report is the result of a panel held at KDD-2006 conference.
This panel will discuss possible exciting and motivating Grand Challenge problems for Data Mining, focusing on bioinformatics, multimedia mining, link mining, text mining, and web mining.
We study an algorithm for feature selection that clusters attributes using a special metric and then makes use of the dendrogram of the resulting cluster hierarchy, to choose the most relevant attributes. The main interest of our technique resides in the improved understanding of the structure of the analyzed data and of the relative importance of the attributes for the selection process.
The rapid and constant growth of databases in business, government, and science has far outpaced our ability to interpret and make sense of this data avalanche, creating a need for a new generation of tools and techniques for intelligent and automated database analysis. These tools and techniques are the subject of the rapidly emerging field of data mining and knowledge discovery in databases (KDD). This paper surveys the state of the art in this field, with a particular focus on the issues and challenges in applying KDD to business databases.
This paper examines the current trends in business applications of data mining and knowledge discovery systems. The focus is on newly emerging “Third Generation” data mining systems, which are solution-oriented and integrate smoothly with existing business systems.
We discuss our effort to date and plan to create a generator of synthetic but realistic microarray datasets, use it to build a comprehensive test suite for to studying the performance of different methods for gene selection, classification, and clustering, and find out which methods perform best under which conditions.
Analyzing gene expression data from microarray devices has many important application in medicine and biology, but presents significant challenges to data mining. Microarray data typically has many attributes (genes) and few examples (samples), making the process of correctly analyzing such data difficult to formulate and prone to common mistakes. For this reason it is unusually important to capture and record good practices for this form of data mining. This paper presents a process for analyzing microarray data, including pre-processing, gene selection, randomization testing, classification and clustering; this process is captured with "Clementine Application Templates". The paper describes the process in detail and includes three case studies, showing how the process is applied to 2-class classification, multi-class classification and clustering analyses for publicly available microarray datasets.
. At the 2001 IEEE International Conference on Data Mining in San Jose, California, on November 29 to December 2, 2001, there was a panel discussion on how data mining research meets practical development. One of the motivations for organizing the panel discussion was to provide useful advice for industrial people to explore their directions in data mining development. Based on the panel discussion, this paper presents the views and arguments from the panel members, the Conference Chair and the Program Committee Co-Chairs. These people as a group have both academic and industrial experiences in different data mining related areas such as databases, machine learning, and neural networks. We will answer questions such as (1) how far data mining is from practical development, (2) how data mining research differs from practical development, and (3) what are the most promising areas in data mining for practical development.