AI is rapidly becoming essential in various industries, raising societal expectations. AI’s societal consequences include impacts on mental health; misinformation; workforce displacement; and economic, regulatory, and law enforcement challenges. Indeed, the regulation of AI usage is on the horizon, with the European Union and China already taking big steps, while the United States drafted its first AI-related bill of rights last year. Professional associations and other nonprofits are also contributing to AI ethics and regulations, increasing the urgency and criticality of this area. In this new context, public services and regulated institutions must ensure responsible AI to avoid biased or inaccurate decision-making. Similarly, companies using AI responsibly can stand out, increase efficiency, and avoid future legal problems. This article highlights the issues and problems that result in many organizations not knowing how to do responsible AI in practice, as they need to identify potential problems, set up safeguards, and conduct ethical impact assessments, among other actions. We present the issues to consider toward a comprehensive approach to responsible AI that should include defining a responsible AI strategy road map; assessing models, processes, and products; and training individuals at different levels. By covering the pressing issues related to the urgent need for adopting responsible AI, we hope to highlight the importance for corporations to seriously consider responsible AI as they rush to adopt this technology for competitive advantage.
Artificial intelligence (AI) represents a novel force in both global and regional developments, transcending geographical, industrial, and academic borders.This article presents a case study in surveying AI challenges and opportunities in the state of Maine, in exploring ways to develop its AI ecosystem, and in fostering collaboration and development aligned with its strengths.It also showcases the potential when academics, government, and industry work together.
Generative AI is all the rage nowadays—primarily driven by ChatGPT capturing the public imagination and attracting hundreds of millions of users in record time, reaching 100 million users in two months. However, there is much ambiguity from the providers about the technology, the methodology, and the way OpenAI makes it work. This compounds the mystique and speculation. I focus on what we know, with a particular emphasis on the aspects that the makers of ChatGPT avoid discussing with the public—namely, the underlying dependence on much manual intervention in training data curation, data labeling, operational interventions by humans, and reinforcement learning. Unfortunately, despite the criticality of these issues to the scientific community, they are hardly discussed. In this article, I attempt to address some of the issues in the hope of stimulating further studies of these less glorified but critical topics.
The 3rd IADSS Workshop on Data Science Standards follows a tradition of two prior KDD workshops and the initial workshop at ICDM-2018. The theme of the 2022 workshop is: Hiring, Assessing and Upskilling Data Science Talent. Organized by the Initiative for Analytics and Data Science (IADSS.org) at KDD, the workshop provides a platform to discuss industry needs and practices around external and internal talent pipeline development in data science. We aim to provide an understanding of the data science job market, and the critical role of collaboration between academic institutions and industry to meet the increasing need for talent. IADSS conducts ongoing research in this domain. In this workshop, we share detailed findings and observations from this research. In addition to contribution from researchers and industry practitioners through an open call for papers, the workshop features several invited presentations and invited speakers. In order to achieve intended aim, interactive panels will discuss topics of interest and feedback from the workshop will be used to produce post-workshop learnings. This workshop is designed as a half-day working meeting with short talks, invited panels and discussion sessions to plan for future steps in the topic. Post conference, learnings from the workshop will be available at the workshop's home page: https://www.iadss.org/kdd2022
Much attention is paid to data science and machine learning as an effective means for getting value out of data and as a means for dealing with the large amounts of data we are accumulating at companies and organizations. This has gained importance with the major waves of digitization we have seen, especially with the COVID-19 pandemic accelerating digital everything. However, in reality, most machine learning models, despite achieving good technical solutions to predictive problems wind up not being deployed. The reasons for this are many and have their origin in data scientists and machine learning practitioners not paying enough attention to issues of deployment in production. The issues range all the way from establishing trust by business stakeholders and users, to failure to explain why models work and when they do not, to failing to appreciate the importance of establishing a robust quality data pipeline, to ignoring many constraints that apply to deployed models, and finally to a lack of understanding the true cost of production deployment and the associated ROI. We discuss many of these problems and we provide what we believe is a pragmatic approach to getting data science models successfully deployed in working environments.
The lack of an agreed-upon classification of job roles related to data science is causing much confusion that is challenging to the industry, educational sector, and practitioners.Prior work in this area has considered different aspects from different fields or points of view and has shown that more detail is needed in subcategorizing data science professionals.However, other prior work has also shown that avoiding the detailed subcategorization leads to challenging problems, for example, the pursuit of the elusive 'data science unicorns.'In this article, we target a simplification of prior work and an anchor categorization of job roles with clear definitions and expectations from each.We achieve this through analysis of survey results, LinkedIn profiles and job descriptions, and in-depth interviews with managing and hiring executives in data science.We also use our judgment as long-term practitioners and employers of data scientists to provide a practice-guided view of the problem.Our analysis has led to a simplification into three key role families with complementary skills: data analyst, data scientist, and data engineer.We believe this anchor categorization helps resolve several problems, including recruiting, forming, training, managing, and retaining effective data science teams.Although we realize there are and will continue to be many variations of these proposed anchor roles, this simplification is an effective tool to bypass the data science unicorn issue, and it can be used as a basis to establish more specialized or domain-specific roles.The combined skills in these role categories converge on the body of knowledge specification from the Initiative for Analytics and Data Science Standards (IADSS) data science knowledge framework (Fayyad & Hamutcu, 2020).The concise and familiar role categories simplify the problem and decompose it into more solvable subchallenges.We describe the essential knowledge required for each role and how, when, and in what ways it can be varied and extended.This description helps align expectations and serves as a step to tackle the pressing issue of training, evaluating, and building effective data science teams.
The growing ubiquity of the Internet and the information overload created a new economy at the end of the twentieth century: the economy of attention. While difficult to size, we know that it exceeds proxies such as the global online advertising market which is now over $300 billion with a reach of 60% of the world population. A discussion of the attention economy naturally leads to the data economy and collecting data from large-scale interactions with consumers. We discuss the impact of AI in this setting, particularly of biased data, unfair algorithms, and a user-machine feedback loop tainted by digital manipulation and the cognitive biases of users. The impact includes loss of privacy, unfair digital markets, and many ethical implications that affect society as a whole. The goal is to outline that a new science for understanding, valuing, and responsibly navigating and benefiting from attention and data is much needed.
How Can We Train Data Scientists When We Can't Agree on Who They Are? 2As the demand for data science talent has exploded, so have the efforts to train data science professionals.There are many programs and formats for training in data science, ranging from short online courses to fulltime undergraduate and graduate degree programs.The article "Statistics Practicum: Placing 'Practice' at the Center of Data Science Education" by Kolaczyk et al. (2020, this issue) presents a great deal of insight into the challenge of designing such a program at one of the prominent academic institutions in the United States.In our opinion, the article makes some distinctions about program design that will surely prove to be very useful for others who are on the journey to building or enhancing their own, particularly the importance of a practicum-type training in statistics education anchored to actual consulting services with real customers and real data.As the title of this discussion article suggests, we will expand on this challenge by posing a critical question that should be top-of-mind in designing education and training programs for data science.As there is not yet an agreed-upon definition of who data scientists are and which skills and knowledge they need to have, designing programs or developing curricula is challenging.On the other hand, organizations in industry are often not able to articulate their expectations from data science talent clearly, which in turn makes hiring, managing, and developing data professionals mostly inefficient and ineffective.In order to start providing answers to the question in the title, we will rely heavily on our work and research at Initiative for Analytics and Data Science Standards 1 (IADSS) and insights from a workshop organized by IADSS at the Knowledge Discovery and Data Mining (KDD) 2020 conference that focused on this exact challenge of training data science professionals.In the second section of the discussion, we will argue that the practicum idea can be even further expanded to a residency-type program with intensive and immersive work on real and current problems, inclusive of problematic data challenges and issues with access and completeness of data.Although our thoughts and findings are driven primarily from an industry perspective, we will try to provide 'student perspective' at the end as we see it is equally challenging for people who are interested in developing their knowledge and skills in data science to find the best path for their learning and career goals. Understanding Roles and Skills in Data ScienceSo really, who are data scientists and what are they expected to know?In the first article (Fayyad & Hamutcu, 2020) authored as part of our IADSS activities to present a framework for the knowledge and skills required in data science, we note that "although 'data scientist' has emerged as a job title, every industry, function, and business appear to be looking for their definition of the role and that universities have responded to the demand for data scientists by creating schools, institutes, and centers and establishing degree programs for relevant disciplines.These suffer from the same confusion-some are housed in business schools and others are established within computer science departments, some are crossdisciplinary, and others are considered specializations of more established disciplines.Some institutes hire a
As the industry is racing to harness the power of data, demand for data science professionals is growing at an increasing rate. However, almost every organization has a unique way of defining roles in data science and associated skills and knowledge. This has resulted in a confusing industry landscape for employers, academic and training institutions, and existing and aspiring data science professionals. This article is the first in a series authored by Initiative for Analytics and Data Science Standards (IADSS). We review the history of data science, which we trace back to 1974, and the emergence of data science as a profession in the industry, followed by a classification of knowledge and skills commonly associated with data science professionals, pointing to a lack of detailed and consistent treatment of the topic. We then present a Data Science Knowledge Framework, that we believe can support industry standardization and building measurement and assessment methodologies for data science professionals.
%% This BibTeX bibliography file was created using BibDesk. %% https://bibdesk. sourceforge.io/ %% Created for jiaqi bao at 2020-02-06 20:34:06 -0800 %% Saved with string encoding Unicode (UTF-8) @url{optics, Author = {Chire}, Date-Added = {2020-03-06 15:10:01 -0800}, Date-Modified = {2020-03-06 15:11:31 -0800}, Lastchecked = {20 October 2011}, Urldate = {https://commons.wikimedia.org/wiki/File:DBSCAN- Illustration.svg}} @url{ae, Author = {Michela Massi}, Date-Added = {2020-03-06 15:07:04 -0800}, Date-Modified = {2020-03-06 15:11:37 -0800}, Lastchecked = {2019}, Urldate = {https://commons.wikimedia.org/wiki/File:Autoencoder_schema.png}} @url{featureset, Date-Added = {2020-02-06 20:32:52 -0800}, Date-Modified …
The Applied Data Science (ADS) Invited Talks Track at KDD-2017 is a continuation of what has now become a "7-year tradition" at KDD conferences. This is the second year the track operates under the ADS name, an evolution from its origins at KDD-2011 as the "Industry Practice Expo". The KDD Conference on Knowledge Discovery and Data Mining (KDD) is the world's first, largest and best conference on Data Science, Data Mining, and Knowledge Discovery. It brings together a healthy mix of academic researchers, industry and government researchers, and practitioners from a wide range of institutions and fields. The primary focus on KDD is on peer-reviewed research contributions and the academic advancement of the field. This is an important goal and in fact the KDD conference is now recognized as the most competitive and prestigious forum for presenting high quality research results. KDD, being fundamentally an applied field, needs the strong representation of applied work of big impact. Over the years of running the conference we observed that our initial speaker-selection approach needed to be re-thought because of the important contributions made to the field outside traditional academic, industrial and government research laboratories. The result of this re-thinking was to create a forum that exposes important contributions to Data Science through Big Data Applications that address strategic problems. We wanted to effectively capture the rising importance of Data Science and Machine Learning especially in the Big Data environment where structured and unstructured data create special challenges, and of course present new opportunities. The goal of the Invited Talks Track is to curate contributions from leaders in our field who have made important contributions through the development of a system, the creation of a new and important business, or the development and market introduction of a product,. Some of these important contributions may never see an academic paper or detailed peer-reviewed paper written about them, yet they are of critical importance to our very applied field. To give you an idea of how rapidly growing this area is, and how this sector of our industry and promises to be highly disruptive across many industries, we cite a couple of articles out of a plethora of such coverage: According to IDC, the global revenues from Big Data and business will grow from $130.1 billion in 2016 to more than $203 billion in 2020, at a compound annual growth rate (CAGR) of 11.7% [1]. Furthermore, to quote from a Forbes article: "Data monetization" will become a major source of revenues, as the world will create 180 zettabytes of data (or 180 trillion gigabytes) in 2025, up from less than 10 zettabytes in 2015.? [2]
This panel aims to address areas that are widely acknowledged to be of critical importance to the success of Data Science projects and to the healthy growth of KDD/Data Science as a field of scientific research. However, despite this acknowledgement of their criticality, these areas receive insufficient attention in the major conferences in the field. Furthermore, there is a lack of actual actions and tools to address these areas in actual practice. These areas are summarized as follows: 1. Ask any data scientist or machine learning practitioner what they spend the majority of their time working on, and you will most likely get an answer that indicates that 80% to 90% of their time is spent on "Data Chasing", "Data Sourcing", "Data Wrangling", "Data Cleaning" and generally what researchers would refer to-often dismissively-as "Data Preparation". The process of producing statistical or data mining models from data is typically "messy" and certainly lacks management tools to help manage, replicate, reconstruct, and capture all the knowledge that goes in 90% of activities of a Data Scientists. The intensive Data Engineering work that goes into exploring and determining the representation of problem and the significant amount of "data cleaning" that ensues creates a plethora of extracts, files, and many artifacts that are only meaningful to the data scientist. 2. The severe lack of Benchmarks in the field, especially ones at big data scale is an impediment to true, objective, measurable progress on performance. The results of each paper are highly dependent on the large degree of freedom an author has on configuring competitive models and on determining which data sets to use (often Data that is not available to others to replicate results on) 3. Monitoring the health of models in production, and deploying models into production environments efficiently and effectively is a black art and often an ignored area. Many models are effectively "orphans" with no means of getting appropriate health monitoring. The task of deploying a built model to production is frequently beyond the capabilities of a Data Scientists and the understanding of the IT team.
These keynotes speeches discuss the following: Big Data, AIl Data, Old Data: Predictive Analytics in a Changing Data Landscape; Relative thinking; Monte Carlo Methods for Big Data and Big Models; Social Computing and Computational Societies: An ACP based Approach for Smart and Parallel Economic Systems; Some Patterns in Online Behavior.
Chapter 5 iSoNTRE: Intelligent Transformer of Social Networks into a Recommendation Engine Environment Rana Chamsi Abu Quba, Rana Chamsi Abu QubaSearch for more papers by this authorSalima Hassas, Salima HassasSearch for more papers by this authorUsama Fayyad, Usama FayyadSearch for more papers by this authorHammam Chamsi, Hammam ChamsiSearch for more papers by this authorChristine Gertosio, Christine GertosioSearch for more papers by this author Rana Chamsi Abu Quba, Rana Chamsi Abu QubaSearch for more papers by this authorSalima Hassas, Salima HassasSearch for more papers by this authorUsama Fayyad, Usama FayyadSearch for more papers by this authorHammam Chamsi, Hammam ChamsiSearch for more papers by this authorChristine Gertosio, Christine GertosioSearch for more papers by this author Book Editor(s):Gérald Kembellec, Gérald KembellecSearch for more papers by this authorGhislaine Chartron, Ghislaine ChartronSearch for more papers by this authorImad Saleh, Imad SalehSearch for more papers by this author First published: 05 December 2014 https://doi.org/10.1002/9781119054252.ch5 AboutPDFPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShareShare a linkShare onEmailFacebookTwitterLinkedInRedditWechat Summary Many works have shown that combining social networks with recommender systems increases the precision of the recommendation. Intelligent Social Network Transformer into Recommendation Engine (iSoNTRE) is a generic engine, which is based on the transformation of social networks into a source which feeds recommendation systems. It aims to transform the implicit actions of users on items and links in social networks into classifications of concepts. Recommendation systems in general and recommendations on social networks in particular use explicit classifications of users. This chapter presents the latest developments and works related to the issue, and also presents the contribution: iSoNTRE, for which the author explains the components. Collaborative Filtering (CF) is a widely used and studied method for making recommendations. The chapter provides a presentation of the experiments carried out and their results. Bibliography Abel F., Gao Q., Houben G.J., et al., " Semantic enrichment of twitter posts for user profile construction on the social web", The Semantic Web:Research and Applications, Springer-Verlag, Lecture Notes in Computer Science, vol. 6644, pp. 375–389, 2011. Abel F., Henze N., Herder E., et al., " Interweaving public user profiles on the web", in P. De Bra, et al. (eds.), User Modeling, Adaptation, and Personalization, Springer-Verlag, Heidelberg, Germany, pp. 16–27, 2010. Ampazis N., " Collaborative filtering via concept decomposition on the Netflix dataset", Proceedings of the 18th European Conference on Artificial Intelligence: Workshop on Recommender Systems (ECAI), IOS Press, Amsterdam, The Netherlands, pp. 26–30, 2008. Bachrach Y., Kosinski M., Graepel T., et al., " Personality and patterns of Facebook usage", ACM Web Sciences, New York, USA, 2012. Breese J.S., Heckerman D., Kadie C., " Empirical analysis of predictive algorithms for collaborative filtering", Proceedings of the 14th Conference on Uncertainty in Artificial Intelligence, pp. 43–52, San Francisco, CA, USA, 1998. Claypool M., Gokhale A., Miranda T., et al., " Combining content-based and collaborative filters in an online newspaper", Proceedings of ACM SIGIR Workshop on Recommender Systems, vol. 60, Berkley, CA, USA, 1999. Cui P., Wang F., Liu S., et al., " Who should share what?: item-level social influence prediction for users and posts ranking", Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 185–194, New York, NY, USA, 2011. Denti L., Barbopuolos I., Nilsson I., et al., " Sweden's largest Facebook study", Göteborg: Gothenburg Research Institute, 2012, available at http://hdl.handle.net/2077/28893. Grimmelmann J., "Facebook and the social dynamics of privacy", Iowa Law Review, vol. 95, no. 4, pp. 1137–1206, 2009. Hofmann T., "Latent semantic models for collaborative filtering", ACM Transactions on Information Systems (TOIS), vol. 22, no. 1, pp. 89–115, 2004. Lee J.-W., Lee S.-G., Kim H.-J., "A probabilistic approach to semantic collaborative filtering using world knowledge", Journal of Information Science, vol. 37, no. 1, pp. 49–66, 2011. Li W.-J., Yeung D.-Y., " Relation regularized matrix factorization", Proceedings of the 21st International Joint Conference on Artificial Intelligence, IJCAI-09, San Francisco, USA, 2009. Ma H., King I., Lyu M.R., " Learning to recommend with social trust ensemble", Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 203–210, New York, USA, 2009. Ma H., Yang H., Lyu M.R., et al., " Sorec: social recommendation using probabilistic matrix factorization", Proceedings of the 17th ACM Conference on Information and Knowledge Management, pp. 931–940, New York, USA, 2008. Ma H., Zhou D., Liu C., et al., " Recommender systems with social regularization", Proceedings of the 4th ACM International Conference on Web Search and Data Mining, pp. 287–296, Kowloon, Hong Kong, 2011. Noel J., Sanner S., Tran K.-N., et al., " New objective functions for social collaborative filtering", Proceedings of the 21st International Conference on World Wide Web, pp. 859–868, New York, USA, 2012. Schmitt D.P., Allik J., Mccrae R.R., et al., "The geographic distribution of big five personality traits patterns and profiles of human self-description across 56 nations", Journal of Cross-Cultural Psychology, vol. 38, no. 2, pp. 173–212, 2007. Shani G., Heckerman D., Brafman R.I., "An MDP-based recommender system", Journal of Machine Learning Research, vol. 6, no. 2, pp. 1265, 2006. Su X., Khoshgoftaar T.M., "A survey of collaborative filtering techniques", Advances in Artificial Intelligence archive, vol. 2009, no. 4, January 2009. Xu B., Bu J., Chen C., Cai D., " An exploration of improving collaborative recommender systems via user-item subgroups", Proceedings of the 21st International Conference on World Wide Web, pp. 21–30, New York, USA, 2012. Yang X., Guo Y., Liu Y., Xiaoyuan S.U., et al., A survey of collaborative filtering techniques, Adv. in Artif. Intell., no. 4, January 2000. Yang S.-H., Long B., Smola A., et al., " Like like alike: joint friendship and interest propagation in social networks", Proceedings of the 20th International Conference on World Wide Web, ACM, New York, pp. 537–546, 2011. Yang X., Steck H., Liu Y., " Circle-based recommendation in online social networks", Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, ACM, New York, pp. 1267–1275, 2012. Recommender Systems ReferencesRelatedInformation
Human is surrounded by a tremendous and scary amount of information on the web. That highlights the continuous need of recommendation systems in the different domains. Unfortunately cold start problem is still an important issue in these systems on new users and new items. The problem becomes more critical in systems that contain resources that lives too shortly like offers on products which stays only for few days (short life resources-SLiR). In this work we highlight how iSoNTRE (the intelligent Social Network Transformer into Recommendation Engine) solves this problem by using users' information in online social networks to overcome the cold start problem on new users, as well as iSoNTRE uses conceptual similarity, this overcomes the problem on new items, and on short life resources also. The work has been evaluated on Twitter on real users and results show that iSoNTRE succeeded in recommending offers to users with 14% of open rate on recommended offers, which is high compared to general open rate in social media, especially when we have nothing about users or offers before.
In August 2013, we held a panel discussion at the KDD 2013 conference in Chicago on the subject of data science, data scientists, and start-ups. KDD is the premier conference on data science research and practice. The panel discussed the pros and cons for top-notch data scientists of the hot data science start-up scene. In this article, we first present background on our panelists. Our four panelists have unquestionable pedigrees in data science and substantial experience with start-ups from multiple perspectives (founders, employees, chief scientists, venture capitalists). For the casual reader, we next present a brief summary of the experts' opinions on eight of the issues the panel discussed. The rest of the article presents a lightly edited transcription of the entire panel discussion.
Human is surrounded by a tremendous amount of information on the web. That highlights the continuous need of recommendation systems in the different domains. Unfortunately cold start problem is still an important issue in these systems on new users and new items. The problem becomes more critical in systems that contain resources that lives too shortly like offers on products which stays only for few days (short life resources - SLiR), or news in a news site. From the other side social networks are very rich with users' information, unfortunately most of the proposed social recommender are applied on domain specific social networks like flickers and epinions which are much less used in the day to day life, because dealing with General Purpose Social Network (GPSN) like Facebook and Twitter needs to transform these GPSN into a useful source of recommendation dealing with them as row, implicit or unary data. In this work we highlight how iSoNTRE (the intelligent Social Network Transformer into Recommendation Engine) addresses this challenge by transforming the GPSN into useful information for recommendation based on middle layer of domain concepts. iSoNTRE overcomes the cold start problem on new users and items. It has been evaluated over Twitter, on new users, recommending offers as a kind of SLiR, results showed that iSoNTRE succeeded in recommending good offers with 14% of click on recommended offers, which is high compared to general open rate in social media, especially when we have nothing about users and we are recommending SLiR resources.
Paul Bradley合作论文数ZirMed26
Andrew Tomkins合作论文数Google3