We study Aggregation Queries over Nearest Neighbors (AQNN), which compute aggregates over the learned representations of the neighborhood of a designated query object. For example, a medical professional may be interested in the average heart rate of patients whose representations are similar to that of an insomnia patient. Answering AQNNs accurately and efficiently is challenging due to the high cost of generating high-quality representations (e.g., via a deep learning model trained on human expert annotations) and the different sensitivities of different aggregation functions to neighbor selection errors. We address these challenges by combining high-quality and low-cost representations to approximate the aggregate. We characterize value- and count-sensitive AQNNs and propose the Sampler with Precision-Recall in Target (SPRinT), a query answering framework that works in three steps: (1) sampling, (2) nearest neighbor selection, and (3) aggregation. We further establish theoretical bounds on sample sizes and aggregation errors. Extensive experiments on five datasets from three domains (medical, social media, and e-commerce) demonstrate that SPRinT achieves the lowest aggregation error with minimal computation cost in most cases compared to existing solutions. SPRinT's performance remains stable as dataset size grows, confirming its scalability for large-scale applications requiring both accuracy and efficiency.
Quiz design is a tedious process that teachers undertake to evaluate the acquisition of knowledge by students. Our goal in this paper is to automate quiz composition from a set of multiple choice questions (MCQs). We formalize a generic sequential decision-making problem with the goal of training an agent to compose a quiz that meets the desired topic coverage and difficulty levels. We investigate DQN, SARSA and A2C/A3C, three reinforcement learning solutions to solve our problem. We run extensive experiments on synthetic and real datasets that study the ability of RL to land on the best quiz. Our results reveal subtle differences in agent behavior and in transfer learning with different data distributions and teacher goals. This was supported by our user study, paving the way for automating various teachers' pedagogical goals.
Reinforcement learning (RL) policies are typically trained for fixed objectives, making reuse difficult when task requirements change. We study inference-time policy reuse: given a library of pre-trained policies and a new composite objective, can a high-quality policy be constructed entirely offline, without additional environment interaction? We introduce lever (Leveraging Efficient Vector Embeddings for Reusable policies), an end-to-end framework that retrieves relevant policies, evaluates them using behavioral embeddings, and composes new policies via offline Q-value composition. We focus on the support-limited regime, where no value propagation is possible, and show that the effectiveness of reuse depends critically on the coverage of available transitions. To balance performance and computational cost, lever proposes composition strategies that control the exploration of candidate policies. Experiments in deterministic GridWorld environments show that inference-time composition can match, and in some cases exceed, training-from-scratch performance while providing substantial speedups. At the same time, performance degrades when long-horizon dependencies require value propagation, highlighting a fundamental limitation of offline reuse.
On October 19 and 20, 2023, the authors of this report convened in Cambridge, MA, to discuss the state of the database research field, its recent accomplishments and ongoing challenges, and future directions for research and community engagement. This gathering continues a long standing tradition in the database community, dating back to the late 1980s, in which researchers meet roughly every five years to produce a forward looking report. This report summarizes the key takeaways from our discussions. We begin with a retrospective on the academic, open source, and commercial successes of the community over the past five years. We then turn to future opportunities, with a focus on core data systems, particularly in the context of cloud computing and emerging hardware, as well as on the growing impact of data science, data governance, and generative AI. This document is not intended as an exhaustive survey of all technical challenges or industry innovations in the field. Rather, it reflects the perspectives of senior community members on the most pressing challenges and promising opportunities ahead.
The ability to reuse trained models in Reinforcement Learning (RL) holds substantial practical value in particular for complex tasks. While model reusability is widely studied for supervised models in data management, to the best of our knowledge, this is the first ever principled study that is proposed for RL. To capture trained policies, we develop a framework based on an expressive and lossless graph data model that accommodates Temporal Difference Learning and Deep-RL based RL algorithms. Our framework is able to capture arbitrary reward functions that can be composed at inference time. The framework comes with theoretical guarantees and shows that it yields the same result as policies trained from scratch. We design a parameterized algorithm that strikes a balance between efficiency and quality w.r.t cumulative reward. Our experiments with two common RL tasks (query refinement and robot movement) corroborate our theory and show the effectiveness and efficiency of our algorithms.
Computer science has been about automating everything including data science pipelines. Data science and humans do not optimize for the same "objective function". Many of us have been claiming that we build human-centric systems. But are we? and if we are, are we doing it properly? This talk will attempt to answer this question at various stages of the data science pipeline, illustrating the essential roles that humans take along the way, as data labelers, as domain experts, and as end-users, and providing recommendations for building data-intensive systems that truly care.
We report progress on automatically classifying written comments that students provide after receiving their performance on tests. The aim of this classification is to help teachers support the development of students' metacognitive skills more effectively. We describe a classification pipeline that seamlessly integrates large or small language models (LLMs or SLMs), leveraging state-of-the-art retrieval augmented generation, and human feedback. We apply our approach to field data from high school physics tests and to a classification scheme derived from a model for self-regulated learning. The best classification accuracies achieved for SLMs are of the order of 0.8, which is comparable to what can be obtained with LLMs. The classification obtained indicates that students in similar classroom contexts have very different perceptions and levels of analysis of their performance on assessments. While some focus solely on the factual interpretation of their quantitative results, others comment on their level of confidence, self-efficacy and learning strategies.
Subjective data, reflecting individual opinions, permeates platforms like Yelp and Amazon, influencing everyday decisions. Upon a user query, collaborative rating platforms return a collection of items ranked in an order that is often not transparent to the users. Then, each item is presented with a collection of reviews in an order that typically is, again, rather opaque. Despite the prevalence of such platforms, little attention has been given to fairness in their context, where groups writing best-ranked reviews for best-ranked items have more influence on users’ behavior. We design and evaluate a fairness assessment pipeline that starts with a data collection phase to gather reviews from real-world platforms, by submitting artificial user queries and iterating through rated items. Following that, a group assignment phase computes and infers relevant groups for each review, based on review content and user data. Finally, the third step assesses and evaluates the fairness of rankings for different user groups. The key contributions are comparing group exposure for different queries and platforms and comparing how popular fairness definitions behave in different settings. Experiments on real datasets reveal insights into the impact of item ranking on fairness computation and the varying robustness of these measures.
We address fairness in the context of sequential bundle recommendation, where users are served in turn with sets of relevant and compatible items. Motivated by real-world scenarios, we formalize producer-fairness, that seeks to achieve desired exposure of different item groups across users in a recommendation session. Our formulation combines naturally with building high quality bundles. Our problem is solved in real time as users arrive. We propose an exact solution that caters to small instances of our problem. We then examine two heuristics, quality-first and fairness-first, and an adaptive variant that determines on-the-fly the right balance between bundle fairness and quality. Our experiments on three real-world datasets underscore the strengths and limitations of each solution and demonstrate their efficacy in providing fair bundle recommendations without compromising bundle quality.
When I was 5 growing up in Algiers, my parents took me to the music conservatory and asked me which instrument I wanted to learn to play. I did not know the answer and I said I wanted to sign up for Ballet dancing. Since then, dancing has been central to my life. When they asked me what I wanted to do after high school, I did not know the answer and I chose computer science because I heard my Math teacher say it was the future. When my husband asked me to marry him, I literally answered ''I am hungry, let's get dinner''. Since then, dinner has been a special moment for us. I was in my mid-career transition when I moved from NYC to Barcelona to Doha and then to Grenoble. I love not knowing the answer and yet making great choices. My mid-career advice is: nurture doubt, develop intuition, and learn to make great choices. The biggest change you will experience when entering your mid-career phase is a widening of your choices. That applies to your collaborators, your academic responsibilities, the conferences you will attend, the projects you will get involved in as a leader or as a partner, the people you want to mentor, the life choices you get to make, the grants you will apply for, the topics you want to work on, the services you get to complete for your research community, and the students and collaborators you will interact with on a daily basis. Choice is a blessing and a responsibility.
This work studies the applicability of expensive external oracles such as large language models in answering top-k queries over predicted scores. Such scores are incurred by user-defined functions to answer personalized queries over multi-modal data. We propose a generic computational framework that handles arbitrary set-based scoring functions, as long as the functions could be decomposed into constructs, each of which sent to an oracle (in our case an LLM) to predict partial scores. At a given point in time, the framework assumes a set of responses and their partial predicted scores, and it maintains a collection of possible sets that are likely to be the true top-k. Since calling oracles is costly, our framework judiciously identifies the next construct, i.e., the next best question to ask the oracle so as to maximize the likelihood of identifying the true top-k. We present a principled probabilistic model that quantifies that likelihood. We study efficiency opportunities in designing algorithms. We run an evaluation with three large scale datasets, scoring functions, and baselines. Experiments indicate the efficacy of our framework, as it achieves an order of magnitude improvement over baselines in requiring LLM calls while ensuring result accuracy. Scalability experiments further indicate that our framework could be used in large-scale applications.
This study explores gender-specific learning patterns among medical students using an online educational tool utilized at the national level in France. Despite increasing numbers of women in medical school admissions, persistent disparities in professional medical careers raise concerns that such inequalities may already be present during university training. In this study, we analyze data from all fourth-year medical students in France for the academic years 2022-2023 and 2023-2024. We employ a Subgroup Discovery algorithm to identify patterns in online engagement and performance within these two national cohorts, focusing on gender as the key demographic variable. Our analysis revealed significant gender differences, with a proportion of male students displaying higher levels of engagement and academic performance. In contrast, female students were more likely to disengage from interactive components of the platform or to perform poorly or averagely in some medical specialties. We hypothesize that these differences may be attributed to the nature of the evaluations, which are predominantly in multiple-choice question format, potentially favoring male students. Additionally, stress and psychological distress, which are known to affect female students, may further exacerbate these performance disparities. These findings underscore the need for targeted interventions in medical education to address early gender disparities, promoting a more equitable learning environment for all students.
We address demographic bias in neighborhood-learning models for collaborative filtering recommendations. Despite their superior ranking performance, these methods can learn neighborhoods that inadvertently foster discriminatory patterns. Little work exists in this area, highlighting an important research gap. A notable yet solitary effort, Balanced Neighborhood Sparse LInear Method (BNSLIM) aims at balancing neighborhood influence across different demographic groups. Yet, BNSLIM is hampered by computational inefficiency, and its rigid balancing approach often impacts accuracy. In that vein, we introduce two novel algorithms. The first, an enhancement of BNSLIM, incorporates the Alternating Direction Method of Multipliers (ADMM) to optimize all similarities concurrently, greatly reducing training time. The second, Fairly Sparse Linear Regression (FSLR), induces controlled sparsity in neighborhoods to reveal correlations among different demographic groups, achieving comparable efficiency while being more accurate. Their performance is evaluated using standard exposure metrics alongside a new metric for user coverage disparities. Our experiments cover various applications, including a novel exploration of bias in course recommendations by teachers’ country development status. Our results show the effectiveness of our algorithms in imposing fairness compared to BNSLIM and other well-known fairness approaches.
This report summarizes the outcomes of the second international workshop on Data Systems Education: Bridging Education Practice with Education Research (DataEd '23). The workshop was held in conjunction with the SIGMOD '23 conference in Seattle, USA on June 23, 2023. The aim of the workshop was to provide a dedicated venue for presenting and and discussing data management systems education experiences and research by bringing together the database and the computing education research communities to share findings, to crosspollinate perspectives and methods, and to shed light on opportunities for mutual progress in data systems education. The program featured two keynote talks, eight research paper presentations, and a discussion session. In this report, we present the workshop's main results, observations, and emerging research directions.
The exploration of large, real-world databases poses major challenges to users due to their volume and complexity. SQL is the preferred language for data exploration. However, the process of iteratively refining SQL queries is tedious and time consuming. We formulate the automation of personalized SQL-based data exploration as the problem of suggesting the most relevant query and accounting for user feedback at each step. We develop an end-to-end solution and a system to assist users in exploring different components of a complex database. We instantiate our solution using Multi-Armed Bandits, a category of algorithms that are suitable for interactive online learning by balancing exploration with exploitation. We design a lightweight algorithm to personalize stepwise SQL recommendations that efficiently discovers the current user preferences in coordination with that user's feedback and what other users prefer. We run extensive experiments that demonstrate the utility of our approach for large-scale data exploration.
Hypothesis testing is a statistical method used to draw conclusions about populations from sample data, typically represented in tables. With the prevalence of graph representations in real-life applications, hypothesis testing on graphs is gaining importance. In this work, we formalize node, edge, and path hypotheses on attributed graphs. We develop a sampling-based hypothesis testing framework, which can accommodate existing hypothesis-agnostic graph sampling methods. To achieve accurate and time-efficient sampling, we then propose a Path-Hypothesis-Aware SamplEr, PHASE, an m-dimensional random walk that accounts for the paths specified in the hypothesis. We further optimize its time efficiency and propose PHASEopt. Experiments on three real datasets demonstrate the ability of our framework to leverage common graph sampling methods for hypothesis testing, and the superiority of hypothesis-aware sampling methods in terms of accuracy and time efficiency.
The Diversity, Equity and Inclusion (DEI) initiative started as the Diversity/Inclusion initiative in 2020 [4]. The current report summarizes our activities in 2023.
Upskilling is a fast-growing segment of the education economy [31]. Yet, there is little algorithmic work that focuses on crafting dedicated strategies to reach high-skill mastery. In this paper, we formalize AdUp, an iterative upskilling problem that combines mastery learning [49] and Zone of Proximal Development [7]. We extend our previous work [9] and design two solutions for AdUp: MOO and MAB. MOO is a multiobjective optimization approach that relies on Hill Climbing to adapt the difficulty of recommended tests to three objectives: learner's predicted performance, aptitude, and skill gap. MAB is a meta approach based on Multi-Armed Bandits to learn the best combination of objectives to optimize at each iteration. We show how these solutions are combined with two common learner simulation models: BKT (KT-IDEM) [47] and Item Response Theory (IRT) [53]. Our simulation experiments demonstrate the necessity of leveraging all three objectives and the need to adapt the optimization objectives to the learner's progression ability as MAB offers a higher mastery rate and a better final skill gain than MOO.
Data Exploration is an incremental process that helps users express what they want through a conversation with the data. Reinforcement Learning (RL) is one of the most notable approaches to automate data exploration and several solutions have been proposed. We first summarize some RL solutions that were built for different applications. In this context, various data exploration operators are leveraged including traditional roll-up and drill-down operations and text-based operations. An RL agent is trained to generate the best policy according to a hand-crafted reward function. The benefit of training RL policies for specific data exploration tasks has been demonstrated more than once for exploring finding a needle in a haystack, for serendipitous galaxy exploration, for helping a customer land on a satisfactory product, for helping a conference chair build a program committee in a stepwise fashion, for summarizing large datasets, etc. With the advent of Large Language Models and their ability to reason sequentially, it has become legitimate to ask the question: would LLMs and AI planning outperform an RL policy in data exploration? More specifically, would LLMs help circumvent retraining for new tasks and striking a balance between specificity and generality? This led us to designing LLM-powered approaches that introduce a new way of thinking about data exploration.
Gerhard Weikum合作论文数Department of Databases and Information Systems, Max-Planck Institute for Informatics8
Motomichi Toyama (遠山元道)合作论文数Keio University5