Large Multimodal Models (LMMs) have recently demonstrated impressive performance on general video comprehension benchmarks. Nevertheless, for broader applications, the robustness of their temporal analysis capability needs to be thoroughly investigated yet predominantly ignored. Motivated by this, we propose a novel temporal robustness benchmark (TemRobBench), which introduces temporal inconsistency perturbations separately at the visual and textual modalities to assess the robustness of models. We evaluate 16 mainstream LMMs and find that they exhibit over-reliance on prior knowledge and textual context in adversarial environments, while ignoring the actual temporal dynamics in the video. To mitigate this issue, we design panoramic direct preference optimization (PanoDPO), which encourages LMMs to incorporate both visual and linguistic feature preferences simultaneously. Experimental results show that PanoDPO can effectively enhance the model's robustness and reliability in temporal analysis.
Accurate text summarization is one of the most common and important tasks performed by Large Language Models, where the costs of human review for an entire document may be high, but the costs of errors in summarization may be even greater. We propose Detecting Errors through Ensembling Prompts (DEEP) - an end-to-end large language model framework for detecting factual errors in text summarization. Our framework uses a diverse set of LLM prompts to identify factual inconsistencies, treating their outputs as binary features, which are then fed into ensembling models. We then calibrate the ensembled models to produce empirically accurate probabilities that a text is factually consistent or free of hallucination. We demonstrate that prior models for detecting factual errors in summaries perform significantly worse without optimizing the thresholds on subsets of the evaluated dataset. Our framework achieves state-of-the-art (SOTA) balanced accuracy on the AggreFact-XSUM FTSOTA, TofuEval Summary-Level, and HaluEval Summarization benchmarks in detecting factual errors within transformer-generated text summaries. It does so without any fine-tuning of the language model or reliance on thresholding techniques not available in practical settings.
Artificial Intelligence (AI) and Extended Reality (XR) have been employed in several foreign language education applications to increase the availability of experiential learning methods akin to international immersion programs. However, research in multi-modal spoken dialogue in L2 combined with immersive technologies and collaborative learning is thin, limiting students' experiences to solo interactions focused mostly on vocabulary and grammar in such settings. We intend to fill this gap as we present the Cognitive Immersive Language Learning Environment (CILLE). The AI in CILLE can hear, see, and understand its users and can engage with them in non-dyadic multimodal conversations. The XR offers students a feeling of being somewhere else without the use of intrusive devices and supports multi-party, multi-modal interactions. Together, AI and XR create naturalistic conversational interactions targeted towards comprehensive foreign language acquisition. We evaluate CILLE as a Chinese-as-a-foreign-language (CFL) education tool through a seven-week, mixed-methods study with university students (N = 10). Results display statistical significance and retained improvement in CFL vocabulary, comprehension, and conversation skills. Coupled with an analysis of student feedback and researcher observations, we show how CILLE is designed and experienced by students to learn CFL.
The past decade has witnessed the rapid expansion of demands for mobile traffic, while the traditional mobile traffic pricing schemes cannot accommodate such demands. Sponsored data plan (SDP), which can increase the revenue of all stakeholders in the market through transferring some of the revenue from content providers (CPs) to end users (EUs), is more suitable. However, existing studies have focused more on Internet service providers (ISPs) and CPs, ignoring the influence of EUs (e.g., the inherent attribute differences of EUs and the interaction among EUs) on the market under SDP. Regarding the difficulty of modeling the abstract property about interaction among EUs, we utilize network congestion as the medium and construct the congestion-aware SDP model based on Stackelberg game. The newly proposed model can not only analyze how network congestion affects SDP mechanism, but also elucidate the impact of interactions among EUs. More specifically, through theoretical analysis, we prove that there is a unique dynamic equilibrium in the interaction among EUs (i.e., the traffic consumption of different EUs). By taking into account network congestion, the newly proposed model also more accurately and realistically describes the optimal strategies and computation methods of all stakeholders in the market. Moreover, simulation experiments demonstrate that the positive effect brought by SDP is not as obvious as before, and EUs influence each other instead of being independent of each other. Overall, this paper emphasizes the non-negligible influence of EUs and promotes a deeper understanding of SDP mechanism, which can guide the relevant stakeholders to optimize their own decision-making details.
This paper proposes a new problem of complementary evidence identification for open-domain question answering (QA). The problem aims to efficiently find a small set of passages that covers full evidence from multiple aspects as to answer a complex question. To this end, we proposes a method that learns vector representations of passages and models the sufficiency and diversity within the selected set, in addition to the relevance between the question and passages. Our experiments demonstrate that our method considers the dependence within the supporting evidence and significantly improves the accuracy of complementary evidence selection in QA domain.
Technological innovations in artificial intelligence and machine learning enable business operators to engage with their customers 24/7 through chatbots. Many customers expect around-the-clock support which puts a strain on human resources; augmenting human resources with a chatbot can reduce costs for an organization and increase customer satisfaction. This work presents the HR Chatbot: a chatbot that answers general questions about human resource topics (i.e. payroll, benefits) for a private university. The research involves a collaboration between computer scientists, user experience researchers, and human resource administration. This work addresses two research questions: What are employees at a private university looking for from a chatbot for human resources?; and what are the appropriate methods to evaluate and measure the success of a chatbot for human resources? The HR Chatbot uses IBM Watson Assistant services, and an initial prototype was designed from a document of 31 frequently asked questions. Three rounds of user testing were conducted with employees of the university. The initial tests revealed that the chatbot was perceived as useful, but many were dissatisfied with the responses, specifically the lack of responses. Errors in the chatbot were classified into different categories; the most common being that the question was not in the content scope for the chatbot. Thus, data from the initial studies informed the scope of the chatbot; the number of unique questions grew to 157 and the total number of questions increased to 463. The HR Chatbot has 90% accuracy and an average sustained usability score of 69.5, surpassing the benchmark score. Following the initial tests, the HR Chatbot was deployed in real-time on the human resources website. This work describes how a chatbot was created, evaluated, and deployed online. We hope that this work inspires and informs others to explore similar use cases with chatbot technologies.
This paper discusses findings in a user study evaluation of a novel digital brainstorming tool situated in a cognitive immersive environment. We compare how the brainstorming tools enable sensemaking by examining thematic units produced by users. The digital tool integrates multimodal inputs that are novel for brainstorming applications, and allows for multiuser collaboration. This user study examines how differences between the analog and digital brainstorming tool formats impacts sensemaking metrics. We find that the digital brainstorming tool allows users to reflect content from a source material more accurately than the analog brainstorming tool. Future work needs to identify the reason for increased idea cohesion within the cognitive immersive room’s digital brainstorming tool.
Abstract Recent advancements in open-domain question answering (ODQA), that is, finding answers from large open-domain corpus like Wikipedia, have led to human-level performance on many datasets. However, progress in QA over book stories (Book QA) lags despite its similar task formulation to ODQA. This work provides a comprehensive and quantitative analysis about the difficulty of Book QA: (1) We benchmark the research on the NarrativeQA dataset with extensive experiments with cutting-edge ODQA techniques. This quantifies the challenges Book QA poses, as well as advances the published state-of-the-art with a ∼7% absolute improvement on ROUGE-L. (2) We further analyze the detailed challenges in Book QA through human studies.1 Our findings indicate that the event-centric questions dominate this task, which exemplifies the inability of existing QA models to handle event-oriented scenarios.
The Design thinking process is a common process across a range of industries. The process, often utilizing sticky notes, has people coming together to collaboratively come up with ideas, aimed at deciding on features, functionality, sale-ability, etc. As part of the process, in its co-located set-up, users are moving around, creating temporary ad-hoc groupings, which help to drive new insights. In this demo, we show a web browser based sticky note application that caters to these co-located ideals, while utilizing ubiquitous interfaces, smartphones and regular TV screens. Users interact with the system through their phone, which gives them options to create notes and then utilize their camera to place or pick-up notes. The large shared screen utilizes a grid-based QR code system that allows for many simultaneous users onto the same screen, detecting where a user's camera is pointing, without blocking any other users using the system simultaneously.
We explore trust in a relatively new area of data science: Automated Machine Learning (AutoML). In AutoML, AI methods are used to generate and optimize machine learning models by automatically engineering features, selecting models, and optimizing hyperparameters. In this paper, we seek to understand what kinds of information influence data scientists' trust in the models produced by AutoML? We operationalize trust as a willingness to deploy a model produced using automated methods. We report results from three studies - qualitative interviews, a controlled experiment, and a card-sorting task - to understand the information needs of data scientists for establishing trust in AutoML systems. We find that including transparency features in an AutoML tool increased user trust and understandability in the tool; and out of all proposed features, model performance metrics and visualizations are the most important information to data scientists when establishing their trust with an AutoML tool.
Competitions that directly pit software agents against one another have proven to be an effective and entertaining way to advance the state of the art in a multitude of AI domains. Less frequently, human-agent competitions have been held to gauge the relative competence of humans vs. agents, or agents vs. agents as measured indirectly by their performance against humans. We are developing a platform that supports a new type of AI competition that involves both agent-agent and human-agent interactions situated in an immersive environment. In this competition, human buyers haggle (in English) with two life-size AI agents that attempt to sell them various goods. We describe several research challenges that arise in this context, present the platform architecture and accompanying technologies, and report on early experiments with simple agents that establish feasibility and suggest that human participants enjoy the experience.
Sponsored Data Plan (SDP) is an emerging pricing model for the wireless data market where the Content Provider (CP) can sponsor the data usage for specific content on behalf of the users. This strategy sheds new light on the data pricing model and receives significant attention from the Internet Service Provider (ISP). However, the existing SDP studies consider traffic price (e.g., sponsorship) as the only factor that affects user decision. The impact of other classic market features, such as the demand for a variety of contents (i.e., love of variety), remains largely unclear. In this paper, we develop a new model to understand the love of variety in the wireless data market under SDPs. Our model has demonstrated that, such variety is important to understand the complex gaming between ISPs, CPs, and users in both short-run and long-run markets. For example, the analysis indicates that the advantage of CPs with higher revenue will be significantly reduced when users have a greater love of variety. Moreover, to help the ISP better adopt the proposed model in the real market, we also develop a practical method to calibrate the related parameters, which can also be applied to quantity the love of variety.
The rapid shift to remote work has forever changed the dynamic of teams, which require the best tools to work collaboratively across any distance. Rensselaer is home to two immersive virtual environments, intelligent rooms with panoramic, human-scale projection screens, spatial audio loudspeaker arrays, and networks of time-of-flight and acoustical tracking sensors. This project seeks to “colocate” teams across both sites such that the experience mimics collaborating within the same room. Simple video conferencing software are rarely well-configured for large groups and do not consider users’ spatial arrangements. This approach captures ultra-low-latency video and audio feeds of each space for presentation at either end, enabling a group at each location to communicate at 1-to-1 scale with the other. A spherical microphone array tracks multiple simultaneous speakers and adjusts their spatial positions across a Wave Field Synthesis loudspeaker array, maintaining audio-visual congruency. Users at each site can use the panoramic displays to present panoramic imagery or immersive data to be explored simultaneously by each group, facilitating a collaborative immersive experience, and interactions with the screen at one site may be echoed at the other to create the appearance of a single space. [Work supported by CISL, NSF No. 1229391, Army DURIP No. 68604-CS-RIP.]
A lot of progress has been made to improve question answering (QA) in recent years, but the special problem of QA over narrative book stories has not been explored in-depth. We formulate BookQA as an open-domain QA task given its similar dependency on evidence retrieval. We further investigate how state-of-the-art open-domain QA approaches can help BookQA. Besides achieving state-of-the-art on the NarrativeQA benchmark, our study also reveals the difficulty of evidence retrieval in books with a wealth of experiments and analysis - which necessitates future effort on novel solutions for evidence retrieval in BookQA.
We explore trust in a relatively new area of data science: Automated Machine Learning (AutoML). In AutoML, AI methods are used to generate and optimize machine learning models by automatically engineering features, selecting models, and optimizing hyperparameters. In this paper, we seek to understand what kinds of information influence data scientists' trust in the models produced by AutoML? We operationalize trust as a willingness to deploy a model produced using automated methods. We report results from three studies -- qualitative interviews, a controlled experiment, and a card-sorting task -- to understand the information needs of data scientists for establishing trust in AutoML systems. We find that including transparency features in an AutoML tool increased user trust and understandability in the tool; and out of all proposed features, model performance metrics and visualizations are the most important information to data scientists when establishing their trust with an AutoML tool.
Augmenting immersive technologies with AI for foreign language learning is a relatively unexplored, multidisciplinary, and complex research paradigm. The Cognitive and Immersive Room at Rensselaer Polytechnic Institute facilitates cultural and foreign language learning beyond a traditional classroom setting. Exploration of virtual environments and real-world renderings occur at human-scale by groups of students and teachers simultaneously, without the need for head-mounted displays. Students travel to true-to-life Panoramic Scenes to learn authentic cultural knowledge and practice vocabulary within the same surroundings they would use such knowledge while traveling abroad. Multimodal input allows students to engage with the system through gesture, voice, and spatial positioning, creating a dynamic language learning experience. This work underwent initial assessment in a six-week summer course, AI-Assisted Immersive Chinese, and preliminary qualitative results find a majority of surveyed students indicate the system is useful, engaging, and fun.
We propose a generative probabilistic model for human motion synthesis. Our model has a hierarchy of three layers. At the bottom layer, we utilize Hidden semi-Markov Model (HSMM), which explicitly models the spatial pose, temporal transition and speed variations in motion sequences. At the middle layer, HSMM parameters are treated as random variables which are allowed to vary across data instances in order to capture large intra- and inter-class variations. At the top layer, hyperparameters define the prior distributions of parameters, preventing the model from overfitting. By explicitly capturing the distribution of the data and parameters, our model has a more compact parameterization compared to GAN-based generative models. We formulate the data synthesis as an adversarial Bayesian inference problem, in which the distributions of generator and discriminator parameters are obtained for data synthesis. We evaluate our method through a variety of metrics, where we show advantage than other competing methods with better fidelity and diversity. We further evaluate the synthesis quality as a data augmentation method for recognition task. Finally, we demonstrate the benefit of our fully probabilistic approach in data restoration task.
This paper discusses a user study conducted on a brainstorming tool for brainstorming in intelligence analysis, situated in a cognitive and immersive system. The cognitive and immersive room is comprised of gesture technology, voice input, tablet input, and a 360 display to allow for group discussion and interaction. The brainstorming tool design is informed by sensemaking theory as defined by Pirolli and Card [1]. The basis for the digital brainstorming tool in our cognitive immersive environment is a structured analytic tool already in use by the intelligence analysis domain. We anticipate that this design will capitalize on cognitive familiarity to enable a more approachable design and facilitate a smoother and more useful tool. The user study (n = 26) is an A/B study conducted to compare the capabilities of the analog brainstorming tool with the digital brainstorming tool, and understand how the multimodal interaction system impacts brainstorming.
Recently, a more challenging state tracking task, Audio-Video Scene-Aware Dialogue (AVSD), is catching an increasing amount of attention among researchers. Different from purely text-based dialogue state tracking, the dialogue in AVSD contains a sequence of question-answer pairs about a video and the final answer to the given question requires additional understanding of the video. This paper interprets the AVSD task from an open-domain Question Answering (QA) point of view and proposes a multimodal open-domain QA system to deal with the problem. The proposed QA system uses common encoder-decoder framework with multimodal fusion and attention. Teacher forcing is applied to train a natural language generator. We also propose a new data augmentation approach specifically under QA assumption. Our experiments show that our model and techniques bring significant improvements over the baseline model on the DSTC7-AVSD dataset and demonstrate the potentials of our data augmentation techniques.
Electroencephalogram (EEG) is a prominent way to measure the brain activity for studying epilepsy, thereby helping in predicting seizures. Seizure prediction is an active research area with many deep learning based approaches dominating the recent literature for solving this problem. But these models require a considerable number of patient-specific seizures to be recorded for extracting the preictal and interictal EEG data for training a classifier. The increase in sensitivity and specificity for seizure prediction using the machine learning models is noteworthy. However, the need for a significant number of patient-specific seizures and periodic retraining of the model because of non-stationary EEG creates difficulties for designing practical device for a patient. To mitigate this process, we propose a Siamese neural network based seizure prediction method that takes a wavelet transformed EEG tensor as an input with convolutional neural network (CNN) as the base network for detecting change-points in EEG. Compared to the solutions in the literature, which utilize days of EEG recordings, our method only needs one seizure for training which translates to less than ten minutes of preictal and interictal data while still getting comparable results to models which utilize multiple seizures for seizure prediction.