The sharing of data between financial institutions is widely recognised as a key component in the industry’s efforts to combat fraud. Broader access to multiple sources of financial data is also critical to the development of high-quality fraud detection mechanisms based on artificial intelligence (AI). Given the challenges relating to sharing real financial data across countries and institutions, the use of synthetic data has recently become critical to enabling the exploration of broader data sharing and supporting open collaboration in AI model development. To generate synthetic data that can substitute for real data, computer algorithms closely mimic the key statistical properties of genuine data, while strictly preserving the privacy and sovereignty of the source data. This paper presents the results of an ongoing exploration into the generation of high-utility synthetic datasets of cross-border payment transactions using transformer models and discusses its application to the development of AI-based fraud prevention solutions.
Compelling public interest is propelling national efforts to advance the evidence base for cancer treatment and control measures and to transform the way in which evidence is aggregated and applied. Substantial investments in health information technology, comparative effectiveness research, health care quality and value, and personalized medicine support these efforts and have resulted in considerable progress to date. An emerging initiative, and one that integrates these converging approaches to improving health care, is "rapid-learning health care." In this framework, routinely collected real-time clinical data drive the process of scientific discovery, which becomes a natural outgrowth of patient care. To better understand the state of the rapid-learning health care model and its potential implications for oncology, the National Cancer Policy Forum of the Institute of Medicine held a workshop entitled "A Foundation for Evidence-Driven Practice: A Rapid-Learning System for Cancer Care" in October 2009. Participants examined the elements of a rapid-learning system for cancer, including registries and databases, emerging information technology, patient-centered and -driven clinical decision support, patient engagement, culture change, clinical practice guidelines, point-of-care needs in clinical oncology, and federal policy issues and implications. This Special Article reviews the activities of the workshop and sets the stage to move from vision to action.
Providing near-term prognostic insight to clinicians helps them to better assess the near-term impact of their decisions and potential impending events affecting the patient. In this work, we present a novel system, which leverages inter-patient similarity for retrieving patients who display similar trends in their physiological time-series data. Data from the retrieved patient cohort is then used to project patient data into the future to provide insights for the query patient. The proposed approach and system were tested using the MIMIC II database, which consists of physiological waveforms, and accompanying clinical data obtained for ICU patients. In the experiments we report the effectiveness of the inter-patient similarity measure and the accuracy of the projection of patients' data. We also discuss the visual interface that conveys the near-term prognostic decision support to the user.
We have made significant progress in automatic speech recognition (ASR) for well-defined applications like dictation and medium vocabulary transaction processing tasks in relatively controlled environments. However, for ASR to approach human levels of performance and for speech to become a truly pervasive user interface, we need novel, nontraditional approaches that have the potential of yielding dramatic ASR improvements. Visual speech is one such source for making large improvements in high...
The paper overviews recent progress and challenges in a number of audiovisual speech processing technologies with main emphasis on the problem of automatic speech recognition. It is well known that visual channel information can improve automatic speech processing for human-computer interaction. To automatically process and incorporate such information into automatic systems, a number of steps are required that are surprisingly similar accross speech technologies. Crucial above all is the issue of feature representation of visual speech and its robust extraction. In addition, appropriate integration of the audio and visual representations is required, in order to ensure improved performance of the bimodal systems over audio-only baselines. These topics are discussed in detail in the talk, with main emphasis on their application to the speech recognition problem in the challenging environments of automobiles and smart rooms.
It is well known that frontal video of the speaker's mouth region contains significant speech information that, when combined with the acoustic signal, can improve accuracy and noise robustness of automatic speech recognition (ASR) systems. However, extraction of such visual speech information from full-face videos is computationally expensive, as it requires tracking faces and facial features. In addition, robust face detection remains challenging in practical human–computer interaction (HCI), where the subject's posture and environment (lighting, background) are hard to control, and thus successfully compensate for. In this paper, in order to bypass these hindrances to practical bimodal ASR, we consider the use of a specially designed, wearable audio-visual headset, a feasible solution in certain HCI scenarios. Such a headset can consistently focus on the speaker's mouth region, thus eliminating altogether the need for face tracking. In addition, it employs infrared illumination to provide robustness against severe lighting variations. We study the appropriateness of this novel device for audio-visual ASR by conducting both small- and large-vocabulary recognition experiments on data recorded using it under various lighting conditions. We benchmark the resulting ASR performance against bimodal data containing frontal, full-face videos collected at an ideal, studio-like environment, under uniform lighting. The experiments demonstrate that the infrared headset video contains comparable speech information to the studio, full-face video data, thus being a viable sensory device for audio-visual ASR.
This paper describes multimodal systems for ad-hoc search constructed by IBM for the TRECVID 2003 benchmark of search systems for broadcast video. These systems all use a late fusion of independently developed speech-based and visual content-based retrieval systems and outperform our individual retrieval systems on both manual and interactive search tasks. For the manual task, our best system used a query-dependent linear weighting between speech-based and image-based retrieval systems. This system has mean average precision (MAP) performance 20% above our best unimodal system for manual search. For the interactive task, where the user has full knowledge of the query topic and the performance of the individual search systems, our best system used an interlacing approach. The user determines the (subjectively) optimal weights A and B for the speech-based and image-based systems, where the multimodal result set is aggregated by combining the top A documents from system A followed by top B documents of system B and then repeating this process until the desired result set size is achieved. This multimodal interactive search has MAP 40% above our best unimodal interactive search system.
We have made significant progress in automatic speech recognition (ASR) for well-defined applications like dictation and medium vocabulary transaction processing tasks in relatively controlled environments. However, ASR performance has yet to reach the level required for speech to become a truly pervasive user interface. Indeed, even in “clean” acoustic environments, and for a variety of tasks, state of the art ASR system performance lags human speech perception by up to an order of magnitude (Lippmann, 1997). In addition, current systems are quite sensitive to channel, environment, and style of speech variations. A number of techniques for improving ASR robustness have met limited success in severely degraded environments, mismatched to system training (Ghitza, 1986; Nadas et al., 1989; Juang, 1991; Liu et al., 1993; Hermansky and Morgan, 1994; Neti, 1994; Gales, 1997; Jiang et al., 2001). Clearly, novel, non-traditional approaches, that use orthogonal sources of information to the acoustic input, are needed to achieve ASR performance closer to the human speech perception level, and robust enough to be deployable in field applications. Visual speech is the most promising source of additional speech information, and it is obviously not affected by the acoustic environment and noise.
Ching-Yung Lin (林清詠)合作论文数IBM T. J. Watson Rsearch Center5