The detection of abnormal operation modes is of fundamental importance for both operational management and predictive maintenance of wind turbines. Anomaly detection approaches in this context should consider the additional information content that probabilistic models can provide. Instead of binary anomaly classification, the probabilistic information is necessary for proper decision making and risk assessment. Common models, such as quantile and distribution regression can provide probabilistic information. While they are appropriate in predicting the cumulative distribution function, they struggle to accurately describe the probability of an event to occur. In this article we present a new, multi-task learning based approach for a continuous distribution regression with deep neural networks. Using real-world data from an offshore wind turbine, we show that with this model we can better reflect the probability of observed events than with conventional methods. While the predicted cumulative distribution function has a similar quality and no significant differences are visible in the continuous ranked probability score, the probability density function will be substantially smoother. This is also reflected in a significantly lower ignorance score.
This paper presents a market model for the EPEX SPOT German continuous intraday market for electric power trading based on the limit order book (LOB). We use the EPEX SPOT M7 order book data, which contains all orders submitted to the German continuous intraday market, to simulate the historic course of the market. Thereby, we reconstruct the complete state of the LOB at every point in (trading) time. We validate our simulation by comparing the transactions that our simulation generated with the actual historical transactions available from a different data set. The LOB based market model can be used to include price volatility risk and illiquidity risk when simulating trading at the EPEX SPOT continuous intraday market. Furthermore, we present all preprocessing steps and decision rules necessary to correctly identify orders from the often ambiguous EPEX SPOT M7 order book data.
This paper presents an approach to quantify the reliability or security level of a pool of flexibility providing controllable energy units (CEUs). In this pool, a backup unit secures the flexibility capacity of the largest flexibility providing unit. The security level is then applied to determine the flexibility capacity of a common pool of wind farms and CEUs, making use of probabilistic wind forecasts. The paper advances previous research conducted in a joint project with industry partners on Frequency Restoration Reserve (FRR) provision of pools of wind farms and controllable power plants.
On last year's EEM conference a method for the dynamic sizing of frequency restoration reserve capacity based on quantile regression was presented. Further research has improved the method and has made it ready for use. It contains the following new features, which will be presented in this paper: an adaptive bias correction function, the allocation of frequency restoration reserves (FRR) to automatic FRR (aFRR) and manual FRR (mFRR) and the calculation of needed reserve capacity for different product lengths. The results will show the advantages of the adaptive bias correction, estimate the needed reserves for aFRR and mFRR, and demonstrate the influence of different product lengths on the needed reserve capacities.
We describe the ICSI-SRI-UW team’s entry in the Spring 2004 NIST Meeting Recognition Evaluation. The system was derived from SRI’s 5xRT Conversational Telephone Speech (CTS) recognizer by adapting CTS acoustic and language models to the Meeting domain, adding noise reduction and delay-sum array processing for far-field recognition, and postprocessing for cross-talk suppression. A modified MAP adaptation procedure was developed to make best use of discriminatively trained (MMIE) prior models. These meeting-specific changes yielded an overall 9% and 22% relative improvement as compared to the original CTS system, and 16% and 29% relative improvement as compared to our 2002 Meeting Evaluation system, for the individual-headset and multiple-distant microphones conditions, respectively.
This thesis proposes several improvements to the correlation-based location features recently used in meeting speaker diarization (answering the question, "Who spoke when?"). The problem of leveraging time delay information is examined for multi-microphone meeting environments, where microphones are placed at unknown, widely spaced, and ad-hoc locations. In addition, conversational speech is challenging because of the many short utterances and speaker overlaps. Finally, assuming no room constraints, the microphone configuration and acoustic environment changes from meeting to meeting. Together, these conditions make it impractical to apply standard localization and beamforming techniques. To address these challenges, we first consider what combination of channel pairs and signal processing to use for location information extraction. Initially, we consider all pairs, then de-emphasizing low quality pairs with feature vector dimension reduction. We also develop an approach for fusing speaker ID information as viewed by different physical processes. Two views are a new time delay estimate and multi-band energy ratios (cues to location) and a third is a vector of mel-warped cepstral coefficients (MFCC's), related to vocal tract characteristics. We find that both MFCC's and energy ratios can improve time delay information when jointly transformed using canonical correlation analysis (CCA). Oracle experiments show that the location feature dimension producing the best diarization error varies with meeting. Therefore, we evaluate automatic methods for determining feature reduction output dimension. In addition, we separately consider reducing the feature dimension by explictly selecting subsets of channel pairs using estimated signal to noise ratio (SNR) and information-theoretic feature selection methods. Location features are also employed to detect speaker overlap, a significant cause of increased speaker diarization error. First, monaural overlap features are developed for a single channel beamformer output. These features are then compared to overlap detector features which make use of location information, but neither type provides good performance due to a high degree of variation across meetings. We also develop a simple, nearest-neighbor overlap processing scheme which, when given accurate overlap detection, improves diarization accuracy. Together, these results underscore the need for dynamic models to handle variable room and recording configurations.
Speaker overlap in meetings is thought to be a significant contributor to error in speaker diarization, but it is not clear if overlaps are problematic for speaker clustering and/or if errors could be addressed by assigning multiple labels in overlap regions. In this paper, we look at these issues experimentally, assuming perfect detection of overlaps, to assess the relative importance of these problems and the potential impact of overlap detection. With our best features, we find that detecting overlaps could potentially improve diarization accuracy by 15% relative, using a simple strategy of assigning speaker labels in overlap regions according to the labels of the neighboring segments. In addition, the use of cross-correlation features with MFCC's reduces the performance gap due to overlaps, so that there is little gain from removing overlapped regions before clustering.
The paper describes our system devised for recognizing speech in meetings, which was an entry in the NIST Spring 2004 Meeting Recognition Evaluation. This system was developed as a collaborative effort between ICSI, SRI, and UW and was based on SRI's 5xRT Conversational Telephone Speech (CTS) recognizer. The CTS system was adapted to the Meetings domain by adapting the CTS acoustic and language models to the Meeting domain, adding noise reduction and delay-sum array processing for far-field recognition, and adding postprocessing for cross-talk suppression for close-talking microphones. A modified MAP adaptation procedure was developed to make best use of discriminatively trained (MMIE) prior models. These meeting-specific changes yielded an overall 9% and 22% relative improvement as compared to the original CTS system, and 16% and 29% relative improvement as compared to our 2002 Meeting Evaluation system, for the individual-headset and multiple-distant microphones conditions, respectively.
We describe the ICSI-SRI-UW team’s entry in the Spring 2004 NIST Meeting Recognition Evaluation. The system was derived from SRI’s 5xRT Conversational Telephone Speech (CTS) recognizer by adapting CTS acoustic and language models to the Meeting domain, adding noise reduction and delay-sum array processing for far-field recognition, and postprocessing for cross-talk suppression. A modified MAP adaptation procedure was developed to make best use of discriminatively trained (MMIE) prior models. These meeting-specific changes yielded an overall 9% and 22% relative improvement as compared to the original CTS system, and 16% and 29% relative improvement as compared to our 2002 Meeting Evaluation system, for the individual-headset and multiple-distant microphones conditions, respectively.
This paper explores packet loss recovery for automatic speech recognition (ASR) in spoken dialog systems, assuming an architecture in which a lightweight client communicates with a remote ASR server. Speech is transmitted with source and channel codes optimized for the ASR application, i.e., to minimize word error rate. Unequal amounts of forward error correction, depending on the data's effect on ASR performance, are assigned to protect against packet loss. Experiments with simulated packet loss in a range of loss conditions are conducted on the DARPA Communicator (air travel information) task. Results show that the approach provides robust ASR performance which degrades gracefully as packet loss rates increase. Transmitting at 5.2 Kbps with up to 200 ms added delay, leads to only a 7% relative degradation in word error rate even under extremely adverse network conditions.
This document contains information, which is proprietary to the “EERA-DTOC” Consortium. Neither this document nor the information contained herein shall be used, duplicated or communicated by any means to any third party, in whole or in parts, except with prior written consent of the “EERA-DTOC” consortium. Report on design tool on variability and predictability Nicolaos A. Cutululis, Luis Mariano Faiella, Scott Otterson, Braulio Barahona, Jan Dobschinski