Photovoltaic (PV) event log data are typically underexploited mainly because of the heterogeneity of the events. To unlock these data, we propose an explorative methodology that overcomes two main constraints: (1) the rampant variability in event labelling, and (2) the unavailability of a clear methodology to traverse the amount of generated event sequences. With respect to the latter constraint, we propose to integrate heterogeneous event logs from PV plants with a semantic model of the events. However, since different manufacturers report events at different levels of granularity and since the finest granularity may sometimes not be the right level of detail for exploitable insights, we propose to explore PV event logs with Multi-level Sequential Pattern Mining. On the basis of patterns that are retrieved across taxonomic levels, several event-related processes can be optimized, e.g. by predicting PV inverter failures. The methodology is validated on real-life data from two PV plants.
Photovoltaic (PV) event log data are typically underexploited mainly because of the heterogeneity of the events. To unlock these data, we propose an explorative methodology that overcomes two main constraints: 1) the rampant variability in event labelling, and 2) the unavailability of a clear methodology to traverse the amount of generated event sequences. With respect to the latter constraint, we propose to integrate heterogeneous event logs from PV plants with a semantic model of the events. However, since different manufacturers report events at different levels of granularity and since the finest granularity may sometimes not be the right level of detail for exploitable insights, we propose to explore PV event logs with Multi-level Sequential Pattern Mining. On the basis of patterns that are retrieved across taxonomic levels, several event-related processes can be optimized, e.g. by predicting PV inverter failures. The methodology is validated on real-life data from two PV plants.
In this paper, we address the general problem of dealing with a complex and interwoven set of influential factors, potentially both linguistic and extralinguistic factors, behind linguistic choices. We do so by investigating the rare phenomenon of swearing with diseases in Dutch, rather than with the more common Western taboo concepts of (among others) religion and sexuality. The standing hypothesis is that swearing with diseases is related to the Calvinistic cultural background of the Dutch. Methodologically speaking, we perform a corpus- based analysis of a large database of location-specific tweets from Flanders and The Netherlands in which we observe "bad language". Since the hypothesized influential factor of Calvinism shows an outspoken geographical pattern, we intuitively expect clear overlap of the area in which we observe a preference for disease-based swearing and the area in which Calvinism is the predominant religion. We do not find a straightforward overlap in the geography of Calvinism and disease-based swearing. Although the locations where disease-based swearing is conspicuously frequent are all within the Calvinistic area of The Netherlands, disease-based swearing is more likely in the highly-urbanized region around Amsterdam. Therefore, we can only conclude that urbanity, socio-economic factors, religious affiliation and nationality play intertwined roles in the choice for disease-based swearing versus swearing with words from other lexical domains. Consequently, we end up with a complex and interwoven set of extra-linguistic factors that influence the lexical choice for a taboo word from a specific domain for swearing.
In this paper, we investigate the properties of Old High German relative clauses. A striking fact is that the finite verb in these constructions may either precede or follow its object(s). We survey different possible factors proposed in the literature that could determine the relative order of the verb and its objects (VO/OV order), such as type, time, and place of origin of the text, information-structural properties of the object of the relative clause, presence of a relative particle, definiteness of the antecedent, specificity of the referent, and type of the relative clause (restrictive or appositive). Our investigation is based on a corpus of nontranslated texts. It reveals that the only factors that have statistically significant influence on word order are the type of the relative clause and some information-structural properties of the object of the relative clause.
Lectometry is a corpus-based methodology that explores how multiple languageexternal dimensions shape language usage in an aggregate perspective. The paper combines this methodology with Semantic Vector Space modeling to investigate lexical variability in written Standard English, as sampled in the original Brown family of corpora (Brown, LOB, Frown and F-LOB). Based on a joint analysis of 303 lexical variables, which are semi-automatically extracted by means of a SVS, we find that lexical variation in the Brown family is systematically related to three lectal dimensions: discourse type (informative versus imaginative), standard variety (British English versus American English), and time period (1960s versus 1990s). It turns out that most lexical variables are sensitive to at least one of these three language-external dimensions, yet not every dimension has dedicated lexical variables: in particular, distinctive lexical variables for the real time dimension fail to emerge.
According to the Third Industrial Revolution, peer-to-peer electricity exchange combined with optimized local storage is the future of our electricity landscape, creating the so-called “smart grid”. Such a grid not only has to rely on predicting electricity production, but also its consumption. A growing body of literature exists on the topic of energy consumption and demand forecasting. Many contributions consist of presenting a methodology, and showing its accuracy. This paper goes beyond this common practice on two levels: first, by comparing two regression techniques to a univariate autoregressive baseline and second, by evaluating the models in term of industrial applicability, in close collaboration with domain experts. It appears that the computationally costly regression models fail to significantly beat the baseline.
Solar plants typically consist of several thousands of passive photovoltaic modules that are connected via thousands of string boxes to hundreds of inverters. In addition, a solar plant has meteo-sensors, power meters and control switches. All these components continuously generate data that is collected by monitoring systems or SCADA systems on-site. From there onwards, this data is pushed to remote analysis servers. The optimal exploitation of this data is hampered by a lack of harmonisation and standardisation in the photovoltaic domain. The data-generating components originate from several different manufacturers, models and versions, and their output is thus not easily commensurable. Crucial for this paper is the fact that conceptually identical failure events are not logged with the same identifying labels. Therefore, every analysis of monitoring system data coming from photovoltaic plants needs an initial integration step to resolve this labeling issue. Our proposal is to facilitate the integration with semantic modelling by means of creating a photovoltaic event ontology with an SWRL reasoning layer.
Lexical sociolectometry considers the aggregated behavior of many lexical variables in different language varieties to identify the multifactorial structure of variation in the lexicon. In this paper, we focus on the problem of generating a large set of lexical variables. Previous sociolectometric studies collected a set of lexical variables by hand, but this manual approach would be unfeasible when it is to be applied to large-scale sociolectometry. Therefore, we propose to automatically model the meaning of words (or a function thereof) with Semantic Vector Space Models. On the basis of these models, candidate lexical variables are automatically generated, facilitating large scale lexical sociolectometry. This methodology is applied to lexical variation in Dutch and it is shown that variation along a register dimension is more outspoken than the pluricentric variation between Dutch as used in Belgium and the Netherlands.
The current paper shows how a sociolectometric approach is needed to disentangle the multidimensional structure of the varieties in a pluricentric language. There are different sociolectometric approaches, i.e. corpus-based methods, perception experiments, or attitude questionaires. Although the focus of a sociolectometric approach is on the varieties, the choice of the variables under analysis is crucial; we focus on lexical variation. Furthermore, in this paper we compare two quantitative corpus-based methods, which differ in their conceptual control of lexical variables: on the one hand, we take a method that ignores the conceptual relationship between the lexemes in the variable set, on the other hand, there is a method that incorporates knowledge about conceptual identity between lexemes. The importance and difficulties of conceptual control when studying variation in the lexicon as a whole is shown by means of a case-study on the pluricentric language Dutch. The pluricentric character of Dutch is now widely accepted: Dutch is used both in Belgium and in the Netherlands, but each nation has its own norm generating center (cf. Clyne, 1992). This is different from the imposed situation in earlier years, especially the sixties, where Dutch in Belgium was supposed to be exogenically modeled on the norms of the Netherlands. Recently, by means of empirical work of e.g. Geeraerts et al. (1999) and experimental work of e.g. Impe et al. (2008), this historical view had to be adjusted to the current view, as described in Auer (2005). Rather than providing further empirical proof of the pluricentric character of the Dutch lexicon, the case-study aims to show the pertinence of a sociolectometric methodology that can aggregate patterns of non-categorical lexical variation while incorporating an appropriate amount of conceptual control — in contrast to a methodology that discards any conceptual knowledge. As such, the study touches upon two general issues in the broader field of variationist linguistics: on the level of words, we look at the problematic status of lexical variation and the difficulty of delineating word meaning; on the level of structure, we run
Although the aggregation of many linguistic variables has provided new insights into the structure of language varieties, aggregation studies have been criticized for obscuring the behavior of individual input variables. Previous solutions to this criticism consisted of extensive post-hoc calculations, simple correlation measures, or highly complex algorithms. We think that these solutions can be improved. Therefore, the current article proposes a creative use of Individual Differences Scaling ( INDSCAL) as an alternative, more straightforward solution. INDSCAL is a branch of Multidimensional Scaling, which is currently the preferred dimension reduction technique for most aggregation studies. The link to the existing methodology and the simplicity of its rationale are the main advantages of INDSCAL. The article introduces INDSCAL by means of a non-linguistic example, a discussion of the mathematical properties, and a case study on the lexical convergence between Belgian and Netherlandic Dutch in a corpus of language from 1950 and 1990. The case study shows how INDSCAL reproduces the results of a typical aggregation study, but elegantly keeps open the possibility of investigating the behavior of individual variables.