Distinguishing between web traffic generated by bots and humans is an important task in the evaluation of online marketing campaigns. One of the main challenges is related to only partial availability of the performance metrics: although some users can be unambiguously classified as bots, the correct label is uncertain in many cases. This calls for the use of classifiers capable of explaining their decisions. This paper demonstrates two such mechanisms based on features carefully engineered from web logs. The first is a man-made rule-based system. The second is a hierarchical model that first performs clustering and next classification using human-centred, interpretable methods. The stability of the proposed methods is analyzed and a minimal set of features that convey the class-discriminating information is selected. The proposed data processing and analysis methodology are successfully applied to real-world data sets from online publishers.
Group decision-making involves selecting among limited options. Experts share their perspectives in a comparative debate but then evaluate alternatives using preference relations, which can result in inconsistencies between their expressions and assessments. To address this, a method is proposed that automates the generation of these relationships from the debate comments, classifying them into positive and negative using sentiment analysis, namely with the Large Language Model. A new operator is introduced that weights these comments to calculate preference relations. Furthermore, modification of the relationships is allowed if the experts so wish. Moreover, another operator is incorporated that adjusts the weight of each expert according to his or her active participation in the discussion, assigning more weight to those who contribute more comments. Finally, this innovative method promotes coherence and equal participation in group decision-making by employing an innovative sentiment analysis detection system.
This book presents ample, richly illustrated account on results and experience from the analysis of data concerning behavior patterns on the Web.
We start with the description of general context and with formulation of the problem that we address. In this manner we set a framework for both the particular issues that we deal with on a technical level, described in the consecutive parts of the book, and for the potential implications thereof, some of them forwarded as more general conclusions or hypotheses.
The case study presented in this book highlights the properties and challenges of distinguishing bot and human traffic using weblogs and compares several solutions to this task. We present here, first, the observations related to the methodological side of the potential problem solution.
Classification is a decision-making problem, in which we aim at the assignment of a correct class label to an observation. In the typical scenario, the set of available class labels is fixed beforehand and remains unchanged. Advanced data processing streams allow for a more flexible definition of this task. Typically, they admit the existence of a “novelty” class, to which some observations may be assigned.
Having characterised the problem that we address and the way of acquiring data that we use, we now turn to the features, which can be extracted from the raw data, and the choice of the possibly good selection of a subset of these variables from the point of view of the problem at hand.
When dealing with challenging, real-world datasets, one may turn to hybrid data processing approaches that join several kinds of data analysis algorithms. The hybrid data analysis techniques are popular in the literature, where we may find fusions of optimization and classification algorithms, clustering and classification algorithms.
Analyzing data from the web is now one of the primary tasks, understood in a variety of manners and solved for a very wide variety of purposes. The talk describes the experience from a project, devoted to analyzing such data while drawing some more general conclusions. The project was aimed at distinguishing artificial ad-related traffic from the genuine one. The rationale is simple: The flow of money depends upon the number of clicks on/views of an ad. If so, fake clicking changes the market to the benefit of some, and to the loss of the other ones. The talk describes the problem and its conceptual framing, as well as a number of technical details, involving the issues and techniques of (1) variable analysis and choice; (2) clustering; (3) classification/classifiers; (4) potential hybrid techniques, along with citations of the most interesting results. These often imply definite general conclusions, some of them quite surprising.
In this chapter, we will look in a more detail at the ways in which data were acquired and processed in the framework of the project in question. We will do so against the background of observations and examples already provided in the preceding section, starting with the first section of the present chapter.
Social choice function or voting procedure is one of the crucial concepts in the domain of political sciences. It maps individuals' preferences over a set of candidates to some subset (possibly one-element) of the candidates who can be thought as the winners of an election procedure. The paper is aimed at applications of formal concept analysis methods to study of social choice functions. We will construct concept lattices over selected set of social choice functions characterized by possessing some properties deemed as important from the point of view of political sciences. We will discuss issues connected with reducibility of both objects and attributes, irreducibility of object concepts as well as attribute concepts and attribute implications. We will discuss also the shape of the constructed concept lattice of social choice functions which in some part is exceptionally regular from the perspective of the lattice theory.
The basic study of fuzzy sets theory was introduced by Lotfi Zadeh in 1965. Many authors investigated possibilities how two fuzzy sets can be compared and the most common kind of measures used in the mathematical literature are dissimilarity measures. The previous approach to the dissimilarities is too restrictive, because the third axiom in the definition of dissimilarity measure assumes the inclusion relation between fuzzy sets. While there exist many pairs of fuzzy sets, which are incomparable to each other with respect to the inclusion relation. Therefore we need some new concept for measuring a difference between fuzzy sets so that it could be applied for arbitrary fuzzy sets. We focus on the special class of so called local divergences. In the next part we discuss the divergences defined on more general objects, namely intuitionistic fuzzy sets. In this case we define the local property modified to this object. We discuss also the relation of usual divergences between fuzzy sets to the divergences between intuitionistic fuzzy sets.
Abstract We propose a new approach to the bipolar database queries, which involve a necessary (required) and optional (desired) conditions, connected with a non-conventional aggregation operator “and possibly”, combined with a context, exemplified by “find houses which are cheap and – with respect to other houses in town – possibly close to a railroad station”. We use our winnow operator based interpretation of the bipolar queries. We assume that the query, posed by the human user, involves terms, which do not directly relate to attributes, and which are then to be decoded using a concept of a query hierarchy, leading to the queries, which involve terms directly related to attribute values. The original query is considered to be of level 0, at the bottom of the precisiation hierarchy, then its required and optional parts are assumed to be bipolar queries themselves, both accounting for context. The precisiation proceeds further, to level 1 queries, level 2, etc. A real estate related example is provided as illustration.
We propose a further extension of a multistage fuzzy control model of stable sustainable regional agriculture development that involves an additional capacity to reflect a stability requirement which addresses a clear preference of the stakeholders for a limited variability of crucial development indicators and parameters. We presented the use of fuzzy dynamic programming for solving the problem in which many crucial aspects, in particular life quality indicators, are subject to objective, by the authorities, and subjective, by the inhabitants, evaluations which are closely related to human perception and cognitive abilities. This model is then augmented with a requirement of a limited variability of crucial development indicators, parameters, etc. For illustration, we have shown a simple example in which the problem is to determine the best (optimal) investment policy under different development scenarios, and subject to objective and subjective evaluations.
We are concerned with 17 Sustainable Development Goals (SDGs) the fulfillment of which is crucial for both the whole world and particular nations. The assessment of attainment of targets included in these SDGs involves human judgments, preferences, intentions, etc. which are human specific and can often be expressed in natural language possibly with a provision for the handling of imprecision. We propose to use fuzzy logic based linguistic data summaries for the assessment and evaluation of both the essence of SDGs and their fulfillment, and such a form is very human consistent because it uses natural language that is the only fully natural means of articulation and communication for the humans. We show the application of linguistic summaries for the analyses of innovation and innovativeness which is explicitly involved in the ninth sustainable development goal (SDG 9). However, as innovations are considered as relevant for all SDGs, then our approach can be used for the analyses of other SDGs too.
Online advertising campaigns are adversely affected by bot traffic. In this paper, we develop and test a method for the estimation of its share, which is necessary for the evaluation of campaign efficiency. First, we present the nature of the problem as well as the underlying business rationale. Next, we describe the essential features of Internet traffic, which ought to be accounted for, and the potential methodologies, which can be used to reach the objective of the project. Finally, some of the results are provided, along with the respective discussion, followed by both technical and also more general conclusions.
The paradigm of the reverse clustering is presented, consisting in attepting to reconstruct, via some clustering procedure, of a definite partition of a given set of data. We give some examples of such real-life problems. The reverse clustering problem is formulated, and its component parts described, i.e. the clustering algorithms considered, their parameters, the weighing or selection of variables, characterising the data items, the definition of distance, and, finally, the measure of similarity of the given definite partition and the one obtained from the reverse clustering procedure. Notation used throughout the book is also introduced. Some comments are given on the applied search procedure.
Adnan Yazıcı合作论文数Department of Computer Engineering
Middle East Technical University2