This paper describes a method of semi-automatic word spotting in minority languages, from one and the same Aesop fable "The North Wind and the Sun" translated in Romance languages/dialects from Hexagonal (i.e. Metropolitan) France and languages from French Polynesia. The first task consisted of finding out how a dozen words such as "wind" and "sun" were translated in over 200 versions collected in the field - taking advantage of orthographic similarity, word position and context. Occurrences of the translations were then extracted from the phone-aligned recordings. The results were judged accurate in 96-97% of cases, both on the development corpus and a test set of unseen data. Corrected alignments were then mapped and basemaps were drawn to make various linguistic phenomena immediately visible. The paper exemplifies how regular expressions may be used for this purpose. The final result, which takes the form of an online speaking atlas (enriching the https://atlas.limsi.fr website), enables us to illustrate lexical, morphological or phonetic variation.
ABSTRACT The relationship between competitive balance or intensity and venue attendance has been extensively tested in the literature. These studies have used complex metrics and econometric testing. This does not favour their understanding and application by managers. This article investigates, over a 10 years period (2009–2019), the influence of competitive intensity on stadium attendance for the French men’s football Ligue 1 through a visualization approach. The use of this method provides visual information via a single chart per season about the influence of competitive intensity on attendance, depending on the moment in the season and the sporting stakes. This means that managers can look at relevant information in one single chart. The article shows the positive influence on attendance of competitive intensity in relation to the different sporting stakes, consistent with previous studies. The visualization approach makes the results easy to grasp for managers, facilitating the identification of managerial implications.
This paper aims at analyzing the changes in the fields of speech and natural language processing over the recent past 5 years (2016–2020). It is in continuation of a series of two papers that we published in 2019 on the analysis of the NLP4NLP corpus, which contained articles published in 34 major conferences and journals in the field of speech and natural language processing, over a period of 50 years (1965–2015), and analyzed with the methods developed in the field of NLP, hence its name. The extended NLP4NLP+5 corpus now covers 55 years, comprising close to 90,000 documents [+30% compared with NLP4NLP: as many articles have been published in the single year 2020 than over the first 25 years (1965–1989)], 67,000 authors (+40%), 590,000 references (+80%), and approximately 380 million words (+40%). These analyses are conducted globally or comparatively among sources and also with the general scientific literature, with a focus on the past 5 years. It concludes in identifying profound changes in research topics as well as in the emergence of a new generation of authors and the appearance of new publications around artificial intelligence, neural networks, machine learning, and word embedding.
HAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. For a mapping of the languages/dialects of Italy and regional varieties of Italian Philippe Boula de Mareüil, Eric Bilinski, Frédéric Vernier, Valentina de Iacovo, Antonio Romano
Nous décrivons ici un atlas linguistique sonore qui prend la forme d’un site web présentant une carte interactive de Belgique, où l’on peut cliquer sur une cinquantaine de points d’enquête pour écouter et lire une même histoire en langues régionales endogènes (romanes et germaniques). Nous avons enregistré la fable d’Ésope « La bise et le soleil » (utilisée par l’Association phonétique internationale pour illustrer nombre de langues du monde) dans le but de mettre en valeur la richesse du picard et du wallon, en particulier.
We evaluate several augmentations to the choropleth map to convey additional information, including glyphs, 3D, cartograms, juxtaposed maps, and shading methods. While choropleth maps are a common method used to represent societal data, with multivariate data they can impede as much as improve understanding. In particular large, low population density regions often dominate the map and can mislead the viewer as to the message conveyed. Our results highlight the potential of 3D choropleth maps as well as the low accuracy of choropleth map tasks with multivariate data. We also introduce and evaluate popcharts, four techniques designed to show the density of population at a very fine scale on top of choropleth maps. All the data, results, and scripts are available from https://osf.io/8rxwg/.
Gamification can be seen as the intentional use of game design elements in non-game tasks, in order to produce psychological outcomes likely to influence behaviour and/or performance. In this respect, we hypothesize that gamification would produce measurable effects on user performance, that this positive impact would be mediated by specific motivational and attentional processes such as flow and that gamification would moderate the social comparison process. In three experimental studies, we examine the effects of gamified electronic brainstorming interfaces on fluency, uniqueness and flow. The first study mainly focuses on time pressure, the second on performance standard and the third one introduces social comparison. The results highlight some effects of the gamified conditions on brainstorming performance, but no or negative effects on flow. All three studies are congruent in that gamification did not occur as a psychological process, which questions popular design trends observed in a number of sectors.
Population information is nowadays available at fine scales and almost always missing from classical choropleth maps. We implemented five techniques to overlay fine-scaled population information on choropleth maps and present the results of an evaluation of three of them. Our results suggest that such overlays do not hinder the understanding of choropleth maps and can help understand the distribution of population on maps, thus possibly avoiding the common pitfalls of choropleth maps.
We describe here a speaking atlas that takes the form of a website presenting interactive maps, where it is possible to click on 260 survey points to listen to as many speech samples and read a transcript of what is said, in regional and minority languages of Hexagonal (i.e. Metropolitan) France and its Overseas Territories. We show how an attractive website enables us to collect more data in underresourced and endangered languages and how these data may be used for phonetic analyses and dialectometry purposes. A one-minute story (“The North Wind and the Sun”) was used, phonetically transcribed automatically by grapheme-to-phoneme converters and forced aligned with the audio signal: a methodology which can be applied to other languages and dialects.
The NLP4NLP corpus contains articles published in 34 major conferences and journals in the field of speech and natural language processing over a period of 50 years (1965-2015), comprising 65,000 documents, gathering 50,000 authors, including 325,000 references and representing approximately 270 million words. This paper presents an analysis of this corpus regarding the evolution of the research topics, with the identification of the authors who introduced them and of the publication where they were first presented, and the detection of epistemological ruptures. Linking the metadata, the paper content and the references allowed us to propose a measure of innovation for the research topics, the authors and the publications. In addition, it allowed us to study the use of language resources, in the framework of the paradigm shift between knowledge-based approaches and content-based approaches, and the reuse of articles and plagiarism between sources over time. Numerous manual corrections were necessary, which demonstrated the importance of establishing standards for uniquely identifying authors, articles, resources or publications.
Multiclass maps are scatterplots, multidimensional projections, or thematic geographic maps where data points have a categorical attribute in addition to two quantitative attributes. This categorical attribute is often rendered using shape or color, which does not scale when overplotting occurs. When the number of data points increases, multiclass maps must resort to data aggregation to remain readable. We present multiclass density maps: multiple 2D histograms computed for each of the category values. Multiclass density maps are meant as a building block to improve the expressiveness and scalability of multiclass map visualization. In this article, we first present a short survey of aggregated multiclass maps, mainly from cartography. We then introduce a declarative model a simple yet expressive JSON grammar associated with visual semantics that specifies a wide design space of visualizations for multiclass density maps. Our declarative model is expressive and can be efficiently implemented in visualization front-ends such as modern web browsers. Furthermore, it can be reconfigured dynamically to support data exploration tasks without recomputing the raw data. Finally, we demonstrate how our model can be used to reproduce examples from the past and support exploring data at scale.
The aim is to show and promote the linguistic diversity of France, through field recordings, a computer program (which allows us to visualise dialectal areas) and an orthographic transcription (which represents an object of research in itself). A website is presented (https://atlas.limsi.fr), displaying an interactive map of France from which Aesopâs fable âThe North Wind and the Sunâ can be listened to and read in French and in 140 varieties of regional languages. There is thus both a scientific dimension and a heritage dimension in this work, insofar as a number of regional or minority languages are in a critical situation.
This chapter presents ongoing research dedicated to augmenting creativity through innovative technologies. Our hypotheses draw on the pros and cons of the brainstorming paradigm to strengthen the former and overcome the latter. The main efficiency factors we are trying to support are Cognitive stimulation, Social comparison, and Group facilitation, while trying to circumvent Production blocking, Social loafing, and Self-censorship. The first technology reviewed is electronic brainstorming systems: It was shown that such devices enable groups, even large ones, to avoid production blocking. However, it may also increase social loafing, which is detrimental to creativity. We then introduce interactive tabletop brainstorming with which groups can conciliate individual reflection, idea sharing, and social setting. We show that this technology reduces social loafing, and we provide interface designs that further support cognitive stimulation, social comparison, and group facilitation. This series of experiments also highlights a new efficiency factor for creativity, namely the Fun factor: The use of innovative technology in itself introduces playfulness, which seems to increase engagement and creative performance. Finally, we report on a recent series of experiments exploring avatar-mediated creativity as a means to counter self-censorship through anonymity and enhance creativity through avatars' appearance. The results confirm that the choice of avatars in virtual brainstorming greatly influences creativity through processes such as self-perception, priming, and social identity. In many respects, avatars and virtual environments offer a new promising tool to support group creativity. We conclude on the potential impact of these findings on real-world innovation challenges.
We have created the NLP4NLP corpus to study the content of scientific publications in the field of speech and natural language processing. It contains articles published in 34 major conferences and journals in that field over a period of 50 years (1965-2015). comprising 65.000 documents. gathering 50.000 authors. including 325.000 references and representing approximately 270 million words. Most of these publications are in English. some are in French. German or Russian. Some are open access. others have been provided by the publishers. In order to constitute and analyze this corpus several tools have been used or developed. Some of them use Natural Language Processing methods that have been published in the corpus. hence its name. Numerous manual corrections were necessary. which demonstrated the importance of establishing standards for uniquely identifying authors. publications or resources. We have conducted various studies: evolution over time of the number of articles and authors. collaborations between authors. citations between papers and authors. evolution of research themes and identification of the authors who introduced them. measure of innovation and detection of epistemological ruptures. use of language resources. reuse of articles and plagiarism in the context of a global or comparative analysis between sources.
By its own nature, the Natural Language Processing (NLP) community is a priori the best equipped to study the evolution of its own publications, but works in this direction are rare and only recently have we seen a few attempts at charting the field. In this paper, we use the algorithms, resources, standards, tools and common practices of the NLP field to build a list of terms characteristic of ongoing research, by mining a large corpus of scientific publications, aiming at the largest possible exhaustivity and covering the largest possible time span. Study of the evolution of this term list through time reveals interesting insights on the dynamics of field and the availability of the term database and of the corpus (for a large part) make possible many further comparative studies in addition to providing a test field for a new graphic interface designed to perform visual time analytics of large sized thesauri.
To address the limitations of traditional line chart approaches, in particular rank charts (RCs) and score charts (SCs), a novel class of line charts called gap charts (GCs) show entries that are ranked over time according to a performance metric. The main advantages of GCs are that entries never overlap (only changes in rank generate limited overlap between time steps) and gaps between entries show the magnitude of their score difference. The authors evaluate the effectiveness of GCs for performing different types of tasks and find that they outperform standard time-dependent ranking visualizations for tasks that involve identifying and understanding evolutions in both ranks and scores. They also show that GCs are a generic and scalable class of line charts by applying them to a variety of different datasets.
Le present travail se propose de renouveler les traditionnels atlas dialectologiques pour cartographier les variantes de prononciation en francais, a travers un site internet. La toile est utilisee non seulement pour collecter des donnees, mais encore pour disseminer les resultats aupres des chercheurs et du grand public. La methodologie utilisee, a base de crowdsourcing (ou « production participative »), nous a permis de recueillir des informations aupres de 2500 francophones d’Europe (France, Belgique, Suisse). Une plateforme dynamique a l’interface conviviale a ensuite ete developpee pour cartographier la prononciation de 70 mots dans les differentes regions des pays concernes (des mots notamment a voyelle moyenne ou dont la consonne finale peut etre prononcee ou non). Les options de visualisation par departement/canton/province ou par region, combinant plusieurs traits de prononciation et ensembles de mots, sous forme de pastilles colorees, de hachures, etc. sont presentees dans cet article. On peut ainsi observer immediatement un /E/ plus ferme (ainsi qu’un /O/ plus ouvert) dans le Nord-Pas-de-Calais et le sud de la France, pour des mots comme parfait ou rose, un /Œ/ plus ferme en Suisse pour un mot comme gueule, par exemple.
According to the Search for Ideas in Associative Memory theory, ideas in a brainstorming session do not come one by one but rather in “trains of thought,” which are rapid accumulations of semantically related ideas. In order to visualize these trains of thought, we developed a brainwriting tabletop interface enabling users to link successive ideas together by means of graphical ropes. To test the effectiveness of this device, 48 participants (in groups of four) brainstormed for 20 min on the tabletop in one of two conditions: either with the train-of-thought interface (with graphical ropes), or without the ropes (control condition). The results show that visualizing the associations between ideas enabled the participants to produce longer trains of thought. We also assessed originality by collecting the unique ideas in the whole corpus of ideas produced by the different groups and observed that the train-of-thought condition produced more original ideas than the control one. One interpretation of this finding is that visualizing trains of thought increases cognitive stimulation, i.e., improves creativity by making others’ ideas more intelligible to the brainstorming partners, in comparison with the classical visualization of ideas as independent items.
Our visualization shows the temporal evolution of " Le Tour de France 2014 " cyclists ranking after each stage. In contrast to standard snapshot ranking tables used to represent this ranking, our visualization clearly shows the temporal evolution of each cyclist's ranking and global trends for each stage. Moreover, our layout algorithm 1) ensures no overlap, tied cyclists at a given stage being adjacent instead of superimposed; and 2) conveys the gap magnitudes between cyclists: the larger the white space between cyclists, the further they are. Additionally, thin gray lines represent gaps of 1 minute. Color encodes cyclist's nationality and stage miniatures provide context on the top of the visualization. For example, we see that the second stage was a key stage, involving many changes. Then the ranking remained stable for two stages, before changing again. We also observe that gap magnitudes increase with time; that flat stages do not impact the rankings neither the gap magnitudes much, whereas stages in mountain strongly impact both; and the expert eye will see that Valverde lost his second position three stages before the end of the race, and that Porte constantly decreased in ranking during high mountain stages, while he was ranked second.
Stéphanie Buisine合作论文数Ecole Nationale Superieure d'Arts et Metiers7
Christian Jacquemin合作论文数3