Assessing changes in community diversity is a central issue with regard to many fundamental and applied aspects in ecology, biogeography and conservation. However, some important features for the assessment of community by diversity measures still need to be considered to quantify spatio-temporal variation, notably their explicit variation through time while integrating network structure and species differences. Here we introduce a new framework based on complex network analysis and on an extension of Hill numbers. It encompasses three main aspects: i) the differences/distances between species (e.g., co-occurrence, function, phylogeny, taxonomy), ii) the topological properties of the network, and iii) the variability along the time dimension by introducing three measures quantifying the overall temporal dynamics of the network. To illustrate how the new framework reveals complementary patterns beyond those identified by traditional approaches in previous studies, we analyze two data sets that vary in both their temporal extent and the types of communities they represent: (i) Mediterranean exploited fish communities sampled during 25 years, (ii) Amazonian bat communities surveyed during 4 years. Our results showed that the fish communities could be classified into five different clusters according to their spatio-temporal behaviors, as well as environmental and fishing forcings. Bat community diversity has higher values in larger forest fragment areas, while habitats with small areas exhibit very high changes through time mainly due to species turnover. By simultaneously incorporating essential features of communities, the new framework enhances identification of their spatio-temporal trends, and helps to identify priority zones of interest for management and conservation.
The development of new methods for covariate model building (CMB) in population pharmacokinetics (popPK) highlights the need for a standardized evaluation framework for method benchmarking. This data paper introduces PMX-CovEval, a framework including a collection of datasets based on 127 distinct scenarios. These scenarios aim to reflect the diversity of models, available clinical studies, and covariates encountered in real-world applications, while remaining limited enough in number to encourage their practical use. To support the evaluation of standard CMB techniques available on PsN and Monolix, model files for NONMEM and Monolix are provided alongside the datasets. Additionally, empirical Bayes estimates (EBEs) from the true base models are included to facilitate the testing of EBE-based regression approaches. By offering pharmacokinetic (PK) datasets, model files, and EBEs in a unified resource, PMX-CovEval provides a standardized, reproducible framework for evaluation and systematic comparison of CMB strategies in popPK. Initial benchmarking results are also provided.
Listwise reranking with large language models (LLMs) is effective for zero-shot information retrieval, but sliding-window inference over large candidate sets remains costly. We study a filter-then-rerank (FtR) design in which the initial retrieved pool is first reduced to a compact shortlist, then reranked in a single listwise LLM pass. We compare embedding-based, cross-encoder, and LLM-based filters combined with several listwise rerankers on the TREC Deep Learning 2019 and 2020 benchmarks. Results show that intermediate filtering yields a stronger effectiveness–efficiency trade-off than sliding-window reranking. Across both benchmarks, shortlists of 10 to 30 passages are sufficient for strong performance. Depending on the benchmark and candidate depth, FtR matches or improves over sliding-window reranking while requiring only one listwise inference call. Using the best-performing configuration for each dataset, FtR achieves an nDCG@10 of 0.756 on TREC-DL 2019 and 0.732 on TREC-DL 2020, while reducing end-to-end online query latency by about 67–68
Thematic maps, like choropleth maps, symbol maps or cartograms, are commonly used to visualize spatial quantitative data. Many studies have been conducted to compare the different approaches and thus define the best strategies to produce suitable and efficient maps. When analyzing spatial data, it is also often necessary to visualize and compare several variables on the same map. Therefore, the question arises of how to best associate two variables in a single representation without one of them prevailing over the other, while avoiding overloading the map and making it difficult to interpret. In this article, we propose a comparison of five types of bivariate maps based on a user study. Participants performed a set of tasks using different maps produced from multiple datasets. Our analysis is based on three approaches: (1) quantitative analysis of user answer accuracy, (2) quantitative analysis of user answer times, and (3) quantitative and qualitative analysis of user feedback. The results suggest that combining symbol and choropleth maps is the most effective approach among those tested, while combining cartograms with any technique is the worst.
Recently, some works have built adaptive systems providing assistance to the user in virtual reality (VR), with little or no knowledge of the user's task. These task-blind help systems can influence behaviours and exploration strategies; however, their ability to significantly improve users' performance on their tasks is still unclear. In this study, we aim to clarify the impact of task-blind help systems on user performance. We also explore two avenues that could provide a better understanding of why these systems can be effective and interesting to study. Our controlled user study involved 56 participants in an immersive analytics environment and compared four VR help-system configurations, including three task-blind systems and a no-assistance baseline. Results showed significant task performance improvements with one task-blind system, highlighting user control as a key factor of efficiency. This work demonstrates the potential of task-blind help systems, offering a flexible framework for adaptive design and raising questions about their broader applications.
The exploration of complex environments (reconstructed locations, immersive data visualisation, etc.) is one of the primary applications of virtual reality (VR) because of the feeling of immersion and the natural interactions that it provides. When exploration is completely free, users easily become disoriented and frustrated due to multiple factors such as task difficulty, interaction techniques, spatial understanding, immersion breaches, etc. Adaptive VR systems aim to overcome these difficulties and increase performance by providing clues to help the user and delivering effective feedback. Current adaptive VR systems are mainly task-based, meaning that they have explicit knowledge about what users must achieve. We assess a new task-blind approach, which aims to enhance users’ attention and recall processes instead of helping directly with their assignment by trying to infer in real time the task a user is performing. We compare three VR help-system configurations (task-based, task-blind, and no-help) through a controlled user study involving 66 participants. The results show that the group assisted by our task-blind system developed a different behavioural pattern while maintaining similar performance scores in comparison to the task-based approach. This paper shows a new way to offer more flexibility in help-system design and opens up the field of task-blind adaptive VR.
The conceptual structures built with Formal Concept Analysis (FCA) and its extensions are appropriate constructs for supporting Exploratory Search (ES). FCA indeed classifies a set of objects described by Boolean attributes in a concept lattice which is prone to (intra-lattice) navigation. Relational Concept Analysis (RCA), for its part, classifies several sets of objects connected through multiple binary relationships by using logical operators (quantifiers) which can be approximate. The output is a set of interconnected concept lattices, thus adding inter-lattice navigation opportunities. In this paper, we describe the web platform RCAviz, which aims to support such intra- and inter-lattice navigation. The user can select a subset of objects and attributes as a starting point for navigation. Then RCAviz shows the associated concept and its close intra- and inter-lattice neighbors. The user can access to the objects and attributes introduced and inherited in a concept. They then can navigate, i.e. zoom and pan the current view, and move from one concept to another. Additional views show the previous and the next conceptual structures, as well as an history which allows the user to browse its navigation. A navigation example is shown on a real dataset to illustrate the potential of RCAviz for ES.
Implication is a core notion of Formal Concept Analysis and its extensions. It provides information about the regularities present in the data. When one considers a relational data set of real-size, implications are numerous and their formulation, which combines primitive and relational attributes computed using Relational Concept Analysis framework, is complex. For an expert wishing to answer a question based on such a corpus of implications, having a smart exploration strategy is crucial. In this paper, we propose a visual approach, implemented in a web platform named FCAvizIR, for leveraging such corpus. Comprised of three interactive and coordinated views and a toolbox, FCAvizIR has been designed to explore corpora of implication rules following Schneiderman's famous mantra "overview first, zoom and filter, then details on demand". It enables metrics filtering, e.g. fixing a minimum and a maximum support value, and the multiple selection of relations and attributes in the premise and in the conclusion to identify the corresponding subset of implications presented as a list and Euler diagrams. An example of exploration is presented using an excerpt of Knomana to analyze plant-based extracts for controlling pests.
PURPOSE:Mapping clinical observations and medical test results into the standardized vocabulary LOINC is a prerequisite for exchanging clinical data between health information systems and ensuring efficient interoperability. METHODS:We present a comparison of three approaches for LOINC transcoding applied to French data collected from real-world settings. These approaches include both a state-of-the-art language model approach and a classifier chains approach. RESULTS:Our study demonstrates that we successfully improve the performance of the baselines using the classifier chains approach and compete effectively with state-of-the-art language models. CONCLUSIONS:Our approach proves to be efficient, cost-effective despite reproducibility challenges and potential for future optimizations and dataset testing.
The analysis of large sets of spatio-temporal data is a fundamental challenge in epidemiological research. As the quantity and the complexity of such kind of data increases, automatic analysis approaches, such as statistics, data mining, machine learning, etc., can be used to extract useful information. While these approaches have proven effective, they require a priori knowledge of the information being sought, and some interesting insights into the data may be missed. To bridge this gap, information visualization offers a set of techniques for not only presenting known information, but also exploring data without having a hypothesis formulated beforehand. In this paper, we introduce Epid Data Explorer (EDE), a visualization tool that enables exploration of spatio-temporal epidemiological data. EDE allows easy comparisons of indicators and trends across different geographical areas and times. It facilitates this exploration through ready-to-use pre-loaded datasets as well as user-chosen datasets. The tool also provides a secure architecture for easily importing new datasets while ensuring confidentiality. In two use cases using data associated with the COVID-19 epidemic, we demonstrate the substantial impact of implemented lockdown measures on mobility and how EDE allows assessing correlations between the spread of COVID-19 and weather conditions.
NDM-1 (New-Delhi-Metallo-β-lactamase-1) is an enzyme developed by bacteria that is implicated in bacteria resistance to almost all known antibiotics. In this study, we deliver a new, curated NDM-1 bioactivities database, along with a set of unifying rules for managing different activity properties and inconsistencies. We define the activity classification problem in terms of Multiple Instance Learning, employing embeddings corresponding to molecular substructures and present an ensemble ranking and classification framework, relaying on a k-fold Cross Validation method employing a per fold hyper-parameter optimization procedure, showing promising generalization ability. The MIL paradigm displayed an improvement up to 45.7 %, in terms of Balanced Accuracy, in comparison to the classical Machine Learning paradigm. Moreover, we investigate different compact molecular representations, based on atomic or bi-atomic substructures. Finally, we scanned the Drugbank for strongly active compounds and we present the top-15 ranked compounds.
Cartographers have long been interested in the representation of various movements such as migration, commercial exchanges and transportation. There are several techniques for visualizing this information; this paper focuses on flow mapping. A flow map shows a set of movements through line symbols connecting an origin to a destination. Each link is associated with a value that corresponds to the volume of the movement. However, once data reach a certain volume, the maps quickly become cluttered and can be difficult to read and understand. Moreover, the values of the movements must be correctly represented to avoid inducing biased interpretations. The objective of this paper is to create flow maps displaying flows of highly variable thicknesses so that the associated values are correctly represented. The technique used to create the flow paths does not create crossings between flows. In order to remove any visual clutter, such as overlaps between flows and geographic features, some areas of the map are distorted. In other words, our method of map distortion adapts the polygon vector base map to the flows, the central information of the visualization, and not the other way around.
CONTEXT We present a post-hoc approach to improve the recall of ICD classification. METHOD The proposed method can use any classifier as a backbone and aims to calibrate the number of codes returned per document. We test our approach on a new stratified split of the MIMIC-III dataset. RESULTS When returning 18 codes on average per document we obtain a recall that is 20% better than a classic classification approach.
In the last years, the scientific community has increasingly studied urban mobility since around 55% of the world population live in urban areas. Thus, individuals living in urban areas have to deal with phenomena like traffic jams, commute time, pollution, among others, which are difficult to understand and solve. Therefore, new innovative approaches such as mobility models, artificial intelligence, or visualization applied to urban mobility analysis problems shed new light on understanding cities' behavior. In this work, we survey the current state of the mathematical and computational tools we have at our disposal to better understand the current situation of urban areas. Our work presents datasets, discusses relevant artificial intelligence and visualization techniques, and reviews mathematical tools to analyze urban data. We hope our work offers a valuable summary of these ideas and provides the base for future investigations.
Machine learning methods are becoming increasingly popular to anticipate critical risks in patients under surveillance reducing the burden on caregivers. In this paper, we propose an original modeling that benefits of recent developments in Graph Convolutional Networks: a patient's journey is seen as a graph, where each node is an event and temporal proximities are represented by weighted directed edges. We evaluated this model to predict death at 24 hours on a real dataset and successfully compared our results with the state of the art.
Nancy Rodriguez合作论文数LIRMM-UM29