This work presents a user-centric recommendation framework, designed as a pipeline with four distinct, connected, and customizable phases. These phases are intended to improve explainability and boost user engagement. We have collected the historical Last.fm track playback records of a single user over approximately 15 years. The collected dataset includes more than 90,000 playbacks and approximately 14,000 unique tracks. From track playback records, we have created a dataset of user temporal contexts (each row is a specific moment when the user listened to certain music descriptors). As music descriptors, we have used community-contributed Last.fm tags and Spotify audio features. They represent the music that, throughout years, the user has been listening to. Next, given the most relevant Last.fm tags of a moment (e.g. the hour of the day), we predict the Spotify audio features that best fit the user preferences in that particular moment. Finally, we use the predicted audio features to find tracks similar to these features. The final aim is to recommend (and discover) tracks that the user may feel like listening to at a particular moment. For our initial study case, we have chosen to predict only a single audio feature target: danceability. The framework, however, allows to include more target variables. The ability to learn the musical habits from a single user can be quite powerful, and this framework could be extended to other users.
Bayesian Networks (BNs) are an important tool for assisting probabilistic reasoning, but despite being considered transparent models, people have trouble understanding them. Further, current User Interfaces (UIs) still do not clarify the reasoning of BNs. To address this problem, we have designed verbal and visual extensions to the standard BN UI, which can guide users through common inference patterns. We conducted a user study to compare our verbal, visual and combined UI extensions, and a baseline UI. Our main findings are: (1) users did better with all three types of extensions than with the baseline UI for questions about the impact of an observation, the paths that enable this impact, and the way in which an observation influences the impact of other observations; and (2) using verbal and visual modalities together is better than using either modality alone for some of these question types.
COVID-19 appeared abruptly in early 2020, requiring a rapid response amid a context of great uncertainty. Good quality data and knowledge was initially lacking, and many early models had to be developed with causal assumptions and estimations built in to supplement limited data, often with no reliable approach for identifying, validating and documenting these causal assumptions. Our team embarked on a knowledge engineering process to develop a causal knowledge base consisting of several causal BNs for diverse aspects of COVID-19. The unique challenges of the setting lead to experiments with the elicitation approach, and what emerged was a knowledge engineering method we call Causal Knowledge Engineering (CKE). The CKE provides a structured approach for building a causal knowledge base that can support the development of a variety of application-specific models. Here we describe the CKE method, and use our COVID-19 work as a case study to provide a detailed discussion and analysis of the method.
The typical phases of Bayesian network (BN) structured development include specification of purpose and scope, structure development, parameterisation and validation. Structure development is typically focused on qualitative issues and parameterisation quantitative issues, however there are qualitative and quantitative issues that arise in both phases. A common step that occurs after the initial structure has been developed is to perform a rough parameterisation that only captures and illustrates the intended qualitative behaviour of the model. This is done prior to a more rigorous parameterisation, ensuring that the structure is fit for purpose, as well as supporting later development and validation. In our collective experience and in discussions with other modellers, this step is an important part of the development process, but is under-reported in the literature. Since the practice focuses on qualitative issues, despite being quantitative in nature, we call this step qualitative parameterisation and provide an outline of its role in the BN development process.
Motivated by the ambiguity of operational case definitions for long COVID and the impact of the lack of a common causal language on long COVID research, in early 2023 we began developing a research framework on this post-acute infection syndrome. We used directed acyclic graphs (DAGs) and Bayesian networks (BNs) to depict the hypothesised mechanisms of long COVID in an agnostic fashion. The DAGs were informed by the evolving literature and subsequently refined following elicitation workshops with domain experts. The workshops were structured online sessions guided by an experienced facilitator. The causal DAGs aim to summarise the hypothesised pathobiological pathways from mild or severe COVID-19 disease to the development of pulmonary symptoms and fatigue over four different time points. The DAG was converted into a BN using qualitative parametrisation. These causal models aim to assist the identification of disease endotypes, as well as the design of randomised controlled trials and observational studies. The framework can also be extended to a range of other post-acute infection syndromes. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This publication is supported by Digital Health CRC Limited ("DHCRC"). DHCRC is funded under the Australian Commonwealth's Cooperative Research Centres (CRC) Program. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethics approval was granted by the Monash University Human Research Ethics Committee (Project ID 26942). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data including source models and dictionaries are available on the Open Science Framework
( Aust N Z J Obstet Gynaecol . 2022;62:813–825) The impact of postpartum hemorrhage (PPH, loss of ≥500 mL blood within 24 h of birth) on women, families, and communities is significant. PPH has been reported with increasing incidence and severity in well-resourced countries and is a major contributor to global maternal morbidity and mortality. Diagnosis can be challenging due to variability of risk factors and underestimation of blood loss leading to delayed response and treatment. Prognostic models to identify patients at risk and to provide resources and facilitate early intervention are not commonly integrated into clinical practice. Few, if any, are considered validated for use in the general obstetric population. This systematic review provides a comprehensive evaluation of current PPH prognostic models to identify best practices for prediction, describe model characteristics, and compare performance.
COVID-19 is a new multi-organ disease causing considerable worldwide morbidity and mortality. While many recognized pathophysiological mechanisms are involved, their exact causal relationships remain opaque. Better understanding is needed for predicting their progression, targeting therapeutic approaches, and improving patient outcomes. While many mathematical causal models describe COVID-19 epidemiology, none have described its pathophysiology. In early 2020, we began developing such causal models. The SARS-CoV-2 virus’s rapid and extensive spread made this particularly difficult: no large patient datasets were publicly available; the medical literature was flooded with sometimes conflicting pre-review reports; and clinicians in many countries had little time for academic consultations. We used Bayesian network (BN) models, which provide powerful calculation tools and directed acyclic graphs (DAGs) as comprehensible causal maps. Hence, they can incorporate both expert opinion and numerical data, and produce explainable, updatable results. To obtain the DAGs, we used extensive expert elicitation (exploiting Australia’s exceptionally low COVID-19 burden) in structured online sessions. Groups of clinical and other specialists were enlisted to filter, interpret and discuss the literature and develop a current consensus. We encouraged inclusion of theoretically salient latent (unobservable) variables, likely mechanisms by extrapolation from other diseases, and documented supporting literature while noting controversies. Our method was iterative and incremental: systematically refining and validating the group output using one-on-one follow-up meetings with original and new experts. 35 experts contributed 126 hours face-to-face, and could review our products. We present two key models, for the initial infection of the respiratory tract and the possible progression to complications, as causal DAGs and BNs with corresponding verbal descriptions, dictionaries and sources. These are the first published causal models of COVID-19 pathophysiology. Our method demonstrates an improved procedure for developing BNs via expert elicitation, which other teams can implement to model emergent complex phenomena. Our results have three anticipated applications: (i) freely disseminating updatable expert knowledge; (ii) guiding design and analysis of observational and clinical studies; (iii) developing and validating automated tools for causal reasoning and decision support. We are developing such tools for the initial diagnosis, resource management, and prognosis of COVID-19, parameterized using the ISARIC and LEOSS databases.
Object-oriented Bayesian networks (OOBNs) allow modellers to construct compositional and hierarchical models, using an inheritance hierarchy of classes ad subclasses, enabling reuse and supporting maintenance. Reasoning with both ordinary Bayesian networks (BNs) and OOBNs requires the important computational task of inference, the computing of new posterior probability distributions given a set of evidence. A widely used inference technique in ordinary BNs involves compiling the BN into a so-called junction tree (JT) before performing the inference; the compilation step is only performed when the model changes. In current OOBN software, the OOBN is first transformed into the underlying BN, so-called flattening, then the standard inference is performed. Researchers have proposed methods for incremental compilation of BNs, rather than recompiling from scratch for each network modification; these can apply to OOBNs also after flattening. Here, we propose a new incremental compilation technique that reuses existing compiled JTs of both embedded components and superclasses, and does not require flattening. We demonstrate through experimental analysis that this can reduce compilation time, and produces compact JTs that are cost-effective for inference.
Bayesian networks (BNs) are a widely used probabilistic modelling tool for reasoning under uncertainty, though scaling them up for complex real-world problems can be challenging. Object-Oriented Bayesian Networks (OOBNs) have been proposed to address this challenge, providing modellers with the ability to define hierarchies of classes and use these classes to construct models with a compositional and hierarchical structure, enabling reuse and supporting maintenance. The object-oriented concept of inheritance supports reuse of existing components, but comes with the challenge of building, and then maintaining, an efficient hierarchy of classes. This paper proposes a supergraph based method that constructs class inheritance hierarchies from a set of OOBN classes; this can be used either to form an initial inheritance hierarchy, or to reform an existing hierarchy into a more efficient one. We also present heuristics to convert a BN to an OOBN class, measures to evaluate a constructed hierarchy and empirical analyses of the proposed approach on synthetic hierarchies and on a real-world OOBN project; results show the algorithm works well in practice.
Our daily life is full of challenges, and the biggest challenge is the unpredictability of many of our significant life events. To deal with this unpredictability, analysing the probability of events has become very important. In particular, the theorem of English statistician Thomas Bayes has been revolutionary. Numerous theories and techniques have been proposed, and many tools have been developed to solve real-life problems based on the theorem, yet it is still very much an area of active research. It still attracts researchers dealing with cutting-edge technologies. One tool that has been used extensively in modelling probabilistic analysis for decades is the Probabilistic Graphical Model (PGM). PGMs have very challenging childhood but glorious youth. The vast applicability of the models in cutting-edge technologies attracts researchers, modellers and scientists of diversified fields. Hence there are numerous models with their respective features, merits and backlogs. To date, there have been very few surveys conducted among the wide range of models and their associated tools. More specifically, those few reviews are highly application and domain focused, and limited to three to four very popular and widely used models and their associated learning and inference algorithms. To the best of our knowledge, this paper is the first that presents the features, limitations, design and implementation platforms, research challenges and applicability of the models based on a common framework that consists of some essential attributes of the popular PGMs and tools for probabilistic analysis. The study helps deciding an appropriate tool as per the perspective of the application and feature of the tool. This paper concludes with future research scope and a non-exhaustive list of applications of PGMs. DUJASE Vol. 6 (2) 82-93, 2021 (July)
Background Postpartum haemorrhage (PPH) remains a leading cause of maternal mortality and morbidity worldwide, and the rate is increasing. Using a reliable predictive model could identify those at risk, support management and treatment, and improve maternal outcomes. Aims To systematically identify and appraise existing prognostic models for PPH and ascertain suitability for clinical use. Materials and Methods MEDLINE, CINAHL, Embase, and the Cochrane Library were searched using combinations of terms and synonyms, including ‘postpartum haemorrhage’, ‘prognostic model’, and ‘risk factors’. Observational or experimental studies describing a prognostic model for risk of PPH, published in English, were included. The Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies checklist informed data extraction and the Prediction Model Risk of Bias Assessment Tool guided analysis. Results Sixteen studies met the inclusion criteria after screening 1612 records. All studies were hospital settings from eight different countries. Models were developed for women who experienced vaginal birth ( n = 7), caesarean birth ( n = 2), any type of birth ( n = 2), hypertensive disorders ( n = 1) and those with placental abnormalities ( n = 4). All studies were at high risk of bias due to use of inappropriate analysis methods or omission of important statistical considerations or suboptimal validation. Conclusions No existing prognostic models for PPH are ready for clinical application. Future research is needed to externally validate existing models and potentially develop a new model that is reliable and applicable to clinical practice.
In the absence of an established gold standard, an understanding of the testing cycle from individual exposure to test outcome report is required to guide the correct interpretation of severe acute respiratory syndrome-coronavirus-2 reverse transcriptase real-time polymerase chain reaction (RT-PCR) results and optimise the testing processes. Bayesian network models have been used within healthcare to bring clarity to complex problems. We use this modelling approach to construct a comprehensive framework for understanding the real-world predictive value of individual RT-PCR results. We elicited knowledge from domain experts to describe the test process through a facilitated group workshop. A preliminary model was derived based on the elicited knowledge, then subsequently refined, parameterised and validated with a second workshop and one-on-one discussions. Causal relationships elicited describe the interactions of pre-testing, specimen collection and laboratory procedures and RT-PCR platform factors, and their impact on the presence and quantity of virus and thus the test result and its interpretation. By setting the input variables as 'evidence' for a given subject and preliminary parameterisation, four scenarios were simulated to demonstrate potential uses of the model. The core value of this model is a deep understanding of the total testing cycle, bridging the gap between a person's true infection status and their test outcome. This model can be adapted to different settings, testing modalities and pathogens, adding much needed nuance to the interpretations of results.
In many complex, real-world situations, problem solving and decision making require effective reasoning about causation and uncertainty. However, human reasoning in these cases is prone to confusion and error. Bayesian networks (BNs) are an artificial intelligence technology that models uncertain situations, supporting probabilistic and causal reasoning and decision making. However, to date, BN methodologies and software require significant upfront training, do not provide much guidance on the model building process, and do not support collaboratively building BNs. BARD (Bayesian ARgumentation via Delphi) is both a methodology and an expert system that utilises (1) BNs as the underlying structured representations for better argument analysis, (2) a multi-user web-based software platform and Delphi-style social processes to assist with collaboration, and (3) short, high-quality e-courses on demand, a highly structured process to guide BN construction, and a variety of helpful tools to assist in building and reasoning with BNs, including an automated explanation tool to assist effective report writing. The result is an end-to-end online platform, with associated online training, for groups without prior BN expertise to understand and analyse a problem, build a model of its underlying probabilistic causal structure, validate and reason with the causal model, and use it to produce a written analytic report. Initial experimental results demonstrate that BARD aids in problem solving, reasoning and collaboration.
In many complex real-world situations, problem solving and decision making require effective reasoning about causation and uncertainty. However, human reasoning in these cases is notoriously prone to confusion and error [Kahneman et al., 1982]. One way to support better reasoning is to employ Bayesian networks (BNs) [Pearl, 1988] to model and represent uncertain situations clearly for the user, and make complex calculations quickly and accurately on demand. BNs have been deployed for this purpose in diverse domains such as medicine, education, engineering, surveillance, the law, weather forecasting, and the environment.
Background ASD and ADHD are prevalent neurodevelopmental disorders that frequently co-occur and have strong evidence for a degree of shared genetic aetiology. Behavioural and neurocognitive heterogeneity in ASD and ADHD has hampered attempts to map the underlying genetics and neurobiology, predict intervention response, and improve diagnostic accuracy. Moving away from categorical conceptualisations of psychopathology to a dimensional approach is anticipated to facilitate discovery of data-driven clusters and enhance our understanding of the neurobiological and genetic aetiology of these conditions. The Monash Autism-ADHD Genetics and Neurodevelopment (MAGNET) Project is one of the first large-scale, family-based studies to take a truly transdiagnostic approach to ASD and ADHD. Using a comprehensive phenotyping protocol capturing dimensional traits central to ASD and ADHD, the MAGNET Project aims to identify data-driven clusters across ADHD-ASD spectra using deep phenotyping of symptoms and behaviours; investigate the degree of familiality for different dimensional ASD-ADHD phenotypes and clusters; and map the neurocognitive, brain imaging, and genetic correlates of these data-driven symptom-based clusters. Methods The MAGNET Project will recruit 1,200 families with children who are either typically developing, or who display elevated ASD, ADHD, or ASD-ADHD traits, in addition to affected and unaffected biological siblings of probands, and parents. All children will be comprehensively phenotyped for behavioural symptoms, comorbidities, neurocognitive and neuroimaging traits and genetics. Conclusion The MAGNET Project will be the first large-scale family study to take a transdiagnostic approach to ASD-ADHD, utilising deep phenotyping across behavioural, neurocognitive, brain imaging and genetic measures.
Bayes Nets (BNs) are extremely useful for causal and probabilistic modelling in many real-world applications, often built with information elicited from groups of domain experts. But their potential for reasoning and decision support has been limited by two major factors: the need for significant normative knowledge, and the lack of any validated methods or software supporting collaboration. Consequently, we have developed a web-based structured technique – Bayesian Argumentation via Delphi (BARD) – to enable groups of domain experts to receive minimal normative training and then collaborate effectively to produce high-quality BNs. BARD harnesses multiple perspectives on a problem, while minimising biases manifest in freely interacting groups, via a Delphi process: solutions are first produced individually, then shared, followed by an opportunity for individuals to revise their solutions. To test the hypothesis that BNs improve due to Delphi, we conducted an experiment whereby individuals with a little BN training and practice produced structural models using BARD for two Bayesian reasoning problems. Participants then received 6 other structural models for each problem, rated their quality on a 7-point scale, and revised their own models if they wished. Both top-rated and revised models were on average significantly better quality (scored against a gold-standard) than the initial models, with large and medium effect sizes. We conclude that Delphi – and BARD – improves the quality of BNs produced by groups. Further, although rating cannot create new models, rating seems quicker and easier than revision and yielded significantly better models – so, we suggest efficient BN amalgamation should include both.
Gaussian graphical models (GGM) express conditional dependencies among variables of Gaussian-distributed high-dimensional data. However, real-life datasets exhibit heterogeneity which can be better captured through the use of mixtures of GGMs, where each component captures different conditional dependencies a.k.a. context-specific dependencies along with some common dependencies a.k.a. shared dependencies. Methods to discover shared and context-specific graphical structures include joint and grouped graphical Lasso, and the EM algorithm with various penalized likelihood scoring functions. However, these methods detect graphical structures with high false discovery rates and do not detect two types of dependencies (i.e., context-specific and shared) together. In this paper, we develop a method to discover shared conditional dependencies along with context-specific graphical models via a two-level hierarchical Gaussian graphical model. We assume that the graphical models corresponding to shared and context-specific dependencies are decomposable, which leads to an efficient greedy algorithm to select edges minimizing a score based on minimum message length (MML). The MML-based score results in lower false discovery rate, leading to a more effective structure discovery. We present extensive empirical results on synthetic and real-life datasets and show that our method leads to more accurate prediction of context-specific dependencies among random variables compared to previous works. Hence, we can consider that our method is a state of the art to discover both shared and context-specific conditional dependencies from high-dimensional Gaussian heterogeneous data.
Nowadays, the global energy system is in a transition phase, in which the integration of renewable energy is among the main requirements for attenuating climate change. Wind power is a major alternative to supply clean energy; hence, its widespread penetration is being pursued in all end-use sectors. In particular, it is currently noteworthy to analyze the feasibility of deploying small-scale wind power technology to provide cleaner and cheaper energy in the residential sector. As a first step, a technical assessment must be carried out to provide crucial information to intensive energy consumers, providers of small-scale wind power technology, electric energy distribution utilities, and any other party, to help them decide whether or not to deploy small-scale wind turbines. With this aim, we propose to perform such an analysis using a suitable probabilistic paradigm to solve complex decision-making problems with uncertainty, namely Bayesian Intelligence, since wind resources and energy demands are intermittent variables, properly characterized by probability distribution functions. Then, the problem of determining the technical feasibility can be formulated as an investigation into whether or not small-scale wind turbine technology can produce enough energy to cover the excess demand of intensive energy residential consumers to get off high-priced tariffs. For this purpose, we introduce a novel model based on probabilistic reasoning to assess the suitability of small-scale wind turbine technology to produce the said energy, taking into consideration the availability of wind resources and the energy pricing structure. To demonstrate the usefulness and performance of the proposed model, we consider a case study of deploying 5 and 10 kW wind turbines and analyze the feasibility of their implementation in Mexico, where the energy pricing structure and scattered wind resource availability pose difficult challenges.
The United Nations sustainable development goals (SDGs) were ratified with much enthusiasm by all UN member states in 2015. However, subsequent progress to meet these goals has been hampered by a lack of data available to measure the SDG indicators (SDIs), and a lack of evidence-based insights to inform effective policy responses. We outline an interdisciplinary program of research into the use of artificial intelligence techniques to support measurement of the SDIs, using both machine learning methods to model SDI measurements and explainable AI techniques to present the outputs in a human-friendly manner. As well as addressing the technical concerns, we will investigate the governance issues of what forms of evidence, methods of collecting that evidence and means of its communication will most usefully inform effective policy development. By addressing these fundamental challenges, we aim to provide policy makers with the evidence needed to take effective action towards realising the Sustainable Development Goals.