Multiple Kernel Learning (MKL) models combine several kernels in supervised and unsupervised settings to integrate multiple data representations or sources, each represented by a different kernel. MKL seeks an optimal linear combination of base kernels that maximizes a generalized performance measure under a regularization constraint. Various norms have been used to regularize the kernel weights, including l1, l2 and lp, as well as the "elastic-net" penalty, which combines l1- and l2-norm to promote both sparsity and the selection of correlated kernels. This property makes elastic-net regularized MKL (ENMKL) especially valuable when model interpretability is critical and kernels capture correlated information, such as in neuroimaging. Previous ENMKL methods have followed a two-stage procedure: fix kernel weights, train a support vector machine (SVM) with the weighted kernel, and then update the weights via gradient descent, cutting-plane methods, or surrogate functions. Here, we introduce an alternative ENMKL formulation that yields a simple analytical update for the kernel weights. We derive explicit algorithms for both SVM and kernel ridge regression (KRR) under this framework, and implement them in the open-source Pattern Recognition for Neuroimaging Toolbox (PRoNTo). We evaluate these ENMKL algorithms against l1-norm MKL and against SVM (or KRR) trained on the unweighted sum of kernels across three neuroimaging applications. Our results show that ENMKL matches or outperforms l1-norm MKL in all tasks and only underperforms standard SVM in one scenario. Crucially, ENMKL produces sparser, more interpretable models by selectively weighting correlated kernels.
Recent advances in artificial intelligence (AI)—including generative approaches—have resulted in technology that can support humans in scientific discovery and forming decisions, but may also disrupt democracies and target individuals. The responsible use of AI and its participation in human–AI teams increasingly shows the need for AI alignment, that is, to make AI systems act according to our preferences. A crucial yet often overlooked aspect of these interactions is the different ways in which humans and machines generalize. In cognitive science, human generalization commonly involves abstraction and concept learning. By contrast, AI generalization encompasses out-of-domain generalization in machine learning, rule-based reasoning in symbolic AI, and abstraction in neurosymbolic AI. Here we combine insights from AI and cognitive science to identify key commonalities and differences across three dimensions: notions of, methods for, and evaluation of generalization. We map the different conceptualizations of generalization in AI and cognitive science along these three dimensions and consider their role for alignment in human–AI teaming. This results in interdisciplinary challenges across AI and cognitive science that must be tackled to support effective and cognitively supported alignment in human–AI teaming scenarios. Ilievski et al. examine differences and similarities in the various ways human and AI systems generalize. The insights are important for effectively supporting alignment in human–AI teams.
Personalised education is one of the domains that can greatly benefit from the most recent advances in Artificial Intelligence (AI) and Large Language Models (LLM). However, it is also one of the most challenging applications due to the cognitive complexity of teaching effectively while personalising the learning experience to suit independent learners. We hypothesise that one promising approach to excelling in such demanding use cases is using a society of minds. In this chapter, we present TrueReason, an exemplar personalised learning system that integrates a multitude of specialised AI models that can mimic micro skills that are composed together by a LLM to operationalise planning and reasoning. The architecture of the initial prototype is presented while describing two micro skills that have been incorporated in the prototype. The proposed system demonstrates the first step in building sophisticated AI systems that can take up very complex cognitive tasks that are demanded by domains such as education.
Human-AI coevolution, defined as a process in which humans and AI algorithms continuously influence each other, increasingly characterises our society, but is understudied in artificial intelligence and complexity science literature. Recommender systems and assistants play a prominent role in human-AI coevolution, as they permeate many facets of daily life and influence human choices through online platforms. The interaction between users and AI results in a potentially endless feedback loop, wherein users' choices generate data to train AI models, which, in turn, shape subsequent user preferences. This human-AI feedback loop has peculiar characteristics compared to traditional human-machine interaction and gives rise to complex and often “unintended” systemic outcomes. This paper introduces human-AI coevolution as the cornerstone for a new field of study at the intersection between AI and complexity science focused on the theoretical, empirical, and mathematical investigation of the human-AI feedback loop. In doing so, we: (i) outline the pros and cons of existing methodologies and highlight shortcomings and potential ways for capturing feedback loop mechanisms; (ii) propose a reflection at the intersection between complexity science, AI and society; (iii) provide real-world examples for different human-AI ecosystems; and (iv) illustrate challenges to the creation of such a field of study, conceptualising them at increasing levels of abstraction, i.e., scientific, legal and socio-political.
Computational prediction of the interaction of T cell receptors (TCRs) and their ligands is a grand challenge in immunology. Despite advances in high-throughput assays, specificity-labeled TCR data remain sparse. In other domains, the pre-training of language models on unlabeled data has been successfully used to address data bottlenecks. However, it is unclear how to best pre-train protein language models for TCR specificity prediction. Here, we introduce a TCR language model called SCEPTR (simple contrastive embedding of the primary sequence of T cell receptors), which is capable of data-efficient transfer learning. Through our model, we introduce a pre-training strategy combining autocontrastive learning and masked-language modeling, which enables SCEPTR to achieve its state-of-the-art performance. In contrast, existing protein language models and a variant of SCEPTR pre-trained without autocontrastive learning are outperformed by sequence alignment-based methods. We anticipate that contrastive learning will be a useful paradigm to decode the rules of TCR specificity. A record of this paper's transparent peer review process is included in the supplemental information.
Human-AI coevolution, defined as a process in which humans and AI algorithms continuously influence each other, increasingly characterises our society, but is understudied in artificial intelligence and complexity science literature. Recommender systems and assistants play a prominent role in human-AI coevolution, as they permeate many facets of daily life and influence human choices through online platforms. The interaction between users and AI results in a potentially endless feedback loop, wherein users' choices generate data to train AI models, which, in turn, shape subsequent user preferences. This human-AI feedback loop has peculiar characteristics compared to traditional human-machine interaction and gives rise to complex and often "unintended" systemic outcomes. This paper introduces human-AI coevolution as the cornerstone for a new field of study at the intersection between AI and complexity science focused on the theoretical, empirical, and mathematical investigation of the human-AI feedback loop. In doing so, we: (i) outline the pros and cons of existing methodologies and highlight shortcomings and potential ways for capturing feedback loop mechanisms; (ii) propose a reflection at the intersection between complexity science, AI and society; (iii) provide real-world examples for different human-AI ecosystems; and (iv) illustrate challenges to the creation of such a field of study, conceptualising them at increasing levels of abstraction, i.e., scientific, legal and socio-political.
Post-disaster emergency restoration (ER) has emerged as a promising approach to enhancing the resilience of critical infrastructure systems (CISs). However, devising optimal ER plans immediately after real-world disasters is inherently challenging due to the vast state space of modern CISs. To address this challenge, this paper proposes a range of strategies to guide these campaigns, with a particular focus on road networks (RNs) under earthquake disasters. Initially, a set of heuristic-based, easy-to-interpret strategies has been developed. Building on prior research, this study investigates the integration of lookahead search, examining its potential to refine and adapt these heuristics to meet diverse optimization objectives of ER campaigns. To operationalize these strategies, a multi-agent-based model (MABM) is established, wherein each restoration group is modelled as an autonomous agent, guided by the proposed planning strategies. The applicability of the model is demonstrated by its implementation in a real-world RN under catastrophic earthquake scenarios. The impact of various strategies on the effectiveness of the ER campaign is examined and elucidated. Notably, the planning strategy that combines a newly developed, accessibility-based heuristic with lookahead proves effective in sequentially balancing the trade-off between the accessibility and criticality of collapsed bridges. Based on the case study result, this approach consistently fulfils diverse optimization objectives across a range of earthquake scenarios, establishing it as the benchmark planning strategy of the post-shock ER of RNs.
This paper presents four theoretical contributions that improve the usability of risk certificates for neural networks based on PAC-Bayes bounds. First, two bounds on the KL divergence between Bernoulli distributions enable the derivation of the tightest explicit bounds on the true risk of classifiers across different ranges of empirical risk. The paper next focuses on the formalization of an efficient methodology based on implicit differentiation that enables the introduction of the optimization of PAC-Bayesian risk certificates inside the loss/objective function used to fit the network/model. The last contribution is a method to optimize bounds on non-differentiable objectives such as the 0-1 loss. These theoretical contributions are complemented with an empirical evaluation on the MNIST and CIFAR-10 datasets. In fact, this paper presents the first non-vacuous generalization bounds on CIFAR-10 for neural networks.
Decision makers may suffer from uncertainty induced by limited data. This may be mitigated by accounting for epistemic uncertainty, which is however challenging to estimate efficiently for large neural networks. To this extent we investigate Delta Variances, a family of algorithms for epistemic uncertainty quantification, that is computationally efficient and convenient to implement. It can be applied to neural networks and more general functions composed of neural networks. As an example we consider a weather simulator with a neural-network-based step function inside - here Delta Variances empirically obtain competitive results at the cost of a single gradient computation. The approach is convenient as it requires no changes to the neural network architecture or training procedure. We discuss multiple ways to derive Delta Variances theoretically noting that special cases recover popular techniques and present a unified perspective on multiple related methods. Finally we observe that this general perspective gives rise to a natural extension and empirically show its benefit.
Principal component analysis(PCA) is a popular method for dimension reduction and has attracted an unfailing interest for decades. More recently, kernel PCA (KPCA) has emerged as an extension of PCA, but despite its use in practice, a sound theoretical understanding of KPCA is missing. We contribute several empirical generalisation bounds on the efficiency of KPCA, involving the empirical eigenvalues of the kernel Gram matrix. Our bounds are derived through the use of probably approximately correct (PAC)-Bayes theory and highlight the importance of some desirable properties of datasets, expressed as variance-typed terms, to attain fast rates, achievable for a wide class of kernels.
With the advancement and utility of Artificial Intelligence (AI), personalising education to a global population could be a cornerstone of new educational systems in the future. This work presents the PEEKC dataset and the TrueLearn Python library, which contains a dataset and a series of online learner state models that are essential to facilitate research on learner engagement modelling.TrueLearn family of models was designed following the "open learner" concept, using humanly-intuitive user representations. This family of scalable, online models also help end-users visualise the learner models, which may in the future facilitate user interaction with their models/recommenders. The extensive documentation and coding examples make the library highly accessible to both machine learning developers and educational data mining and learning analytics practitioners. The experiments show the utility of both the dataset and the library with predictive performance significantly exceeding comparative baseline models. The dataset contains a large amount of AI-related educational videos, which are of interest for building and validating AI-specific educational recommenders.
Background: Applying survival analysis techniques to epidemiological inference within research into ageing offers opportunities to estimate the association between exposure and outcome in longitudinal data. This study used Cox regression to investigate how socioeconomic inequality in mortality can be explained by exposure to various factors including smoking, diet, alcohol and physical activity. This study seeks to complement and extend previous work which found that the contribution of the socioeconomic gradient to inequalities in health was underestimated by baseline analysis. Methods: Data was obtained from Whitehall II, a British longitudinal cohort study, which investigated social determinants of health. Analysis is based on 11 waves of data collected over 32 years on 10,308 civil servants aged between 35 and 90. Socioeconomic position was defined by baseline employment grade (1-3). During the follow-up 2,427 participants died. Extensive experimental analysis was conducted using a vast number of health behaviours. Cox regression produced an age-and-sex-adjusted hazard ratio for the socioeconomic inequality in mortality. Health behaviours (smoking, physical activity, alcohol consumption, and diet) were then added as covariates to determine the extent to which they statistically explain this inequality, and how this differed from the last similar analysis from 2009. This was done at baseline and longitudinally. The health behaviours were then combined linearly, nonlinearly and new health behaviours were added. Results: Adding the above health behaviours as covariates statistically explained the socioeconomic gradient in mortality at baseline from 42% to 2009, to 51% to 2021. Longitudinal consideration increased the explanatory power, when all health behaviours were added as time-varying covariates, from 51% to 87%. Adding more variables in the form of a more comprehensive diet score statistically explained the gradient further, to 91%. The nonlinear model of smoking and exercise most accurately predicted mortality and had a 13% higher explanatory power when explaining the gradient compared to the linear model in longitudinal data. Conclusion: In the Whitehall II study, socioeconomic position and mortality showed an association. There is a gain in explanatory power of the set of health behaviours at baseline when follow-up is extended by 12 years, from 42% to 51%. When changes in behaviour over the 32 years of follow-up were also accounted for, this association was now significantly explained by over 90%, compared with 51% when considered at baseline. We suggest that reverse causation is partly responsible for the almost complete explanation of the social gradient in mortality by health behaviours. These results would therefore lead us to question why health behaviours are socially patterned in the way that has been observed, which would be significant for targeting health behaviours in lower socioeconomic statuses. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study did not receive any funding ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: UCL Research Ethics Committee (85/0938) of University College London gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data is protected in Dementias Platform UK and available upon application with reasonsable request.
In an effort to reduce pesticide use, agronomists and computer scientists have joined forces to develop site-specific weed detection and classification systems. These systems aim to recognize and locate weed species within a crop field, using precision equipment to apply required herbicides timely and only where needed, with the objective of reducing the sprayable surface required to eliminate the given weed and protect the crop, with both economic and environmental benefits. Yet, with climate change on the rise, common weeds are expected to undergo some changes to adapt to their environment, possibly with new or invasive weeds spreading to areas where they did not exist before. These changes (often morphological) as well as new invasions need to be taken into account by future classifiers and detection algorithms to ensure system robustness and adaptation to new habitats/climate dynamics. This paper proposes a set of experiments evaluating the use of transfer learning and zero-shot learning for weed classification using our novel TomatoWeeds dataset. Residual networks of variable depth, pretrained on the Imagenet and/or DeepWeeds datasets were evaluated. A ResNet50 pretrained on both datasets and fine-tuned on the TomatoWeeds dataset performed best, returning a holdout set accuracy of 77.8%, showing the advantageous use of transfer learning in this domain. Zero-shot learning, using both embeddings of images and morphological and habitat text-based descriptions, is implemented to test the ability of machine learning pipelines of recognising unseen classes at test time (which may arise e.g. due to changing climate dynamics), a learning task in which the field (and our experiments) are still far from satisfactory results. Further research could benefit from larger weed-specific datasets for transfer learning as well as deeper network architectures to improve model performance. The projection-based ZSL could also benefit from larger datasets and new zero-shot learning architectures in hope that unseen classes are accurately projected.
Artificial Intelligence (AI) in Education claims to have the potential for building personalised curricula, as well as bringing opportunities for democratising education and creating a renaissance of new ways of teaching and learning. Millions of students are starting to benefit from the use of these technologies, but millions more around the world are not, due to the digital divide and deep pre-existing social and educational inequalities. If this trend continues, the first large-scale delivery of AI in Education could lead to greater educational inequality, along with a global misallocation of educational resources motivated by the current techno-solutionist narrative, which proposes technological solutions as a quick and flawless way to solve complex real-world problems. This work focuses on posing questions about the future of AI in Education, intending to initiate the pressing conversation that could set the right foundations (e.g., inclusion and diversity) for a new generation of education that is permeated with AI technology. The main goal of our opinion piece is to conceptualise a sustainable, large-scale and inclusive AI for the education ecosystem that facilitates equitable, high-quality lifelong learning opportunities for all. The contribution starts by synthesising how AI might change how we learn and teach, focusing on the case of personalised learning companions and assistive technology for disability. Then, we move on to discuss some socio-technical features that will be crucial to avoiding the perils of these AI systems worldwide (and perhaps ensuring their success by leveraging more inclusive education). This work also discusses the potential of using AI together with free, participatory and democratic resources, such as Wikipedia, Open Educational Resources and open-source tools. We emphasise the need for collectively designing human-centred, transparent, interactive and collaborative AI-based algorithms that empower and give complete agency to stakeholders, as well as supporting new emerging pedagogies. Finally, we ask what it would take for this educational revolution to provide egalitarian and empowering access to education that transcends any political, cultural, language, geographical and learning-ability barriers, so that educational systems can be responsive to all learners’ needs.
Current PAC-Bayes generalisation bounds are restricted to scalar metrics of performance, such as the loss or error rate. However, one ideally wants more information-rich certificates that control the entire distribution of possible outcomes, such as the distribution of the test loss in regression, or the probabilities of different mis classifications. We provide the first PAC-Bayes bound capable of providing such rich information by bounding the Kullback-Leibler divergence between the empirical and true probabilities of a set of M error types, which can either be discretized loss values for regression, or the elements of the confusion matrix (or a partition thereof) for classification. We transform our bound into a differentiable training objective. Our bound is especially useful in cases where the severity of different mis-classifications may change over time; existing PAC-Bayes bounds can only bound a particular pre-decided weighting of the error types. In contrast our bound implicitly controls all uncountably many weightings simultaneously.
Artificial intelligence (AI) has been attracting increased attention from researchers, entrepreneurs, investors and policy makers on all continents. Innovative national and international development collaboration mechanisms, such as micro-funding and social impact bonds, are being tested, refined and implemented to assist AI researchers, contribute to scientific excellence and scale to market. This essay looks at emerging networks of excellence in the Global South, particularly AI4D Africa. It examines how bottom-approach, small-scale investments resulted in significant research on different scientific and non-scientific, engineering and educational topics, including a profile of African languages.
Kitsuchart Pasupa合作论文数Faculty of Information Technology, King Mongkut's Institute of Technology Ladkrabang7
Mario Marchand合作论文数Departement d'informatique6