We design a large-language-model (LLM) agent that extracts causal feedback fuzzy cognitive maps (FCMs) from raw text. The causal learning or extraction process is agentic both because of the LLM's semi-autonomy and because ultimately the FCM dynamical system's equilibria drive the LLM agents to fetch and process causal text. The fetched text can in principle modify the adaptive FCM causal structure and so modify the source of its quasi-autonomy–its equilibrium limit cycles and fixed-point attractors. This bidirectional process endows the evolving FCM dynamical system with a degree of autonomy while still staying on its agentic leash. We show in particular that a sequence of three finely tuned system instructions guide an LLM agent as it systematically extracts key nouns and noun phrases from text, as it extracts FCM concept nodes from among those nouns and noun phrases, and then as it extracts or infers partial or fuzzy causal edges between those FCM nodes. We test this FCM generation on a recent essay about the promise of AI from the late diplomat and political theorist Henry Kissinger and his colleagues. This three-step process produced FCM dynamical systems that converged to the same equilibrium limit cycles as did the human-generated FCMs even though the human-generated FCM differed in the number of nodes and edges. A final FCM mixed generated FCMs from separate Gemini and ChatGPT LLM agents. The mixed FCM absorbed the equilibria of its dominant mixture component but also created new equilibria of its own to better approximate the underlying causal dynamical system.
We automatically generate feedback causal fuzzy cognitive maps (FCMs) from text by teaching large-language-model agents to break the text into overlapping chunks of text. Convex mixing of these chunk FCMs gives a representative cyclic FCM knowledge graph. The text chunks can have different levels of overlap. The chunk FCMs still mix to form a new FCM causal knowledge graph. The mixing technique scales because it uses light computation with sparse causal chunk matrices. The mixing structure allows an operator-level type of Bayesian inference that produces "de-chunked" or posterior-like FCMs from the mixed FCM. These de-chunked FCMs are useful in their own right and allow further iterations of Bayesian updating. We demonstrate these mixing techniques on the essay text of Allison's "Thucydides Trap" model of conflict between a dominant power such as the United States and a rising power such as China. The FCM dynamical systems predict outcomes as they equilibrate to fixed-point or limit-cycle attractors. Seven out of 8 FCM knowledge graphs predicted a type of war when we stimulated them by turning on and keeping on the concept node that stands for the rising power's ambition and entitlement. Gemini 3.1 LLMs served as the chunking AI agents.
An adaptive multiexpert mixture of feedback causal models can approximate missing or phantom nodes in large-scale causal models. The result gives a scalable form of big knowledge. The mixed model approximates a sampled dynamical system by approximating its main limit-cycle equilibria. Each expert first draws a fuzzy cognitive map (FCM) with at least one missing causal node or variable. FCMs are directed signed partial-causality cyclic graphs. They mix naturally through convex combination to produce a new causal feedback FCM. Supervised learning helps each expert FCM estimate its phantom node by comparing the FCM's partial equilibrium with the complete multi-node equilibrium. Such phantom-node estimation allows partial control over these causal hallucinations and helps approximate the future trajectory of the dynamical system. But the approximation can be computationally heavy. Mixing the tuned expert FCMs gives a practical way to find several phantom nodes and thereby better approximate the feedback system's true equilibrium behavior.
A large language model (LLM) can map a feedback causal fuzzy cognitive map (FCM) into text and then reconstruct the FCM from the text. This explainable AI system approximates an identity map from the FCM to itself and resembles the operation of an autoencoder (AE). Both the encoder and the decoder explain their decisions in contrast to black-box AEs. Humans can read and interpret the encoded text in contrast to the hidden variables and synaptic webs in AEs. The LLM agent approximates the identity map through a sequence of system instructions that does not compare the output to the input. The reconstruction is lossy because it removes weak causal edges or rules while it preserves strong causal edges. The encoder preserves the strong causal edges even when it trades off some details about the FCM to make the text sound more natural.
We present the new bidirectional variational autoencoder (BVAE) network architecture. The BVAE uses a single neural network both to encode and decode instead of an encoder-decoder network pair. The network encodes in the forward direction and decodes in the backward direction through the same synaptic web. Simulations compared BVAEs and ordinary VAEs on the four image tasks of image reconstruction, classification, interpolation, and generation. The image datasets included MNIST handwritten digits, Fashion-MNIST, CIFAR10, and CelebA-64 face images. The bidirectional structure of BVAEs cut the parameter count by almost 50% and still slightly outperformed the unidirectional VAEs.
We introduce new soft diamond regularizers that both improve synaptic sparsity and maintain classification accuracy in deep neural networks. These parametrized regularizers outperform the state-of-the-art hard-diamond Laplacian regularizer of Lasso regression and classification. They use thick-tailed symmetric alpha-stable (S alpha S) bell-curve synaptic weight priors that are not Gaussian and so have thicker tails. The geometry of the diamond-shaped constraint set varies from a circle to a star depending on the tail thickness and dispersion of the prior probability density function. Training directly with these priors is computationally intensive because almost all S alpha S probability densities lack a closed form. A precomputed look-up table removed this computational bottleneck. We tested the new soft diamond regularizers with deep neural classifiers on the three datasets CIFAR-10, CIFAR-100, and Caltech-256. The regularizers improved the accuracy of the classifiers. The improvements included 4.57% on CIFAR-10, 4.27% on CIFAR-100, and 6.69% on Caltech-256. They also outperformed L-2 regularizers on all the test cases. Soft diamond regularizers also outperformed L-1 lasso or Laplace regularizers because they better increased sparsity while improving classification accuracy. Soft-diamond priors substantially improved accuracy on CIFAR-10 when combined with dropout, batch, or data-augmentation regularization.
Non-uniform prior probabilities between hidden layers improved deep neural classifiers trained with bidirectional backpropagation. The resulting Bayesian bidirectional backpropagation algorithm jointly maximizes the forward and backward network likelihoods along with the weight priors. The backward direction exploits a hidden regression that ordinary unidirectional backpropagation ignores. Simulations compared Laplacian, Gaussian, Cauchy, and the new sinc-squared hidden priors on the CIFAR-10 and CIFAR-100 balanced image data sets. These hidden priors improved the classification accuracy of deep neural classifiers compared with default uniform priors and default unidirectional backpropagation. They did so at little extra computational cost. Sinc-squared and Cauchy multivariate priors often had the best classification accuracy. Cauchy hidden priors gave sparse hidden weights similar to the Laplacian priors associated with sparse lasso regression.
A bidirectional autoencoder learns or approximates an identity mapping as it trains a single network with a version of the new bidirectional backpropagation algorithm. Ordinary unidirectional autoencoders find many uses in image processing and in large language models. But they use separate networks for encoding and decoding. Bidirectional auto encoders use the same synaptic weights for encoding and decoding. The forward pass encodes while the backward pass decodes. Bidirectional auto encoders improved network performance and significantly reduced memory usage and used fewer parameters. Simulations compared unidirectional with bidirectional autoencoders for image compression and de noising. The models trained on the MNIST handwritten-digit and CIFAR-IO image datasets. The performance measures were the peak signal-to-noise ratio and the index of structural similarity. Bidirectional autoencoders outperformed unidirectional autoencoders and still reduced the number of trainable synaptic parameters by about 50%.
A statistical formulation of recurrent backpropagation (RBP) allows direct noise boosting for time-varying classification and regression. The noise boost reduces training iterations and improves accuracy. The injected noise is just that noise that makes the current signal more probable. This noise-boost result extends the two recent results that backpropagation is a special case of the generalized expectation maximization (EM) algorithm and that careful noise injection can always speed the average convergence of the EM algorithm to a local maximum of the log-likelihood surface. The noise-benefit conditions differ for additive and multiplicative noise in RBP. We tested noise-boosted RBP classifiers on 11 classes of sports video clips and tested RBP regressors on predicting the dollar-rupee exchange rate. Injecting noisy-EM (NEM) noise outperformed injecting blind noise or injecting no noise at all. Additive NEM noise usually outperformed multiplicative noise. The best case of NEM noise injection with RBP training of a recurrent neural classification model speeded up its training by 60% and improved its classification accuracy by 9.51% compared with noiseless RBP training and accuracy. The best performance of the NEM noise with the RBP training of a recurrent neural regression model yielded a 38% speedup in training and also reduced the squared error by 49.3%. The injection of the additive NEM noise in the output and hidden neurons performed best.
We can combine expert knowledge by combining the probability mixtures that represent the if-then rules of the experts. Fuzzy rules define a generalized probability mixture whose moments describe a fuzzy system and its uncertainty. The mixture’s Bayesian structure gives a complete posterior probability description of the if-then fuzzy-set rules as they fire. A new theorem extends the uniform convergence of a fuzzy system’s mixture to the uniform convergence of the sequence of expert mixtures that represent any number of combined fuzzy systems as they each converge to a target function. A mixture of just two normal bell curves exactly represents the target function in the scalar case and serves as the probabilistic target of the converging mixture sequence. A sampled deep neural network can serve as the target function. Then the mixture defines a proxy system that gives a probabilistic form of explainable AI. The uniform convergence result extends to any continuous transformation of the converging fuzzy systems and further extends to the uniform mixture convergence of any continuous function of the combined systems and their continuous transformations.
A random foam defines a modular rule-based ontology after sampling from a neural or other input-output system. A random foam combines several rule-based systems and averages the systems. It gives a Bayesian posterior over the subsystems. It also gives separate Bayesian posteriors over the rules of each subsystem. The shape of the rules controls how well the random-foam ontology performs in classification and regression. We found that a heterogenous ontology that mixes different rule shapes can perform better than a homogenous ontology based on a single Gaussian or other rule type. Random foams are also universal function approximators. So they can train on a neural black box and act as its explainable proxy system. We prove this uniform approximation theorem for the important case of bump-function random foams with throughput combination. Random foams also measure their output’s uncertainty through the conditional variance. Bump function rules performed better than Cauchy rules at classification while Cauchy rules performed better at regression. Gaussian rules performed best in both classification and regression. A homogeneous Gaussian random foam that trained on a 96.7% accurate neural classifier was itself 95.96% accurate on the MNIST data set. A heterogeneous random foam with two-thirds Gaussian rules and one-third Laplacian rules did better than did the all-Gaussian foam ontology.
The new NoVa hidden neurons have outperformed ReLU hidden neurons in deep classifiers on some large image test sets. The NoVa or nonvanishing logistic neuron additively perturbs the sigmoidal activation function so that its derivative is not zero. This helps avoid or delay the problem of vanishing gradients. We here extend the NoVa to the generalized perturbed logistic neuron and compare it to ReLU and several other hidden neurons on large image test sets that include CIFAR-100 and Caltech-256. Generalized NoVa classifiers allow deeper networks with better classification on the large datasets. This deep benefit holds for ordinary unidirectional backpropagation. It also holds for the more efficient bidirectional backpropagation that trains in both the forward and backward directions.
We show that training neural classifiers with Bayesian bidirectional backpropagation improves the performance of the network. Bidirectional backpropagation trains a deep network for both forward and backward recall through the same layers of neurons and with the same weights. It maximizes the network's joint forward and backward likelihood. Bayesian bidirectional backpropagation combines prior probabilities at the input and output layers with the likelihood structure of the layers. It maximizes the posterior probability of the network. It differs from other forms of neural Bayesian estimation because it uses the bidirectional likelihood of the network instead of the unidirectional likelihood. Bayesian bidirectional backpropagation outperformed classifiers trained with both unidirectional and bidirectional backpropagation. The networks trained on the CIFAR-10 and CIFAR-100 image test sets. A Laplacian or Lasso-like prior outperformed both Gaussian and uniform priors.
The new bidirectional backpropagation algorithm helps blocking networks learn and recall large numbers of image patterns. Bidirectional backpropagation exploits backward-pass learning that ordinary unidirectional backpropagation ignores. The backward pass reveals a hidden regressor in classifiers since the input neurons are identity units. Blocking networks allow deep classifiers to learn and accurately recognize more patterns than the older classifiers that use softmax neurons at the output classification layer. Blocking networks use logistic neurons at the output layer of a block. They use random bipolar coding from the vertices of a hypercube rather than from the vertices of the simplex embedded in it as with l-in-K encoding. Bidirectional deep sweeps improved classification accuracy on the CIFAR-100 image data base and did so at little extra computational cost.
A random foam trains several fuzzy-rule-foam function approximators and then combines them into a single rule-based approximator. The foam systems train independently on bootstrapped random samples from a trained neural classifier. The foam systems convert the neural black box into an interpretable set of rules. The fuzzy rule-based systems have an underlying probability mixture structure that gives rise to an interpretable Bayesian posterior over the rules for each input. A rule foam also measures the uncertainty in its outputs through the conditional variance of the generalized probability mixture. A random foam combines the learned additive fuzzy systems by averaging their throughputs or rule structure. A random foam is also interpretable in terms of its rules, its posterior of its rules, and its conditional variance. Thirty 1000-rule foams trained on random subsets of the MNIST digit data set. Each such foam system had about 93.5% classification accuracy. The random foam that averaged throughputs achieved \(96.80\%\) accuracy while the random foam that averaged only their outputs achieved 96.06% accuracy. The throughput-averaged random foam also slightly outperformed a standard random forest that output-averaged 30 classification trees. Thirty 1000-rule foams also trained on a deep neural classifier that had 96.26% accuracy. The random foam that averaged these foam throughputs was itself 96.14% accurate. The random foam that averaged their outputs was just 95.6% accurate. The appendix proves a Gaussian combined-foam version of the uniform approximation theorem for additive fuzzy systems.
Bidirectional associative memories (BAMs) pass neural signals forward and backward through the same web of synapses. Earlier BAMs had no hidden neurons and did not use supervised learning. They tuned their synaptic weights with unsupervised Hebbian or competitive learning. Two-layer feedback BAMs always converge to fixed-point equilibria for threshold or threshold-like neurons. Every rectangular connection matrix is bidirectionally stable. These simpler BAMs extend to arbitrary hidden layers with supervised learning if the resulting bidirectional backpropagation algorithm uses the proper layer likelihood in the forward and backward directions. Bidirectional backpropagation lets users run deep classifiers and regressors in reverse as well as forward. Bidirectional training exploits pattern and synaptic information that forward-only running ignores.
A random rule foam grows and combines several independent fuzzy rule-based systems by randomly sampling input-output data from a trained deep neural classifier. The random rule foam defines an interpretable proxy system for the sampled black-box classifier. The random foam gives the complete Bayesian posterior probabilities over the foam subsystems that contribute to the proxy system's output for a given pattern input. It also gives the Bayesian posterior over the if-then fuzzy rules in each of these constituent foams. The random foam also computes a conditional variance that describes the uncertainty in its predicted output given the random foam's learned rule structure. The mixture structure leads to bootstrap confidence intervals around the output. Using the Bayesian posterior probabilities to prune or discard low-probability sub-foams improves the system's classification accuracy. Simulations used the MNIST image data set of 60,000 gray-scale images of ten hand-written digits. Dropping the lowest-probability foams per input pattern brought the pruned random foam's classification accuracy nearly to that of the neural classifier. Posterior pruning outperformed simple accuracy pruning of a random foam and outperformed a random forest trained on the same neural classifier.
The new NoVa (nonvanishing) logistic neuron activation allows deeper neural networks because its derivative is positive. So it helps mitigate the problem of vanishing gradients in deep networks. Deep neural classifiers with NoVa hidden units had better classification accuracy on the CFAR-10, CFAR-100, and Caltech-256 image databases compared with threshold-linear ReLU hidden units. Still simpler identity hidden units also outperformed ReLU hidden units in deep classifiers but usually had less classification accuracy than NoVa networks. NoVa hidden neurons also outperformed ReLU hidden neurons in deep convolutional neural networks.
The probability mixture structure of additive fuzzy systems allows uniform convergence of the generalized probability mixtures that represent the if-then rules of one system or of many combined systems. A new theorem extends this result and shows that it still holds uniformly for any continuous function of such fuzzy systems if the underlying functions are bounded. This allows fuzzy rule-based systems to approximate a far wider range of nonlinear behaviors for a given set of sample data and still produce an explainable probability mixture that governs the rule-based proxy system.
A rule foam converts a neural black-box classifier into a probabilistic rule-based ontology where a fresh Bayesian posterior describes the relative rule firings for each input pattern. The rules define a generalized probability mixture that in turns yields the Bayesian posterior over the rules or subsystems. The rules and posterior explain each observed input-output pair and further give a confidence measure of the predicted output in terms of the mixture’s conditional variance. A random foam creates and combines several foams by randomly sampling the neural classifier. It’s mixture structure gives a Bayesian posterior over the constituent foam systems and a finer-grained Bayesian posterior over each system’s rules. We illustrate the foam technique on a deep neural classifier trained on the CIFAR-10 image dataset.