Two disjoint sets of entities and their relationship can be modelled as a bipartite graph. Real-life examples include drug-target interaction in biological networks, user-item relationships in e-commerce networks, etc. Motif-based analysis is essential for understanding the structure of large-scale networks, and bipartite graphs are no exception. In contrast to unsigned graphs, motif analysis in signed bipartite graphs has received limited attention. The smallest non-trivial motif in a signed bipartite graph is a balanced (2,2)-biclique, often called a balanced butterfly, which captures only local patterns and cannot reveal higher-order relationships. Bipartite motifs have been studied in the literature in the context of signed bipartite graphs, such as maximal biclique, bitruss, and so on. None of these works addresses bipartite motifs with fixed-sized vertex sets, which are often relevant in practical situations. In this work, we study the balanced (p,q)-biclique counting problem for small values of p and q. As a baseline, we first adapt and extend the state-of-the-art BCList++ algorithm for unsigned bipartite graphs to incorporate edge signs, which we call SBCList++. We then propose two efficient algorithms: BBWC, a wedge-centric approach that enforces balance constraints during enumeration, and BBVP, a vertex-based pruning approach that directly enumerates feasible vertex sets. Extensive experiments on large real-world datasets demonstrate that the vertex-based pruning algorithm, BBVP, significantly outperforms the baseline, achieving an average speedup of 636× over SBCList++ (where p=q=3).
Balanced butterfly counting, corresponding to counting balanced (2, 2)-bicliques, is a fundamental primitive in the analysis of signed bipartite graphs and provides a basis for studying higher-order structural properties such as clustering coefficients and community structure. Although prior work has proposed an efficient CPU-based serial method for counting balanced (2, k)-bicliques. The computational cost of balanced butterfly counting remains a major bottleneck on large-scale graphs. In this work, we present the highly parallel implementations for balanced butterfly counting for both multicore CPUs and GPUs. The proposed multi-core algorithm (M-BBC) employs fine-grained vertex-level parallelism to accelerate wedge-based counting while eliminating the generation of unbalanced substructures. To improve scalability, we develop a GPU-based method (G-BBC) that uses a tile-based parallel approach to effectively leverage shared memory while handling large vertex sets. We then present an improved variation, G-BBC++, which integrates dynamic scheduling to mitigate workload imbalance and maximize throughput. We conduct an experimental assessment of the proposed methods across 15 real-world datasets. Experimental results exhibit that M-BBC achieves speedups of up to 71.13x (average 38.13x) over the sequential baseline BB2K. The GPU-based algorithms deliver even greater improvements, achieving up to 13,320x speedup (average 2,600x) over BB2K and outperforming M-BBC by up to 186x (average 50x). These results indicate the substantial scalability and efficiency of our parallel algorithms and establish a robust foundation for high-performance signed motif analysis on massive bipartite graphs.
Butterflies, or 4-cycles in bipartite graphs, are crucial for identifying cohesive structures and dense subgraphs. We propose distributed agent-based algorithms for Butterfly Counting in a bipartite graph G((A,B),E), where the agents first determine their partition for which they construct a spanning tree and elect a leader in O(n log lambda) rounds with O (log lambda) bits of memory per agent. A novel meeting mechanism between adjacent agents enhances efficiency and removes the need for prior graph knowledge, requiring only the highest agent ID (lambda) among.. agents. Building on these foundations, agents count butterflies per node in O(Delta) rounds and compute the total butterfly count of G in O(Delta + min{|A|, |B |}) rounds.
This study introduces a novel methodology for assessing ice-jam flood hazards along river channels. It employs empirical equations that relate non-dimensional ice-jam stage to discharge, enabling the generation of an ensemble of longitudinal profiles of ice-jam backwater levels through Monte-Carlo simulations. These simulations produce non-exceedance probability profiles, which indicate the likelihood of various flood levels occurring due to ice jams. The flood levels associated with specific return periods were validated using historical gauge records. The empirical equations require input parameters such as channel width, slope, and thalweg elevation, which were obtained from bathymetric surveys. This approach is applied to assess ice-jam flood hazards by extrapolating data from a gauged reach at Fort Simpson to an ungauged reach at Jean Marie River along the Mackenzie River in Canada’s Northwest Territories. The analysis further suggests that climate change is likely to increase the severity of ice-jam flood hazards in both reaches by the end of the century. This methodology is applicable to other cold-region rivers in Canada and northern Europe, provided similar fluvial geomorphological and hydro-meteorological data are available, making it a valuable tool for ice-jam flood risk assessment in other ungauged areas.
Machine-learning algorithms have been employed in river ice research for flood estimation. This study aimed to introduce a machine learning-based model for predicting ice jam floods. An ice-jam dataset was created using a stochastic modelling approach in which thousands of possible scenarios were simulated. This approach integrated a hydrodynamic model, RIVICE, into a Monte Carlo Analysis (MOCA) framework. The set of parameters, boundary conditions, and associated backwater level elevations was then applied to a machine-learning algorithm to implement a preliminary ice-jam flood prediction model, combining decision tree regressors (DTR) with an adaptive boosting (AdaBoost) regressor. Shapley Additive explanations (SHAP) were then applied in the preliminary model to identify the most influential parameters of ice-jam flooding. Identified variables from SHAP were then used to construct a simple ice-jam flood hazard prediction model with fewer variables. The Athabasca River in Fort McMurray, Canada, is a test site for this modelling framework.
Butterflies, or 4-cycles in bipartite graphs, are crucial for identifying cohesive structures and dense subgraphs. While agent-based data mining is gaining prominence, its application to bipartite networks remains relatively unexplored. We propose distributed, agent-based algorithms for \emph{Butterfly Counting} in a bipartite graph $G((A,B),E)$. Agents first determine their respective partitions and collaboratively construct a spanning tree, electing a leader within $O(n \log λ)$ rounds using only $O(\log λ)$ bits per agent. A novel meeting mechanism between adjacent agents improves efficiency and eliminates the need for prior knowledge of the graph, requiring only the highest agent ID $λ$ among the $n$ agents. Notably, our techniques naturally extend to general graphs, where leader election and spanning tree construction maintain the same round and memory complexities. Building on these foundations, agents count butterflies per node in $O(Δ)$ rounds and compute the total butterfly count of $G$ in $O(Δ+\min\{|A|,|B|\})$ rounds.
In northern regions, river ice- jam flooding can be more severe than open-water flooding causing property and infrastructure damages, loss of human life and adverse impacts on aquatic ecosystems. Very little has been performed to assess the risk induced by ice-related floods because most risk assessments are limited to open-water floods. The specific objective of this study is to incorporate ice-jam numerical modelling tools (e.g. RIVICE, Monte-Carlo simulation) into flood hazard and risk assessment along the Peace River at the Town of Peace River (TPR) in Alberta, Canada. Adequate historical data for different ice-jam and open-water flooding events were available for this study site and were useful in developing ice-affected stage-frequency curves. These curves were then applied to calibrate a numerical hydraulic model, which simulated different ice jams and flood scenarios along the Peace River at the TPR. A Monte-Carlo analysis was then carried out to acquire an ensemble of water level profiles to determine the 1:100-year and 1:200-year annual exceedance probability flood stages for the TPR. These flood stages were then used to map flood hazard and vulnerability of the TPR. Finally, the flood risk for a 200-year return period was calculated to be an average of $32/m(2)/a ($/m(2)/a corresponds to a unit of annual expected damages or risk). Copyright (c) 2016 John Wiley & Sons, Ltd.
Ice-jam flooding is a prevalent extreme event that impacts flood hazard and vulnerability. We introduce a conceptual model framework for Dynamic Ice-jam Flood Risk Assessment (DIFRA). DIFRA integrates ice-jam flood hazard, ice-jam flood risk, and human adaptation. Using agent-based modeling, we captured top-down (artificial breakup) and bottom-up (flood-proofing) adaptive behavior. Our study in Fort McMurray, Canada, shows the complex interaction between micro-level behaviors and macro-level phenomena over time. Our variance-based global sensitivity analysis shows the role of dynamic adaptive behavior in ice-jam flood risk, where the artificial breakage by the government can lead to a regime shift and a decrease in the ice-jam flood risk. However, it can also decrease the number of newly adapted residents to flood-proofing and the role of residents in ice-jam flood risk. DIFRA offers a comprehensive approach to understanding and managing ice-jam flood risk, with potential applications to similar riverine communities in cold regions.
Analysis of large-scale networks for different structural patterns (also called motifs) remains an active area of research in the domain of graph data management and mining. In the past three decades, research has led to a large volume of literature in this area. However, the literature in the context of a signed bipartite graph is limited. In this paper, we study the problem of counting the balanced motifs in a signed bipartite graph. Specifically, We design an efficient algorithm BB2K for counting balanced (2, k)-bicliques. Previous studies for balanced bicliques in a signed bipartite graph focus on a set enumeration-based approach. We observe that the set enumeration-based approaches are expensive as they need to discard a large number of structures, which are not balanced. We take a different approach where we systematically group the symmetric and asymmetric wedges and count the balanced bicliques by counting on those wedges in such a way that we eliminate the generation of unbalanced structures completely. We conducted experiments with nine real-life datasets, and the experimental results demonstrate that our algorithm BB2K is more than 100× faster than the baseline SBCList++ - an adaptation of the state-of-the-art algorithm BCList++ combined with the filtering technique to filter out unbalanced bicliques. We have also shown the scalability of the proposed algorithm by using datasets of different sizes in our experiments.
Triangle counting in graphs is a fundamental problem with a diverse application domain. In this paper, we propose a solution to the triangle counting problem in an anonymous graph using autonomous mobile agents. We further use the triangle count to address the Truss Decomposition problem which involves finding maximal sub-graphs with strong interconnections. Truss decomposition helps in identifying maximal, highly interconnected sub-graphs, or trusses, within a network. Additionally, the triangle count is also used to compute two important metrics - Triangle Centrality and Local Clustering Coefficient for the nodes of the graph. Our goal is to devise algorithms that effectively solve these problems minimizing both the overall time complexity and the memory usage at each agent.
Graphlet counting is an important problem as it has numerous applications in several fields, including social network analysis, biological network analysis, transaction network analysis, etc. Most of the practical networks are dynamic. A graphlet is a subgraph with a fixed number of vertices and can be induced or non-induced. There are several works for counting graphlets in a static network where graph topology never changes. Surprisingly, there have been no scalable and practical algorithms for maintaining all fixed-sized graphlets in a dynamic network where the graph topology changes over time. We are the first to propose an efficient algorithm for maintaining graphlets in a fully dynamic network. Our algorithm is efficient because (1) we consider only the region of changes in the graph for updating the graphlet count, and (2) we use an efficient algorithm for counting graphlets in the region of change. We show by experimental evaluation that our technique is more than 10x faster than the baseline approach.
River ice-jams can create severe flooding along many rivers in cold regions. While ice-jams often form during the spring breakup, the midwinter breakup can cause ice-jamming and flooding. Although many studies have already been focused on forecasting spring ice-jam flooding, studies related to forecasting mid-winter breakup jamming and flooding severity are sparse. The main purpose of this research is to develop a stochastic framework to forecast the severity of mid-winter ice-jam flooding along the transborder (New Brunswick/Maine) Saint John River of North America. A combination of hydrological (MESH) and hydraulic model (RIVICE) simulations was applied to develop the stochastic framework. A mid-winter breakup along the river that occurred in 2018 has been hindcasted as a case study. The result shows that the modelling framework can capture the real-time ice-jam severity. The results of this research will help to improve the capacity of ice-jam flood management in cold regions.
Bipartite graphs offer a powerful framework for modeling complex relationships between two distinct types of vertices, incorporating probabilistic, temporal, and rating-based information. While the research community has extensively explored various types of bipartite relationships, there has been a notable gap in studying Signed Bipartite Graphs, which capture liking / disliking interactions in real-world networks such as customer-rating-product and senator-vote-bill. Balance butterflies, representing 2 x 2 bicliques, provide crucial insights into antagonistic groups, balance theory, and fraud detection by leveraging the signed information. However, such applications require counting balance butterflies which remains unexplored. In this paper, we propose a new problem: counting balance butterflies in a signed bipartite graph. To address this problem, we adopt state-of-the-art algorithms for butterfly counting, establishing a smart baseline that reduces the time complexity for solving our specific problem. We further introduce a novel bucket approach specifically designed to count balanced butterflies efficiently. We propose a parallelized version of the bucketing approach to enhance performance. Extensive experimental studies on nine real-world datasets demonstrate that our proposed bucket-based algorithm is up to 120x faster over the baseline, and the parallel implementation of the bucket-based algorithm is up to 45x faster over the single core execution. Moreover, a real-world case study showcases the practical application and relevance of counting balanced butterflies.
In the higher latitudes of the northern hemisphere, ice jam related flooding can result in millions of dollars of property damages, loss of human life and adverse impacts on ecology. Since ice-jam formation mechanism is stochastic and depends on numerous unpredictable hydraulic and river ice factors, ice-jam associated flood forecasting is a very challenging task. A stochastic modelling framework was developed to forecast real-time ice jam flood severity along the transborder (New Brunswick/Maine) Saint John River of North America during the spring breakup 2021. Modélisation environnementale communautaire—surface hydrology (MESH), a semi-distributed physically-based land-surface hydrological modelling system was used to acquire a 10-day flow forecast. A Monte-Carlo analysis (MOCA) framework was applied to simulate hundreds of possible ice-jam scenarios for the model domain from Fort Kent to Grand Falls using a hydrodynamic river ice model, RIVICE. First, a 10-day outlook was simulated to provide insight on the severity of ice jam flooding during spring breakup. Then, 3-day forecasts were modelled to provide longitudinal profiles of exceedance probabilities of ice jam flood staging along the river during the ice-cover breakup. Overall, results show that the stochastic approach performed well to estimate maximum probable ice-jam backwater level elevations for the spring 2021 breakup season.
Assessing the impact of future climate on the severity of ice jam floods (IJFs) is an essential component of a flood mitigation strategy for many ice jam-prone northern communities. The general circulation model (GCM) outputs are used to derive hydrological conditions under future climate scenarios. Although GCMs are often downscaled to the point of interest, there can still be significant differences between modelled climate scenarios and historically observed climate scenarios. Therefore, the changes between the model simulated baseline and future scenarios are applied to observed baseline values to derive projected future values. In IJF modelling, such projected values are provided as input (boundary conditions) to assess the frequency and severity of future IJF events. Different methods can be used to calculate the climate change signal (difference between model-simulated future vs model-simulated historical conditions), and depending upon the method employed, results can vary. In this study, we evaluate the impact of using different delta change methods (i.e., absolute vs relative) to directly bias-correct the hydrological model output on the assessment of frequency and severity of IJFs under future climate. The method was tested in the Athabasca River at Fort McMurray in western Canada, an ice jam-prone location, to assess the IJF probabilities and intensities in the 2041–2070 period. Our results indicate that there is a notable difference in the projected frequency and severity of IJFs between absolute and relative delta change approaches, suggesting the methods should be carefully selected and results cautiously interpreted.
Projection of the impact of future climate on ice-jam flood intensity is an essential component of a flood mitigation strategy for many northern communities. General Circulation Model (GCM) outputs are used to derive hydrological conditions under future climate scenarios. Although GCMs are often downscaled to a point of interest, there can still be significant differences between modelled climate scenarios and historically observed climate scenarios. Therefore, the model-indicated changes between baseline and future values of climatic scenarios are applied to observed baseline values to estimate projected future values. This can be carried out by using the delta change method which is an approach for adjusting GCM output. This study evaluates the impact of the delta change method on the frequency and severity of ice-jam flooding under a future climate scenario. The Athabasca River at Fort McMurray is presented as the test site. Streamflow conditions were derived from a physically-based hydrological model, Modélisation Environnementale communautaire-Surface Hydrology (MESH), by forcing the Canadian Regional Climate Model (CRCM) driven by the Third Generation Coupled Climate Model (CGCM3) for both baseline (1971–2000) and future (2041–2070) periods. Streamflow under future climatic conditions was developed based on the delta change method for both absolute and relative changes. The adjusting streamflow was then used in a fully dynamic river ice hydraulic model, RIVICE, to project future ice-jam scenarios using a stochastic modelling framework. Finally, the impact of the delta changes on the frequency and severity of simulated ice-jam flooding was assessed by producing ice-jam stage-frequency distributions (SFDs) under future climatic conditions. The results indicate that there is a notable difference in the projected frequency and severity of ice-jam flooding between absolute and relative change approaches.
Abstract Ice‐jam flood risk management requires new approaches to reduce flood damages. Although many structural and non‐structural measures are implemented to reduce the impacts of ice‐jam flooding, there are still many challenges in identifying appropriate strategies to reduce the ice‐jam flood risk along northern rivers. The main purpose of this study is to provide a novel methodological framework to assess the feasibility of various ice‐jam flood mitigation measures based on risk analysis. A total of three ice‐jam flood mitigation measures (artificial breakup, sediment dredging and dike installation) were examined using a stochastic modelling framework for the potential to reduce the ice‐jam flood risk along the Athabasca River at Fort McMurray. An ensemble of hundreds of backwater level profiles was used to construct ice‐jam flood hazard maps to estimate expected annual damages, using depth‐damage curves for structural and content damages, within the downtown area of Fort McMurray. The results show that, while sediment dredging may be able to reduce a certain level of expected annual damages in the town, and artificial breakup and a dike with a crest elevation of 250 m a.s.l. can be the most effective measures to reduce the amount of expected annual damages.
Quasi-cliques are dense incomplete subgraphs of a graph that generalize the notion of cliques. Enumerating quasi-cliques from a graph is a robust way to detect densely connected structures with applications in bioinformatics and social network analysis. However, enumerating quasi-cliques in a graph is a challenging problem, even harder than the problem of enumerating cliques. We consider the enumeration of top-k degree-based quasi-cliques and make the following contributions: (1) we show that even the problem of detecting whether a given quasi-clique is maximal (i.e., not contained within another quasi-clique) is NP-hard. (2) We present a novel heuristic algorithm KernelQC to enumerate the k largest quasi-cliques in a graph. Our method is based on identifying kernels of extremely dense subgraphs within a graph, followed by growing subgraphs around these kernels, to arrive at quasi-cliques with the required densities. (3) Experimental results show that our algorithm accurately enumerates quasi-cliques from a graph, is much faster than current state-of-the-art methods for quasi-clique enumeration (often more than three orders of magnitude faster), and can scale to larger graphs than current methods.