
Graph databases like Neo4j are powerful for analysing interconnected data, but performing advanced analyses such as spectral clustering often requires exporting data to external tools, incurring overhead and breaking the in-database workflow. While Neo4j offers various analytical capabilities, a dedicated, comprehensive plug-in for in-database spectral clustering has been a notable gap. To address this, we introduce SimKit1, a Java-based plug-in for Neo4j that enables true in-database spectral clustering by providing a suite of Cypher-callable procedures for end-to-end spectral clustering, including similarity graph construction (fully connected, epsilon-neighbourhood, k-nearest neighbourhood, and mutual k-nearest neighbourhood), Laplacian eigendecomposition of graph Laplacians, subsequent K-Means clustering of the resultant eigenvectors, and evaluation using silhouette coefficient and the Adjusted Rand Index. We evaluated SimKit on a range of feature-based datasets (Iris, Madelon, 20 Newsgroups) and native graph datasets (CORA, PubMed, CiteSeer). The results demonstrate that SimKit facilitates practical, end-to-end clustering workflows entirely within the Neo4j database system, achieving clustering quality comparable to established Python implementations.
With the widespread adoption of cloud computing and containerization technologies, large-scale clusters now host increasingly complex applications, yet the fragmentation of multidimensional resources (CPU, GPU, memory) has significantly degraded resource utilization efficiency. To address this challenge, this study investigates scheduling strategies for multidimensional resource fragmentation optimization. First, a unified quantitative metric is established by extending a task-statistics-based fragmentation measurement approach to threedimensional resource constraints (CPU, memory, GPU), enabling objective assessment of fragmentation levels. Subsequently, a multidimensional resource fragmentation optimization scheduling strategy is proposed based on the FGD framework. This approach employs TOPSIS to holistically address multidimensional resource fragmentation optimization. Building upon this methodology, we design and implement both two-dimensional and three-dimensional scheduling strategies. Comprehensive evaluation using an extended Kubernetes scheduler simulator with Alibaba production dataset reveals that: 1) All TOPSIS-based multi-dimensional strategies achieve performance comparable to or exceeding baseline approaches in their target dimensions while demonstrating synergistic effects.; 2) Two-dimensional strategies achieve optimal performance in their target dimensions but may degrade others 3) The three-dimensional optimization strategy TOPSIS-FGD-CGM can overcome the limitations of two-dimensional optimization strategies and demonstrates robust comprehensive performance.
Three-dimensional (3D) reconstruction based on aerial images of unmanned aerial vehicles (UAVs) is important for surveying and monitoring natural resources, supporting land cover analysis and geological hazard mapping. However, the performance of a fixed UAV docking station and manual deployments in large areas is limited by slow data acquisition and transmission. To accelerate data acquisition, emerging mobile UAV docking stations can collaboratively capture images in wide areas. Moreover, by leveraging edge computing power on UAVs and docking stations, the computations of 3D reconstruction can be performed locally, eliminating the need for data transmission. This paper builds a mathematical model for the entire process. The model can be decomposed into three associated problems: area partitioning, computation offloading, and mobile docking station path planning. The optimization objective is to minimize the execution time under the constraint of UAV battery power. A suboptimal solution is first derived using an enumeration-genetic algorithm and then fine-tuned using deep reinforcement learning. The experimental results validate the feasibility of using mobile docking stations for 3D reconstruction. In addition, the numerical results indicate that our proposed solution reduces execution time compared to the benchmark solution.
The growing disparity between computing speed and data access times poses a significant challenge to the management of data storage in large-scale supercomputers. To address these challenges, storage systems have evolved into hierarchical architectures with multiple levels that accommodate various hardware technologies. Each level offers a unique combination of performance, cost, and capacity. In this work, we introduce a paradigm shift in data management across the storage tiers by transitioning from block-level to file-level granularity to optimize data placement. Our online prediction model uses file characterization (input, log, checkpoint, work, and output) to predict file reuse with 97% accuracy, categorizing each file during usage and enabling proactive data management strategies. Using three real applications and a benchmark, we validate our work and demonstrate significant improvements in data hit rates, reaching 55% and 57% for LQCD and the benchmark applications, respectively. Even with lower performance improvements, our approach remains competitive with NEMO and NAMD, achieving a 3% gain over both LRU and LFU on NEMO and matching LRU performance while being 1% lower than LFU on NAMD.
The occurrence of geological hazards is not only affected by natural factors, but is also closely related to human engineering activities. In order to systematically integrate information related to multi-source and heterogeneous geological hazards and to improve the ability of hazard awareness and risk assessment, this paper proposes a multi-level and multiscale geological hazards knowledge graph construction method. The method constructs a multilevel geohazard ontology model containing elements such as hazard types, hazard-causing factors, hazard events, and hazard-bearing bodies, and introduces a multiscale spatio-temporal coding mechanism, integrating geological and human activity data at different spatial and temporal scales, to achieve the structured expression and organization of the hazard knowledge.
Traditional CNN and Transformer show good performance in medical image segmentation, but both have problems. CNN capture local features with convolution operations but have trouble modeling long-range dependencies. Transformers use self-attention to understand global context but are very computationally expensive. In medical images, lesions have big changes in shape and size, requiring models to include both general shapes and specific local boundaries. This makes it important to combine local and global features well. To solve these problems, we suggest the Multi-Feature Fusion Mamba model (MF-Mamba). The model uses a Multi-Scale Channel Fusion Network (MCFN) to get object features at different scales and capture local context, helping the model segment lesions of different sizes. It also adds a Direction Perception Attention (DPA) module to capture long-range context, improving the network's capability to model long-range interdependencies. Tests on the ISIC2017, ISIC2018, and Synapse public datasets demonstrate that MF-Mamba works much better in skin lesion segmentation tasks, proving it has a strong edge.
Modern LLM decoding has the drawbacks of high cost and slow speed, and speculative decoding has been shown to be an effective solution to this problem. However, the inference latency still poses a significant challenge to maintaining service level objectives (SLOs) in systems that employ multiple draft models for speculative decoding. The verification phase in such systems if reliant on tree attention often constitutes a bottleneck especially when draft sequences lack common prefixes and substantially underutilizes GPU parallelism while increasing end-to-end latency. We introduce SuperSpec, an end-to-end speculative decoding system designed to co-optimize verification, sampling and draft generation. SuperSpec integrates three pivotal innovations: an Efficient Batch Verifier, which substitutes treebased flattening with batch parallel validation and layer-wise KV Cache replication; a Global Optimal Sampler, which assesses all candidate sequences within a batch to ascertain the longest valid path, thereby circumventing the local optima frequently encountered in tree-based rejection sampling; and a Dynamic Adaptive Multi-Drafter, which dynamically modulates the speculative length (K) for each drafter predicated on real-time idleness metrics and acceptance rates. Empirical evaluations of Qwen2.5-72B and the OPT-66B on various datasets show that SuperSpec improves average acceptance rate by 6.4% to 30.2%, and the end-to-end inference acceleration ratio by 7.12% to 62.06%, when compared to the state-of-the-art tree-based speculative decoding system SpecInfer. These improvements were achieved without compromising the quality of text generation, making SuperSpec an effective solution for accelerating LLM inference.
The growing demand for sustainable energy solutions has elevated photovoltaic (PV) systems as a practical and environmentally viable alternative to conventional power sources in residential settings. This study presents a systematic methodology for the design and performance analysis of a grid-connected PV system tailored for a residential area, with particular emphasis on local meteorological conditions using the PVsyst simulation software. The research encompasses the interpretation of climatic data, system sizing, and performance evaluation. A detailed PV system model was developed in PVsyst to assess key performance metrics, including nominal array energy output, system losses illustrated via Sankey diagrams, effective energy at the array output, and energy delivered to the grid. The proposed system, designed with a nominal capacity of 11.3 kWp, comprises 20 JASolar photovoltaic modules and a 20 kW Huawei inverter. Simulation results indicate an annual energy yield of $\mathbf{2 0, 9 1 8}$ kWh and an average performance ratio of 75.62%. This methodological approach provides a valuable reference for engineers and researchers aiming to adopt climate-responsive design principles for residential PV system deployment.
Urbanisation has led to urban population growth affecting the economy and the environment, including degrading air quality via pollution. Air pollution has been linked to a variety of conditions and health risks including heart disease, stroke, asthma, Alzheimer's and neurodevelopmental disorders. However, it is difficult for a citizen to find precise air pollution data at a particular location. Smart City strategies usually stipulate that city councils should focus on delivering platforms for active citizen participation using existing technology. Existing civic data hubs such as the London Datastore, Open Data Bristol etc., provide air pollution data but lack elaborate representations for user-defined locations. Existing air quality initiatives such as the Smart Citizen platform and Sensor.Community provide more advanced graphical representations. However, they restrict themselves to showing data coming from their respective devices. The paper presents the Open City Air Quality Platform (OpenCAQP), a development that merges a wide range of data sources and air pollution parameters into a single platform. The OpenCAQP allows citizens, environmentalists, data analysts, and developers to access and visualise data. The proposed solution contributes to two key objectives: i) analysis of the air pollution data sources available in a city; ii) a replicable scalable, modular open source capability aggregating and visualising air pollution data from multiple sources. Its effectiveness has been evaluated by measuring quality, usability and increased awareness of users through a feedback questionnaire.
With the rapid development of artificial intelligence (AI) technology, high-quality data has become increasingly critical, as data quality has emerged as a key bottleneck limiting AI applications. Existing data evaluation methods face three main issues: (1) static evaluation models fail to match dynamic training needs, (2) criteria primarily focus on data attributes while ignoring operational and training effectiveness, (3) AI developers and users hold differing perspectives. To effectively address these challenges, we propose a comprehensive three-dimensional evaluation framework encompassing input quality, operation quality, and training quality. This framework integrates methods such as entropy analysis, RoBERTa-based toxicity detection, Markov chain monitoring, and PSI, covering the full data life cycle. Its core innovation lies in incorporating AI-specific factors such as toxicity risk, and distribution stability, while establishing a two-way link between data quality and model performance. Practical applications show this approach improves model training and offers a standardized, reproducible method for data evaluation and governance in AI.
Transformer-based Neural Machine Translation (NMT) models have attained state-of-the-art performance within the machine translation community, automatically acquiring translation knowledge from bilingual corpora through the attention mechanism. This contrasts with the approach of human translators, who heavily rely on prior knowledge during sentence translation. Inspired by this disparity, we propose a Word Translation Augmentation (WTA) method to enhance Transformer-based NMT models. Our methodology comprises three key steps: initially, we construct word alignment rules based on the training set; subsequently, we generate translation rules for source words in accordance with these alignment rules; finally, we integrate potential translation candidates for each source word into the NMT model during both training and testing phases. Additionally, the WTA method introduces the concept of Mixup for augmenting translation candidates of source words and employs two augmentation strategies to enrich the encoder. Experimental results on the WMT14 (German $\leftrightarrow$ English) and AIChallenger 2018 (Chinese $\leftrightarrow$ English) translation tasks demonstrate the effectiveness of our proposed method compared to strong baseline models.
In this paper, we present the Rubus tool suite, with a focus on its static (offline) real-time scheduler. The schedules generated by the scheduler are executed by a real-time operating system certified according to the ISO 26262 safety standard. The Rubus tool suite and its scheduler have been utilized in the vehicle industry for model- and component-based software development of resource-constrained embedded systems for over $\mathbf{2 5}$ years. Since its introduction in 1998, the scheduler has evolved significantly, transitioning from pure Earliest Deadline First (EDF) heuristics to incorporating priorities, data dependencies, and the ability to distribute the schedule over the entire hyperperiod of the software application, among other heuristics. We provide an in-depth discussion on the mechanisms and algorithms that constitute the Rubus offline scheduler. Moreover, we provide an example of its application in generating an offline schedule for a part of software architecture in an industrial setting.
The COVID-19 pandemic posed significant challenges to global healthcare systems, with pronounced effects on critical treatments such as lung cancer care. This study investigates the impact of the pandemic on the public healthcare network in Ceará, Brazil, utilizing complex network analysis. By examining 5,467 Hospitalization Authorization Records (AIH) from DATASUS (2018-2023), the research identifies structural shifts and regional disparities across three phases: pre-pandemic, pandemic, and post-pandemic. Key findings reveal increased centralization in Fortaleza city, network fragmentation in Sobral, and resilience in Cariri. The analysis highlights the necessity for decentralized strategies to enhance future healthcare resilience.
Efficient processing and sharing of geological environmental data are crucial for sustainable development and informed decision-making. However, current analysis methods struggle with low efficiency and resource utilization, especially in complex computational tasks. This paper proposes a parallel processing framework based on container orchestration that systematically improves the efficiency of geological environment data analysis by integrating container technology and complex task processing optimization strategies. Leveraging containerized processing, we established a standardized packaging and deployment mechanism for geological environmental data analysis algorithms, enabling flexible encapsulating and management of multiple models. In addition, we proposed a complex task decomposition method for pipeline parallelism and realized multicontainer collaborative geological environment data processing in a distributed environment based on container orchestration. For enhancing the efficiency of complex task processing purposes, this paper proposed a task scheduling optimization strategy based on the dynamic merging of directed acyclic graph, which improves resource utilization and processing speed through task merging. Experimental results demonstrate that the proposed framework enhances processing efficiency by over 50% in typical geological environmental data analysis scenarios, while improving resource utilization by $\mathbf{4 8 \% - 6 9 \%}$. It exhibits strong reliability and scalability, offering technical support for intelligent analysis and service sharing of geological environmental data.
Single-camera markerless systems have emerged as a robust methodology for human motion capture and rehabilitation applications. Traditional methodologies typically necessitate multiple strategically positioned cameras or special equipment, including sensors to capture patient ambulatory motion, requiring preliminary calibration and synchronization procedures, which may incur significant costs. This paper presents a wireless single-camera markerless framework for rehabilitation applications that leverages advanced deep learning (DL) architectures to estimate and extract three-dimensional skeletal coordinates from monocular camera views of ambulatory patients. The extracted skeletal representation is subsequently transmitted across wireless communication channels. Then, the rendering technique has been applied for displaying virtual movement of the patient for privacy enhancement. Simulations demonstrate the effectiveness of the framework while maintaining motion assessment capabilities, presenting opportunities for deployment in remote healthcare monitoring scenarios.
Generative Adversarial Networks (GANs) have proven to be effective in generating synthesized data. However, their use often requires large amounts of data and sufficient computing power for model training. This presents significant challenges for GAN training in edge devices, which typically lack the necessary resources for GAN training. Additionally, the data collected by these devices often contain sensitive information that requires privacy preservation. The data volume may also be insufficient for GAN training, and the data collected by individual edge devices may have both common and unique features across the network. To address these challenges, we propose TSGAN, a lightweight GAN framework for Edge-Cloud computing. Coordinated by a Cloud server, TSGAN allows a network of resource-limited edge devices to train GAN models for privacy-preserving data generation. We propose a teacher-student model to enable edge devices to generate high-quality data. We also propose a novel deployment mechanism that facilitates effective distributed learning across edge devices and the Cloud server, while preventing the Cloud server generating the synthetic data from accessing the data collected by edge devices. Finally, we introduce a joint restraint learning function that enhances the effectiveness of learning unique features from data on individual edge devices. We conducted extensive experiments. The results have verified the effectiveness of TSGAN.
This paper provides a description of an alternative to compiled code and GUI peripheral configuration tools in the field of embedded systems design and rapid prototyping. This is explored in the form of an interpreted language loaded onto a microcontroller, named NiTTY. Interaction with NiTTY is via a read-evaluate-print loop (REPL) over serial communication. NiTTY is used to demonstrate reliable reconfiguration of peripheral register functions without the need for recompilation or external programs, while also allowing for interfacing with high level languages. NiTTY was designed to imitate the simplicity of shell languages and implemented using a tree-walking parsing technique. A simple data structure containing searchable data about the permutation of registers, their configuration, and status of the enabled peripherals is included. The interpreter was tested using a unit test of the scanner & parser, along with four tests to verify the function and reliability of the program. This paper demonstrates a proof of concept of interfacing platform independence, reliability, and compactness. Questions about the performance, features, and optimisation of the program are raised, and reconciliation of these before effective use for rapid prototyping or commercial development is outlined.
To address the challenges of manual dependency and complex management in railway data governance, and to promote data value realization, we propose DGMAS, a Data Governance Multi-Agent System based on Large Language Models (LLMs). The architectural design adopts a system comprising a mother system and three subsystems. The DGMAS comprises four critical phases: initialization, execution, re-planning, and failure attribution. Through this system, the process of railway data governance can be simulated, to help stakeholders anticipate potential issues and develop solutions, thereby supporting improved railway data governance.
Task scheduling decisions for multimodal large language model (MLLM) applications in multilayer computing power networks present a significant challenge, as they simultaneously balance system delay, carbon emissions, and model accuracy requirements while adapting to network conditions and varying energy availability. Thus, in this paper, we formulate the joint optimization problem of MLLM task scheduling, resource allocation, and green energy utilization to minimize system delay and carbon emissions while meeting accuracy requirements. We propose Green-Aware MAPPO, a novel approach that integrates graph attention networks (GAT) with multi-agent proximal policy optimization (MAPPO) for distributed decision-making in multilayer computing power networks. By modeling the problem as a partially observable Markov decision process (POMDP), our algorithm enables agents to capture complex resource dependencies through relation-specific attention mechanisms while maintaining high performance with limited local observations. Experiments in various network configurations demonstrate that Green-Aware MAPPO significantly outperforms baseline algorithms.