Mainframe systems, written in legacy languages such as COBOL, PL/I, and JCL, continue to support missioncritical applications across various industries. Their complexity and limited documentation hinder maintenance and modernization, especially in regions where localized explanations are essential for accurate understanding. However, existing approaches predominantly generate English-only outputs and rely on resource-intensive models unsuitable for secure, on-premises environments. This study explores multilingual explanation generation for mainframe programs using lightweight language models suitable for constrained enterprise settings. We evaluate two strategies-(a) direct generation in the target language and (b) translation-based generation from English—across five languages: Japanese, French, German, Spanish, and Portuguese. Explanation quality is assessed using BLEU, ROUGE-L, METEOR, and semantic similarity. Preliminary results show that lightweight models can produce semantically adequate multilingual explanations. Translation-based generation generally yields higher lexical and structural quality across languages and models, while direct generation shows promise in specific scenarios. These findings demonstrate the feasibility of deploying multilingual explanation systems in enterprise environments and highlight opportunities to refine generation strategies based on language and code characteristics.
Recent advances in Large Language Model (LLM) based Generative AI techniques have made it feasible to translate enterprise-level code from legacy languages such as COBOL to modern languages such as Java or Python. While the results of LLM-based automatic transformation are encouraging, the resulting code cannot be trusted to correctly translate the original code, making manual validation of translated Java code from COBOL a necessary but time-consuming and labor-intensive process. In this paper, we share our experience of developing a testing framework for IBM Watsonx Code Assistant for Z (WCA4Z) [5], an industrial tool designed for COBOL to Java translation. The framework automates the process of testing the functional equivalence of the translated Java code against the original COBOL programs in an industry context. Our framework uses symbolic execution to generate unit tests for COBOL, mocking external calls and transforming them into JUnit tests to validate semantic equivalence with translated Java. The results not only help identify and repair any detected discrepancies but also provide feedback to improve the AI model.
Recent advances in Large Language Model (LLM) based Generative AI techniques have made it feasible to translate enterprise-level code from legacy languages such as COBOL to modern languages such as Java or Python. While the results of LLM-based automatic transformation are encouraging, the resulting code cannot be trusted to correctly translate the original code. We propose a framework and a tool to help validate the equivalence of COBOL and translated Java. The results can also help repair the code if there are some issues and provide feedback to the AI model to improve. We have developed a symbolic-execution-based test generation to automatically generate unit tests for the source COBOL programs which also mocks the external resource calls. We generate equivalent JUnit test cases with equivalent mocking as COBOL and run them to check semantic equivalence between original and translated programs. Demo Video: https://youtu.be/aqF_agNP-lU
Registering altered domain names with the purpose of confusing users and conducting malicious activities is one of the most widespread types of attacks on the Web, conforming a family of techniques known as domain squatting. Detecting these domains is a difficult task, given the large a mount of combinations and the massive and heterogeneous nature of the Web. In this work, we propose a set of models to firstly learn the distributional regularities from detected squatted domains, and from that, automatically generate realistic modified domains. Our goal is to proactively guide the generation of squatted domains towards malicious domains that exists but have not been detected yet. We conducted an empirical study for both typo-squatting and combo-squatting generation approaches against strong baselines on real world data, showing their feasibility and providing insights to support for proactive defense in the context of cloud security.
It is crucial to ensure that server rack doors are properly closed in data centers in order to prevent serious dangers caused by the doors suddenly opening in disasters and causing accidents that hurt maintenance workers and damage facilities. In this paper, we propose multi-factor-based motion detection for server rack doors that have been left open. Our approach recognizes the status of server rack doors on the basis of maintenance workers' motion and prevents a worker from forgetting to close the doors without using any additional devices, only smartphones, which maintenance engineers already carry for work.
In software development, bug localization is the process finding portions of source code associated to a submitted bug report. This task has been modeled as an information retrieval task at source code file, where the report is the query. In this work, we propose a model that, instead of working at file level, learns feature representations from source changes extracted from the project history at both syntactic and code change dependency perspectives to support bug localization. To that end, we structured an end-to-end architecture able to integrate feature learning and ranking between sets of bug reports and source code changes. We evaluated our model against the state of the art of bug localization on several real world software projects obtaining competitive results in both intra-project and cross-project settings. Besides the positive results in terms of model accuracy, as we are giving the developer not only the location of the bug associated to the report, but also the change that introduced, we believe this could give a broader context for supporting fixing tasks.
In this paper, we describe our proposal for the task of Semantic Extraction from Cybersecurity Reports. The goal is to explore if natural language processing methods can provide relevant and actionable knowledge to contribute to better understand malicious behavior. Our method consists of an attention-based Bi-LSTM which achieved competitive performance of 0.57 for the Subtask 1. In the due process we also present ablation studies across multiple embeddings and their level of representation and also report the strategies we used to mitigate the extreme imbalance between classes.
In the cloud based service provisioning industry, one of the main challenges that providers face involves keeping existing tenants engaged while attracting new ones. To address this, providers need to gain insights about customer satisfaction. In that regards, support ticket data, understood as the main way of communication between both parties, can be mined to obtain an estimation of customer satisfaction by means of the polarity of the sentiment extracted from the report descriptions. To that end, in this work we propose a model which can learn a feature representation for sentiment polarity changes, from the sequence of tickets emitted by a given customer during the period associated with the service subscription term. Then, that resulting feature representation, combined with other handcrafted features related to contract and ticket data, is passed to a classifier which estimates the likelihood of service subscription renewal by the customer. Experiment results using real data from a service provider shows that learned representation of sentiment polarity changes from support ticket data in combination with other handcrafted features improves the accuracy in predicting subscription renewals. Moreover, our architecture is flexible enough incorporate and integrate several feature representations and give more expressive power to the prediction.
We propose to study the generation of descriptions from source code changes by integrating the messages included on code commits and the intra-code documentation inside the source in the form of docstrings. Our hypothesis is that although both types of descriptions are not directly aligned in semantic terms —one explaining a change and the other the actual functionality of the code being modified— there could be certain common ground that is useful for the generation. To this end, we propose an architecture that uses the source code-docstring relationship to guide the description generation. We discuss the results of the approach comparing against a baseline based on a sequence-to-sequence model, using standard automatic natural language generation metrics as well as with a human study, thus offering a comprehensive view of the feasibility of the approach.
Log-based business operation analysis is getting more and more attention from enterprise decision makers. However, at the very first step of the analysis service, two primary obstacles are always encountered: how to process a wide variety of local event log formats and how to handle personally identifiable information weaved in an event log. Due to these obstacles, typical business analysts who do not have programming skills have lost business opportunities at the early stages. We propose a privacy-preserving data curation specification language, BELAS, for such analysts and present experimental results that show how most of a real-life event log could be processed in process analysis services.
Extended abstract of a paper presented at Microscopy and Microanalysis 2013 in Indianapolis, Indiana, USA, August 4 – August 8, 2013.
Reducing the energy used in Cloud Computing is an important issue a sustainable society. There are many existing approaches for reducing energy use in data centers, but new approaches are needed in case of energy management of Cloud Computing. Cloud Computing involves decentralized data centers, so new flexible way for collecting energy consumption data becomes quite important. Many current approaches focus on reducing energy consumption by air handling equipment, however energy from IT resources also need to be optimized for a total energy management of Cloud Computing. For advanced energy management for Cloud Computing, we developed a Cloud energy management system with sensor management functions, with an optimized VM allocation tool to minimize energy consumption at multiple data centers. Our evaluations showed more than a 30% energy savings for the servers in our experimental environment. Our system can be extended to optimize energy usage from various perspectives, such as for minimizing electricity bills or carbon emissions.
Companies around the world are increasingly expected to report their greenhouse gas emissions. Currently there are various formulas to calculate emissions, and there are different reporting formats. Most of the reporting formats are paper-based or non-readable-by-machine formats. The emissions of companies will influence their accounting results due to ‘cap & trade’ systems or environmental taxes. Analyses of financial impacts are important for management decisions and corporate evaluations by interested third parties. A standardized reporting format for GHG (greenhouse gas) emissions is critical for reliable analysis of the impact of emissions on finances. This paper proposes an XBRL (eXtensible Business Markup Language) format as the foundation for standardizing the emissions reporting formats, and provides a preliminary XBRL taxonomy for emissions reporting. XBRL makes it possible to combine the financial reports and the emissions reports. Evaluations of the emissions impact are easier for both managers of the company and external parties, even if a large number of emissions reports must be analyzed.
Companies are being required to manage and reduce their emissions to address the climate change problem. It is quite important to correctly monitor and calculate emissions from business activities for reliable emissions trading. Some programs for monitoring or calculations have been developed, but we believe that the IT system needs to manage emissions considering the entire energy management cycle, not just monitoring. We propose a Cloud-based energy management system that allows us to manage emissions according to the full cycle and output emissions reports in the XBRL format. These emissions results have been used as new evaluation indices with financial numbers. Applying XBRL to emissions reporting makes it easy to combine emissions data and financial numbers, and hence it is especially beneficial for interested third parties. We give an example of advanced analysis using emissions and financial data extracted from XBRL. Our approach can also be extended to other kinds of environmental resource management beyond emissions, such as water management.
The configuration of non-functional requirements, such as security, has become important for SOA applications, but the configuration process has not been discussed comprehensively. In current development processes, the security requirements are not considered in upstream phases and a developer at a downstream phase is responsible for writing the security configuration. However, configuring security requirements properly is quite difficult for developers because the SOA security is cross-domain and all required information is not available in the downstream phase. To resolve this problem, this chapter clarifies how to configure security in the SOA application development process and defines the developer’s roles in each phase. Additionally, it proposes a supporting technology to generate security configurations: Model-Driven Security. The authors propose a methodology for end-to-end security configuration for SOA applications and tools for generating detailed security configurations from the requirements specified in upstream phases model transformations, making it possible to configure security properly without increasing developers’ workloads.
An application based Service-Oriented Architecture(SOA) consists of an assembly of external services and the application is called as a composite service. Acomposite service could be implemented by other composite services hence the application could have a recursive structure, which is one of the features of SOA application. Securing an SOA application is an important non-functional requirement. However, specifying a security policy of a composite service is not so easy because the policy should keep the consistency with other policies of external services which are invoked in the process. We need the way to assure the consistency of policies, but the concrete way is not developed yet to specify a consistent policy for a composite service. Therefore, this paper proposes a security policy composition mechanism from existing policies of external services. Our contribution is creating a security policy of a composite service automatically based on predicate logic, with support for two approaches of policy composition: bottom-up and top-down. Also, we focus on three kinds of security policies, such as a Data Protection Policy, an Access Control Policy, and a Composite Process Policy, and propose the policy composition rules for each policy. Our mechanism makes it possible to validate the consistency of policies by inference without increasing a developer's workload, even if a composite service has a recursive structure.
Web Services Security (WS-Security) is a technology to secure the data exchanges in SOA applications. The security requirements for WS-Security are specified as a security policy expressed in Web Services Security Policy (WS-SecurityPolicy). The WS-I Basic Security Profile (BSP) describes the best-practices security practices for addressing the security concerns of WS-Security. It is important to prepare BSP-conformant security policies, but it is quite hard for developers to create valid security polices because the security policy representations are complex and difficult to fully understand. In this paper, we present a validation technology for security policy conformance with WS-Security messages. We introduce an Internal Representation (IR) representing a security policy and its validation rules, and a security policy is known to be valid if it conforms to the rules after the policy is transformed into the IR. We demonstrate the effectiveness of our validation technology and evaluate its performance on a prototype implementation. Our technology makes it possible for a developer without deep knowledge of WS-Security and WS-SecurityPolicy to statically check if a policy specifies appropriate security requirements.
Kugamoorthy Gajananan合作论文数National Institute of Informatics, The Graduate University for Advanced Studies (SOKENDAI)4
Michiaki Tatsubori合作论文数IBM Research - Tokyo2