We demonstrate 128 Gbps/port (8-λ×16 Gbps/λ) natively error-free transmission across eight optical ports using a 8-port, 8-λ/port WDM remote laser source and a pair of monolithically integrated CMOS optical I/O chiplets with 4.96-5.56 pJ/bit optical Tx+Rx chiplet energy efficiency.
Outsourcing enterprise IT service management is an increasingly challenging business. On one hand, service providers must deliver with respect to customer expectations of service quality and innovation. On the other hand, they must continuously seek competitive reductions in the costs of service delivery and management. These targets can be achieved with integration of innovative service management tools, automation, and advanced analytics. In this paper, we focus on service analytics, the subset of analytics problems and solutions concerning specific service delivery and management performance and cost optimization. The paper reviews various service analytics methods and technologies that have been developed and applied to enhance IT service management. We use our industrial experience to highlight the challenges faced in the development and adoption of service analytics, and we discuss open problems.
Automated Teller Machine (ATM) service providers are increasingly challenged with improving the quality of customer service while reducing the cost of cash flow management. Effectively balancing the need to have enough cash in the ATMs to avoid out-of-cash incidents as well as to reduce the cash interest cost and the cash refill cost challenges the most experienced cash flow management teams. In this paper we propose an optimization framework for managing the ATM cash flow network. The interactions among various constraints and cost factors are included in the framework to allow decision-making regarding the optimal cash refill amount and schedule. We demonstrate the effectiveness of the proposed approach using sample data from a large commercial bank.
Extractive text or speech summarization manages to select a set of salient sentences from an original document and concatenate them to form a summary, enabling users to better browse through and understand the content of the document. A recent stream of research on extractive summarization is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing speech summarization in an unsupervised fashion. However, one of the major challenges facing the LM approach is how to formulate the sentence models and accurately estimate their parameters for each sentence in the document to be summarized. In view of this, our work in this paper explores a novel use of recurrent neural network language modeling (RNNLM) framework for extractive broadcast news summarization. On top of such a framework, the deduced sentence models are able to render not only word usage cues but also long-span structural information of word co-occurrence relationships within broadcast news documents, getting around the need for the strict bag-of-words assumption. Furthermore, different model complexities and combinations are extensively analyzed and compared. Experimental results demonstrate the performance merits of our summarization methods when compared to several well-studied state-of-the-art unsupervised methods.
Ticket annotation and search has become an essential research subject for the successful delivery of IT operational analytics. Millions of tickets are created yearly to address business users' IT related problems. In IT service desk management, it is critical to first capture the pain points for a group of tickets to determine root cause; secondly, to obtain the respective distributions in order to layout the priority of addressing these pain points. An advanced ticket analytics system utilizes a combination of topic modeling, clustering and Information Retrieval (IR) technologies to address the above issues and the corresponding architecture which integrates of these features will allow for a wider distribution of this technology and progress to a significant financial benefit for the system owner. Topic modeling has been used to extract topics from given documents; in general, each topic is represented by a unigram language model. However, it is not clear how to interpret the results in an easily readable/understandable way until now. Due to the inefficiency to render top concepts using existing techniques, in this paper, we propose a probabilistic framework, which consists of language modeling (especially the topic models), Part-Of-Speech (POS) tags, query expansion, retrieval modeling and so on for the practical challenge. The rigorously empirical experiments demonstrate the consistent and utility performance of the proposed method on real datasets.
Recent years have seen a major increase in the application of predictive analytics to the service delivery domain as more and more service providers rely on such analytics for proactive risk management. At the pre-contract stage, identifying potential project risks accurately is of vital importance since it allows service providers to avoid profit erosion through proactive risk management. This paper describes a data-driven approach to project failure prediction of complex information technology (IT) projects. We introduce a novel theoretical framework of Latent Trait Analysis (LTA), whose original form was first developed in psychometrics. We take as the input questionnaire data of risk assessment reviews in the quality assurance (QA) process of IT projects before contract signing, and attempt to predict the project health in the delivery phase after contract signing. The idea is to explicitly capture the human cognitive process through LTA, and estimate the latent project failure tendency hidden behind the questionnaire answers collected by QA experts. Using real QA data of an IT service provider, we demonstrate that our approach outperforms existing approaches in project failure prediction while providing practical information on the usefulness of individual question items.
The task of extractive speech summarization is to select a set of salient sentences from an original spoken document and concatenate them to form a summary, facilitating users to better browse through and understand the content of the document. In this paper we present an empirical study of leveraging various supervised discriminative methods for effectively ranking important sentences of a spoken document to be summarized. In addition, we propose a novel margin-based discriminative training (MBDT) algorithm that aims to penalize non-summary sentences in an inverse proportion to their summarization evaluation scores, leading to better discrimination from the desired summary sentences. By doing so, the summarization model can be trained with an objective function that is closely coupled with the ultimate evaluation metric of extractive speech summarization. Furthermore, sentences of spoken documents are embodied by a wide range of prosodie, lexical and relevance features, whose utilities are extensively compared and analyzed. Experiments conducted on a Mandarin broadcast news summarization task demonstrate the performance merits of our summarization method when compared to several well-studied state-of-the-art supervised and unsupervised methods.
Statistical language modeling (LM) that purports to quantify the acceptability of a given piece of text has long been an interesting yet challenging research area.In particular, language modeling for information retrieval (IR) has enjoyed remarkable empirical success; one emerging stream of the LM approach for IR is to employ the pseudo-relevance feedback process to enhance the representation of an input query so as to improve retrieval effectiveness.This paper presents a continuation of such a general line of research and the main contribution is threefold.First, we propose a principled framework which can unify the relationships among several widely-used query modeling formulations.Second, on top of the successfully developed framework, we propose an extended query modeling formulation by incorporating critical query-specific information cues to guide the model estimation.Third, we further adopt and formalize such a framework to the speech recognition and summarization tasks.A series of empirical experiments reveal the feasibility of such an LM framework and the performance merits of the deduced models on these two tasks.
Ticket annotation and search has become an important research subject in the IT service desk delivery. Millions of tickets are created yearly to address business users' IT related problems. In IT service desk management, it is critical to first capture the pain points for a group of tickets to determine root cause; secondly, to obtain the respective distributions in order to layout the priority of addressing these pain points. An advanced ticket analytics system utilizes a combination of topic modeling and clustering to address the above issues and the integration of these features into information architecture will allow for a wider distribution of this technology and progress to a remarkable financial impact for IT industry. Topic modeling has been used to extract topics from given documents; each topic is represented by unigram distributions. However, it is not clear how to interpret the results. Due to the inadequacy to render top concepts, in this paper, we propose a probabilistic framework, which integrates topic models, POS tags, query expansion and so on, for the practical challenge. The rigorously empirical experiments demonstrate the consistent and utility performance of the proposed method on real datasets.
The adoption of the cloud computing model continues to be dominated by startups seeking to build new applications that can take advantage of the cloud's pay-as-you-go pricing and resource elasticity. In contrast, large enterprises have been slow to adopt the cloud model, partly because migrating legacy applications to the cloud is technically non-trivial and economically prohibitive. Both challenges arise, in part, from the difficulty in discovering the complex dependencies that these legacy applications have on the underlying IT environment. In this paper, we introduce a novel Kullback-Leibler (KL) divergence based method that can systematically discover the complex server-to-server and application-to-server relationships. We evaluate our method using live real datasets from large enterprise migration efforts. Our results demonstrate that our new method is capable of finding critical application correlations; it performs better than traditional approaches, such as Bayesian or mutual information models. Additionally, by cleverly subdividing the sample space, we are able to uncover intriguing phenomena in different subspaces. These analyses aid migration engineers in a variety of tasks ranging from migration planning to failure mitigation, and can potentially lead to significant cost reduction in migration to cloud.
Ticketing is a fundamental management process of IT service delivery. Customers typically express their requests in the form of tickets related to problems or configuration changes of existing systems. Tickets contain a wealth of information which, when connected with other sources of information such as asset and configuration information, monitoring information, can yield new insights that would otherwise be impossible to gain from one isolated source. Linking these various sources of information requires a common key shared by these data sources. The key is the server names. Unfortunately, due to historical as well as practical reasons, the server names are not always present in the tickets as a standalone field. Rather, they are embedded in unstructured text fields such as abstract and descriptions. Thus, automatically identifying server names in tickets is a crucial step in linking various information sources. In this paper, we present a statistical machine learning method called Conditional Random Field (CRF) that can automatically identify server names in tickets with high accuracy and robustness. We then illustrate how such linkages can be leveraged to create new business insights.
A part and parcel of any automatic speech recognition (ASR) system is language modeling (LM), which helps to constrain the acoustic analysis, guide the search through multiple candidate word strings, and quantify the acceptability of the final output hypothesis given an input utterance. Despite the fact that the n-gram model remains the predominant one, a number of novel and ingenious LM methods have been developed to complement or be used in place of the n-gram model. A more recent line of research is to leverage information cues gleaned from pseudo-relevance feedback (PRF) to derive an utterance-regularized language model for complementing the n-gram model. This paper presents a continuation of this general line of research and its main contribution is two-fold. First, we explore an alternative and more efficient formulation to construct such an utterance-regularized language model for ASR. Second, the utilities of various utterance-regularized language models are analyzed and compared extensively. Empirical experiments on a large vocabulary continuous speech recognition (LVCSR) task demonstrate that our proposed language models can offer substantial improvements over the baseline n-gram system, and achieve performance competitive to, or better than, some state-of-the-art language models.
IT service delivery relies on intelligent data-driven insights to make strategic decisions. It is a highly complex business with many sub-organizations that focus on different aspects of delivery operations. High-level business insights that emerge from understanding the collective value of all these viewpoints are invaluable to achieving excellent service quality and solid profit margin. However, this is hindered by the inability to integrate data models and taxonomies across business components such as asset management, configuration management, and incident management. Innovative solutions are necessary to effectively "connect-the-dots", bridging the gaps between available content and higher-level business insights. In this paper, we describe several real-world business decisions in service delivery and logistics that suffer from this content-model gap. We propose a unified approach to bridge this gap, with an information system component called Business-Knowledge Discovery Component. We discuss key challenges, architectural framework and the text analytic techniques that are involved.
Transliteration is the process of proper name translation based on pronunciation. It is an important process in many multilingual natural language tasks. A common and essential component of transliteration approaches is a verification mechanism that tests if the two names in different languages are translations of each other. Although many transliteration systems have verification as a component, verification as a stand-alone problem is relatively new. In this paper, we propose a simple, effective and robust training framework for the task of verification. We show the many applications of the verification techniques. Our proposed method can operate on both phonemic and orthographic inputs. Our best results show that a simple, straightforward orthographic representation is sufficient and no complex training method is needed. It is effective because it achieves remarkable accuracies. It is robust because it is language-independent. We show that on Chinese and Korean our technique achieves equal error rate well below 1% and around 1% for Japanese using 2009 and 2010 NEWS transliteration generation share task dataset. Our results also show that the orthographic system outperforms the phonemic system. This is especially encouraging because the orthographic inputs are easier to generate and secondly, one does not need to resort to more complex training algorithm to achieve excellent results. This approach is integrated for proper name based cross lingual information retrieval without translation. 1
Query-by-example information retrieval provides users a flexible but efficient way to accurately describe their information needs. The query exemplars are usually long and in the form of either a partial or even a full document. However, they may contain extraneous terms that would have potential negative impacts on the retrieval performance. In order to alleviate those negative impacts, we propose a novel term-based query reduction mechanism so as to improve the informativeness of verbose query exemplars. We also explore the notion of term discrimination power to select a salient subset of query terms automatically. Experiments on the TDT Chinese collection show that the proposed approach is indeed effective and promising.