In the era of Big Data Analytics, information dissemination, data integrity, and identifying unique records from large pool of data poses a big challenge for analysts in entity matching and linking scenarios. Data ingestion from multiple sources of same real-world entity exhibits several data quality issues like redundancy, incorrectness, variations, etc. Also, there are data input errors like typographical/spelling mistakes as well as missing fields. In order to achieve entity resolution, uniqueness and eradicate data redundancy and improve the data quality issues, deduplication is the solution. India being a multi-lingual and multi-cultural country with vast demographic variations, there is a need to develop India-centric model for handling deduplication on various Indian structured data held by various authorities. This research proposes a novel approach catering to India-centric demographic variations, region-specific naming conventions, address standardization using a highly customizable and scalable deep learning approach, by customizing DeepMatcher algorithm along with a synthetic data generation tool reckoning Indian variations of names and addresses in a region-specific manner.
The automobile and transport industries have always tried to reduce downtime by the way of preventive maintenance. In the past decade, electromechanical sensors have become more accurate along with novel innovations in IoT and machine learning, and automobiles have leveraged this. In this paper, an end-to-end open-source predictive maintenance solution is presented to predict the severity of faults in a car using onboard historical and real-time sensor data using IoT and machine learning. Sensor data is collected from a Suzuki Swift VXi Model, and classifiers like logistic regression, random forest, and gradient boosting trees are used to train the data with imputed faults. F1 score and AUC are used as evaluation metrics. An end-to-end onboard diagnostics (OBD) data to the user dashboard pipeline is proposed with final predicted faults visible on a real-time dashboard.
The widespread use of image-based memes on socioeconomic or political issues has witnessed a booming effect unparallel to any form of media in the recent years. The ability to go viral on social media in seconds and the popularity of memes on online platforms give a wide scope and pathway for research as it will help in understanding the usage patterns of the public and in turn be used for analyzing their sentiment toward a specific topic/event. In this paper, initially gap analysis on the features used for sentiment extraction on memes is presented. Exploring the correlation of image based and textual features, this paper gives a novel approach (correlating the facial features along with the text in the meme itself) for the extraction of sentiment from image-based memes. This paper also addresses the challenges faced in this relatively new area of sentiment extraction on memes. Finally, this paper concludes with insightful results.
In the age of big data analytics, text mining on large sets of digital textual data does not suffice all the analytic purposes. Prediction on unstructured texts and interlinking of the information in various domains like strategic, political, medical, financial, etc. is very pertinent to the users seeking analytics beyond retrieval. Along with analytics and pattern recognition from the textual data, there is a need to formulate this data and explore the possibility of predicting future event(s). Event prediction can best be defined as the domain of predicting the occurrence of an event from the textual data. This paper spans over two main sections of surveying technical literature and existing tools/technologies functioning in the domain of prediction and analytics based on unstructured text. The survey also highlights fundamental research gaps from the reviewed literature. A systematic comparison of different technical approaches has been listed in a tabular form. The gap analysis provides future scope of more optimized algorithms for textual event prediction.
Word Sense Disambiguation (WSD) is considered as one of the pivotal problems of Semantic classification among polysemous words that can be addressed using Natural Language Processing (NLP) for identifying the sense of the ambiguous word in a particular context. The application areas of WSD pertain to machine translation, information extraction and retrieval (IE-IR), dialogue systems, and automatic summarization kind of NLP solutions. This paper presents a survey on WSD approaches in major AI-NLP methods by comparing different approaches for WSD in supervised, unsupervised, and knowledge based algorithms. This paper also aims at providing gap analysis in surveyed WSD systems comparing strengths and weaknesses of various surveyed systems and their accuracy. Based on the findings, a future hybrid approach synergizing rule-based and machine learning based methods are contemplated. The findings of this survey are envisaged through an ongoing research on WSD based Meta-Search algorithm under C-DAC purview for an Intelligent NLP based system to detect the actual sense of search queries and providing semantic classification of news headlines and snippets containing ambiguous words.
In the domain of psychological practice, experts follow different methodologies for the diagnosis of psychological disorders and might change their line of treatment based on their observations from previous sessions. In such a scenario, a standardized clinical decision support system based on big data and machine learning techniques can immensely help professionals in the process of diagnosis as well as improve patient care. The technology proposed in this paper, attempts to understand psychological case studies by identifying the psychological disorder they represent along with the severity of that particular case, with the help of a Multinomial Naive Bayes model for disorder identification and a regular expression based severity processing algorithm. A knowledge base is created based on the knowledge of human experts of psychology. Psychological disorders however need not possess distinct symptoms to easily differentiate between them. Some are very closely connected with a variety of overlapping symptoms between them. Our work, in this paper, focuses on analyzing the performance of such psychological disorders represented in the form of case studies in a decision support system, with an aim of understanding this gray area of psychology.
Named Entity Recognition (NER) is a tool based on principles of Artificial Intelligence (AI) and Natural Language Processing (NLP) for automatically tagging Named Entities from unstructured text. Named Entities, which are generally proper nouns, can be the name of a person, organization, location etc. Named Entity Recognition is a very crucial tool of Natural Language Processing. Some of the application areas of NER include- Information Extraction and Retrieval, Machine Translation, Text Summarization etc. This paper analyses different approaches used in the NER of Indian languages, with emphasis on Hindi. The study compares the different approaches for NER viz. Machine Learning (ML), Rule-based and Hybrid. This study is orchestrated to provide gap analysis in available NER systems for Indian Languages, especially Hindi. It is because existing NER systems are evaluated for accuracies based on predefined datasets, which are not providing universal results on any dataset. Also, there is no standard dataset available for fair comparison of accuracies of Indian Languages. Furthermore, we concentrated our scope to Hindi as it is the official language of India and representative to other Indo-Aryan languages having similar structure, making it a generalized solution. Moreover, the scope of Named Entity tags are limited. Whereas, there is a scope of further classification of Named Entities. As future work, we would like to develop an efficient system, in terms of accuracy, and catering to more Named Entity tags than available presently in Hindi NER tools.
International Journal of Computer Sciences and Engineering (A UGC Approved and indexed with DOI, ICI and Approved, DPI Digital Library) is one of the leading and growing open access, peer-reviewed, monthly, and scientific research journal for scientists, engineers, research scholars, and academicians, which gains a foothold in Asia and opens to the world, aims to publish original, theoretical and practical advances in Computer Science,Information Technology, Engineering (Software, Mechanical, Civil, Electronics & Electrical), and all interdisciplinary streams of Computing Sciences. It intends to disseminate original, scientific, theoretical or applied research in the field of Computer Sciences and allied fields. It provides a platform for publishing results and research with a strong empirical component. It aims to bridge the significant gap between research and practice by promoting the publication of original, novel, industry-relevant research.
International Journal of Computer Sciences and Engineering (A UGC Approved and indexed with DOI, ICI and Approved, DPI Digital Library) is one of the leading and growing open access, peer-reviewed, monthly, and scientific research journal for scientists, engineers, research scholars, and academicians, which gains a foothold in Asia and opens to the world, aims to publish original, theoretical and practical advances in Computer Science,Information Technology, Engineering (Software, Mechanical, Civil, Electronics & Electrical), and all interdisciplinary streams of Computing Sciences. It intends to disseminate original, scientific, theoretical or applied research in the field of Computer Sciences and allied fields. It provides a platform for publishing results and research with a strong empirical component. It aims to bridge the significant gap between research and practice by promoting the publication of original, novel, industry-relevant research.
In the current scenario, there is no computing system in practice which understands or provides analytical insights for the language used in the psychological documents maintained by psychiatrists. Most importantly, there is no standardized language or methodology used in psychiatric practice to diagnose a psychological disorder. The main aim of our system is to provide a standardized platform for an assistance in diagnosis using Natural Language Processing techniques and Sentiment Inference. The system would suggest to the psychiatrist, the psychological disorder identified and its severity from the document text. The severity is obtained via a gradient score, which is computed by the semantic based gradient processing module - the core of our research. The psychiatrist can approve/correct the system diagnosis if necessary and reassign the gradient values for phrases contributing towards disorder severity. In an automated process, model gets re-trained after a particular threshold is reached, continually improving its performance.
In today's digital age, in the dawning era of big data analytics, it is not the information but the linking of information through entities and actions, which defines the discourse. Any textual data either available on the Internet or off-line (like newspaper data, Wikipedia dump, etc) is basically connected information which cannot be treated isolated for its wholesome semantics. There is a need for an automated retrieval process with proper information extraction to structure the data for relevant and fast text analytics. The first big challenge is the conversion of unstructured textual data to structured data. Unlike other databases, graph databases handle relationships and connections very elegantly. Our project aims at developing a graph based information extraction and retrieval system.
With recent developments in Deep Learning, neural network methods have proved to be useful for QA (Question Answering) systems. Researchers have begun to tackle open-domain QA, in which the model is given a question and access to a large, diverse corpus (e.g. Wikipedia). Open-domain QA is complex, as it requires large-scale search for relevant passages by an Information Retrieval component, combined with a Reading Comprehension component that reads the passages to generate an answer to the question. In our paper, we propose a novel open-domain system with two components viz., Paragraph Retriever and Paragraph Reader. Traditionally, retrieval systems only retrieve the most relevant documents to the given question. Our Retriever component retrieves the most relevant paragraphs with respect to the question. In the Reader component, we propose a modification to the Dynamic Coattention Network model, by augmenting it with other useful linguistic features like Parts of Speech of words, Named Entities, Exact Match between question and paragraph, Wh question word (what, where, when, etc.) identification, for extracting the answer from a paragraph.
Natural Language Generation (NLG) is a challenging problem in the field of Artificial Intelligence. The difficulty stems from the natural language’s flexibility to convey the same message in different ways. Psycholinguists have always believed that learning simpler sentences early can lead to complex sentence creation using the same knowledge.This is also the intuition behind curriculum learning. Thus, in this paper, we investigate the use of curriculum learning for natural language generation. We show that curriculum learning is a promising training methodology for deep learning systems for NLG. We show this by reporting improvements obtained using i) a particular curriculum strategy and ii) augmenting data using curriculum logic. We use TGen, a deep learning based NLG system, for experimentation. We use 5 metrics for NLG evaluation and 8 metrics for readability evaluation. Our quantitative and qualitative evaluation shows that the system trained using curriculum methodology produces better quality text as compared to the sys-tem trained on normal data.
Natural Language text is not bound by a fixed structure. For a machine to understand the language, the challenge lies in resolving the ambiguities and capturing innovativeness. Due to its unstructured nature, discourse linking, required for understanding and generating text by a machine is a challenging task. Also, dealing with sentences varied in nature and changing them into generic structure is an additional challenge. This paper presents a way to create a discourse linking tool for English language. This tool is based on specially created lexicon and hand-crafted rules suited for discourse linking purpose. Currently available lexicons such as dictionaries, WordNet, etc. are not suited as they lack domain knowledge. Therefore, an ontology is developed for political news domain which serves as a lexicon and its inherent relations are used for discourse linking purpose.