Context retrieval and ranking have always been an area of interest for researchers around the world. The ranking provides significance to the data that has to be presented in front of users but it also consumes time if the ranking architecture is not organized. The retrieval is dependent upon the co-relation among the data attributes that are supplied against a class label also referred to as ground truth and the ranking depends upon the sensing polarity that indicates the hold of the outcome towards asked information. This paper illustrates an ontological architecture that involves two phases namely context retrieval and ranking. The ranking phase is composed of three different algorithm architectures namely k-means, Support Vector Machines (SVM), and Deep Neural Networks (DNN). The DNN is tuned to fit and work as per the availability of a total number of samples. The proposed work has been evaluated for both quantitative and qualitative parameters in different sets and scenarios. The proposed work has also been compared with other state of art techniques and is illustrated in the paper itself.
Social media platforms, namely Instagram, Facebook, Twitter, YouTube, etc. have gained a lot of attention as users used to share their views, and post videos, audio, and pictures for social networking. In near future, understanding the meaning and analyzing this enormously rising volume and size of online data will become a necessity in order to extract valuable information from them. In a similar context, the paper proposes an analysis model in two phases namely the training and the sentiment classification using the reward-based grasshopper optimization algorithm. The training architecture and context analysis of the tweet are presented for the sentiment analysis along with the ground truth processing of emotions. The proposed algorithm is divided into two phases namely the exploitation and the exploration part and creates a reward mechanism that utilizes both phases. The proposed algorithm also uses cosine similarity, dice coefficient, and euclidean distance as the input set and further processes using the grasshopper algorithm. Finally, it presents a combination of swarm intelligence and machine learning for attribute selection in which the reward mechanism is further validated using machine learning techniques. The comparative performance in terms of precision, recall, and F-measure has been measured for the proposed model in comparison to existing swarm-based sentiment analysis works. Overall, simulation analysis showed that the proposed work based on grasshopper optimization outperformed the existing approaches for Sentiment 140 by 5.93% to 10.05%, SemEval 2013 by 6.15% to 12.61%, and COVID-19 tweets by 2.72% to 9.13%. Thus, demonstrating the efficiency of the context-aware sentiment analysis using the grasshopper optimization approach.
Analyzing the multiple relevant documents returned in reply to an end-user request by an information retrieval system is challenging. It is very time-consuming and less efficient to find analogous web pages without applying the clustering. Clustering of web pages arranges a large number of web documents into relevant small clustered groups. In this paper, a novel similitude degree computation technique is proposed to provide the web documents related to the context in which multiple related web documents are the members of the same cluster. The clustering module results in web documents’ arrangement with their associated topic and corresponding computed similitude or similarity score. This provides the user clusters containing equivalent web documents related to the issue of desire. This context-based grouping of web documents reduces the time taken for searching relevant data and improves the results in response to a user request. Moreover, the comparison and analysis of the proposed technique are done with different existing similarity measures on the basis of performance metrics purity and entropy. It has shown the proposed scheme provides better results to the user.
Computers and the Internet have changed how we live our lives in ways that are often difficult to grasp, therefore it is extremely helpful to utilise computers and the Internet for the purposes of making computers and the Internet easier to use. An ontology-based search engine will assist in more accurate searches by utilising artificial intelligence to improve natural language processing and analysing keywords to more precisely target data (NLP). A general evaluation of some of the existing search engine design methodologies was introduced in this document, including with some information about the difficult areas of development. Text analysis with artificial intelligence NLP techniques could be used to build an engine that generates intelligent responses based on user requests. This research is all about intelligent ontology-based semantic search engines and evaluating their research in terms of creating methodologies. Furthermore, elements of artificial intelligence technology that are increasingly difficult to deal with, including such deep learning and machine learning, are investigated to influence development of search engines. These publications demonstrate considerable progress in the decade-long drive towards AI-infused ontology-driven search engine via use of ontologies.
This paper aims to select the appropriate node(s) to effectively destabilize the terrorist network in order to reduce the terrorist group’s effectiveness. Considerations are introduced in this literature as fuzzy soft sets. Using the weighted average combination rule and the D–S theory of evidence, we created an algorithm to determine which node(s) should be isolated from the network in order to destabilize the terrorist network. The paper may also prove that if its power and foot soldiers simultaneously decrease, terrorist groups will collapse. This paper also proposes using entropy-based centrality, vote rank centrality, and resilience centrality to neutralize the network effectively. The terrorist network considered for this study is a network of the 26/11 Mumbai attack created by Sarita Azad.
Globally, the rising level of air pollution is becoming a cause of serious concern. The situation has reached such an alarming situation that knowing the extent of air pollution well in advance has become an absolute necessity before we step out of our home. An advance prediction can help the urban travellers to know the possibility of enhanced pollution ahead of time at strategic locations of a city and thereby be useful in planning a less polluted route. Different cities of the world have pollutants levels with varied dispersion pattern. As a result, a generic prediction model is needed that can cater to all types of pollutants and which can offer better forecasts of pollution data irrespective of their dispersion levels. In this paper, the authors modelled recurrent neural network (RNN)-based bidirectional long short-term memory (Bi-LSTM) that forecasts pollutants concentration with least RMSE. The model is trained and tested on the pollution data of Delhi, the most polluted capital city in the world for the second consecutive year in 2019. Air pollutants like PM 10 , PM 2.5 , NO 2 , O 3, and CO are considered that have a varied dispersion pattern. To test the efficacy of our predictions, the model is tested on real data obtained from the Central Pollution Control Board (CPCB) of India. The performance metrics are also generated to evaluate the performance of the proposed model.
People have realized the importance of finding and archiving information with the computer advents for thousands of years, and storing of large amount of information became possible. It is actually not related to the fetching of the documents, it informs the user on the whereabouts and existence of the documents. In this paper, hybrid model has been used in which the document is classified using the support vector machine (SVM) classifier, and after the condition is applied, if it is satisfied, the extraction of the matched paragraph and the sentence is responsible for the generation of relevant answer. The knowledge base gets updated if condition does not match, and new updated answer will be generated. Finally, the best answer is displayed after ranking by using the PSO optimization. Word2vector is applied for feature extraction. In this paper, comparison of RankSVM, RankPSO and RankHSVM + PSO for the implementation of IR ranking is considered. Here, first SVM is used as a classifier for dividing most relevant and non-relevant results, and afterward PSO is used for the optimization of the result means extraction of the best answer or document. Selection of appropriate parameters is difficult in case of simple SVM, but for the ranking of the answers it gives potential solutions. PSO is used for optimization which has global search capability and is easy to implement and thus to optimize the ranking of document retrieval. We propose the RankHSVM + PSO model to find the fitness function. This technique improves the performance of the system as comparative to other techniques. The result shows that the algorithm applied here improves the value of performance evaluation by 4–5%. TREC 2004 QA DATA dataset is used which contains my datasets. It has a question answering track since 1999. The task was defined in each track. Retrieval of true equivalent test collection for standard retrieval is an open problem. In a retrieval test collection, the unit that is judged the document has a unique identifier.
Now a day’s many research works are going on in the field of Information Retrieval Ranking. Retrieval and Ranking of information from the huge database of internet world is become a most useful and interesting task. Machine learning plays the major role now days. In this paper we have worked for monolingual and cross lingual information retrieval ranking in which the retrieval of document within the same language and retrieval of document in different language has been done. We take English language for monolingual IR ranking and for cross lingual we take Hindi queries and get documents in English. TREC 2008 QA Dataset and FIRE 2011 ADHOC Dataset has been taken in this work respectively. Finally the performance Evaluation has been done for SVM, PSO and SVM+PSO and the results has been compared. The results clearly explain that SVM+PSO model gives best result in compare to other two.
Translation ranking is inherently of great significance for machine translation (MT), as it allows the comparison of performances of multiple MT systems as well as for its efficient training. This article demonstrates a mechanism that is used for ranking the translation outputs generated by the MT systems from best to worst. To implement this approach, the system exploits a supervised learning algorithm trained over existing manual ranking by using various features obtained after the linguistic analysis of both source and target side sentences without relying on the reference translation.
This article having a significant analysis of the Mumbai terrorist attack held on November 26, 2011, using a decision based on multi-criteria which are a potential approach for the social network analysis. The data source for this analysis is a dossier report on 26/11 Mumbai attack submitted to the Ministry of External Affairs that was published in the year 2009 and many more articles. This report gives the complete details about this tragic event consisting number of terrorist involved in India as well as from Pakistan, points where they had done operation, their communications in between, number of casualties, etc. When law enforcement agencies want to analyze any terrorist attack concerning key players involved in the attack, investigators should consider more than one criterion or factors that may also inconsistent and contradictory. Therefore, key player selection and ranking of terrorist nodes are a multi-criteria decision-making issue. AHP resolves this multi-criteria decision-making issue. This study reveals the key players and ranked them accordingly, involved in 26/11 Mumbai terrorist attack using the analytic hierarchy process (AHP).
With the web getting greater and acclimatizing information about various ideas and areas, it is getting extremely hard for straightforward information base driven applications to catch the information for a space. Accordingly engineers have come out with ontology based frameworks which can store enormous measure of data and can apply thinking and produce ideal data. Subsequently encouraging compelling information the board. Despite the fact that this methodology has made our carries on with simpler, and yet has offered ascend to another issue. Two unique ontologies acclimatizing same information will in general utilize various terms for similar ideas. This makes disarray among information architects and laborers, as they don't realize which is a superior term then the other, and also the language issue, human and machine interaction is in different language. So we need to merge the two cross lingual ontology by using cross lingual ontology matching. This paper shows the development of cross lingual ontology matcher at two levels:1) at String level and 2) Semantic level. We have utilized a Graph Matching Technique which works at the center of the framework. We have additionally assessed the framework and have tried its presentation with its antecedent which works just on string coordinating. Hence current methodology delivers better outcomes. General Terms: Ontology Matching, Ontology Alignment.
In recent years the Terrorist network research community has collected huge information for the appearance of composite and disparate communication patterns in complex terrorist social network. The identification of most centralized node(s) is crucial to understand the information flow and their communicable spread. Strategies for identifying influential nodes in complex terrorist network have been of interest. Techniques have been proposed from different perspectives, each with its own particular favourable circumstances and deficiency. Current study based on Formal Concept Analysis with Weight, under Terrorist Network Mining to identify the node(s) in a Terrorist Network, who influenced other members of the terrorist group. It may help to destabilize the terrorist network more effectively by removing or destroying the all communication link of higher ranking node(s). First we construct the adjacency matrix of 26/11 Mumbai Terrorist Attack network as the formal context. Next, we calculate the all possible concepts of the formal context. Subsequently the weight of every node within the terrorist network calculated and ranked each node accordingly. A comparison has been made of result with other well-known centrality algorithms like closeness centrality, node betweenness centrality, flow betweeness, PageRank, Katz, Reach centrality and PN centrality
The quality of air that we breathe is one of the more serious environmental challenges that the government faces all around the world. It is a matter of concern for almost all developed and developing countries. The National Air Quality Index (NAQI) in India was first initiated and unveiled by the central government under the Swachh Bharat Abhiyan (Clean India Campaign). It was launched to spread cleanliness, and awareness to work towards a clean and healthy environment among all citizens living in India. This index is computed based on values obtained by monitoring eight types of pollutants that are known to commonly permeate around our immediate environment. These are particulate matter PM10; particulate matter PM2.5; nitrogen dioxide; sulfur dioxide; carbon monoxide; lead; ammonia; and ozone. Studies conducted have shown that almost 90% of particulate matters are produced from vehicular emissions, dust, debris on roads, and industries and from construction sites spanning across rural, semi-urban, and urban areas. While the State and Central governments have devised and implemented several schemes to keep air pollution levels under control, these alone have proved inadequate in cases such as the Delhi region of India. Internet of Things (IoT) offers a range of options that do extends into the domain of environmental management. Using an online monitoring system based on IoT technologies, users can stay informed on fluctuating levels of air pollution. In this paper, the design of a low-price pollution measurement kit working around a dust sensor, capable of transmitting data to a cloud service through a Wi-Fi module, is described. A system overview of urban route planning is also proposed. The proposed model can make users aware of pollutant concentrations at any point of time and can also act as useful input towards the design of the least polluted path prediction app. Hence, the proposed model can help travelers to plan a less polluted route in urban areas.
Analysis of the terrorist network is a process to analyze or deriving Useful information from the available network data. Ranking The Terrorist nodes within a terrorist network in identifying the most influential node is essential for the elaboration of Covert network mining. The purpose of this paper is to implement an approach of two dimensional criteria weight determination along with logarithmic concept implementation for vital node investigation in term of their influential ability. Betweenness, Closeness, Eigenvector, Hub, In-degree, Inverse closeness, Out-degree and Total degree considered as criteria and terrorist involved in 9/11 terrorist attack considered as alternatives used to formulate a decision problem. Although an integrated approach of Fuzzy based subjective-Aggregation concept based objective criteria weight determination and Ranking alternatives by the implementation of the logarithmic concept is used to solve multi criteria decision problems in order to show the application of most centralized node identification process which can be obtained easily by classification and selection problem solution using multiple criteria and alternatives.
Storing of information that is large in amount became possible after the realization of the importance of finding and archiving with the computer by the people. In this paper we are using query in Hindi and using corpus dictionary the query is translated into English and related to the extracted keywords the relevant document or answer is classified using SVM and answer will be generated using PSO in English. SVM is used as a classifier which classifies the answers or documents .If condition doesn’t match the knowledge base gets updated and new updated answer will be generated .Finally using the PSO optimization the best answer is displayed after ranking. For feature extraction Word2Vector is applied. We have used PSO, SVM and SVM-PSO for the Implementation of CLIR and compare the results between them. The performance of the system thus improves 4 to 5 percent by SVM-PSO. FIRE 2011 ADHOC dataset is used for the performance evaluation of CLIR. Keywords : CLIR, Ranking, PSO, SVM, Machine Learning
This submission describes the study of linguistically motivated features to estimate the translated sentence quality at sentence level on English-Hindi language pair. Several classification algorithms are employed to build the Quality Estimation (QE) models using the extracted features. We used source language text and the MT output to extract these features. Experiments show that our proposed approach is robust and producing competitive results for the DT based QE model on neural machine translation system.
In Natural Language Parsing, in order to perform sequential labeling and segmenting tasks, a probabilistic framework named Conditional Random Field (CRF) have an advantage over Hidden Markov Models (HMMs) and Maximum Entropy Markov Models (MEMMs). This research work is an attempt to develop an efficient model for shallow parsing which is based on CRF. For training the model, around 1,000 handcrafted chunked sentences of Hindi language were used. The developed model is tested on 864 sentences and evaluation is done by comparing the results with gold data. The accuracy is measured by precision, recall, and F-measure and is found to be 98.04, 98.04, and 98.04, respectively.
Rapid urbanization combined with an almost insatiable need for energy has spawned various forms of pollution. Researchers have found air pollution to be at the top of list of factors that cause the most fatalities among urban dwellers today. Scoping urban areas that harbor the most air pollutants and contaminants can help an urban dweller identify comparatively less polluted routes. However, processing of information related to air pollutants is time intensive. As such, temporal forecasting takes preeminence in designing a system that can provide information well in advance of concentration levels of air pollutants at any given time in day or night. In this paper, the authors approach problems related timely forecasts for predicting and tracing air pollution levels across major thoroughfares in urban environments, using fog computing and Internet of Things (IoT). The objective of the research and proposed method is to offer a time-sensitive forecasting to enable citizens to adopt a more agile route-planning approach at any given point of time. In the wake of rising deaths owing to air-borne pollutants and chemicals, results of the research conducted indicate an object-oriented approach toward building a smarter city.
Translation is the technique in which system translate text from source natural language to target natural language, so that the original message is retained in target language. Deep Neural Networks are capable models that achieved malicious achievement on challenging learning tasks such as visual object recognition and speech recognition and work well whenever large amount of training sets are available. This paper represent Hindi to English machine translation at Hindi-English parallel corpus in which supervised learning algorithm applied with attention model and in which one Recurrent Neural Network map the input sequence to a vector in fixed dimensionality, and another Recurrent Neural Network decode the target sequence from the vector and show how neural machine translation is better way to translate the data from source language to target language.
The paper represents the advanced NLP learning resources in context of Indian languages: Hindi and Urdu. The research is based on domain-specific platforms which covers health, tourism, and agriculture corpora with 60 k sentences. With these corpora, some NLP-based learning resources such as stemmer, lemmatizer, POS tagger, and MWE identifier have been developed. All of these resources are connected in sequential form, and they are beneficial in information retrieval, language translation, handling word sense disambiguation, and many other useful applications. Stemming is first and foremost process of root extraction from given input word, but sometimes it does not produce valid root word. So the problem of stemming has been resolved by developing Lemmatizer, which produces the exact root by adding some rules in stemmed output. Eventually, statistical POS tagger has been designed with the help of Indian Government (TDIL) tagset (Indian Govt. Tagset, [1]). With this POS-tagged file, MWE identifier was developed. However, for developing MWE identifier, some rules are created for MWE tagset and then MWE-tagged file has been developed which in turn produces the automatic extraction of the MWEs from tagged corpora using CRF $${+}{+}$$ tool. Moreover, evaluation of learning resources has been performed to calculate the accuracy, and as a result, the output of corresponding proposed resources such as stemmer, lemmatizer, POS tagger, and MWE identifier are 77.0, 86.8, 73.20, and 43.50% for Hindi and 74.0, 85.4, 84.97 and 47.2% for Urdu, respectively.