Many Natural Language Processing (NLP) applications involve Named Entity Recognition (NER) as an important task, where it leads to improve the overall performance of NLP applications. In this paper the Deep learning techniques are used to perform NER task on Hindi text data as it found that as compared to English NER, Hindi language NER is not sufficiently done. This is a barrier for resource-scarce languages as many resources are not readily available. Many researchers use various techniques such as rule based, machine learning based and hybrid approaches to solve this problem. Deep learning based algorithms are being developed in large scale as an innovative approach now a days for the advanced NER models which will give the best results out of it. In this paper we devise a Novel architecture based on residual network architecture for preferably Bidirectional Long Short Term Memory (BiLSTM) with fasttext word embedding layers. For this purpose we use pre-trained word embedding to represent the words in the corpus where the NER tags of the words are defined as the used annotated corpora. BiLSTM Development of an NER system for Indian languages is a comparatively difficult task. In this paper, we have done the various experiments to compare the results of NER with normal embedding and fasttext embedding layers to analyse the performance of word embedding with different batch sizes to train the deep learning models. Here we present a state-of-the-art results with said approach F1 Score measures.
Background: Big data describe volume amount of structured and unstructured data. Big data sources like social media, Facebook collected large number of unstructured data. When unstructured data collected from different sources, maintain the quality of data is also important. Big data sources provide insights to businesses for improve business decision making. The proposed system improves the quality of business data which are collected for business decision making. Data quality Assessment is a way for practitioners to understand the scope of how poor data quality effects on business process and develop a business case for data quality management. Methods: This paper contributes to providing a solution by introducing new assessment model to evaluate and manage the quality of social media data. Sentiment analysis is used for monitoring real-time data. Generate the new rules and attributes to assess the quality of data. Apply quality attributes on input data and assess only those data which are fit into the quality attribute dimensions. Evaluate data quality using large data set. Result: The system provides the visualized data and generates a report based on sentiment analysis. Conclusion: The proposed system improve solution by provide real time and validate data for the user.
File is a super abstraction of hypothetically huge volume of information stored in containers (disk blocks) which is persists for very long time (memory abstraction). File is linear Array of bites and blocks of bytes accessed by multiple clients which is controlled by File system. File System function is disk seek, read maximum information from container and maps file names and offsets to disk blocks for better read and write operations. File system is logic to device how Information is stockpiled and how effectively to retrieve it. File system manages references (pointers) to memory block and assist to seek information in feasible manner. Without which data would be placed in large heaps of memory at location unknown to operating system to manage resources effectively. File System is structure and sense rules accustomed in managing clusters of data is termed as file system. Current age is age of High performance information processing and intelligent knowledge generation which urges for better Scalable File System. Current File System come with Numerous limitations as architecture design pattern limitations, I/O Operational, component failure, reliability issue, data decomposition, fault Tolerance etc. This manuscript give a first Survey first step towards synchronized research.This is our First research article on Scalable File System which facilities view map of new Scalable File System. Beta Survey Methodology is been incorporated with pattern of Abstract methodology and scope Writing for every research article related to Scalable file system Survey. The Key point answer of Conclusion is Luster a scalable file system which has been used as core in development of world top 100 supercomputers which fulfills key space to be scalable file SystemFuture research work would be implementation of Luster File System and Evaluation on parameters of global name space IOPS and Luster File Size, No of clients and OST's used with Throughput MBPS and read/write pattern evaluation.
Information Retrieval Systems and Search engines lack capability to Map Human perception, as words have limited expression power come up with ambiguity in different contexts and concepts. A picture or image is bigger broader and best way to express thing. An image is concept that represents information urge in more relevant and desired answer. Even though an image would represent a set of Thousands of keywords and phrases it give rise to image ambiguity just Word Sense Disamguity (WSD). It's very challenging to map user keyword query to retrieve image as answer, as relevance depends on user perception and intent. Web-Image search engines work on principle on keyword as queries and likewise work on surrounding information like tags annotation to find Images depicting user perception. Search engines development comes with fist challenges to map correctly keywords in relevant classes of Image. Visual attributes most of time cannot co-relate with image class signature which interprets conceptual meaning of user keyword or phrase search. Relevance Feedback Research Technique incorporates Image re-ranking as proficient Approach to enhance results of web image search. This principle methodology is been implemented by most popular and commercial search engines Google and Bing. Asking user feedback as implicit is best feedback mechanism In corporation of user feedback (i.e. one click feedback) to search results with re-ranking and mapping search results in accordance has proved best method for improved search in case of text based and image based retrieval which has been incorporated by www search engines for image(Image Re-ranking). Input a Keyword based Query group of image are retrieved by search engine. Taking in one click from client images are re-arranged and ranked by mapping visual similarity of similar images to clicked image. But a major problem Resemblances of visual parameters do not fine relate with images Semantic sense that construe client's' image search goal. On supplementary side, learning an entire visual semantic space to distinguish vastly dissimilar pictures from www is problematic and inefficient. This research propose an inventive image re-ranking design, which inevitably offline acquires dissimilar visual semantic spaces for diverse keyword based queries through keyword enlargements (expansion). Visual structures of pictures are projected into their associated visual semantic area to acquire sense (semantic) signatures. At online phase, pictures are re-ranked by matching their semantic signs acquired from visual semantic area specified by keyword query. This newfangled methodology significantly increases both accurateness and efficiency of image re-ranking. The unique visual features of 1000's of aspects are been projected to semantic signs as tiny as 25 extents. Investigational outcomes display that maximum 40% comparative progress has been attained on re-ranking precisions equated with state of art methodologies. Automated indexing and text alignment with similar image clustering adds improved technique to IIR (image information retrieval). The research further implements incremental learning framework. Semi-supervised methodology is been implements which always stood better than supervised and unsupervised methodology. Furthermore audio and video or crowd motion datasets re-ranking adds to novelty of research. The multimedia text-image corpus generation facilitates additional contribution of research area.
Decision making involves comparing solution with each other in Decision support but it also necessary recommend which objects are comparable and in what way. This is challenging Question in data and knowledge processing, which urges for better pattern mining. Today's web is web of document where we find reviews,complaints, feedbacks posted on blogs, e-commerce websites and social networks which are rich source of knowledge for pattern mining. Research presents analyzing comparable question and then extraction of information for two objects as comparable and if not recommendation on objects comparable, if comparable answers. We propose a Decision support Engine that Answers queries asked by finding comparable objects if not comparable identifying user search intent find comparable object and answer with key values comparing them.Current state of art research system check if objects are comparable if not they don't provide recommendation to user for comparable objects. Supervised methods are limited to set of input and expected output, whereas unsupervised system output at times is false positive.In order to overcome this limitation semi-supervised methodology is used to develop algorithm for mining. This article is outcome of Methodical summary of literature on current research scope in data mining and NLP. Precisely it is analysis on abstract methodology and research scope on 24 appropriate manuscripts retrieved as per our research domain. The search contributes to field of Information retrieval and web search by solving five Research Question and major issues and challenges with procedures.Types of patterns that have been extracted in previous approaches with new learned method to extract complex patterns. In General review consequences demonstrate as many scholars have worked on pattern mining and decision support system there is need of precision and accuracy in pattern mining. Evaluation of research needs to be tested with various parameters this research evaluates decision engine with Mean average precision (MAP) and feedback rating of user to answers produced by decision engine, with regular evaluation of precision and recall.
Information Present in Different language and Structure gives Rise to language as barrier in information retrieval. Informative Document on queen Elizabeth is been writing by foreign language English which makes its difficult for a Marathi reader to understand and seek History of England, on Similar lines Literature Work on Shivaji is mostly documented in Marathi which makes foreign Historians difficult to gain know, in both case user is at times unknown of facts due to language and may lose interest on information. Vital information on current happing on village and taluka level are been published in newspaper with local language which setback information spread among other Masses. Many Time government documents and forms are been presented in English Language where a lay man from Marathi language background finds difficulty to understand information and even avoid such procedures therefore it highly urges for need of automated Software based Translation system which would assist in cross Domain information Retrieval. Machine Translation assist to translate Information presented in one language to other language. Information can be present in form of text, speech and image translating this information helps for sharing of information and ultimately information gain. A lot of work has been done on Translation of English to Hindi, Tamil Bangla and other foreign languages also.Machine Translation is challenging Research Area with numerous issues due to language ambiguity like grammar, Structure and even fluency of use. Numerous Methodologies have been proposed and developed which have uplifts and downfalls also, Although statistical and rule based at core with each having limitations. rule based produce accurate mapped translation and are trainable system but costly, whereas statistical produce fluent translation but lack accuracy and sense. Hybrid is combine approach which integrated approach and helps to optimize translation output.The research manuscript we present hybrid machine Translator for English to Marathi language which translated Web pages, text Documents on Agriculture (crops fruits for farmer), Medical reports in Marathi and tourism related information. Proposed System consists of Parallel Multi-Engines which process statistical and rule based Translation for same input document and produce a optimized result by performing statistical over rule based which give fluent language sense outputs. Mapper algorithm is been used in rule based Translation, with Agriculture corpus, medical and tourism corpus for statistical evaluation. Marathi wordnet has been implemented to enhance dictionary and incorporate better translation resultCurrently System has been proposed for text Document which can be extended to speech and voice. Comparative analysis for point view in one dimension of only limited set of Queries is done with Google Translator. Holding hybrid approach as better methodology.A Systematic survey of only 10 key articles used in research has been done. This research article is extension of our previous research surveys and partial implementations. And innovative Smeasure has been new parameter proposed and evaluated by our research team.
In the era of web new information is upcoming day by day. Researches add their work for their research domains. Detecting of originality of research work is in hype. In Academic sector students researchers bring in innovative ideas, algorithms stating that their work outperforms prior research. They may implement NULL Hypothesis or alternative Hypothesis, detecting their effort is a challenge. By means of plagiarism detectors such academic efforts can be evaluated or graded. This reflects the essence of research in the field of Plagiarized content detection and grading. Some of our research issue highlights to technical scenario to design an algorithm which is adaptable to changing nature of dataset. The dataset grows, as new research work is added in due course of time. Data extraction from unstructured information is challenging, as no standard pattern is yet defined. Such patterns vary from research to research and are domain specific. A document in question i.e plagiarized or not? Is a join of one or more sentences that originate by the authors research or referenced from previous publications. Authors to prove originality use paraphrasing which may have semantic similarity, also some of the contents act as metaphor for upcoming research work. It is complex task point out such an activity.Methodology states that a document in question is a join of sentences, whereas each sentence is a join of terms. Thus we conclude by fork and join operations; plagiarism detection is possible in effective way. Document in question is split to produce a sentence vector. A term vector is generated by forking sentence to terms for each sentence in sentence vector. Mapper is implemented that maps term to sentence and sentence to source document. To enhance the accuracy of the model a Multi Agent Based System MAS frame is recommended to adapt varying similarity functions. Achieve parallelism in system and adaptability of new similarity measures as well remove one which are not suitable any more to the task.
This paper discusses and revolves around providing an agent-based text-mining document search-engine architecture and implementing two algorithms one of which will be used for document-weighting and other which can be used for document-ranking. The main challenge lies in defining the factor on which the algorithms work and hence the „weight of the document‟ has been defined as the deciding factor. While browsing the internet, people use different Search Engines such as Google, Yahoo, and Bing etc. In India Internet access is not easily available everywhere especially to the person who knows a little about computers. But there will be people who are especially seeking proper and genuine information related to their queries in a particular domain such as cancer diseases in medical science, data mining techniques in computer science, etc. Hence effort is made here is to develop a web-browser based document search-engine which incorporates novel algorithms for document-weighting and document-ranking and hence provide a good way to find out relevant documents with ease. With incorporating of Quality assurance activities in the software development phases ensures all required issues are addressed and implemented to achieve quality product development.
Software quality cannot be improved simply by following industry standards which require adaptive/upgrading of standards or models very frequently. Quality Assurance (QA) at the design phase, based on typical design artifacts, reduces the efforts to fix the vulnerabilities which affect the cost of product. For this different design metrics are available, based on its result design artifacts can be modified. But to modify or make changes in artifacts is not an easy task because these artifacts are designed by rigorous study of requirements. The purpose of this research work is to automatically find out software artifacts for the system from natural language requirement specification as forward engineering and from source code as reengineering, to generate formal models specification in exportable form that can be used by UML compliment tool to visually represent the model of system. This research work also assess these design models artifacts for quality assurance and suggest alternate designs options based on primary constraints given in requirement specification. Following problems are resolved in this research work 1. Automatic generation of design phase class model from natural language input 2. Automatic generation of design phase class model from already developed source code 3. Generation of secure validated deign from above generated class models with different level of security as high low and medium with the help of different software metrics To resolve these problem there is need of automated environment which will assess generated design artifacts from natural language as forward engineering and from source code as reengineering and finally suggest and validates alternate designs options for better quality assurance.
Software quality cannot be improved simply by following industry standards which require adaptive/upgrading of standards or models very frequently. Quality Assurance (QA) at the design phase, based on typical design artifacts, reduces the efforts to fix the vulnerabilities which affect the cost of product. Different design metrics are available, based on their results design artifacts can be modified. Modifying or making changes in artifacts is not an easy task as these artifacts are designed by rigorous study of requirements. The purpose of this research work is to automatically find out software artifacts for the system from natural language requirement specification as forward engineering and from source code as reengineering, to generate formal models specification in exportable form that can be used by UML compliment tool to visually represent the model of system. This research work also assess these design models artifacts for quality assurance and suggest alternate designs options based on primary constraints given in requirement specification. To analyze, extract and transform the hidden facts in natural language to some formal model has many challenges and obstacles. To overcome some of these obstacles in software analysis there should be some mean or a technique which aims to generate software artifacts to build the formal models such as UML class diagrams. Initially, the proposed technique converts the NL business requirements into a formal intermediate representation to increase the accuracy of the generated artifacts and models. Next, it focuses on identifying the various software artifacts to generate the analysis phase models. Finally it provides output in the format understood by model visualizing tool. The re-engineering process to find out design level artifacts and model information about the previous version of software system from available source code with easy layout is a very difficult task. Performing this task manually has many problems as the ability of human brains to deal with the complexity and security of large software systems is limited. To overcome this difficulty there is need of automated environment which will assess generated design artifacts from natural language as forward engineering and from source code as reengineering and finally suggest and validates alternate designs options for better quality assurance.
Wireless Sensor Networks (WSNs) consists of numerous small sensors. These sensors are wirelessly connected to each other for perfor m- ing same task collectively such as monitoring weather conditions or specifically parameters like temperature, pressure, sound and vibrations etc. For all applications partial or full time synchronization is required and the message exchanged by sensor nodes for data fusion must be time stamped by each sensor's local clock. This helps to achieve a common notion of time in wireless sensor networks. This paper contains a survey, relative study and analy- sis of existing time synchronization protocols for wireless sensor networks, based on various parameters. No single protocol is optimal and sufficient in all aspects for designing a clock synchronization system. So the comparative study and design considerations will help a lot to the designer for designing a scheme which may or may not be application specific.
Wireless sensor networks were initially deployed for military applications. Gradually researchers found them to be very useful in applications like weather monitoring, target tracking, agriculture, industrial applications, and recently smart homes and kindergartens. All the WSN applications need partial or full time synchronization. Applications like acoustic ranging, target tracking or monitoring need a common notion of time. Every data is time stamped sensor nodes local clock. Two main approaches to time synchronization are receiver-receiver synchronization and sender-receiver synchronization. In this paper we analyze the receiver-receiver synchronization and discuss the results of simulation in network simulator. This study, design considerations and simulation methodology will help a lot to the designer for designing a time synchronization scheme or system.
Extensive research efforts are going on to contribute to the Internet of Things (IoT) application development. IoT will create network of “Things” capable of communicating and sharing information with one another. The main goal of IoT is to make the physical environment more intelligent. IoT plays an important role in smart cities and smart homes. The goals of this research paper are six-fold: (i) serve as a guideline for researchers who are new to the Internet of Things (IoT) and want to contribute to this research area, (ii) analyze problems and challenges identified in the implementation of middleware for IoT, (iii) provides a brief overview of the sensor network in Internet of Things (IoT) for building smart cities, (iv) depicts challenges on technologies and applications from India’s perspective, (v) proposes a general IoT architecture to meet the architecture challenge, and (vi) provides further research directions required into the Internet of Things (IoT) middleware and software architectures.