As technological innovations rise, the amount of information generated by the healthcare sector is enormous and requires efficient statistical analysis procedures to extract maximum information. The research paper presents clear and simple instructions on the way to choose the most appropriate statistical analysis tool in healthcare analysis data. Although numerous research studies have employed different statistical tools in the healthcare sector, in most cases they merely apply any of the statistical tools without giving much consideration to the statistical tool that suits better. The current paper explores SPSS, R, STATA, and SAS to provide the data to a researcher when choosing the most appropriate statistical tool. This paper also utilizes statistical tools on a heart disease dataset and analyzes their performance. This empirical evaluation provides additional quantitative evidence supporting the theoretical comparison discussed earlier. It is also the first attempt at having a comprehensive discussion and comparison of various statistical tools that can be employed in healthcare with distinct results that can guide the researchers in making informed decisions related to the tool to be used.
Object Detection is a fundamental, but challenging problem in computer vision. However, the rise of deep learning techniques has led to creation of numerous detection frameworks that strike an optimal trade-off between speed and accuracy. These models are split into two categories: single stage and two-stage detectors. This study evaluates six prominent object detection models: four single-stage detectors (YOLOv3, YOLOv8n, YOLOv8x, and YOLOv11), and two two-stage detectors (Fast R-CNN and Faster R-CNN). Performance of these models is evaluated using mean Average Precision (mAP) and inference time. Two-stage methods, such as Fast R-CNN and Faster R-CNN, use region proposals to enhance detection accuracy, especially in visually complex scenes. Faster R-CNN was shown to achieve the highest detection accuracy with an mAP score of 85.72
Lately, there has been a remarkable surge in online movie reviews in Hindi with the advent of the UTF-8 standard. Movie reviews are an excellent source of sentiments; therefore, Hindi movie review classification is one of the exciting and demanding tasks of NLP, as it helps the viewers decide whether a film/movie is worth watching. Much work in movie reviews sentiment classification has been done mainly for resource-affluent languages. Still, preliminary work is being done in Hindi due to its complex nature and scarce resources like adequate-labeled datasets. This paper aims to develop a machine learning-based solution for performing binary sentiment classification on movie reviews in Hindi (Devanagari Script). To this end, a primary binary polarity dataset, namely, Movie Reviews in Hindi (MRH) consisting of 5K reviews, is made. Apart from MRH, the Hindi IIT-P movie and product review datasets are also deployed in this work. Firstly, all three datasets are prepared for further processing using the preprocessing steps, and the features used are unigram, bigram, and trigram, along with TF-IDF. Second, various state-of-the-art classifiers are applied to all three datasets. Further, we proposed and used a stacked model of classifiers for performing binary sentiment classification on Hindi reviews. Experimental results on all three datasets prove that the proposed stacking ensemble based on the employed features compared favorably to all the baseline classifiers applied and achieved reasonably high performance. Therefore, it indicates the efficacy of the proposed stacked model for sentence level movie reviews sentiment classification in a resource-scarce scenario.
Motivation: Recognizing and studying DNA patterns is crucial for improving knowledge of illnesses, cell function, and gene control. Motifs determine which transcription factor a protein may bind to, leading to a better unraveling of gene expression. Advancements in the fields of deep learning and high-throughput sequencing have made possible the exploration of motif discovery anew, with greater accuracy and performance. Methodology: In this paper, a novel deep learning framework (XDeMo – Transformer-based Deep Motifs) for DNA motif mining using Transformer models is proposed. Furthermore, a hybrid encoding scheme is also introduced, called ‘blended’ encoding specifically designed for use with deep learning transformer models that are trained using DNA sequences. Results: Our proposed transformer-based framework for DNA motif discovery augmented by blended encoding outperforms many state-of-the-art deep learning models on many baseline performance metrics when trained on the standard datasets. Our models demonstrated robust performance in predicting motifs with high discriminative power, precision, recall, and F1 score. Conclusion: The model’s ability to capture intricate sequence patterns and long-range dependencies led to the discovery of biologically meaningful motifs that were verified from known transcription factor binding motif databases. This shows that our novel framework can be effectively used to find DNA motifs and therefore, aid in further downstream analyses for biomedical and biotechnological applications. XDeMo’s practical implications span the realms of gene regulation research, genomics tool development, molecular biology, and diagnostic applications. It offers a robust foundation for further advancements in genomic analysis, with the potential to accelerate discoveries in gene regulation and the development of novel therapeutic strategies.
With evolving technology, Hindi web content is growing considerably and catching popularity as the larger audience feels more connected and heard using their native language. A volcanic growth in online movie reviews (MRs) in Hindi has been observed lately; manually analyzing them is impossible. Hence, the research problem of automatic organization and classification of Hindi reviews is apparent as this can help viewers decide whether a movie is worth watching or not. This work focuses on developing a deep learning-based system for bi-polar sentiment classification of MRs for resource deficient language – Hindi. To this end, a primary Hindi movie review (MR) corpus is made and manually annotated with binary polarity class labels - positive or negative. The corpus is preprocessed using the preprocessing steps, and Random Word Embeddings (WEs) are utilized for feature extraction. This paper proposes an ensemble CNN_BiGRU, which is an integration of 1D CNN with BiGRU for the bipolar classification of Hindi MRs. To prove our ensemble’s efficacy; other widely used mainstream deep learning models (DLMs) such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN) based models – Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), are also applied and compared with proposed ensemble using average classification accuracy. Empirical results show the effectiveness of the proposed ensemble in achieving a reasonably good average accuracy of 89.366%. The results indicate that proposed CNN_BiGRU compares favorably to the state-of-the-art DLMs applied and hence gives an effective solution for sentence-level bi-polar classification of MRs in a resource deficient scenario.
There is a pressing need for ongoing vigilance and genetic scrutiny of the SARS-CoV-2 virus that devastated humanity globally since its advent in December 2019. Particular focus is needed on the less-explored 5′ and 3′ untranslated regions. These regions, often overlooked in previous studies, hold immense potential for shedding light on critical aspects of viral behavior. Specifically, their analysis has unveiled potential consequences for viral transmissibility and disease severity, contributing to the understanding of this dynamic virus. In this research, a comprehensive genomic computational analysis of significant strains is carried out with an in-depth examination of mutation frequencies, robust phylogenetic analysis, important mutations in the untranslated regions, etc. Intriguingly, this investigation demonstrates how the choice of alignment model can influence estimated mutation rates, particularly concerning insertions and deletions. Nevertheless, overarching patterns of mutation rates across different variants remain generally consistent between the alignment models employed. The findings reveal a spectrum of mutation rates within SARS-CoV-2, ranging from approximately 1.2 × 10−3–1.1 × 10−2 per base. Notably, the Omicron variants exhibit relatively high substitution rates compared to their counterparts, indicative of a substantial accumulation of mutations. These subvariants are distinctive, forming a separate cluster in the phylogenetic tree, characterized by a more recent origin and an accelerated evolutionary pace. These attributes could potentially account for their heightened transmissibility and enhanced immune evasion capabilities. Crucially, this research highlights the pivotal role of mutations in the untranslated regions of the viral genome, particularly in the Delta strain and Omicron subvariants. These mutations significantly contribute to the virus’s ability to propagate rapidly within the human host. Their impact underscores the importance of comprehensive wet lab experimentation to discern the precise elements governing regulatory pathways in SARS-CoV-2 variants. This study not only advances the comprehension of SARS-CoV-2’s genetic landscape but also underscores the critical need for sustained surveillance and genetic analysis. By addressing the research gaps illuminated by this study, governments can better inform and tailor public health interventions to mitigate the impact of this evolving virus effectively.
The current study aims to develop robust contextual knowledge of deep-learning methodology for DNA/RNA motif sequence identification and recognition of correct transcription factor-binding sites (TFBS) for gene regulatory mechanisms in humans. Knowledge of the exact sequence specificities of DNA- and RNA-binding to particular transcriptional factors (TF) seems to be an excellent strategy to develop unique deep-learning models for gene regulatory processes. But uncertainty in the sequence specificity of genomic sequences to a particular TFBS is a big issue. It may be possible to resolve this issue using deep-learning techniques, and thus, it will be helpful to gain generalizable domain knowledge of deep-learning architectures, which offers researchers to know better, their performance to select a unified computational approach for the discovery of a selective kind of motif pattern. This scoping review serves to synthesize evidence for DNA/RNA motif sequences binding with transcriptional factor sites using the PRISMA-ScR guidelines (Preferred Reporting Items for Systematic reviews and Meta-Analyses of Scoping Reviews) to better understand and further assessment of the scope of literature on DNA or RNA motif mining using deep-learning methods. A deep-learning architecture literature survey for DNA and RNA sequence specificity for human ChIP-seq (Chromatin Immuno-Precipitation sequence), DNase-seq (DNase hypersensitive site sequence), CLIP-seq (Cross Linking Immuno-Precipitation sequence), ATAC-seq (Assay for Transposase-Accessible Chromatin sequence), etc. datasets, common motif pattern, and their corresponding TF-DNA/RNA-binding site affinities are included in this study. Deep-learning (DL) models have been used to find selective motifs and have been demonstrated to be more reproducible than traditional methods. As per our literature survey, 33 DL models exist to detect DNA/RNA motifs that have varied framework designs and implementation styles. Through literature survey and PRISMA-ScR reporting guidelines, it is easy to analytically evaluate the performances of each DL model in terms of model size, automatic calibration ability, tool selection, and training set, and it has been found that the DESSO (DEep Sequence and Shape mOtif), DeepFinder, and DeepBind are the selective DL models that are appropriate to study the true biological relationship, especially concerning gene expression patterns and sequence analysis. This study concludes that the application of existing deep-learning methods in the field of motif discovery is the faster way to process complex data relevant to genomic sequences. Through the PRISMA-ScR reporting guidelines and literature survey analysis, more than 30 existing deep-learning models are compared, and it is concluded that complex DL models are preferred over simpler DL models in terms of performance and scalability evaluation. Selective selection of a DL model architecture can be made to understand the complex behavior of motifs and their associated regulatory mechanism at the gene level.
Generally, global pairwise alignments are used to infer homology or other evolutionary relationships between any two sequences. The significance of such sequence alignments is vital to determine whether an alignment algorithm is generating the said alignment as evidence of homology or by random chance. Gauging the statistical significance of a sequence alignment obtained through the application of a global pairwise alignment algorithm is a difficult task, and research in this direction has only provided us with nebulous solutions. Moreover, the case of nucleotide alignments with gaps has been scarcely explored. Very little literature exists on the statistical significance of gapped global alignments employing affine gap penalties. This manuscript aims to provide insights into how the statistical significance of gapped global pairwise alignments that may be inferred using Monte Carlo techniques.
This paper introduces the method and algorithm for improvement of detection system based on CNN. due to some poor resolution and conflict background system faces problem of small object detection. So this paper proposed hybrid method for small object detection. Using improved loss function on IOU for bounding box regression and solve the positioning deviations by using ROI pooling operation and also use the multi-scale convoluted fussed features for more information. This algorithm gives good performance on small object like knife, cutter, wallet, screwdriver etc. Algorithm recall rate reach 85% and accuracy reach 80%. So, the detection is significantly better than other method and more effective to detect small object.
Lately, distinct classifiers have been seen to complement each other in classification performance. This gave rise to the belief of using an ensemble of many simple models instead of using one complex fine-tuned model for classification tasks. Movie review mining or movie review sentiment classification or sentiment analysis (SA) of movie reviews is the process of abstracting the attitude and feelings of the writer from movie reviews text, meaning thereby whether sentiment expressed in written reviews is of positive valence or negative valence. It is known that ensemble models are a promising way of solving various classification problems. These days, researchers have proposed many models for Hindi movie review mining. However, existing research investigations neglect the potency and competence of ensemble models in Hindi movie review sentiment analysis. The voting ensemble is a very popularly used ensemble technique in which many machine learning classifiers (MLCs) can be combined for yielding better predictive performance. This paper proposes a majority voting ensemble model of classifiers for doing binary sentiment classification of Hindi movie reviews. The proposed voting ensemble model has been compared with baseline models. Experimental results on our Hindi movie reviews unfold that our proposed voting ensemble model delivers better acceptable performance when compared to other stated models for sentiment classification on Hindi movie reviews.
In recent technologies, object detection is considered as an effective tool for diagnosing the anonymous activities of a particular location. We can recognize the specific object from images and videos. Therefore, we can obtain essential information for developing a highly secure framework. The ODEF technique is developed for enhanced object detection and classification process. By utilizing the proposed technique, the anonymous activities in a specific region can be detected through the video. This technique detected the objects with bounding boxes; therefore, malicious activities can be shown in distinct visual. Then, the utilization of DRCNN technique provides better platform to classify the object.
Object detection had gained importance in previous decade due to large amount of data that is being generated throughout the world by cameras, mobile phones, satellite imaginary, medical image, social media, UAV etc. As hardware cost to render these images had been reduced significantly and we have access to plethora of algorithms, framework to detect the object and use this information to solve day to day problems. The object detection is most researched area but it still fails to detect and recognize small objects as detecting large objects had got more focus. But small object detection had got less attention and the algorithms and methodology developed for detecting large object does not yield the desired results and accuracy. In this paper we attempt to detect small objects by using state of art algorithm yolov7 and roboflow and try to evaluate the robustness of object detection with scarcity of data in dataset.
The objective of this study is to supply an overview of research work based on video-based networks and tiny object identification. The identification of tiny items and video objects, as well as research on current technologies, are discussed first. The detection, loss function, and optimization techniques are classified and described in the form of a comparison table. These comparison tables are designed to help you identify differences in research utility, accuracy, and calculations. Finally, it highlights some future trends in video and small object detection (people, cars, animals, etc.), loss functions, and optimization techniques for solving new problems.
Sentiment analysis has significantly progressed in English, whereas Hindi research is still nascent. Despite being the third most spoken language worldwide, Hindi remains an RRL. Movie reviews are a treasure trove of opinionated content fueled by people’s passionate engagement with film industry. The proliferation of great use of Hindi in writing reviews has catalyzed our endeavor to devise an approach for bipolar sentiment classification of movie reviews. We compiled and manually annotated a Hindi Language Movie Review (HLMR) dataset comprising 10K reviews for experiments, and challenges associated with Hindi have also been identified. In addition to HLMR, two publicly available IIT-P movie and product review datasets are used. Following dataset preprocessing, we explored TF-ISF with word-level N-gram features for text representation. Studies suggest that performance of machine learning approaches can be enhanced by hyperparameter tuning and ensemble learning. Several baseline classifiers were initially applied, and their parameters were hyper-tuned using Grid search. Subsequently, ensemble-based classifiers were applied individually. Lastly, we propose a simplistic yet powerful stacked ensemble-based architecture (SEBA), which effectively classifies Hindi reviews by leveraging the strengths of both approaches. Comprehensive experiments were conducted on all deployed datasets. Empirical results demonstrate that SEBA outperformed individual baselines and exhibited superior performance with unigrams and TF-ISF as features across deployed datasets. SEBA achieved an accuracy, precision, and recall of 0.808% and an F1-score of 0.807% on the HLMR dataset. These findings strongly advocate for the effectiveness of proposed solution and indicate its suitability for online deployment in binary review classification tasks.
Artificial intelligence includes various technologies: machine learning, neural network, deep learning, robotic, etc. Machine learning (ML) is a part of artificial intelligence in which machine is trained to learn without being explicitly programmed, and artificial neural network (ANN) is a popular model that is used for machine learning. The balanced combination of weight and bias plays a vital role in artificial neural network for error prediction. This paper is intended to provide a detailed study of different supervised artificial neural networks: feedforward ANN, multilayer perceptron, PatternNet, cascade feedforward network based on the different combination of weight and bias for link prediction. A comparative study has been done which highlights that the performance of ANN gets better as weight and bias initialization changes.
Machine Learning techniques have been widely popular in the field of healthcare and have numerous biomedical applications owing to technological advancements in computing technology, processing power, and medical imaging. Since the COVID-19 pandemic began in 2019, researchers and doctors all over the globe are looking for new ways to tackle this crisis. To date, we are unsure of many aspects of this disease including its viral origin and intermediate host (if any), the role of many of its transcribed proteins and other genetic components, the effectiveness of guidelines that may reduce chances of transmission, etc. Efforts are ongoing to contain the pandemic and machine learning models are being increasingly employed to handle many aspects of the disease, like drug discovery, vaccine efficacy, prediction of disease progression, understanding the SARS-CoV-2 protein structure and function, etc. The global impact of COVID-19 on the socio-economic front, the impact of social distancing policies, and vaccinations are also being modeled using artificial intelligence and machine learning algorithms. Much data concerning COVID-19 is now readily available online and research articles related to this disease have been mostly made open for access by all, making it easier for researchers worldwide to get access to valuable information for analysis. Machine learning algorithms are data-hungry in the sense that a lot of data is required by these complex computing algorithms to mine meaningful knowledge. However, once proper data is available, these algorithms learn features from this data and can gauge interesting patterns found. In this manuscript, we aim to look at how machine learning has helped us in the realm of coronavirus epidemiological aspects and how these techniques can be applied in the future, to predict and contain such epidemics and pandemics.