
The rapid growth of digital multimedia sharing platforms increases the demand for secure video copyright protection and unauthorized content tracking. Conventional video watermarking approaches often suffer from low robustness, limited embedding capacity, and poor resistance against geometric and signal-processing attacks. Existing methods also exhibit inadequate detection accuracy under compressed and noisy transmission environments. These limitations create significant challenges in multimedia authentication, copyright verification, and secure video communication applications. This study presents a Hybrid Deep Neural Video Watermarking Framework that integrates attention-driven watermark embedding with intelligent tamper detection mechanisms for robust multimedia security. The proposed method combines Convolutional Neural Networks (CNN), Bidirectional Long Short-Term Memory (BiLSTM), and Transformer-based attention modules to achieve adaptive watermark insertion and reliable extraction. Initially, video frames are decomposed into spatial and temporal components through adaptive feature learning. The embedding model identifies perceptually significant regions using an attention-guided encoder, where encrypted watermark information is inserted with minimal visual distortion. Subsequently, a dual-stage decoder network performs watermark recovery and tamper localization through deep residual feature analysis. The framework also incorporates adversarial training and adaptive noise filtering to improve resilience against compression, frame dropping, Gaussian noise, rotation, and scaling attacks. Experimental evaluation demonstrates that the proposed framework achieves a PSNR of 48.7 dB, SSIM of 0.986, NC value of 0.994, tamper detection accuracy of 98.4%, and watermark recovery accuracy of 96.7% under diverse multimedia attack environments. The framework also maintains stable extraction performance under compression, scaling, Gaussian noise, rotation, and frame-dropping attacks. Comparative analysis indicates that the proposed model improves robustness by 19.8%, perceptual quality by 12.4 dB, and watermark reconstruction accuracy by 18.5% compared with conventional DWT-SVD and CNN-based watermarking approaches.
Computer Assisted Diagnosing (CAD) plays a vital role in healthcare education, surgical training and treatment planning. The backbone of CAD is medical images. The 2D medical images will lack the ability to interpret the information. 3D model construction helps to abolish the hurdle. The 2D medical images can be corrupted by Intensity In-Homogeneity (IIH) or bias. Lesser Intensity In-Homogeneity is hardly noticeable but can cause organ loss while reconstructing a 3D model. The purpose of this research work is to present a novel multi-layered approach that accurately removes the barely perceptible bias, segments the image organ and reconstructs a 3D model from the bias-influenced brain MR images. A Benchmark data set with different bias spectra is used for this research. The formulated 3D model is precise in its illustration without any organ loss.
Multimedia systems have faced persistent challenges in maintaining perceptual quality under dynamic network and computational constraints. Traditional optimization techniques have struggled to preserve visual fidelity while adapting to heterogeneous content distributions. A hybrid deep learning framework was proposed to address these limitations by combining Generative Adversarial Networks (GANs) with attention-based feature learning. The proposed method, named Attention-Guided Generative Adversarial Multimedia Optimization Network (AG-GAMON), was designed to enhance spatial-temporal feature representation and improve adaptive quality control. The GAN component has been utilized to generate high-fidelity reconstructed frames, while the attention mechanism has been employed to selectively focus on semantically important regions. The discriminator has been trained to distinguish between reconstructed and original multimedia samples, ensuring improved perceptual consistency. The framework has integrated a reinforcement-based adaptive weighting strategy that has dynamically adjusted loss contributions across content types. Multimedia systems have faced persistent challenges in maintaining perceptual quality under dynamic network and computational constraints. A hybrid deep learning framework has been proposed to address these limitations by combining Generative Adversarial Networks (GANs) with attention-based feature learning. The proposed method, Attention-Guided Generative Adversarial Multimedia Optimization Network (AG-GAMON), has enhanced spatial-temporal feature representation and adaptive quality control. Experimental evaluation has demonstrated improved performance with 35.2 dB PSNR, 0.94 SSIM, 0.013 MSE, and 94 VMAF compared to CNN, LSTM, and GAN baselines.
Multimedia transmission systems have faced significant performance degradation due to dynamic network conditions and heterogeneous user demands. The background of adaptive multimedia optimization has remained critical in supporting real time quality of service requirements across modern communication systems. The problem has been observed in inefficient allocation of bandwidth and inability of conventional methods to adapt to spatiotemporal variations. To address this issue a Spatiotemporal Deep Q Network based Reinforcement Learning framework has been proposed named STDRL for adaptive multimedia quality optimization. The framework has integrated convolutional feature extraction with temporal dependency modeling using recurrent structures that capture evolving network states. The model has been trained using reward driven policy optimization that balances latency throughput and perceptual quality metrics. Experimental evaluation demonstrates that the proposed method achieves 40.7 dB PSNR compared to 34.0 dB in baseline DQN Adaptive Streaming. SSIM improves to 0.98 compared to 0.91 in conventional methods. Latency reduces to 60 ms compared to 100 ms in baseline approaches, indicating faster adaptive response. Throughput increases to 16.5 Mbps compared to 13.2 Mbps, showing improved bandwidth utilization efficiency. QoE stabilizes at 5.0, indicating optimal user satisfaction. The reinforcement learning agent learns adaptive policies through reward driven optimization that balances visual quality, latency, and smoothness.
Estimation of Depth (z-coordinate) of 3D face from 2D (x-, y-) face images based on similarity transform is an optimization problem. In this work Correlation Scale Factor Differential Evolution (csDE) is proposed and used to estimate the optimal depth values which represent the z-coordinate. The different correlations considered to compute the scale factor of Differential Evolution are Spearman’s Rho, Kendall’s Tau and Pearson Correlation Coefficient. The proposed algorithm is implemented in MATLAB and empirical study is conducted on 2D images taken in lab and 3D Bosphorus database images. Similarity measure is computed between the estimated depth values and Candide Face model depth values for the 2D images captured in lab. In the case of 3D Bosphorus database similarity measure is computed between the estimated and true depth values provided with the database. The similarity measure obtained using Correlation Scale Factor Differential Evolution for the sample images of 3D Bosphorus database is compared with other similar estimation algorithms.
Alzheimer’s disease (AD) is a neurological condition that causes memory loss and cognitive impairment and is gradual and irreversible. Timely intervention and a better quality of life for a patient are possible only through early and accurate diagnosis. This paper seeks to model abnormal neural circuits with brain networks that forecast a sound and timely detection of AD via cutting-edge neuroimaging methods. The Decoupling Generative Adversarial Network (DecGAN) is suggested to identify aberrant neural networks associated with AD. A decoupling module in the model separates brain networks into (i) a sparse network of the brain that contains circuits of great significance, and (ii) an additional network of trivial disease contribution. A conflicted learning scheme guarantees the emphasis on the features that are related to the disease, whereas a sparse capacity loss operation maintains the inherent topographical arrangement of neural networks. The model is trained and tested on DTI and rs-fMRI data of ADNI. Performance is evaluated using ROC-AUC, accuracy, precision, recall, F1-score, and training & validation loss. The proposed DecGAN had 92.1% accuracy, 89.2% precision, 86.5% recall, and 87.1% F1-score, with a ROC-AUC of 93.8% and a final validation loss of 0.0040 and it was significantly better than the current baseline and advanced classification methods. Superior discriminative performance for early-stage AD detection is indicated by a higher ROC-AUC, while better convergence, greater generalization, and less overfitting are demonstrated by lower validation loss. In this work, a sparse capacity loss function that maintains the neural circuit’s topological distribution during decoupled graph reconstruction is introduced for the first time. Proposed methodology allows robust detection of AD-related aberrant circuits even under moderate changes in brain network structure by explicitly restricting sparsity and topology combined. Previous AD classification approaches based on GANs were unable to capture this capability.
Digital media supply chain systems have experienced rapid expansion due to increasing demand for real-time content distribution and adaptive forecasting mechanisms. However, the variability in cross-modal data streams has created significant uncertainty in predicting demand, resource allocation, and delivery efficiency. Traditional forecasting models have struggled to capture nonlinear dependencies across heterogeneous media sources, leading to inconsistent performance in dynamic environments. This study proposed a Neuro-Fuzzy Evolutionary Cross-Modal Forecasting (NFECF) framework to address these limitations. The framework integrated neural network learning capabilities with fuzzy inference reasoning and evolutionary optimization strategies to enhance predictive accuracy. The neuro component modeled nonlinear relationships across multimodal datasets, while the fuzzy layer handled uncertainty in data interpretation. Evolutionary optimization refined model parameters through iterative selection and adaptation. Experimental results demonstrate that the proposed NFECF framework achieved 0.41 MAE, 0.60 RMSE, 4.5% MAPE, 0.97 R², and 96% forecasting accuracy, outperforming LSTM, NFIS, and GA Regression significantly in cross-modal digital media supply chain forecasting tasks.
Pavement crack detection needs to be done to identify and assess cracks in road surfaces. Detecting the crack and its measurement by manual methods is extremely time-consuming and requires a lot of manpower. This process is crucial for maintaining road safety and infrastructure integrity. By detecting cracks early, authorities can prioritize repairs and prevent further damage, ultimately extending the lifespan of roads and reducing maintenance costs. Additionally, crack detection helps improve driving conditions and safety for motorists by enabling timely repairs to be made. Overall, pavement crack detection plays a vital role in ensuring the durability, safety, and efficiency of road networks. Some factors, such as non-uniform intensity, complexity, and irregular patterns of cracks, complicate the process, and the accuracy of the results may be affected. The aim of this study is to develop a practical crack segmentation method for real-time maintenance. In this paper, two models are proposed based on U-Net architecture and feature pyramidal network (FPN) architecture. To verify the superiority and generalizability of the proposed method, two publicly available CRACK500 and CFD datasets are used. Metrics such as AIU (Average Intersection over Union) and ODS (Overall Dice Similarity) measure are used to evaluate the performance. These metrics indicate that the proposed method effectively segments cracks in pavement images, demonstrating its potential for use in real-world applications.
The quality of onions (Allium cepa), a vegetable that is consumed worldwide, is essential for food safety and agricultural economics. This study suggests a deep learning-based approach that uses convolutional neural networks (CNNs) to categorize images of onions as either healthy or unhealthy. The technique uses a proposed architecture and a Customized Dataset (CDS) and publically accessible sources like Fruit360, Onion-det, and Vegetable360. The model, which outperformed the others, reached a peak accuracy of 98.33% on the CDS dataset. The study also highlights the importance of color information in onion disease categorization, with models trained on RGB images performing better than monochrome counterparts. The model’s classification skills are further confirmed using confusion matrices. The CNN architecture has a lot of potential for automated onion quality evaluation, outperforming conventional pre-trained models regarding resilience and dependability, delivering high classification accuracy and great generalization across various datasets and images circumstances.
Detecting red lesions in color retinal fundus images is essential for preventing vision loss and blindness in people with diabetic retinopathy (DR). Among these lesions, microaneurysms (MAs) are the earliest and most common indicators of DR, making their identification particularly important for effective large-scale screening programs. However, accurately spotting MAs is challenging due to low contrast and varying image quality across different imaging conditions. To overcome these challenges, computer-aided diagnostic (CAD) systems powered by deep learning have shown immense potential for supporting timely and precise diagnosis. In this study, we propose a comprehensive CAD framework that combines advanced deep learning models to improve both detection and classification of retinal abnormalities. Our method begins by enhancing image quality—reducing noise, improving clarity, and standardizing image size to ensure consistent input for downstream analysis. We then differentiate between healthy and DR-affected retinas using a VGG-16 network enhanced with a Spatial Pyramid Pooling (SPP) layer to extract rich and meaningful features. These features are then fed into an Extreme Gradient Boosting (XGBoost) classifier, which separates normal from diseased cases. Next, to locate potential microaneurysms, we employ a Residual U-Net architecture with atrous depthwise separable convolutions (RAS-UNet). This model consists of an encoder, an atrous convolution module, and a decoder. The atrous module combines cascaded and parallel operations to capture features at multiple scales, enabling more reliable detection of MAs of different sizes. Finally, we refine the results by passing candidate regions through a Convolutional Neural Network with MGCA (CNN-MGCA) to distinguish true microaneurysms from false positives. We evaluated our system using a range of performance metrics, including accuracy, AUC, sensitivity, specificity, positive predictive value (PPV), F1-score, and FROC analysis. Overall, our experimental results demonstrate that the proposed approach outperforms existing methods reported in the literature, offering a promising tool for large-scale automated diabetic retinopathy screening and early intervention.
Biometric technology protects valuable assets and digital data by utilizing human physical and behavioral traits. Among these traits, palmprint recognition is recognized as particularly effective due to its unique characteristics. This paper introduces a new method called the Hybrid Approach of Compressed Contour Texture Analysis for Palmprint Recognition System with Supervised Deep Learning Classifier (HCCTA-SDLCNet) to enhance security in biometric systems. This method integrates the Hybrid approach of Compressed Contour Texture Analysis (HCCTA) for attribute extraction with the Supervised Deep Learning Classifier (SDLCNet). The process begins with the data normalization of a Two Dimensional Palmprint Region of Interest (2DPROI), followed by the capture of the Contour Pre-Processed Image of 2D-PROI Sample (CPI) using the Progressing Amalgamation of Conventional Compression (PACC) and the Canny edge detection model. The PACC combines the Discrete Wavelet Transform (DWT) and Principal Component Analysis (PCA) for sample compression, final in a Compressed Contour Preprocessed 2DPROI Image (CCPI). The HCCTA method focuses on extracting essential texture features from the CCPI. These extracted features are then analyzed using the SDLCNet classification model to identify individuals. The research utilizes 2DPROI data from the POLYU database at Hong Kong Polytechnic University. The intended system demonstrates a high benchmark, achieving 99.5% recognition accuracy, which is superior to existing biometric approaches.
Human Activity Recognition (HAR) systems have demonstrated excellent performance in controlled laboratory settings, but their performance tends to vary significantly when applied to diverse users, devices, and environments across domains. To overcome the above-mentioned limitation, this paper presents DA-MMHAR, a domain-adaptive multimodal artificial intelligence framework for cross-user and cross-environment HAR. The proposed framework combines visual, motion, and contextual modalities in a single end-to-end architecture. Modality-driven encoders learn complementary spatiotemporal features, which are then projected into a shared latent space and adaptively combined using an attention-based mechanism to dynamically modulate the contribution of each modality. To reduce the distribution differences between source and target domains, an adversarial domain adaptation approach is employed to promote the learning of domain-invariant feature representations. Furthermore, a multimodal data processing pipeline is constructed to coordinate the heterogeneous inputs, and consistency regularization is used to stabilize the predictions and enhance generalization. Comprehensive experiments are carried out on popular HAR benchmarks, NTU RGB+D, UTD-MHAD, and PAMAP2, following cross-user and cross-environment HAR evaluation settings. The experimental results clearly show that DA-MMHAR outperforms the state-of-the-art unimodal and traditional multimodal HAR methods in terms of recognition accuracy and robustness, while maintaining comparable inference efficiency. These results confirm the potential of the proposed framework for reliable real-world HAR applications in dynamic and heterogeneous environments.
Social media platforms have generated massive visual content streams, where sentiment interpretation remains crucial for adaptive marketing and decision making. Problem: existing models have struggled to integrate spatial visual cues with temporal dependencies, which has limited forecasting reliability under noisy and dynamic environments. Method: this study has proposed an Evolutionary Fuzzy CNN-BiLSTM Fusion (EFCBF) framework that has combined convolutional feature extraction, fuzzy logic-based uncertainty handling, and BiLSTM temporal modeling optimized through an evolutionary algorithm. Results: the proposed system has demonstrated improved sentiment classification stability and forecasting accuracy across benchmark social media datasets, outperforming baseline deep learning and hybrid models in precision, recall, and F1-score metrics. The fuzzy inference layer has reduced ambiguity in visual sentiment interpretation, while the evolutionary optimization has enhanced parameter selection efficiency and convergence behavior. The integrated CNN-BiLSTM architecture has captured both local spatial patterns and sequential dependencies effectively, which has strengthened predictive consistency under diverse content distributions. The proposed model achieves 91.2% accuracy, 90.3% precision, 89.9% recall, 90.1% F1-score, and 0.94 AUC-ROC, which outperform CNN-LSTM Hybrid, Fuzzy Logic Classifier, and Evolutionary CNN across all evaluation steps. The fuzzy layer improves uncertainty handling, while evolutionary optimization enhances parameter stability and convergence behavior. The CNN-BiLSTM fusion effectively captures spatial and temporal dependencies in social media image sequences.
Handwritten text recognition in degraded documents remains a major challenge in document image analysis due to factors such as noise, handwriting variability, uneven lighting, faded ink, and physical distortions in historical or low-quality scans. Traditional OCR methods often perform poorly under these circumstances. To improve recognition accuracy, this research proposes a strong architecture that combines a fixed Convolutional Recurrent Neural Network (CRNN)with refined image prose. The preprocessing pipeline includes grayscale normalization, adaptive thresholding, noise filtering (e.g., median and Gaussian smoothing), and morphological operations like dilation and erosion to improve image clarity while preserving critical handwriting features. These refined images are then processed by a CRNN architecture, comprising convolutional layers for spatial feature extraction, bidirectional recurrent layers (LSTM) for sequence modelling, and a Connectionist Temporal Classification (CTC) loss for transcription without character-level segmentation. The addition of preprocessing models reduces the rate of transmitter rate (CER) and word error speed (WER), increasing training stability and flexibility. Our study forms a base line to detect handwriting, such as creating old manuscripts digital and analysing multilingual documents under complex, real -world conditions, which are both effective and expandable.
Pose variation in facial imagery presents a persistent challenge for automated face recognition systems, particularly in uncontrolled environments such as surveillance, access control, and mobile device authentication. This paper introduces an approach based on Conditional Generative Adversarial Network (cGAN) for synthesizing photorealistic frontal views from single profile images. The proposed architecture concatenates a spatially replicated noise vector with the input profile, enabling generation diversity while retaining subject identity. A composite loss function integrating adversarial, L1, and L2 losses is employed to enhance both global realism and pixel-level fidelity. The model is trained on a custom dataset comprising 4,682 images of 44 subjects, each with a single frontal view and multiple side profiles. Training is performed incrementally to improve stability and convergence. Qualitative results indicate that the method produces visually convincing frontal images with preserved identity details. This work establishes a foundation for future extensions involving perceptual loss, identity-preserving regularization, and large-scale evaluations.
Video restoration has remained an important task in multimedia processing because visual data captured in real environments often contain noise, motion artifacts, and resolution degradation. The demand for high-quality video has increased with the growth of surveillance systems, streaming platforms, and intelligent vision applications. Traditional denoising and super-resolution approaches have relied on spatial filtering and convolutional neural networks. However, these techniques have faced limitations in modeling long-range temporal dependencies across frames. As a result, inconsistent textures, motion blur, and temporal flickering have frequently appeared in restored videos. The present study has addressed these challenges by introducing a Recurrent Optical Flow Transformer (ROFT), a recurrent transformer architecture that has integrated optical flow estimation with temporal attention for joint video denoising and super-resolution. The proposed framework has utilized a recurrent transformer module that has captured temporal correlations between adjacent frames while maintaining spatial consistency. An optical flow estimation unit has guided the alignment of frames, which has reduced motion distortion and misalignment during reconstruction. In addition, a temporal attention mechanism that has analyzed contextual dependencies across multiple frames has enhanced feature representation for dynamic regions. The network has processed sequential frames through recurrent connections that have preserved temporal memory and improved reconstruction stability. Experiments have been conducted on benchmark video restoration datasets that contained noisy and low-resolution sequences. The experimental evaluation demonstrates that the proposed ROFTT framework achieves superior performance compared with existing approaches. The model produces a PSNR value of 35.8 dB and an SSIM value of 0.97, which indicate improved reconstruction quality and structural preservation. The reconstruction error decreases to 0.005 MSE, while the temporal consistency error reduces to 0.007, which confirms stable frame transitions across video sequences. Furthermore, the model achieves an FSIM value of 0.995, which indicates strong preservation of perceptual texture features. These results demonstrate that the proposed architecture effectively integrates optical flow alignment and temporal transformer attention that enhances both spatial detail recovery and temporal coherence in restored video frames.
Surgical video analysis has become an essential component in computer-assisted interventions and clinical documentation. The rapid growth of minimally invasive surgery has produced large volumes of surgical recordings that require detailed frame-level annotations for training intelligent systems. Manual annotation of surgical videos remains a labor-intensive and time-consuming process that often requires expert knowledge. As a result, the development of automated annotation systems has become a critical research direction in medical image analysis. Existing segmentation and annotation approaches have faced limitations in handling complex surgical scenes, instrument occlusions, illumination variations, and tissue deformation. Conventional deep learning models often rely on large labelled datasets, whereas surgical datasets usually remain limited due to the difficulty of manual labeling. This challenge has reduced the reliability and scalability of automated surgical video segmentation systems. To address these issues, this study has proposed an Active Deep Ensemble Segmentation Network (ADES-Net) for automated surgical video segmentation annotation. The framework has integrated an ensemble of convolutional segmentation models with an active learning strategy that has selectively identified informative frames for annotation. The ensemble architecture has combined multiple deep segmentation networks that have captured diverse spatial representations from surgical frames. An uncertainty-driven active sampling mechanism has prioritized frames that required expert labeling, which has reduced redundant annotations. Feature representations that were extracted from each model have contributed to robust segmentation predictions, while iterative learning cycles have refined the annotation quality. The experimental evaluation demonstrates that the proposed ADES-Net framework achieves superior segmentation performance across multiple metrics. The model achieves a Dice similarity coefficient of 0.93, an IoU of 0.86, precision of 0.93, recall of 0.91, and an F1 score of 0.92 when trained with twenty-five annotated frames. These results indicate that the active ensemble mechanism effectively captures spatial and contextual features, reduces false positives, and improves boundary delineation. Compared with baseline methods such as U-Net, Attention U-Net, and DeepLabV3+, the proposed framework achieves improvements of 5–10% across all metrics, demonstrating enhanced segmentation reliability, efficiency, and robustness in automated surgical video annotation tasks.
Sentiment analysis in social media has gained substantial attention due to the rapid growth of multimedia content across digital platforms. Traditional sentiment analysis techniques primarily relied on textual information, which has limited the capability of capturing the rich emotional cues that appear in audio signals and visual expressions. Social media posts frequently contain videos that integrate speech, facial expressions, and textual captions. These heterogeneous modalities carry complementary emotional information that conventional unimodal models have struggled to interpret effectively. The inability of earlier systems to integrate multimodal information has created limitations in sentiment classification accuracy and contextual understanding. To address this challenge, the present study has introduced a Fusion Transformer for Multimodal Sentiment Analysis (FTMSA), which has integrated audio, visual, and textual modalities into a unified representation framework. The proposed architecture has utilized transformer based attention mechanisms that have captured inter modal relationships among speech tone, facial features, and textual semantics. A feature extraction module has processed textual embeddings through contextual language representation, while acoustic descriptors have represented speech characteristics and visual encoders have captured facial emotional cues. These heterogeneous features have been fused through a cross modal attention transformer that has learned correlations among modalities. The training procedure has employed supervised learning that has optimized sentiment classification performance across multimodal inputs. Experimental evaluation has demonstrated that the proposed FTMSA model has achieved improved sentiment recognition accuracy when compared with conventional unimodal and early fusion techniques. The experimental evaluation demonstrates that the proposed FTMSA achieves a maximum accuracy of 93.2%, precision of 92.3%, recall of 91.3%, F1 score of 91.8%, and specificity of 92.7%, outperforming existing methods such as MAN, RMNN, and TBMM. The model maintains superior performance across varying training epochs and dataset sizes, validating the effectiveness of the cross modal attention mechanism in capturing textual, acoustic, and visual sentiment cues for accurate prediction.
The rapid growth of multimedia communication has significantly increased the demand for efficient video compression techniques. Conventional video coding standards often rely on fixed or globally optimized rate–distortion strategies that inadequately adapt to spatial content variations across video frames. As a result, regions with complex textures or motion frequently experience quality degradation, while smoother areas unnecessarily consume coding resources. This imbalance has created challenges in achieving optimal compression efficiency without sacrificing perceptual quality. Therefore, an adaptive mechanism that intelligently allocates coding resources across spatial regions has remained an important research requirement. To address this limitation, this study has proposed a novel neural compression framework termed Spatially Variable Rate–Distortion Neural Coding (SVRD-NC). The framework has utilized a deep neural encoder–decoder architecture that has integrated spatial attention modules and adaptive rate–distortion optimization strategies. Within the architecture, a content-aware feature extractor has analyzed spatial characteristics of video frames, including texture density, motion intensity, and structural complexity. These extracted features have guided a spatial weighting module that has dynamically adjusted the rate–distortion trade-off for different regions of each frame. The optimization mechanism has employed a learning-based distortion estimator that has predicted perceptual reconstruction errors across spatial segments. This prediction has enabled selective bitrate allocation to visually important regions while maintaining efficient compression in smoother areas. The neural entropy model that has been incorporated within the framework has further enhanced coding efficiency by modeling spatial probability distributions of latent representations. Experimental evaluation has been conducted on widely used video datasets that include diverse motion patterns and scene complexities. Experimental evaluation demonstrates that the proposed SVRD-NC framework achieves significant improvements in neural video compression performance. The method achieves a maximum PSNR value of 37.1 dB, which exceeds the Deep Convolutional Autoencoder Compression model that produces 34.2 dB under similar complexity conditions. The structural similarity evaluation indicates that the proposed framework reaches 0.98 SSIM, while the attention-based compression method achieves 0.97. The bitrate analysis shows that the proposed method reduces the transmission requirement to 620 kbps, compared with 720 kbps that appears in the convolutional autoencoder model. The compression ratio improves to 25.1, while the existing approaches remain between 21.2 and 23.6. The reconstruction accuracy also improves because the Mean Squared Error decreases to 0.006, compared with 0.010 that appears in the baseline compression model. These results demonstrate that the spatially adaptive rate–distortion mechanism effectively improves compression efficiency while preserving the perceptual quality of reconstructed video frames.
The rapid growth of visual data has increased the demand for efficient image representation techniques that reduce storage and computational requirements while preserving structural information. Neural networks have provided powerful mechanisms for learning compact image representations, yet conventional backpropagation often struggles with local minima, slow convergence, and inefficient parameter optimization when handling highly compressed visual features. These limitations have created challenges for developing scalable learning frameworks that maintain reconstruction accuracy and representation efficiency. This study has proposed a hybrid meta-heuristic optimization framework for learning compressive neural image representations. The framework has integrated a Compressive Backpropagation Neural Network with a hybrid search mechanism that has combined Particle Swarm Optimization and Differential Evolution strategies. The hybrid mechanism has guided weight initialization and adaptive parameter tuning during training, which has improved the exploration and exploitation balance within the optimization space. The compressive representation module has transformed high-dimensional image data into compact latent vectors that preserved essential spatial patterns. The neural network has then reconstructed the images from these compressed representations through iterative backpropagation that has minimized the reconstruction loss. The meta-heuristic component has refined network parameters that ensured stable convergence and prevented premature stagnation. The experimental results show that the proposed framework achieves a peak PSNR of 37.4 dB, SSIM of 0.97, MSE as low as 0.009, and a compression efficiency of 24.2. The model converges rapidly within 58 epochs and 126 seconds, outperforming existing methods such as the Convolutional Neural Representation Model, Sparse Autoencoder Representation Model, and Particle Swarm Optimized Neural Network Model. These results indicate that the proposed hybrid meta-heuristic framework effectively balances high reconstruction accuracy with efficient compressive learning.