
In this paper, a blind image quality assessment (BIQA) method is proposed, which leverages both a visual neuron model and visual attention mechanism. The research of BIQA is aimed at evaluating the perceptual quality of images without access to reference images. In order to make the quality scores of the BIQA methods more in line with the visual perception, modeling the human visual system (HVS) is an effective approach. The neurons in the visual cortex are the key components of the HVS, and their response characteristics can be described by scene statistics models. Therefore, we design a visual neuron model to simulate cortical responses. This model is implemented using kernel principal component analysis (KPCA) to simulate the complexity of neural responses, thereby enabling the extraction of perceptually relevant statistical features. The attention mechanism of the HVS allows it to focus on the most salient regions of an image. The visual attention network (VAN), utilizing large-kernel convolutions, effectively captures long-range dependencies and global information within an image, making it suitable for extracting attention features. The extracted statistical and attention features are then fused and input into a regression network to predict the image quality score. Experimental results on multiple benchmark datasets demonstrate that the proposed method outperforms state-of-the-art IQA models.
The heterogeneity of the network at the edge creates unequal bandwidth availability for clients connected through diverse networks and devices. This variability can lead to bandwidth fluctuations and slow response times, resulting in buffering or degraded streaming performance. As such, effective resource allocation and bandwidth utilization are crucial for optimizing video quality. Live events, which often experience sudden spikes in traffic, present additional challenges in terms of resource management. These challenges are further intensified when malicious users exploit shared bandwidth, placing additional strain on the system. This study is aimed at enhancing adaptive streaming by balancing both security and quality. We integrated watermarking techniques with the analysis of user streaming behavior and traffic patterns in real live streaming platforms within a multi-CDN (Content Delivery Network) architecture. While Digital Rights Management (DRM) primarily controls access to content, watermarking technologies go a step further by embedding unique identifiers to track content distribution. This capability enables precise monitoring of content, even when sourced from unknown origins, though it comes with added complexity and challenges. By leveraging real-time system data analytics, we extract critical insights that drive decision-making, strengthen fraud detection, and optimize resource allocation. Our empirical results show that, along with providing secure live video streaming, having a comprehensive overview of network conditions and traffic distribution can improve Key Performance Indicators (KPIs) in practice by up to 17%.
With the continuous development of information technology, multimedia information faces many security threats. To realize efficient multimedia information processing under the premise of guaranteeing information security, this research analyzes and extracts the parameters of different kinds of information in view of the fact that each of them has unique attributes and complex structures. The study discusses in depth the encryption technology and confidentiality system of multimedia information. On this basis, the study analyzes the coding format, changes the signal-to-noise ratio and sampling frequency, and performs a balance test on the parameters according to the changes in audio quality for multimedia audio, which is a representative type of multimedia information. The parameters that pass the test are screened as sensitive audio frames, and the encryption scheme based on the parameter sensitivity test is designed. The experimental results revealed that the variation of the peak signal-to-noise ratio of the three types of audio after encryption ranged from 0 to 25 dB. The average audio encryption elapsed time was shortened by 5.3%, 6.2%, and 10.4%, respectively, compared to the traditional method. The encryption time share of Horn, Inst, Popm, and Bass did not exceed 5% overall. The results indicated that the designed multimedia information encryption scheme based on parameter sensitivity was effective in securing multimedia information and improving encryption efficiency. The research results have guiding significance for the security control of multimedia information.
The rapid growth of the intelligent media industry has led to an increasing demand for high-quality video in fields such as video production, digital content exchange, and the film industry, especially for 4K, 8K, and ultra-high-definition (UHD) content. However, existing video compression methods, such as HEVC and JPEG-XS, often struggle to balance compression efficiency and visual quality, particularly with color-rich and dark-field-rich videos. This paper proposes a novel research on cloud compression codec technology for remote production, to address the compression challenges of high-resolution, color-rich, and dark-field-rich videos. By incorporating advanced block partitioning techniques and enhanced encoding tools, the algorithm improves compression efficiency while minimizing visual quality loss, especially under HDR and low-light conditions. Using the AVS3 video encoding standard, it achieves significant gains in compression ratio and real-time encoding performance, making it suitable for UHD video streaming, VR, AR, and other real-time applications. Experimental results show that the proposed algorithm outperforms existing methods in both compression ratio and visual quality, offering a promising solution for future high-quality video applications.
This study explores the implementation challenges of data science in Indonesian digital-native newsrooms by focusing on two leading platforms: Detik.com and Kumparan. Rather than generalizing to the entire media ecosystem, this research provides a focused comparative case study to capture institutional variations in how data science is adopted and managed across different newsroom models. The comparison between Detik.com and Kumparan is significant because both are among Indonesia ' s most data-driven digital-native media, yet they differ in organizational structure, technological resources, and innovation strategy-making them ideal cases for examining the broader dynamics of data science adoption in emerging media systems. Using a mixed-method approach that includes in-depth interviews, field observation, and directed content analysis, the study examines how data science is integrated into newsroom practices for content planning, audience engagement, and predictive analytics. Interviews with key personnel provide insight into internal decision-making and organizational constraints, while content analysis of 120 news items reveals the extent and patterns of data-driven editorial output. The findings indicate that both platforms face similar difficulties in data quality, fragmented system integration, and limited human resource capacity in analytics. Although predictive tools are being tested, their implementation remains experimental due to infrastructure and talent gaps. The study contributes to the literature on data-driven journalism in emerging media environments by highlighting the tension between technological potential and organizational readiness.
History education represents a multifaceted concept, encompassing a visionary idea, a transformative educational reform movement, and a dynamic process aimed at fostering equitable learning experiences for all students. Despite its essential role in shaping historical awareness, national development, and identity, history education is often perceived as boring and less desirable in many parts of the world. To address this, teachers should develop teaching materials focused on themes like the archipelago and independence, drawing from books and other resources. This paper is aimed at creating a mobile game application to enhance educational resources for children, employing the seven stages of game development according to Pickell. The research findings indicate that the mobile game application is effective and significantly increases students’ history learning because it is fun. Therefore, it is recommended that elementary schools adopt this mobile game application to improve history learning.
Predicting origin-destination (OD) flow presents a significant challenge in intelligent transportation due to the intricate dynamic correlations between starting points and destinations. Although existing OD prediction methodologies leveraging graph neural networks have demonstrated commendable performance, they often struggle to address the complexities inherent in two-sided correlations. To address this gap, this paper introduces a novel approach, the 2D spatiotemporal hypergraph convolution network (2D-HGCN), designed specifically for forecasting OD traffic flow. Our proposed model employs a two-stage architecture. Initially, temporal characteristics of traffic flow between OD pairs are captured using a 1D convolution neural network (1D-CNNs). Subsequently, a 2D hypergraph convolutional network is introduced to uncover spatial correlations in OD flow patterns. The unique aspect of our 2D-HGCN lies in its dynamic hypergraph, which evolves over time, enabling the model to adaptively learn changing spatial dependencies. Experimental evaluation conducted on real-world datasets highlights the efficiency of our suggested model for predicting OD flows. Our results demonstrate a promising predictive performance, showcasing the ability of the 2D-HGCN to effectively capture the intricate dynamics of OD traffic flow.
Artistic image transformation is a computer technique widely applied in art creation, design, entertainment, and cultural heritage by converting images into artistic styles. It offers innovative ways for artists to express themselves, provides designers with more choices and inspiration, enhances visual esthetics, and enables creative implementations in movies, games, and virtual reality. Additionally, it aids in the restoration and preservation of ancient artworks, allowing a deeper appreciation of classical art. Traditional image transformation methods, though effective for simple effects, lack the flexibility and expressiveness of deep learning–based approaches. To enhance the effectiveness and efficiency of artistic image transformation, this paper employs generative adversarial networks (GANs), which utilize an adversarial training mechanism between a generator and a discriminator to produce high-quality and realistic image transformations. This study introduces spectral normalization (SNGAN) to further improve GAN performance by constraining the spectral norm of the discriminator’s weight matrix, preventing gradient issues during training, thus improving convergence and image quality. Experimental results on the CHAOS dataset indicate that the proposed SNGAN model achieves the lowest mean absolute error (MAE) of 0.3420, the highest peak signal-to-noise ratio (PSNR) of 32.1423, and a structural similarity index (SSIM) of 0.6696, closely matching the best result. Additionally, the SNGAN model demonstrates the shortest training time, highlighting its efficiency. These results confirm that the proposed method achieves more realistic and efficient artistic image transformations compared to traditional methods and other deep learning algorithms.
Nowadays, the strong development of the economy and society has driven the increase in traffic participation, making traffic management increasingly difficult. To effectively address this issue, AI applications are being applied to improve urban traffic management and operations. Therefore, we propose a smart system to detect and monitor vehicles across multiple surveillance cameras. Our system leverages data collected from traffic surveillance cameras and harnesses the power of deep learning technology to detect and track vehicles smoothly. To achieve this, we use the YOLO model for detection in conjunction with the DeepSORT algorithm for precise vehicle tracking on each camera. Furthermore, our system uses a ResNet backbone model for feature extraction of objects within each camera’s frame. It utilizes cosine distance to identify similar objects in other cameras, facilitating multicamera tracking. To ensure optimal performance, our system is implemented using the NVIDIA DeepStream SDK, enabling it to achieve an impressive speed of 21 fps on each camera and an average of precision approximately 85% for three modules. The results of our study affirm the system’s suitability and its potential for practical applications in the field of urban traffic management.
Conventional data acquisition systems face challenges in achieving high acquisition speeds and rapid storage of large data volumes using microcontrollers. In contrast, field-programmable gate arrays (FPGAs) offer numerous advantages, including high clock frequencies, minimal internal delays, fast operational speeds, abundant internal RAM resources, and simplified control of complex peripheral circuits. This study presents the design of an FPGA-based multimedia remote monitoring system for information technology server rooms. The proposed system utilizes an FPGA as the primary controller and incorporates environmental sensors, electrical energy sensors, carbon monoxide sensors, smoke sensors, and A/D converter modules to monitor multiple locations within the server room. Simultaneously, the FPGA transmits the collected data from each monitoring point via a serial port to an LCD serial screen for display. An alarm is triggered if any environmental anomalies are detected, indicating abnormal statuses. Additionally, the system employs the fast Fourier transform algorithm and butterfly operations to derive voltage and current AC quantities, while utilizing relative temperature differences to identify equipment faults within the server room. The system is evaluated in terms of functionality, user-friendliness, and reliability. Experimental results demonstrate that all performance measures align with expectations, fulfilling the initial design objectives and highlighting the potential applicability and relevance of FPGA technology in the monitoring field. In addition, a comparison between FPGA and traditional microcontroller systems is performed, showcasing the superior processing speed and performance of the FPGA-based system. This comparative analysis further validates the advantages of using FPGA technology in high-speed data acquisition and monitoring applications.
Postgraduates being as the backbone of the innovative talent team, the enhancement of postgraduates’ innovative behavior has become the focus of higher education reform in the new era. Postgraduate students’ innovative behaviors are influenced by multiple factors, and mentorship plays a crucial role in the cultivation of postgraduate students’ innovative behaviors. This study is based on social cognitive theory and empirical analysis from the perspective of individual postgraduate students and mentorship cultivation. Multisource data were obtained from a research team attending a teacher training college in southwest China, and a questionnaire survey was conducted on 362 postgraduate students based on an online approach. The empirical study was conducted using SPSS software and AMOS software combined with hierarchical regression analysis and structural analysis of covariance to examine the mechanism of the effect of transformational tutoring style on postgraduate students’ innovative behavior, using creative self-efficacy as a mediating variable. The results showed that the transformational tutoring style had a significant positive effect on the innovation behavior of postgraduate students, and the creative self-efficacy partially mediated the effect of the transformational tutoring style on the innovation behavior. According to the findings of the study, the creative self-efficacy of postgraduate students is enhanced through the collaboration of “multiple” subjects; the “integrated” cultivation model is built to create a transformative tutor team; a mentoring community is established. The study is aimed at providing a reference for the cultivation of innovation ability of master students.
Digital technology offers numerous advantages, such as preserving the authenticity, replicating reality, and facilitating dissemination. It enables the preservation of intangible cultural heritage (ICH) in its original form and allows for the creation of comprehensive graphic, audio, and visual databases. Among these technologies, holographic technology holds promise for protecting ICH and promoting its dissemination. This paper focuses on interactive holographic technology and presents the design and implementation of a dynamic holographic display system that combines digital hologram (DH) and computer-generated hologram (CGH) to showcase 3D images consisting of both virtual and real objects. Real-time loading of DH into a spatial light modulator enables the optical reproduction of real objects, while the loading of two CGHs into other spatial light modulators facilitates the optical reproduction of virtual objects. Computational holography allows for the addition of virtual information, such as coordinate text, and the fusion of the three reconstructed images in space, resulting in an augmented reality experience and enhanced 3D display of real objects. An experimental setup employing three liquid crystal on silicon (LCOS) devices confirms the validity of the proposed method. Compared to other techniques, this approach demonstrates improved image signal-to-noise ratio, reduced alignment errors, and wider coverage of light traversal for laser 3D reconstruction images. The holographic technology presented in this paper enables the fusion display of real and virtual scenes and real-time two-way interaction between the audience and virtual images. This research holds significant practical value in promoting the effective dissemination and protection of ICH.
In the context of the rapid development of multimedia and information technology, machine translation plays an indispensable role in cross-border e-commerce between China and Japan. However, due to the complexity and diversity of natural languages, a single neural machine translation model tends to fall into local optimality, leading to poor accuracy. To solve this problem, this paper proposes a general multimodal machine translation model based on visual information. First, visual information and text information are used to generate a visual representation of perceptual text information. Then, the two modal information are encoded separately, and the proportion of visual information in the whole multimodal information is controlled by a gating network. Finally, experiments are conducted on the image description datasets MSCOCO, Flickr30k, and video dataset VATEX, respectively. The results show that the algorithm in this paper achieves the best performance on both the BLEU and METEOR evaluation metrics.
The digital multimedia network data traffic is huge, and many end devices are applied to the digital multimedia network; therefore, the digital media network communication delay transmission must be analysed in order to improve the performance of the digital multimedia network communication transmission. Due to the presence of a large number of data packets in the communication network used, some random delays in data can occur during use. Conventional data transmission methods rely solely on extended data storage capacity to achieve this, making it difficult to achieve intelligent scheduling, resulting in delays and optimized performance degradation during communication data transmission. The aim of this paper is to enhance the communication quality of digital multimedia networks by developing an algorithmic model that rapidly mitigates communication delays. The proposed approach employs a delay cancellation algorithm in a network predictive control system. Simulation techniques are utilized to simulate and compensate for delays in multimedia communication networks. The efficacy of the network control strategy is demonstrated through simulation tests on a large-scale multimedia network communication system. The research findings indicate that the network control prediction model and elimination algorithm can substantially reduce the packet loss rate and quickly eliminate communication delays in digital multimedia networks.
In order to solve the bottleneck problem of public service advertising, an analysis method of creative design and design and production technology of animation public service advertising based on the ant colony optimization algorithm was proposed. In the current creative design process of animation public service advertising, along with the improvement of people’s aesthetic requirements, it is necessary to carry out a comprehensive innovative design of its design methods and content. This paper mainly describes how to use the ant colony optimization algorithm in the actual design process to carry out the corresponding design and analysis and then improve the overall communication of advertising. The research results show that under the ant colony algorithm, the customer satisfaction rate of public service advertising design is 89%, the general rate is 10%, and the dissatisfaction rate is 1%, which is superior to the traditional algorithm. Using Flash, 3DS Max, Premiere, and other software as well as “3D simulation,” Easy Mocap motion capture, and other technologies, the production quality of animation public service advertising films has been effectively improved.
With the development of film and television industry and the rise of new media short video, film and television postproduction process has higher and higher requirements for image quality. In film and television postproduction, optimizing the image quality can enhance the resolution and make the image more vivid and detailed. High-quality image can fully embody the value of film and television, as well as promote the development of new media short videos. This paper optimizes image quality by improving image processing technology, thus improving the quality and value of film and television and new media short videos. In this paper, a convolutional neural network combined with a nonlinear activation function is used to establish an improved image processing technology model to efficiently extract image features. This technology can enhance the ability to extract image features, improve the accuracy of image feature extraction, and then improve the image resolution and details, thus improving the image quality in the process of film and television postproduction. The results show that the average value of PSNR is 30.29. The average value of PSNR of the proposed algorithm is higher than that of other algorithms, indicating that the error between the image processed by the proposed algorithm and the original image is small. The average SSIM of the algorithm in this paper is 0.903, which is closer to 1. Compared with other algorithms, the structure processed by the algorithm in this paper is more similar to the original structure, resulting in a better graph. The algorithm in this paper has the best performance on both the peak signal-to-noise ratio (PSNR) and the structural similarity (SSIM). The improved image processing technology proposed in this paper can effectively improve the accuracy of image feature extraction, making film and television or new media short video images of higher quality.
With the change of the times, under the leadership of big data, cloud computing, network technology, mobile Internet, and other high technology, education has gradually broken free from the traditional teaching mode. At the same time, advanced technology has also lifted the confinement of traditional teaching mode in space and time and opened the skylight of intelligent thinking mode. From the current situation of Chinese language fever around the world, the demand of Chinese language learners for learning materials and learning resources will continue to increase in the future. With the continuous development of Internet education, the combination of Internet technology and international Chinese education will definitely become the key direction for the development of international Chinese education. In this context, online teaching resources will become an important basis for the development of “Internet+Chinese international education”, so it is necessary to investigate and study the online teaching resources of “Internet+Chinese international education.” As a matter of fact, with the development of society and the advancement of technology, the era of informationization has come. As a result, the main theme of education in the information age is to provide suitable education for students who want to learn Chinese and to promote the active and lively development of each student. In the past, the one-size-fits-all education model of classroom teaching could not meet the individual development of students and the demand of society for diversified talents. Therefore, the traditional teaching of Chinese as a foreign language is in urgent need of change, and personalized teaching is attracting attention. With the emergence of technologies such as cloud computing, Internet of Things, and big data in education, personalized teaching has received technical support. This paper is aimed at exploring how to apply the new technologies in the teaching process to help teachers personalize teaching, stimulate students’ interest in learning, meet their individual needs, break through the traditional teaching methods, and make it possible to teach according to their abilities. As one of the main learning resources, the quality of multimedia courseware will be the criterion to measure whether it meets the teaching and learning needs. Consequently, the use of big data, multimedia, and multimodal technologies to improve the Chinese language material library and develop special software for teaching Chinese as a foreign language, as well as the production of high-quality multimedia courseware that is highly compatible with the teaching materials, will become the trend of future research.
Blockchain technology is widely used in the field of digital right protection technology. The traditional digital right protection scheme is not only inefficient and highly centralized but also has the risk of being modified. Due to its own characteristics, blockchain cannot completely store all the original files of digital resources. In this paper, a convolutional neural network algorithm based on visual priority rule is proposed (CNNVP). This algorithm can recognize facial expressions in the original files of digital resources (for short video of face class). The algorithm extracts facial expression features accurately and makes these features form log files that can represent the original files of digital resources. Then, the paper proposes a short video copyright storage algorithm based on blockchain and facial expression recognition and stores the log file into the blockchain. The above methods not only improve the efficiency of short video copyright storage, reduce the degree of storage centralization, and eliminate the risk that copyright is easy to be modified. Moreover, the computing operation of deep learning technology on short video not only ensures the privacy of storage certificate information but also ensures the possibility of blockchain storage of video information. Experiments show that the algorithm proposed in this paper is more efficient than the traditional copyright storage method. Moreover, the algorithm proposed in this paper can provide technical support to the media resource management department.
Compression is an essential process to reduce the amount of information by reducing the number of bits; this process is necessary for uploading images, audio, video, storage services, and TV transmission. In this paper, image compressions with losses from this action will be shown for some common patterns. The compression process uses different mathematical equations that have different methods and efficiencies, so some common mathematical methods for each style are presented taking into consideration the pros and cons of each method. In this paper, it is demonstrated that there is a quality improvement by applying anisotropic interpolation to edge enhancement for its ability to satisfy the dispersed data of the propagation process, which leads to faster compression due to concern for optimum quality rather than fast algorithms. The test images for these patterns showed a discrepancy in the image resolution when the compression coefficient was increased, as the results using three types of image compression methods proved a clear superiority when using “partial differential equations (PDE)”.
For multiuser millimeter wave (mmWave) massive multiple-input multiple-output (MIMO) systems, the key factor that causes the system sum-rate to change is its interference plus noise, corresponding to the number of base station (BS) antennas and signal-to-noise ratio (SNR) variation causing the system sum-rate changes, which in turn affects the user’s communication quality. Based on this, this paper proposes an improved low-complexity hybrid beamforming grouping sum-rate maximization (HBG-SRM) algorithm to achieve system sum-rate maximization under the premise that both BS and user side use hybrid beamforming architecture and the channel state information (CSI) of the downlink channel is perfect. The algorithm first predefines a relevant threshold value, which is used to group multiusers; then, the user uses the maximum likelihood (ML) criterion to identify the optimal beam and estimate its beamforming gain within each group, and finally, the user compares all the candidate optimal beam gains between each group to confirm the optimal beamforming vector. The simulation results also verify the superiority of the proposed algorithm’s sum-rate over other algorithms.