Aiming at the problem that remote sensing images are difficult to effectively restore spectral details in spectral super-resolution reconstruction due to the complexity and diversity of surface feature spectral characteristics or the low spectral resolution of original images, this paper proposes a super-resolution network model based on multi-attention residual blocks as the generator. Meanwhile, the Histogram Transformer Block and Non-Local Spectral Attention Module are introduced to optimize the spectral feature expression from the levels of local details, long-distance correlation, and global dependence, thereby generating hyperspectral super-resolution remote sensing images with high-quality spectral details. The experimental results show that the average peak signal-to-noise ratio of the proposed method at 8x magnification is 23.136 dB, the average structural similarity index is 0.54294, the average relative of error and the average perceptual similarity are 1.25890 and 0.48436, respectively. Compared with some existing methods, the proposed method performs better in the above four evaluation indicators and can generate remote sensing images with higher spectral details and closer to real hyperspectral resolution.
Few-shot knowledge graph completion refers to inferring missing entity using limited instances. A key challenge lies in entity representation, which is complicated by diverse neighbor attributes. Although the entity’s neighborhood topology holds potential to address this, its significance is overlooked in current research. In this paper, we propose a structure-aware graph attention network for few-shot knowledge graph completion. Firstly, to enhance entity representations, we design a structure-aware graph attention encoder to capture the graph’s structural features of nodes, generating embedding for entity pairs. Secondly, a semantic prototype matching network is employed to compute the prediction score. Experiments on the NELL-One and Wiki-One datasets show that our proposed model outperforms the best baseline models by 0.021, 0.026, 0.039, 0.032 and 0.016, 0.064, 0.043, 0.040 in terms of MRR, Hits@10, Hits@5, and Hits@1 metrics, respectively. This demonstrates that our model can effectively leverage neighborhood topological information to improve the accuracy of knowledge completion, and achieve a better generalization.
为了提高说话人识别系统的性能,提出基于改进语谱图的深度学习说话人识别算法.语谱图当中包含了语音的内容、情绪、语种以及说话人身份等多种信息,在以往的说话人识别算法中,往往没有考虑到说话人身份特性,采用直接提取语音中的语谱图作为网络输入,而说话人识别系统中需要提取语谱图中表征身份的信息,因此需要在原始语谱图的基础上进行改进.在语谱图中,基音频率以及共振峰等信息最能表现说话人的身份特征,从而提出根据语音信号中每一帧的基音频率进行自适应梳状滤波,得到改进后的语谱图,再通过卷积神经网络提取说话人特征,从而达到提升识别准确率的效果.网络模型采用MobileNetv2神经网络,该网络模型具有模型参数少、收敛速度快、识别速度快等优点,有利于实际应用.在对照实验结果中,该方法相对于原始语谱图的准确率分别提高了2.3%、5.2%、3%.
为了更好地保护音乐作品的版权,维护作品所有者的权益,提出一种基于质心和改进奇异值分解(Ameliorate Singular Value Decomposition,ASVD)的鲁棒音频水印算法.首先对水印信息应用交织编码技术充分发挥汉明码的纠错能力,选择质心所在子带作为水印嵌入子带.然后对水印嵌入子带进行离散小波变换(Discrete Wavelet Transform,DWT)和ASVD,利用DWT多尺度分解和ASVD奇异值相较于传统奇异值分解更加稳定的特性,实现水印的嵌入和提取.实验结果表明,采用该算法嵌入水印对音频的质量几乎没有影响,同时水印信息被攻击后具有很强的鲁棒性.
To solve the problem of the low identification rate of language identification in a noisy environment, a language identification method based on the Gammatone-scale power-normalized coefficients spectrograms is proposed, which is obtained by extracting coefficients as features based on the suppression of noise in power and the auditory features of the Gammatone filter-banks. The coefficients are then transformed into images as spectrograms. Then the dark channel prior algorithm and automatic color scale algorithm are applied to enhance and denoise the images. Finally, the residual neural network is used for training and identification. Experiment results show that the identification rate of the proposed method is improved by 39.1%, 12.3%, 19.0%, 5.5%, 28.2% and 28.5% relative to the linear gray-scale spectrograms under the conditions of the signal-to-noise ratio is 0 dB and noise sources are white noise, volvo noise, pink noise, high frequency channel noise, babble noise and factory floor noise respectively. The identification rate under other signal-to-noise ratios is also improved.
为了解决在音频中添加水印信息后如何保持音频质量以及嵌入的水印信息在遭受攻击后的安全问题,提出了一种基于平稳小波变换(SWT)和离散余弦变换(DCT)的音频水印算法.首先,利用Lorenz混沌系统生成密钥对原始水印信息进行加密.然后对原始音频分帧,通过音频帧的能量特征和过零率特征确定水印嵌入帧,对水印嵌入帧进行三级平稳小波变换后,将其三级近似分量平均分成两个一维矩阵,分别进行离散余弦变换,同时计算这两个一维矩阵幅度绝对值的平均值,通过不同的水印信息修改离散余弦变换系数嵌入水印.通过实验选取3种不同类型音乐(classical、hip-hop、rock)测试该算法,结果表明,3种音乐的音频的信噪比分别为25.4517、22.2963、25.2431,高于国际标准;在经过各种攻击后其误码率均在0.02以下;相关系数都在0.98以上.
To solve the issue of low accuracy of language identification in a noisy environment, a language identification method is proposed by combining Mel-scale frequency cepstral coefficients and Gammatone frequency cepstral coefficients. First, the Mel-scale frequency cepstral coefficients and Gammatone frequency cepstral coefficients of speech are extracted, and the feature dimensions are screened based on the language contribution. Then, the feature is mapped in the spatial coordinate system composed of the Mel domain-Gammatone domain to obtain the Mel Gammatone cepstral coefficients(MGCC). Finally, the fusion feature is input into the deep bottleneck network. The experimental results show that the identification accuracy and speed of the proposed method are much higher than those of the single acoustic feature and other features. The accuracy can reach 99.38% in the clean corpus, and can still reach more than 89% under the-5 dB environment, which fully proves the effectiveness and robustness of the proposed method.
语音信号回声隐写后其倒谱系数会在回声延迟出产生峰值,传统回声隐写分析主要采用倒谱系数的统计特征作为隐写检测特征,然而在低回声幅度时隐写信号倒谱系数的峰值并不明显,基于统计特征的方法检测性能并不理想.本文将倒谱分析与图像识别技术结合,提出了一种基于倒谱图像的语音回声隐写分析方法,对语音信号分帧加窗后进行倒谱计算,然后以时间为横轴,倒谱序列点为纵轴,倒谱系数幅值为灰度级生成倒谱图像,将生成的倒谱图像作为隐写检测的输入,采用残差神经网络作为分类器进行回声隐写分析.实验结果表明,在3种经典回声隐写算法上低回声幅度时检测准确率分别达到98.2%、98.6%和96.1%,本文方法在低回声幅度时检测准确率相较传统回声隐写分析方法有较大提升,解决了传统回声隐写分析方法在低回声幅度检测效果不佳的问题.
针对短时语音时长过短以及训练语音和测试语音时长不等,导致语种识别性能大幅度下降的问题,提出了一种可变时长的短时广播语音多语种识别模型(Variable?Duration-Language?Identification,?VD-LID).?首先,对不同时长的语音进行时长规整;然后,对规整后的短时语音进行特征提取,提取其对数功率谱包络图作为语种特征;最后,将语种特征输入到残差神经网络中进行分类.?实验结果表明,相比于传统特征输入,对数功率谱包络图特征将短时语音的语种识别准确率提高到了82.4%;相比于没有引入时长规整层的语种识别模型,VD-LID在测试语音时长为5?s和10?s的实验中,语种识别准确率分别提升了27.9%和37.7%.
针对广播语种识别问题,提出一种语音时域滤波方法,用gammatone时域函数与预处理后的语音信号进行卷积滤波,再分帧加窗并求对数化能量得到时域GF(gammatone filterbank)特征.将特征参数图像化表示,然后通过VGG19和Resnet34分类网络进行语种识别实验.同时,也使用自动色阶算法对加噪语音的图像化特征参数进行去噪,并对比不同维数的特征参数以及不同噪声类型和信噪比对语种识别率的影响.结果表明,采用该特征参数的广播语种识别准确率高于使用传统的GFCC特征、GFCC-D-A特征、GFCC-SDC特征及Fbank特征,且在不同噪声类型和不同信噪比的广播语音识别场景下,语种识别准确率均有一定提升.
针对广播音频语种识别中与语种识别无关的特征对识别结果产生影响的问题,提出一种基于伽马频率倒谱系数的改进特征参数的语种识别方法.通过提取每帧信号的能量谱包络,去除部分与说话人相关的特征,采用Gammatone滤波器组滤波,经离散余弦变换后再进行倒谱提升,得到改进的伽马频率倒谱系数特征参数.将广播音频信号提取特征参数输入隐Markov模型中进行训练测试,得到的语种识别结果表明,该方法有效提升了广播音频语种识别的准确率,优于目前使用的伽马频率倒谱系数特征及其衍生方法.
基于普通图像的数字音频水印算法,鲁棒性弱以及自动化检测复杂.针对这种情况,本文在此基础上提出了一种新的盲水印嵌入算法.利用二维码(QR码)自身的纠错能力,将QR码作为待嵌入的水印图像.首先将水印图像进行分块处理,以Arnold变换为基础,通过添加密匙来提升数字音频水印的安全性.其次对原始音频信号完成分帧预处理后,首先对每帧信号应用3级离散小波变换(DWT),选取低频分量进行离散余弦变换(DCT),然后把得到的一维信号进行奇异值分解.最后通过对奇异值的量化,并利用重复码的特点,将水印信息进行嵌入.实验结果表明,本文算法对噪声、低通滤波、压缩等常见的攻击方法具有良好的鲁棒性.
对于常见的基于离散小波变换-奇异值分解的水印算法应对常见的攻击鲁棒性较差的情况,提出一种基于奇异值比(Singular Value Ratio,SVR)的混合域音频水印算法.先将音频分帧,利用音频特性选择适合水印嵌入的帧,然后对水印嵌入帧进行离散小波变换和离散余弦变换,将离散余弦变换系数分为4段,选取中频系数进行奇异值分解,计算奇异值比.根据音频的质量最优修改奇异值,嵌入水印.同时利用Arnold变换和Logistic混沌序列提升其安全性,利用汉明码增强其纠错能力.实验表明所提算法安全性高,对添加高斯噪声、低通滤波、重采样、重量化和压缩具有良好的鲁棒性.
Based on the spectral peak point characteristics of Chinese speech, this study proposes a syllable matching algorithm to improve the matching effect of Chinese speech syllables in noisy environments. First, a discrete cosine transform is used to extract the speech signal envelope spectrogram, and the human ear masking effect is used for spectral energy judgment to obtain the extreme value points of spectral energy in each frame. Then, the syllable signal is corresponded to a binary sequence by performing binary quantization in the logarithmic frequency range. Finally, the syllable matching result is determined based on the template comparison of the binary sequence. The results show that the proposed algorithm outperforms the conventional methods for matching syllables in the noiseless Chinese speech. Additionally, it has a high matching accuracy at low signal-to-noise ratios.
通过建立一种基于人际关系的传染病传播仿真模型对传染病传播过程以及预测在相关防控措施下疫情发展的趋势进行研究.基于人际关系描述个体间的接触与交互,以个体为单位建立仿真模型,根据中国卫健委平台收集武汉地区COVID-19疫情的初期数据调整模型参数,估算基本再生数(R0)验证模型,并模拟不同疫情防控手段的场景,探讨不同干预措施下疫情传播的趋势.建立的基于人际关系的传染病传播模型首先模拟了武汉疫情初期的传播过程,估算武汉封城前COVID-19的R0;然后对扬州疫情发展趋势进行了初步预测,发现疫情已进入可控阶段.通过探讨在人口密集接触场所(以学校为例)中不同的干预措施对疫情发展的影响,针对学生秋季开学提出相关防控意见,讨论了个体在社交接触网络中的社交移动距离对疫情传播的影响.
Aiming at the problem of low language recognition rate under low signal-to-noise ratio, a language recognition method based on fractional wavelet transform was proposed.Firstly, the adaptive filtering algorithm was used to filter the noise of the noisy signal, so as to reduce the influence of noise on the feature extraction and improve the processing ability of the system for non-stationary signals.Secondly, the motion of the signal on the basilar membrane of the cochlea was simulated, and then the signal was compressed by a nonlinear power function.Finally, the improved CFCC were extracted by simulating the human hearing process.Experiments show that compared with the traditional CFCC, the language recognition rate is significantly improved, and the language recognition rate is increased by 11.1% on average under the 0 dB signal-to-noise ratio, which verifies the effectiveness and robustness of the proposed algorithm.
目的 为了充分考虑种群中个体差异和人员流动对传染病发展的影响,构建一种基于个体行为的传染病模型来揭示传染病传播的动力学特征,对疫情防控措施的有效性进行评价.方法 以个体为基本研究对象建立仿真模型,为个体属性赋予不同数值以体现个体间的差异性,将影响传染病传播的因素以参数形式引入模型,通过改变个体的属性值来反映传染病流行期间个体的状态变化.结果 通过对相关参数的设置,模型仿真结果与疫情发展趋势高度吻合,并以武汉地区新型冠状病毒肺炎疫情的基本再生数为指标,验证了模型的有效性.在此基础上,进一步讨论了个体社交活跃度对疫情发展的影响,模拟了不同防控措施下疫情的发展趋势.结论 该模型参数灵活,适用于传染病在多种情况下的趋势分析,能对防控措施的有效性进行科学评价,为疫情防控提供科学指导.
In mathematical models of microwave heating with infinite-dimensional characteristics, it is difficult to use traditional numerical methods to improve computational efficiency. In this work, we propose a fast and accurate method to calculate the temperature distribution of materials under microwave heating. First, we analysed the relationship between the choice of model order and the solution accuracy by downscaling the infinite-dimensional heat conduction partial differential equation (PDE) model into a finite-dimensional ordinary differential equation (ODE) model. Additionally, the effect of different boundary conditions on the global temperature distribution was analysed. Second, the equilibrium conversion matrix was calculated using the singular value decomposition (SVD) truncation method under homogeneous boundary conditions. Using this matrix, a lower-dimensional microwave heating ODE model was further obtained. Finally, the numerical simulation results showed that the root mean square error (RMSE) was only 0.07 and the maximum relative error was only -0.85%. The computation time of the equilibrium conversion matrix was 2.12 similar to 3.00 ms, and the model calculation time was reduced by 97.78%. We compared the calculated temperature rise curves with those obtained using the conventional COMSOL model. The SVD truncation method achieved an efficient and accurate solution for the microwave heating model.
In the language identification system, the interference of silent segments and the inconsistency of voice decibel range leads to a decline in language identification. Additionally, algorithms using spectrograms for language identification cannot effectively show the information of its low-frequency part, which results in performance failure. To mitigate this, we proposed a language identification method based on joint voice activity detection and dynamic range control. First, we extracted the first dimension coefficient of the Mel-scale frequency cepstral coefficients. Second, we applied median filtering to smooth the feature parameters and perform voice activity detection to remove the silent segment of the voice. Next, we used the dynamic range control to adjust the decibel range of different voices. Finally, we put the log scale spectrogram into the convolutional neural network for classification. The experimental results show that the proposed algorithm improved performance by 7. 16 percentage points as compared with the traditional language identification algorithm using spectrogram in the VoxForge public corpus under the ResNeSt network. Additionally, under the same experimental settings, the recognition performance of the log scale spectrogram showed superiority over other mainstream features, which fully validates the effectiveness and superiority of the proposed algorithm and features.
目前,汉语并列结构的研究对标注语料的依赖较强,无法利用未标注语料中的语义信息,且未引入半监督学习方法.该文以条件随机场为基本框架,提出了一种基于半监督学习的并列结构识别方法.从未标注语料中训练出词向量继而提取无监督特征,同时引入语言学特征进行对比实验,考察不同特征对并列结构识别效果的影响.实验表明,无监督特征的融入能提高并列结构的识别效果,使F值达到85.75%,语言学特征和无监督特征结合后的F值为85.77%.说明语言学特征对结果的影响甚微,而无监督特征的引入可以减少人工选取特征的工作量,并将语义信息以较简洁的方式融入识别模型中.