There can be numerous electronic components on a given PCB, making the task of visual inspection to detect defects very time-consuming and prone to error, especially at scale. There has thus been significant interest in automatic PCB component detection, particularly leveraging deep learning. However, deep neural networks typically require high computational resources, possibly limiting their feasibility in real-world use cases in manufacturing, which often involve high-volume and high-throughput detection with constrained edge computing resource availability. As a result of an exploration of efficient deep neural network architectures for this use case, we introduce PCBDet, an attention condenser network design that provides state-of-the-art inference throughput while achieving superior PCB component detection performance compared to other state-of-the-art efficient architecture designs. Experimental results show that PCBDet can achieve up to 2$\times$ inference speed-up on an ARM Cortex A72 processor when compared to an EfficientNet-based design while achieving $\sim$2-4\% higher mAP on the FICS-PCB benchmark dataset.
The percussion response of long bone has the potential to be used as a measure of bone strength for Osteoporosis detection. Modelling the vibration response requires describing the shape of the long bone which can have several features. An overly simplistic model of the shape does not give enough insight into their influence on the vibration response. This paper identifies the key features of the shape of a tibia and femur bone (cross-sectional shape, twist, and scale of the ends) and investigates their individual effects on the eigenfrequencies using finite element modelling. A femur and tibia model are dissected at the thicker ends and length adjusted to isolate the influence of the proximal and distal ends on the eigenfrequencies. Selected cross-sectional shapes are investigated to simplify the modelling and compared to real bone cross-sections and results. The twist is added across the longitudinal axis of the model producing an inline twist to the cross-section and resulting in a 1.5–2.5% decrease in frequencies per 20° of twist. The scale of the cross-sections at the ends of the model are increased along a set length of the bone to emulate the larger proximal and distal end of the long bones. The results show that any model for the vibro-acoustic response of long bones needs to include asymmetry in the cross-section as well as the scaling of the ends.
Osteoporosis is a prevalent but asymptomatic condition that affects derly population, causing bone f fracture. Several methods have been developed and are of general hospitals to assess the bone quality and diagnose osteoporosis; however these are a cessed via referrals of general practitioners healthcare settings such as general practitioner or family doctors' clinics and cost constraints. This paper describes a new method that uses a medical exert testing stimuli, i.e. tapping, an electronic stethos and intelligent signal processing based rosis.
An advantageous approach to DSP equalization of loudspeakers is proposed in this paper adopting spatial averages of complex responses acquired from 3D balloon measurements. Alignment of the off-axis impulses responses with the on-axis impulse responses are accomplished using a cross-correlation technique prior to spatial averaging to attain meaningful statistics of magnitude and phase responses. This is performed over a pre-defined listening window from the complete loudspeaker response balloons (both magnitude and phase). The resulted average of the complex response within a suitably defined listening window is used to obtain, via the least mean square adaptive technique, an inverse filter that corrects the linear behaviour of the loudspeaker.
In the field of audio classification, audio signals may be broadly divided into three classes: speech, music and events. Most studies, however, neglect that real audio soundtracks can have any combination of these classes simultaneously. In this study, a novel feature, “Entrocy”, is proposed for the detection of music both in pure form and overlapping with the other audio classes. Entrocy is defined as the variation of the information (or entropy) in an audio segment over time. Segments which contain music were found to have lower Entrocy since there are fewer abrupt changes over time. We have also compared Entrocy with existing music detection features and the entrocy showing a promising performance. Keywords—Music detection, audio content analysis, audio indexing, Entropy, real world audio classification.
The speech transmission index (STI) is one of the most widely used and standardized methods for objective prediction of speech intelligibility of transmission channels. The original verification of the relationship between the STI and the intelligibility for the English language was published in 1987. The methodology employed then for the listening tests and the different input spectrum recommended today by the current STI method suggest that the relationship STI vs. speech intelligibility needs to be verified for the English language. This paper presents a new verification of the current STI for the English language with binaural listening and the speech materials presented to the listeners in a real room from a real sound source. Two hundred and ten subjects participated in speech intelligibility tests designed to replicate real-life listening conditions. The speech materials were contaminated with natural reverberation, noise, band pass limiting and echoes, and the listeners were not familiarized with the speech materials before the tests, in order to better replicate everyday situations. Results showed lower intelligibility scores than the earlier verification presented in Annex E of the current STI standard (IEC60268–16:2011) for most of the scenarios that were investigated. The correlation between STI and speech intelligibility was also investigated for the English language with a newly proposed male speech spectrum. The accuracy of STI intelligibility prediction with the new male spectrum was found higher than that attained with the current IEC specified male spectrum. These new findings give new and additional insights into the interpretation of intelligibility using the STI values.
Wind induced noise is one of the major concerns of outdoor acoustic signal acquisition. It affects many field measurement and audio recording scenarios. Filtering such noise is known to be difficult due to its broadband and time varying nature. In this paper, a new method to mitigate wind induced noise in microphone signals is developed. Instead of applying filtering techniques, wind induced noise is statistically separated from wanted signals in a singular spectral subspace. The paper is presented in the context of handling microphone signals acquired outdoor for acoustic sensing and environmental noise monitoring or soundscapes sampling. The method includes two complementary stages, namely decomposition and reconstruction. The first stage decomposes mixed signals in eigen-subspaces, selects and groups the principal components according to their contributions to wind noise and wanted signals in the singular spectrum domain. The second stage reconstructs the signals in the time domain, resulting in the separation of wind noise and wanted signals. Results show that microphone wind noise is separable in the singular spectrum domain evidenced by the weighted correlation. The new method might be generalized to other outdoor sound acquisition applications.
Osteoporosis is a prevalent but asymptomatic condition that affects a large population of the elderly, resulting in a high risk of fracture. Several methods have been developed and are available in general hospitals to indirectly assess the bone quality in terms of mineral material level and porosity. In this paper we describe a new method that uses a medical reflex hammer to exert testing stimuli, an electronic stethoscope to acquire impulse responses from tibia, and intelligent signal processing based on artificial neural network machine learning to determine the likelihood of osteoporosis. The proposed method makes decisions from the key components found in the time-frequency domain of impulse responses. Using two common pieces of clinical apparatus, this method might be suitable for the large population screening tests for the early diagnosis of osteoporosis, thus avoiding secondary complications. Following some discussions of the mechanism and procedure, this paper details the techniques of impulse response acquisition using a stethoscope and the subsequent signal processing and statistical machine learning algorithms for decision making. Pilot testing results achieved over 80% in detection sensitivity.
Osteoporosis is an asymptomatic bone condition that affects a large proportion of the elderly population around the world, resulting in increased bone fragility and increased risk of fracture. Previous studies had shown that the vibroacoustic response of bone can indicate the quality of the bone condition. Therefore, the aim of the authors’ project is to develop a new method to exploit this phenomenon to improve detection of osteoporosis in individuals. In this paper a method is described that uses a reflex hammer to exert testing stimuli on a patient’s tibia and an electronic stethoscope to acquire the impulse responses. The signals are processed as mel frequency cepstrum coefficients and passed through an artificial neural network to determine the likelihood of osteoporosis from the tibia’s impulse responses. Following some discussions of the mechanism and procedure, this paper details the signal acquisition using the stethoscope and the subsequent signal processing and the statistical machine learning algorithm. Pilot testing with 12 patients achieved over 80% sensitivity with a false positive rate below 30% and accuracies in the region of 70%. An extended dataset of 110 patients achieved an error rate of 30% with some room for improvement in the algorithm. By using common clinical apparatus and strategic machine learning, this method might be suitable as a large population screening test for the early diagnosis of osteoporosis, thus avoiding secondary complications.
Semantic data can be extracted from soundtracks to enable automated audio content analysis and effective metadata generation, which enables scene analysis, indexing and search. Existing methods classify audio at a high level into speech, music or event sound segments exclusively, discarding content information in overlapped parts. Taking into account the fact that many real soundtracks include overlapped types of audio components, nonexclusive segmentation and indexing are essential pre-processors for reliable and non-lossy audio information mining. This paper argues the importance of nonexclusive segmentation, identifies the challenges imposed by overlapped classes, and proposes the use of meta features and random forests to detect and segment overlapped audio. Proposed algorithms are presented and validation results discussed. Estimated statistical distributions about zero crossing, RMS values, discrete cosine transform (DCT) coefficients and entropies calculated from consecutive 20 ms analytical frames over a 1 second period seem to be adequate for suitable classifiers to discriminate and identify the content in question. DCT spectra of entropy further improve the performance but only slightly.
Reliability of Speaker Recognition (SR) is crucial for critical applications, especially in adverse acoustic conditions. Ambient noises and their variations represent a significant challenge for such applications. In this paper, a new technique is proposed to address the issue of performance degradation in noisy environments. Based on the estimation of the signal to noise ratio (SNR) and profile of the ambient noise from input signals, the proposed method re-trains the enrolment model for the claim speaker to generate new noisy models that adapt to the noise profile. This technique is termed "training on the fly". Evaluation results show notable enhancement in performance in terms of the reduction of equal error rates over a range of SNRs and different types of noise.
This Soundtracks are information rich; knowledge can be extracted and discovered from them. Audio signals may be broadly divided into three classes: speech, music and events. Dedicated recognition algorithms are typically used to extract semantic data for further knowledge discovery. A preprocessor to classify and segment audio into these three classes is essential. Current practice often neglects that audio data from the real world can have any combination of these classes simultaneously. This can result in information loss, thus compromising the knowledge discovery. Singular Spectrum Analysis (SSA) is further developed to reduce the degree of overlap between classes of audio content, in order to mitigate such information losses and improve the performance of knowledge discovery. In particular the SSA method serves to mitigate the overlapping ratio between speech and music in the mixed soundtracks by generating two new soundtracks with a lower level of overlapping. Next, feature space is calculated for the output audio streams, and these are classified using random forests into either speech or music. The classification performance of overlapped soundtracks is effectively improved and singular spectrum analysis has been found to be an efficient way to discriminate speech/music in mixed soundtracks.