Brain-Computer Interface (BCI) applications provide a direct way to map human brain activity onto the control of external devices, without a need for physical movements. These systems, crucial for medical applications and also useful for non-medical applications, predominantly use EEG signals recorded non-invasively, for system control, and require algorithms to translate signals into commands. Traditional BCI applications heavily depend on algorithms tailored to specific behavioral paradigms and on data collection using EEG systems with multiple channels. This complicates usability, comfort, and affordability. Moreover, the limited availability of extensive training datasets limits the development of robust models for classifiying collected data into behavioral intents. To address these challenges, we introduce an end-to-end EEG classification framework that employs a pre-trained Convolutional Neural Network (CNN) and a Transformer, initially designed for image processing, applied here for spatiotemporal representation of EEG data, and combined with a custom developed automated EEG channel selection algorithm to identify the most informative electrodes for the process, thus reducing data dimensionality, and easing subject comfort, along with improved classification performance of EEG data onto subjects intent. We evaluated our model using two benchmark datasets, the EEGmmidb and the OpenMIIR. We achieved superior performance compared to existing state-of-theart EEG classification methods, including the commonly used EEGnet. Our results indicate a classification accuracy improvement of 7% on OpenMIIR and 1% on EEGmmidb, reaching averages of 81% and 75%, respectively. Importantly, these improvements were obtained with fewer recording channels and less training data, demonstrating a framework that can support a more efficient approach to BCI tasks in terms of the amount of training data and the simplicity of the required hardware system needed for brain signals. This study not only advances the field of BCI but also suggests a scalable and more affordable framework for BCI applications. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study did not receive any funding ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced are available online at openmiir and physionet
Voice data is increasingly being used in modern digital communications, yet there is still a lack of comprehensive tools for automated voice analysis and characterization. To this end, we developed the VANPY (Voice Analysis in Python) framework for automated pre-processing, feature extraction, and classification of voice data. The VANPY is an open-source end-to-end comprehensive framework that was developed for the purpose of speaker characterization from voice data. The framework is designed with extensibility in mind, allowing for easy integration of new components and adaptation to various voice analysis applications. It currently incorporates over fifteen voice analysis components - including music/speech separation, voice activity detection, speaker embedding, vocal feature extraction, and various classification models. Four of the VANPY's components were developed in-house and integrated into the framework to extend its speaker characterization capabilities: gender classification, emotion classification, age regression, and height regression. The models demonstrate robust performance across various datasets, although not surpassing state-of-the-art performance. As a proof of concept, we demonstrate the framework's ability to extract speaker characteristics on a use-case challenge of analyzing character voices from the movie "Pulp Fiction." The results illustrate the framework's capability to extract multiple speaker characteristics, including gender, age, height, emotion type, and emotion intensity measured across three dimensions: arousal, dominance, and valence.
The unprecedented growth in video conferencing usage is accompanied by multiple security and privacy threats. Importantly, protecting users' privacy is not always in their own hands. Posting meeting images affects all participants, leading to an easy collection of personal data including age, gender and linkage with participation in other meetings. Here, we explored privacy issues that may be at risk by attending virtual meetings. We extracted private information from collage images of meeting participants that are publicly posted online. We used image processing, text recognition tools, as well as social network analysis to explore our curated dataset of over 15 700 collage images, and over 142 000 face images of meeting participants. We demonstrate that video conference users are facing prevalent security and privacy threats. Our results indicate that it is relatively easy to collect thousands of publicly available images of video conference meetings and extract personal information about the participants, including their face images, age, gender, usernames, and even full names. This type of data can vastly and easily jeopardize people's security and privacy both in the online and real-world, affecting not only adults but also more vulnerable segments of society, such as children and older adults. Finally, we show that cross-referencing facial image data with social network data may put participants at additional privacy risks they may not be aware of and that it is possible to identify users that appear in several video conference meetings, thus providing a potential to maliciously aggregate different sources of information about a target individual.
In the last decades, global awareness toward the importance of diverse representation has been increasing. The lack of diversity and discrimination toward minorities did not skip the film industry. Here, we examine ethnic bias in the film industry through commercial posters, the industry’s primary advertisement medium for decades. Movie posters are designed to establish the viewer’s initial impression. We developed a novel approach for evaluating ethnic bias in the film industry by analyzing nearly 125,000 posters using state-of-the-art deep learning models. Our analysis shows that while ethnic biases still exist, there is a trend of reduction of bias, as seen by several parameters. Particularly in English-speaking movies, the ethnic distribution of characters on posters from the last couple of years is reaching numbers that are approaching the actual ethnic composition of the US population. An automatic approach to monitoring ethnic diversity in the film industry, potentially integrated with financial value, may be of significant use for producers and policymakers.
In recent years, video conferencing (VC) popularity has skyrocketed for a wide range of activities. As a result, the number of VC users surged sharply. The sharp increase in VC usage has been accompanied by various newly emerging privacy and security challenges. VC meetings became a target for various security attacks, such as Zoombombing. Other VC-related challenges also emerged. For example, during COVID lockdowns, educators had to teach in online environments struggling with keeping students engaged for extended periods. In parallel, the amount of available VC videos has grown exponentially. Thus, users and companies are limited in finding abnormal segments in VC meetings within the converging volumes of data. Such abnormal events that affect most meeting participants may be indicators of interesting points in time, including security attacks or other changes in meeting climate, like someone joining a meeting or sharing a dramatic content. Here, we present a novel algorithm for detecting abnormal events in VC data. We curated VC publicly available recordings, including meetings with interruptions. We analyzed the videos using our algorithm, extracting time windows where abnormal occurrences were detected. Our algorithm is a pipeline that combines multiple methods in several steps to detect users' faces in each video frame, track face locations during the meeting and generate vector representations of a facial expression for each face in each frame. Vector representations are used to monitor changes in facial expressions throughout the meeting for each participant. The overall change in meeting climate is quantified using those parameters across all participants, and translating them into event anomaly detection. This is the first open pipeline for automatically detecting anomaly events in VC meetings. Our model detects abnormal events with 92.3% precision over the collected dataset.
Video conferencing (VC) has become increasingly popular, bringing new challenges in privacy and security, one notable example is of Zoombombing. Furthermore, other issues related to VC usage have emerged, such as keeping students involved. Identifying abnormal segments in VC meetings in vast data is a challenging task. Here, we introduce a novel algorithm to detect such anomalies in VC automatically. By analyzing publicly available VC recordings, our algorithm tracks and analyzes changes in participants’ facial expressions to identify and quantity overall meeting climate changes. We demonstrate performance of 92.3% precision in anomaly detection on the collected dataset. Our model offers a pioneering solution for recognizing abnormal events in VC meetings.
provide value insight that policymakers could use in considering where to lead European CS when distributing budgets, by either encouraging leading fields and collaborations or strengthening those that fall behind. In the last decade, 30% of worldwide CS publications were of European origin. For comparison, North America leads with 33% and Asia provides 30% of worldwide CS publications. Interestingly, in terms of worldwide attention, Europe also holds 30% of worldwide CS citations, while North America impressively approaches nearly half (47%) of them, Asia following third Publication volumes are important, but do not necessarily imply quality or impact. One problem of increasing global concern is that publication volume has been dramatically increasing, affected by the race to publish. We therefore analyzed separately high/low impact publications (papers with fewer than five citations considered low impact). The spatial distribution of citations is similar to volume mapping, with some smaller countries (like Switzerland and Netherlands) of high impact, and East-West impact gaps, possibly resulting from historical separation. Since results may be affected by countries’ size, reflected also by numbers of publishing institutes, we computed average citations per institute in each country.d This approach
The spread of the Red Palm Weevil has dramatically affected date growers, homeowners and governments, forcing them to deal with a constant threat to their palm trees. Early detection of palm tree infestation has been proven to be critical in order to allow treatment that may save trees from irreversible damage, and is most commonly performed by local physical access for individual tree monitoring. Here, we present a novel method for surveillance of Red Palm Weevil infested palm trees utilizing state-of-the-art deep learning algorithms, with aerial and street-level imagery data. To detect infested palm trees we analyzed over 100,000 aerial and street-images, mapping the location of palm trees in urban areas. Using this procedure, we discovered and verified infested palm trees at various locations.
The COVID-19 pandemic outbreak, with its related social distancing and shelter-in-place measures, has dramatically affected ways in which people communicate with each other, forcing people to find new ways to collaborate, study, celebrate special occasions, and meet with family and friends. One of the most popular solutions that have emerged is the use of video conferencing applications to replace face-to-face meetings with virtual meetings. This resulted in unprecedented growth in the number of video conferencing users. In this study, we explored privacy issues that may be at risk by attending virtual meetings. We extracted private information from collage images of meeting participants that are publicly posted on the Web. We used image processing, text recognition tools, as well as social network analysis to explore our web crawling curated dataset of over 15,700 collage images, and over 142,000 face images of meeting participants. We demonstrate that video conference users are facing prevalent security and privacy threats. Our results indicate that it is relatively easy to collect thousands of publicly available images of video conference meetings and extract personal information about the participants, including their face images, age, gender, usernames, and sometimes even full names. This type of extracted data can vastly and easily jeopardize people's security and privacy both in the online and real-world, affecting not only adults but also more vulnerable segments of society, such as young children and older adults. Finally, we show that cross-referencing facial image data with social network data may put participants at additional privacy risks they may not be aware of and that it is possible to identify users that appear in several video conference meetings, thus providing a potential to maliciously aggregate different sources of information about a target individual.
Brain computer interface applications, developed for both healthy and clinical populations, critically depend on decoding brain activity in single trials. The goal of the present study was to detect distinctive spatiotemporal brain patterns within a set of event related responses. We introduce a novel classification algorithm, the spatially weighted FLD-PCA (SWFP), which is based on a two-step linear classification of event-related responses, using fisher linear discriminant (FLD) classifier and principal component analysis (PCA) for dimensionality reduction. As a benchmark algorithm, we consider the hierarchical discriminant component Analysis (HDCA), introduced by Parra, et al. 2007. We also consider a modified version of the HDCA, namely the hierarchical discriminant principal component analysis algorithm (HDPCA). We compare single-trial classification accuracies of all the three algorithms, each applied to detect target images within a rapid serial visual presentation (RSVP, 10 Hz) of images from five different object categories, based on single-trial brain responses. We find a systematic superiority of our classification algorithm in the tested paradigm. Additionally, HDPCA significantly increases classification accuracies compared to the HDCA. Finally, we show that presenting several repetitions of the same image exemplars improve accuracy, and thus may be important in cases where high accuracy is crucial.
We examined the effects of aging on visuo-spatial attention. Participants performed a bi-field visual selective attention task consisting of infrequent target and task-irrelevant novel stimuli randomly embedded among repeated standards in either attended or unattended visual fields. Blood oxygenation level dependent (BOLD) responses to the different classes of stimuli were measured using functional magnetic resonance imaging. The older group had slower reaction times to targets, and committed more false alarms but had comparable detection accuracy to young controls. Attended target and novel stimuli activated comparable widely distributed attention networks, including anterior and posterior association cortex, in both groups. The older group had reduced spatial extent of activation in several regions, including prefrontal, basal ganglia, and visual processing areas. In particular, the anterior cingulate and superior frontal gyrus showed more restricted activation in older compared with young adults across all attentional conditions and stimulus categories. The spatial extent of activations correlated with task performance in both age groups, but the regional pattern of association between hemodynamic responses and behavior differed between the groups. Whereas the young subjects relied on posterior regions, the older subjects engaged frontal areas. The results indicate that aging alters the functioning of neural networks subserving visual attention, and that these changes are related to cognitive performance.
In complex natural environments, auditory and visual information often have to be processed simultaneously. Previous functional magnetic resonance imaging (fMRI) studies focused on the spatial localization of brain areas involved in audiovisual (AV) information processing, but the temporal characteristics of AV information flow in these regions remained unclear. In this study, we used fMRI and a novel information–theoretic approach to study the flow of AV sensory information. Subjects passively perceived sounds and images of objects presented either alone or simultaneously. Applying the measure of mutual information, we computed for each voxel the latency in which the blood oxygenation level-dependent signal had the highest information content about the preceding stimulus. The results indicate that, after AV stimulation, the earliest informative activity occurs in right Heschl's gyrus, left primary visual cortex, and the posterior portion of the superior temporal gyrus, which is known as a region involved in object-related AV integration. Informative activity in the anterior portion of superior temporal gyrus, middle temporal gyrus, right occipital cortex, and inferior frontal cortex was found at a later latency. Moreover, AV presentation resulted in shorter latencies in multiple cortical areas compared with isolated auditory or visual presentation. The results provide evidence for bottom-up processing from primary sensory areas into higher association areas during AV integration in humans and suggest that AV presentation shortens processing time in early sensory cortices.
A new approach for analysis of event-related fMRI (BOLD) signals is proposed. The technique is based on measures from information theory and is used both for spatial localization of task-related activity, as well as for extracting temporal information regarding the task-dependent propagation of activation across different brain regions. This approach enables whole brain visualization of voxels (areas) most involved in coding of a specific task condition, the time at which they are most informative about the condition, as well as their average amplitude at that preferred time. The approach does not require prior assumptions about the shape of the hemodynamic response function (HRF) nor about linear relations between BOLD response and presented stimuli (or task conditions). We show that relative delays between different brain regions can also be computed without prior knowledge of the experimental design, suggesting a general method that could be applied for analysis of differential time delays that occur during natural, uncontrolled conditions. Here we analyze BOLD signals recorded during performance of a motor learning task. We show that, during motor learning, the BOLD response of unimodal motor cortical areas precedes the response in higher-order multimodal association areas, including posterior parietal cortex. Brain areas found to be associated with reduced activity during motor learning, predominantly in prefrontal brain regions, are informative about the task typically at significantly later times.
To what extent do all brains work alike during natural conditions? We explored this question by letting five subjects freely view half an hour of a popular movie while undergoing functional brain imaging. Applying an unbiased analysis in which spatiotemporal activity patterns in one brain were used to “model” activity in another brain, we found a striking level of voxel-by-voxel synchronization between individuals, not only in primary and secondary visual and auditory areas but also in association cortices. The results reveal a surprising tendency of individual brains to “tick collectively” during natural vision. The intersubject synchronization consisted of a widespread cortical activation pattern correlated with emotionally arousing scenes and regionally selective components. The characteristics of these activations were revealed with the use of an open-ended “reverse-correlation” approach, which inverts the conventional analysis by letting the brain signals themselves “pick up” the optimal stimuli for each specialized cortical area.
Synaptic transmission between pairs of excitatory neurones in layers V (N= 38) or IV (N= 6) of somatosensory cortex was examined in a parasagittal slice preparation obtained from young Wistar rats (14–18 days old). A combined experimental and theoretical approach reveals two characteristics of short‐term synaptic depression. Firstly, as well as a release‐dependent depression, there is a release‐independent component that is evident in smaller postsynaptic responses even following failure to release transmitter. Secondly, recovery from depression is activity dependent and is faster at higher input frequencies. Frequency‐dependent recovery is a Ca2+‐dependent process and does not reflect an underlying augmentation. Frequency‐dependent recovery and release‐independent depression are correlated, such that at those connections with a large amount of release‐independent depression, recovery from depression is faster. In addition, both are more pronounced in experiments performed at physiological temperatures. Simulations demonstrate that these homeostatic properties allow the transfer of rate information at all frequencies, essentially linearizing synaptic responses at high input frequencies.
Spike-frequency adaptation in neocortical pyramidal neurons was examined using the whole cell patch-clamp technique and a phenomenological model of neuronal activity. Noisy current was injected to reproduce the irregular firing typically observed under in vivo conditions. The response was quantified by computing the poststimulus histogram (PSTH). To simulate the spiking activity of a pyramidal neuron, we considered anintegrate-and-fire model to which an adaptation current was added. A simplified model for the mean firing rate of an adapting neuron under noisy conditions is also presented. The mean firing rate model provides a good fit to both experimental and simulation PSTHs and may therefore be used to study the response characteristics of adapting neurons to various input currents. The models enable identification of the relevant parameters of adaptation that determine the shape of the PSTH and allow the computation of the response to any change in injected current. The results suggest that spike frequency adaptation determines a preferred frequency of stimulation for which the phase delay of a neuron's activity relative to an oscillatory input is zero. Simulations show that the preferred frequency of single neurons dictates the frequency of emergent population rhythms in large networks of adapting neurons. Adaptation could therefore be one of the crucial factors in setting the frequency of population rhythms in the neocortex.
Synaptic transmission in the neocortex is dynamic, such that the magnitude of the postsynaptic response changes with the history of the presynaptic activity. Therefore each response carries information about the temporal structure of the preceding presynaptic input spike train. We quantitatively analyze the information about previous interspike intervals, contained in single responses of dynamic synapses, using methods from information theory applied to experimentally based deterministic and probabilistic phenomenological models of depressing and facilitating synapses. We show that for any given dynamic synapse, there exists an optimal frequency of presynaptic spike firing for which the information content is maximal; simple relations between this optimal frequency and the synaptic parameters are derived. Depressing neocortical synapses are optimized for coding temporal information at low firing rates of 0.5-5 Hz, typical to the spontaneous activity of cortical neurons, and carry significant information about the timing of up to four preceding presynaptic spikes. Facilitating synapses, however, are optimized to code information at higher presynaptic rates of 9-70 Hz and can represent the timing of over eight presynaptic spikes.
We compared the spike activity of individual neurons in the Aplysia abdominal ganglion with the movement of the gill during the gill-withdrawal reflex. We discriminated four populations that collectively encompass approximately half of the active neurons in the ganglion: (1) second-order sensory neurons that respond to the onset and offset of stimulation of the gill and are active before the movement starts; (2) neurons whose activity is correlated with the position of the gill and typically have a tonic output during gill withdrawal; (3) neurons whose activity is correlated with the velocity of the movement and typically fire in a phasic manner; and (4) neurons whose activity is correlated with both position and velocity. A reliable prediction of the position of the gill is achieved only with the combined output of 15-20 neurons, whereas a reliable prediction of the velocity depends on the combined output of 40 or more cells.
Misha Tsodyks合作论文数Department of Neurobiology
Weizmann Institute of Science5
Amir B. Geva合作论文数Electrical and Computer Engineering Department
Ben-Gurion University of the Negev1