Offensive speech covers many different ways someone may harm another such as humiliation, criticism, yelling, or verbal abuse. Such behavior results in negative consequences with possible dangerous impacts like biasing people’s thoughts, spreading racism, discrimination and even physical violence among citizens. While offensive speech takes place online and offline and in all forms of social interaction, our work contributes at filtering the alarming diffusion of such behaviors in televisiondebate programs by proposing a speech-based offensive behavior detection system. To do so, we investigated the use of the Vera Am Mittag (VAM) corpus which is a collection of recordings taken from a German TV talk show. From these, a mixture of Mel Frequency Cepstral Coefficients (MFCC) and Stationary Wavelet Transform (SWT) based features was extracted and several feature selection techniques were applied for capturing the most relevant features. Finally, both of K-Nearest Neighbors (KNN) and deep learning Convolutional Neural Networks (CNNs) were employed for classification. Results highlight that the best performance is associated with CNNs whith feature selection, reaching a classification accuracy of $97.21 \%$.