This paper presents a generic event detection system evaluated in the Surveillance Event Detection (SED) task of TRECVID 2011 campaign. We investigate a generic statistical approach with spatio-temporal features applied to seven event classes, which were defined by the SED task. This approach is based on local spatio-temporal descriptors, which is named as MoSIFT and generated by pair-wise video frames. Visual vocabularies are generated by cluster centers of MoSIFT features, which were sampled from the event part video clips. We also estimated the spatial distribution of actions by over generated person detection and background subtraction. Different slide window sizes and steps were adopted for different events by events’ duration prior. Several sets of one-against-all action classifiers were trained using cascade non-linear SVMs and Random Forest, which could improve the classification performance in unbalanced data just like the SED datasets. 9 runs results were presented with variations in i) Slide window size ii) step size of BOW, iii) classifier threshold and iv) classifiers. The performance shows improvement over last year on the event detection task.