End-to-end Multi-modal Low-resourced Speech Keywords Recognition Using Sequential Conv2D Nets | AMiner