We address the content-identification problem by modeling it as a multi-class classification problem. The goal is to pave the way and establish a general framework to incorporate the powerful algorithms of the machine learning literature in learning from data into this problem. Through this end, a particular successful approach, linked with the coding theory known as ECOC is considered and studied. We argue that the conventional codings used in this approach are suboptimal by analyzing the problem from an information-theoretic viewpoint. We then advise the use of our recently proposed method for this problem. The ECOC approach converts the multi-class problem to several binary problems. We consider these equivalent binary classification tasks in more details and use the Gaussian Mixture Models instead of SVM’s. This latter brings significant reduction in complexity by having an assumption on the distributions.