One of the biggest problems of the World Wide Web and pub has probably, several times, been experiencing different, but almost identical, or nearly identical content results in a simple Google search. In the case of publishers, duplicate content represents plagiarism or copyright infringement, therefore large financial losses. Google is seriously dealing with content duplication and may penalize duplicate sites or pages as a result. There are many software packages and services that identify and/or solve this problem, but they are not free. The current paper describes how to remove duplicate, or nearly identical, content from a database, using utilities that implement specific algorithms. One of these utilities is Natural Language Toolkit (NLTK) written in Python. NLTK is a platform for creating programs in Python, to work with data in human language. The way to detect duplicate content is described for the applications running on open source, Linux and Windows operating systems.