Data preparation modules are ubiquitous and are used to perform, amongst other things, operations such as record retrieval, format transformation, data combination to name a few. To assist scientists in the task of discovering suitable modules, semantic annotations can be leveraged. Experience suggests, however, that while such annotations are useful in describing the inputs and outputs of a module, they fail in crisply describing the functionality performed by the module. To overcome this issue, we outline in this poster paper a solution that utilizes semantic annotations describing the inputs and outputs of modules together with data examples that characterize modules’ behavior as ingredients for querying data preparation modules. Data examples are constructed using retrospective provenance of module executions. The discovery strategy that we devised is iterative in that it allows scientists to explore existing modules by providing feedback on data examples.