Thanks to the popularity of smartphones with high-quality cameras and social media platforms, an exceptional amount of image data is generated and shared daily. This visual data can provide unprecedented insights into daily life and can be used to help answer research questions in psychology. However, the traditional methods used to analyze visual data are burdensome and are either time-intensive (e.g., content analysis) or require technical training (e.g., developing and training deep learning models). Zero-shot learning, where a pretrained model is used without any additional training, requires less technical expertise and may be a particularly attractive method for psychology researchers aiming to analyze image data. In this tutorial, we aim to provide an overview and step-by-step guide on how to analyze visual data with zero-shot learning. Specifically, we demonstrate how to use two popular models (Contrastive Language-Image Pretraining and Large Language and Vision Assistant) to identify a beverage in an image from a data set where we manipulated the type of beverage present, the setting, and the prominence of the beverage in the image (foreground, midground, background). To guide researchers through this process, we provide open code and data on GitHub and as a Google Colab notebook. Finally, we discuss how to interpret and report accuracy, how to create a validation data set, what steps need to be taken to implement the models with new data, and discuss future challenges and limitations of the method. To conclude, zero-shot learning requires less technical expertise and may be a particularly attractive method for psychology researchers aiming to analyze image data. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
更多