Multi-label data stream classification has emerged as a critical learning paradigm for real-world applications in which data arrive continuously, labels are not mutually exclusive, and the underlying data distribution evolves over time.However, this task faces problems inherent to dynamic environments, such as the continuous arrival of data at high speed and volume, changes in data distribution (concept drift), the emergence of new labels (concept evolution), and the latency in the arrival of ground-truth labels.This systematic literature review presents an in-depth analysis of proposals for multi-label data stream classification.We characterize the latest methods published between 2016 and 2025, provide a comprehensive overview, construct a thorough hierarchy, and discuss how each proposal addresses each problem.Furthermore, we discuss the adopted evaluation strategies and analyze the asymptotic complexity and resource consumption of the methods.Finally, we identify the main gaps and offer recommendations for future research directions in the field.The review reveals a strong methodological focus on concept drift adaptation, often at the expense of other critical challenges.In particular, label latency and evolving label spaces are rarely addressed in a principled or scalable manner, despite their relevance in real-world deployments.Overall, the field shows methodological maturity in drift-aware learning but lacks integrated solutions that jointly address delayed supervision, evolving labels, and efficiency constraints.