Variable stars play a crucial role in advancing knowledge of stellar evolution, Galactic structure, and cosmological distance scales. The advent of massive datasets from modern time-domain surveys (e.g., Zwicky Transient Facility, All-Sky Automated Survey for Supernovae, or Gaia) has made the classification of these stars increasingly complex. Traditional methods depend extensively on manual feature engineering, a process that is both labor-intensive and difficult to scale, thereby impeding timely scientific progress. To address these limitations, we propose a novel classification framework that treats photometric light curves as structured sequences and fine-tunes a pretrained large language model (LLM) for variable star classification with only a small amount of labeled data. We design a unified tokenization scheme that integrates both time-series measurements and physical parameters (period, parallax, and color index) into a single textual prompt, thereby streamlining traditional feature engineering. When evaluated on a multisurvey dataset comprising 19,580 labeled variable stars across nine classes, the optimal model achieves an overall accuracy of 0.98 and a weighted F _1 score of 0.98. We also verify the complementary contributions of each physical parameter through ablation studies. Our study indicates that LLM-based sequence modeling constitutes an accurate, scalable, and flexible alternative to conventional methods, with strong potential for integration into future large-scale classification of variable stars.
更多
查看译文
关键词
Light curve classification,Time series analysis,Variable stars,Stellar types