The conversion between visibilities and images is a fundamental yet computationally expensive operation in radio interferometric imaging. Although existing algorithms combining high-precision convolution kernels with w-stacking have achieved imaging precision on the order of 10-12, the efficient processing of massive datasets from next-generation arrays such as the SKA telescope remains a formidable challenge, constrained by the hardware limitations of heterogeneous platforms. In this work, we introduce the baseline separation paradigm (BSP), an innovative imaging framework designed to optimize workload distribution. The core strategy of BSP is to partition the dataset into two regimes: utilizing the GPU for the massive volume of short-baseline data to bypass the memory bottleneck, while leveraging the CPU for long baselines to preserve high-resolution details, with both regimes utilizing grids strictly minimized to their respective spatial frequency limits. This strategy effectively resolves the conflict between the limited memory capacity of accelerators and the large grid sizes required for wide-field imaging. We implemented a prototype to evaluate the algorithm performance. Experimental results show that, for SKA1-Low-scale simulations, BSP achieves a precision of 10-11, comparable to ducc0.wgridder, while providing a substantial speedup. This approach provides a scalable solution for future exascale astronomical data processing.
Chemically peculiar (CP) stars offer a unique observational sample for exploring key processes in stellar evolution. Leveraging the massive amounts of data obtained from large-scale sky surveys, deep learning technology has significantly advanced research in identifying special celestial objects in large datasets in recent years. In this study, we developed a hybrid deep learning approach that integrates a long short-term memory network with a convolutional neural network to search for CP1/CP2/CP3 stars in low-resolution spectra of LAMOST DR12 (v1.0). We first trained a binary classification model to distinguish CP stars from non-CP stars, achieving an accuracy of 95.92%. Applying this model to the LAMOST DR12 (v1.0) dataset, we identified 79,216 CP star candidates. We then constructed a three-class classification model to distinguish between CP1, CP2, and CP3, achieving an accuracy rate of 97.34%. This model further identified 59,333 CP1 stars, 15,179 CP2 stars, and 1424 CP3 stars from the CP star candidates. After cross validation, 51,241 stars were consistent with previously reported CP stars based on crossmatching results, with 24,695 newly identified CP1/CP2/CP3 candidates, including 16,432 new CP1 candidates, 7132 CP2 candidates, and 1131 CP3 candidates. By combining visual inspection with the MKCLASS code, we ultimately identified 2899 CP1 stars, 4536 CP2 stars, and 747 CP3 stars.
Be stars are rapidly rotating B-type stars that exhibit Balmer emission lines in their optical spectra, which makes them crucial for studying stellar evolution and circumstellar disk structures. In this study, we performed a systematic identification of Be stars using low-resolution spectroscopic data from Large Sky Area Multi-Object Fiber Spectroscopic Telescope Data Release 11. We constructed a dataset and developed a hybrid identification model combining long short-term memory networks and convolutional neural networks, achieving a testing accuracy of 97.83%. The trained model was applied to spectra with signal-to-noise ratios greater than 10 in the g band, yielding 55,667 B-type candidates. After further validation using the MKCLASS tool, manual verification, and removing duplicated observational data, 26,925 B-type candidates were confirmed. These B-type candidates were subsequently crossmatched with published H alpha emission-line catalogs to identify emission-line B-type objects, yielding a sample of 5384 Be star candidates. By comparing these samples with existing Be star databases, we identified 2881 previously reported objects and 2503 newly discovered Be stars. Based on infrared color criteria, 1658 of these new detections were classified as classical Be stars and 32 as Herbig Be stars. The remaining 813 objects were categorized as inconclusive, as they either exhibited ambiguous classifications satisfying only partial infrared criteria or lacked sufficient photometric information for definitive categorization.
This study used Bayesian analysis alongside various statistical methods to examine the young open cluster NGC 7067, utilizing Gaia DR3 and Two Micron All Sky Survey data. We identified probable members of the cluster using density-based spatial clustering of applications with noise and used Bayesian Analysis for Stellar Evolution with nine variables software to fit the age, [Fe/H], distance and Av of the cluster. Then we used the fitting results to simulate the mass of each member star and identified binary stars and carried out mass analysis, reddening analysis, and orbit analysis. We estimate the log(Age) of the cluster to be about 7.11-0.058+0.059 , [Fe/H] to be -0.082-0.031+0.034 , distance to be 5267.4-74.61+85.53pc , and Av to be 2.565-0.022+0.023 . The dynamical relaxation time of the cluster in log(Age) is 7.52. This makes NGC 7067 a good object for studying the causes of mass segregation and the evolution of star formation. The cluster's orbit indicates that it was born very close to the Galactic disk.
Very metal-poor (VMP; [Fe/H] < -2) stars are critical tracers for understanding early star formation and Galactic chemical evolution. However, identifying these rare objects from the massive datasets generated by the Dark Energy Spectroscopic Instrument (DESI) presents significant challenges in efficiency and precision due to the scarcity of high-fidelity labels and low signal-to-noise ratios in the metal-poor regime. To address this, we propose a novel dual-model deep learning framework that integrates a 1D-ResNet binary classifier with a specialized parameter regression model. Leveraging a transfer learning strategy with high-quality labels from APOGEE and the Large Sky Area Multi-object Fiber Spectroscopic Telescope, we optimized the framework for DESI spectra. The classification model achieves an accuracy of 97.87%, while the regression model predicts stellar metallicity with an rms error of 0.093 dex. Applying this framework to the high-quality DESI-HQ-DATA dataset and applying a strict temperature cut (T-eff >= 4500 K) to avoid extrapolation in the cool dwarf regime, we constructed a highly purified catalog of 2569 high-confidence VMP candidates. Internal validation demonstrates a significant improvement in consistency between DESI pipeline measurements, and external crossmatching with Gaia XP spectra confirms the reliability of our metallicity estimates. Finally, compared to existing compilations, this work contributes 1377 new VMP candidates, providing a robust and statistically significant sample for future studies of the Galactic halo and ancient stellar populations.
White dwarfs (WDs) with infrared (IR) excesses probe dusty debris disks, low-mass companions, and the late-stage evolution of planetary and binary systems. Conventional searches usually rely on source-by-source spectral energy distribution (SED) fitting and visual inspection, which become time-consuming for the rapidly growing samples produced by large spectroscopic surveys. We develop a supervised multimodal deep learning framework for scalable preselection of IR-excess WD candidates in the Dark Energy Spectroscopic Instrument (DESI) Data Release 1. The model combines Pan-STARRS1 z - and y -band images, unWISE W1- and W2-band images, and atmospheric, astrometric, and photometric tabular features. On the internal validation set, the model achieved an area under the receiver operating characteristic curve of 0.9765, demonstrating effective separation of literature-reported IR-excess candidates from comparison WDs. Applied to the 10,988 objects in DESI-WD-SEARCH-DATA, the model selected 1886 first-stage candidates for subsequent validation. Image-based screening for Wide-field Infrared Survey Explorer (WISE)-scale contamination and composition-dependent SED validation identified 1041 objects satisfying the adopted excess criteria. Of these, 294 had sufficient photometric coverage for further assessment, and catalog-specific photometric-quality screening yielded a final catalog of 221 candidates, including 204 main-sample and 17 warning candidates. We also provide the complete list of 1886 first-stage candidates with flags recording the outcomes of subsequent screening steps. This catalog provides targets for future high-resolution IR imaging, spectroscopic follow-up, and studies of the physical origins of IR excesses around WDs.
Next-generation radio interferometers, with their enhanced spatial, temporal, and high-frequency resolution, will pose considerable challenges to data processing and storage. Frequency averaging reduces the volume of data, but would cause smearing of the bandwidth. The examination of bandwidth smearing is essential in the processing of radio interferometer data to achieve scientific objectives. Existing analysis methods can quantify parameters such as peak intensity loss in restored images under ideal assumptions. However, determining the shape of the source smearing in dirty images requires using complex observational simulation methods, which limits comprehensive studies of bandwidth smearing from frequency averaging and its optimal application. In this paper, we introduce a semianalytical method (SAM) and validate its accuracy through the simulation method. The SAM efficiently and precisely evaluates the impact of the smearing effect from frequency averaging for various telescope pointing positions, offering a plausible “distorted beam” representation of the non-phase-centered beam shape of point sources in the dirty image. Based on this beam shape, the specific shape of smearing after frequency averaging can be accurately and comprehensively described. This not only enables the determination of the extent of smearing, prediction of accuracy loss, and formulation of effective observation and data processing strategies prior to data processing, but also facilitates the development of more advanced frequency averaging methods.
Hunting for open clusters (OCs) in Gaia data is a hot topic for astronomical big data analysis. Significant progress has been made in searching for OCs in Gaia data using machine learning methods. Based on our previous research, we applied a hybrid unsupervised clustering algorithm (friends-of-friends and pyUPMASK) and a binary classification algorithm (Random Forests) to perform fine-grained blind searching beyond $5\:$kpc on Gaia DR3. After isochrone-fitting, cross-matching, and visual inspection, we obtained 2932 plausible candidate clusters, 436 of which have been published as candidates in other catalogs. We performed a Kings's model profile fitting and a comparative study of theoretical tidal radii and the observed mass-radius correlation to determine the physical reality of these OCs. After physical reality checks, we validated the remaining 872 candidate clusters using statistical analysis and dynamical binding distinction. Statistical analysis shows that the distributions of these candidates' proper motion, OC age, and metallicity are consistent with other related studies. We also analyzed the intrinsic dispersion of morphological features, sky maps, ages, and metallicities of these candidate clusters. Following rigorous verifications and analytical validations, we have ultimately identified 739 OC candidates with high confidence.
Studying open clusters (OCs) is essential for a comprehensive understanding of the structure and evolution of the Milky Way. Many previous studies have systematically searched for OCs near the solar system within 1.2 kpc or 20 degrees of galactic latitude. However, few studies searched for OCs at higher galactic latitudes and deeper distances. In this study, based on a hybrid unsupervised clustering algorithm (Friends-of-Friends and pyUPMASK) and a binary classification algorithm (Random Forest), we extended the search region (i.e., galactic latitude |b|>=20 degrees) and performed a fine-grained blind search of Galactic clusters in Gaia DR3. After cross-matching, the newly discovered cluster candidates are fitted using isochrone fitting to estimate the main physical parameters (age and metallicity) of these clusters. These cluster candidates were then checked using manual visual inspection. Their statistical properties were compared with previously exposed cluster catalogs as well. In the end, we found 1,179 new clusters with considerable confidence within 5kpc.
Studying open clusters (OCs) may help to provide a comprehensive understanding of the structure and evolution of the Milky Way. Many previous studies have systematically searched for OCs near the solar system within 1.2 kpc or 20° of Galactic latitude. However, few studies have searched for OCs at higher Galactic latitudes and deeper distances. In this study, based on a hybrid unsupervised clustering algorithm (friends-of-friends and pyUPMASK) and a binary classification algorithm (random forest), we extended the search region (i.e., Galactic latitude b ≥ 20 ° ) and performed a fine-grained blind search for Galactic clusters in Gaia DR3. After crossmatching, the newly discovered cluster candidates are fitted using isochrone fitting to estimate the main physical parameters (i.e., age and metallicity) of these clusters. These cluster candidates were then checked using manual visual inspection. Their statistical properties were also compared with previously exposed cluster catalogs. We found 1179 new clusters with considerable confidence within 5 kpc.
The rapid development of new generation radio interferometers such as the Square Kilometer Array (SKA) has opened up unprecedented opportunities for astronomical research. However, anthropogenic Radio Frequency Interference (RFI) from communication technologies and other human activities severely affects the fidelity of observational data. It also significantly reduces the sensitivity of the telescopes. We proposed a robust Convolutional Neural Network (CNN) model to identify RFI based on machine learning methods. We overlaid RFI on the simulation data of SKA1-LOW to construct three visibility function datasets. One dataset was used for modeling, and the other two were used for validating the model's usability. The experimental results show that the Area Under the Curve (AUC) reaches 0.93, with satisfactory accuracy and precision. We then further investigated the effectiveness of the model by identifying the RFI in the actual observational data from LOFAR and MeerKAT. The results show that the model performs well. The overall effectiveness is comparable to AOFlagger software and provides an improvement over existing methods in some instances.
Hot subdwarf star is a particular type of star that is crucial for studying binary evolution and atmospheric diffusion processes. In recent years, identifying Hot subdwarfs by machine learning methods has become a hot topic, but there are still limitations in automation and accuracy. In this paper, we proposed a robust identification method based on the convolutional neural network (CNN). We first constructed the dataset using the spectral data of LAMOS DR7-V1. We then constructed a hybrid recognition model including an 8-class classification model and a binary classification model. The model achieved an accuracy of 96.17% on the testing set. To further validate the accuracy of the model, we selected 835 Hot subdwarfs that were not involved in the training process from the identified LAMOST catalog (2428, including repeated observations) as the validation set. An accuracy of 96.05% was achieved. On this basis, we used the model to filter and classify all 10,640,255 spectra of LAMOST DR7-V1, and obtained a catalog of 2393 Hot subdwarf candidates, of which 2067 have been confirmed. We found 25 new Hot subdwarfs among the remaining candidates by manual validation. The overall accuracy of the model is 87.42%. Overall, the model presented in this study can effectively identify specific spectra with robust results and high accuracy, and can be further applied to the classification of large-scale spectra and the search of specific targets.
The execution framework would significantly decrease the operation costs and improve the execution performance, but developing applications of the execution framework is not an easy job. In this chapter, we present a series of examples to demonstrate how to design and develop execution framework technology-based applications.
SKA 科学数据处理产生的数据超出了所有已存在的分布式处理系统的处理能力,如何实现一个分布式执行框架是当前科学数据处理的一个重要研究内容。Spark 是非常成熟的一个商业框架,在互联网应用中被广泛应用,本文根据SKA项目进展要求,重点研究了如何将算法参考库(ARL)中的部分管线移植到Spark上执行。本文对部分实现过程进行了分析讨论,给出了相应的任务流程实现。最终结果表明,移植后代码生成结果符合预期,Spark能够满足部分数据分布式数据的要求,但迫切需要解决自身存在的一系列问题。
Astronomy has become one of the biggest consumers of computing resources in the past 10 years. Therefore, new computational solutions (both hardware and software) are emerging, dedicated to the various fields of science. For example, in radio astronomy, instruments are becoming extremely large, much like the square kilometer array and its precursor telescopes. Thus, it is expected that astronomical experiments will become larger in all dimensions: larger data collections, more accurate data analysis and processing, and more detailed results. The significant increase of these large-size experiments performed by instruments built around the world requires not only huge processing power, but also clever system design. We realize that execution framework technology is becoming a significant component for modern astronomical data processing.
平方千米阵列(Square Kilometre Array, SKA)望远镜建成后将具有超高的灵敏度、超快的巡天速度以及宽视场,进而产生超海量的观测数据。SKA天文台与各国区域数据中心间的海量数据同步传输是当前SKA建设中的一个难点。SKA先导项目使用的下一代归档存储系统(Next Generation Archive System, NGAS)在应用测试中存在效率低下,性能不足等问题。提出了一种基于ZeroMQ的数据存储与同步方法,通过采用更加高效的异步消息机制实现同步传输数据,回避了NGAS原有的采用HTTP协议的局限。实验结果表明,新方法在平均数据归档存储效率方面比原有方法快了近40倍,能够基本满足10 GB带宽的全速传输需要,取得了较好的使用效果。
随着天文技术的发展,天文数据处理软件的需求也不断更迭变化,导致软件运行环境渐趋复杂.对于开发者和使用者,急需提出一种对复杂天文数据处理软件敏捷化封装和部署的方法.我国明安图射电频谱日像仪已进入常规观测,与之配套的数据处理软件也已完成开发并投入使用.由于该软件的部署涉及操作系统环境、 图形处理器运行环境及底层依赖软件等配置问题,导致安装过程既繁琐又容易出错.结合容器技术的特点,提出了一种基于Docker容器对日像仪软件系统进行敏捷封装与部署的方法,并对该方法的设计进行介绍,通过实验验证了其可用性,以及相比于传统虚拟机可获得较优异的性能表现.该方法可为未来天文数据处理软件的封装部署提供参考.可以预见,未来容器技术将成为天文海量数据处理的基础支撑技术.
观测控制系统是当前的一个研究热点,底层通信是观测控制系统架构的一个重要环节。但传统的观测控制系统一般采用裸套接技术实现,缺少统一的传输控制机制,在密集数据通信时经常存在延迟,影响实时控制的需要;同时缺少广播与组播机制,限制了观测控制系统的设计。针对观测控制系统的要求,研究了基于ZeroMQ不同通信模型及对应的天文控制模式,分析了不同的通信控制架构的可用性,在此基础上通过实验对相应的通信模型进行测试,验证其在天文仪器分布控制中的可用性。测试结果进一步说明,ZeroMQ构建观测控制系统的底层通信架构是可行的,能够满足观测控制系统对设备的各种控制需求。
The independent control of telescope is an important part of modern astronomical observation technology. In the current mainstream autonomous control system,the open source RTS2 has the characteristics of modularity, plug-and-play, fast response and stable operation, resulting in its widely application in the autonomous control system. As RTS2 is mainly based on the Linux platform,which is mainly based on the CLI for remote access control, the requirements for the observers are strict. In this paper, the RTS2 system is thoroughly analyzed, and the JSON API is modified to use JSON as the data transmission format. The microcomputer application in the mobile terminal is used as the carrier,and the data access and function call of the RTS2 telescope control system are carried out across different platforms. The control system is transplanted to WeChat Micro-program,so that astronomical researchers can quickly and easily use the mobile terminal in the WeChat platform for telescope remote control and real-time monitoring telescope autonomous control system status. With this model, it can be extended to other autonomous control architectures such as ASCOM,so as to realize a remote control of mobile terminal based on micro-application.