Comparative Analysis of Clustering Methodologies in DNA Storage

Subhasiny Sankar,Yixin Wang, Zhang Jiayu, Nur Sabrina,Erry Gunawan,Yong Liang Guan,Noor-A-Rahim Md.,Chueh Loo Poh

2022 26th International Computer Science and Engineering Conference (ICSEC)(2022)

引用 0|浏览5
暂无评分
摘要
Owing to the significance of DNA storage technology in meeting exponential storage demands and longevity, the challenges caused by bio-molecular errors while reading/sequencing data from DNA molecules must be addressed. By reading redundant copies, data can be reconstructed but with associated cost of sequencing and decoding complexities. Hence, solutions for dealing with both errors and complexities are sought after. The main objective of this work is to study data reconstruction methods for processing sequence readouts at downstream stage of DNA data storage. We investigated applicability of three clustering tools -Starcode, Slidesort, MeShClust, and two algorithms - Majority Nucleotide Selection (MNS), Cooperative Sequence Clustering (CSC) by transforming them into suitable tools for storage application. We observed that for fixed redundancy of 6.3x to 8.6x based on the nature of the dataset, Starcode outperforms other tools with 1% to 40% higher recovery rate. However, it costs the highest decoding complexity whereas MNS and CSC provides the lowest decoding complexity. Moreover, the distribution of the cluster and clustering speed of each tool/method are compared. This is the first comparative analysis study of tools/methods for data reconstruction in DNA data storage.
更多
查看译文
关键词
DNA data storage,Illumina sequencing,Data reconstruction,Clustering
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要