SeSQL: A High-Quality Large-Scale Session-Level Chinese Text-to-SQL Dataset.

Saihao Huang,Lijie Wang,Zhenghua Li,Zeyang Liu,Chenhui Dou, Fukang Yan,Xinyan Xiao,Hua Wu,Min Zhang

NLPCC (1)（2023）

引用 0|浏览14

暂无评分

摘要

As the first session-level Chinese dataset, CHASE contains two separate parts, i.e., 2,003 sessions manually constructed from scratch (CHASE-C), and 3,456 sessions translated from English SParC (CHASE-T). We find the two parts are highly discrepant and incompatible. In this work, we present SeSQL, a high-quality large-scale session-level Chinese text-to-SQL dataset, consisting of 5,028 sessions all manually constructed from scratch. Compared with previous datasets, in order to guarantee data quality, we adopt an iterative annotation workflow to facilitate intense and in-time review of previous-round natural language (NL) questions and SQL queries. Moreover, by completing all context-dependent NL questions, we obtain 27,012 context-independent question/SQL pairs, allowing SeSQL to be used as the largest dataset for single-round text-to-SQL parsing. We conduct benchmark session-level text-to-SQL parsing experiments on SeSQL via employing three competitive session-level parsers, and present detailed analysis.

查看译文

关键词

chinese,high-quality,large-scale,session-level,text-to-sql

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要