Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic
Knowledge
Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023
Main:8 Pages
8 Figures
Bibliography:4 Pages
7 Tables
Appendix:2 Pages
Abstract
In this paper, we study the problem of knowledge-intensive text-to-SQL, in which domain knowledge is necessary to parse expert questions into SQL queries over domain-specific tables. We formalize this scenario by building a new Chinese benchmark KnowSQL consisting of domain-specific questions covering various domains. We then address this problem by presenting formulaic knowledge, rather than by annotating additional data examples. More concretely, we construct a formulaic knowledge bank as a domain knowledge base and propose a framework (ReGrouP) to leverage this formulaic knowledge during parsing. Experiments using ReGrouP demonstrate a significant 28.2% improvement overall on KnowSQL.
View on arXivComments on this paper
