46
6

FPUTS: Fully Parallel UFANS-based End-to-End Text-to-Speech System

Abstract

A Text-to-speech (TTS) system that can generate high quality audios with small time latency and fewer errors is required for industrial applications and services. In this paper, we propose a new non-autoregressive, fully parallel end-to-end TTS system. It utilizes the new attention structure and the recently proposed convolutional structure, UFANS. Different to RNN, UFANS can capture long term information in a fully parallel manner. Compared with the most popular end-to-end text-to-speech systems, our system can generate equal or better quality audios with fewer errors and reach at least 10 times speed up of inference.

View on arXiv
Comments on this paper