The aim of text-to-speech (or TTS, Text-To-Speech) is to automatically calculate the speech signal corresponding to a given text. The text itself can come from a variety of sources: newspapers, books, voice response systems, dialogue or automatic translation systems (interactive terminals, personal assistants), information system databases, video games, e-mails, SMS, documents browsed on the web, or simply text typed on a computer keyboard.
Voice response in its simplest form can be a set of pre-recorded messages (or "prompts"). Text-to-speech synthesis is more ambitious: it automatically calculates the sound samples corresponding to any written statement, which is not known in advance and may be large in size.
The two sides of speech synthesis are, on the one hand, text analysis and interpretation, and on the other, prediction of the acoustic-phonetic parameters of the sound and signal synthesis itself:
Text analysis: the first stage in transforming text into speech involves the ability to analyze and understand the written text, its nuances and connotations, the speech situation and the speech act to be performed. In addition to the text, the context can be specified (speaking style, emotion, attitude, character type, specific voice...);
Signal synthesis: once the text has been analyzed, the aim is to calculate the acoustic signal that best interprets the linguistic content, with a voice that sounds as natural as possible, resembling a particular speaker, and with the nuances of attitude and even emotion that the text calls for. In addition to the audio signal, the synthesizer can provide instructions for synchronizing the lip movements of an avatar or video character, or the movements of a robot.