Refactor into a single package
This commit is contained in:
parent
02c3e4c908
commit
5e866e4fca
94 changed files with 198 additions and 2867 deletions
348
mimic3_tts/README.md
Normal file
348
mimic3_tts/README.md
Normal file
|
|
@ -0,0 +1,348 @@
|
|||
# Mimic 3
|
||||
|
||||
A fast and local neural text to speech system for [Mycroft](https://mycroft.ai/) and the [Mark II](https://mycroft.ai/product/mark-ii/).
|
||||
|
||||
* [Available voices](https://github.com/MycroftAI/mimic3-voices)
|
||||
* [Mimic 3 Architecture](#architecture)
|
||||
|
||||
|
||||
## Command-Line Tools
|
||||
|
||||
|
||||
### mimic3
|
||||
|
||||
|
||||
#### Basic Synthesis
|
||||
|
||||
```sh
|
||||
mimic3 --voice <voice> "<text>" > output.wav
|
||||
```
|
||||
|
||||
where `<voice>` is a [voice key](https://github.com/MycroftAI/mimic3/#voice-keys) like `en_UK/apope_low`.
|
||||
`<TEXT>` may contain multiple sentences, which will be combined in the final output WAV file. These can also be [split into separate WAV files](#multiple-wav-output).
|
||||
|
||||
|
||||
#### SSML Synthesis
|
||||
|
||||
```sh
|
||||
mimic3 --ssml --voice <voice> "<ssml>" > output.wav
|
||||
```
|
||||
|
||||
where `<ssml>` is valid [SSML](https://www.w3.org/TR/speech-synthesis11/). Not all SSML features are supported, see [the documentation](#ssml) for details.
|
||||
|
||||
If your SSML contains `<mark>` tags, add `--mark-file <file>` to the command-line and use `--interactive` mode. As the marks are encountered, their names will be written on separate lines to the file:
|
||||
|
||||
```sh
|
||||
mimic3 --ssml --interactive --mark-file - '<speak>Test 1. <mark name="here" /> Test 2.</speak>'
|
||||
```
|
||||
|
||||
|
||||
#### Long Texts
|
||||
|
||||
If your text is very long, and you would like to listen to it as its being synthesized, use `--interactive` mode:
|
||||
|
||||
```sh
|
||||
mimic3 --interactive < long.txt
|
||||
```
|
||||
|
||||
Each input line will be synthesized and played (see `--play-program`). By default, 5 sentences will be kept in an output queue, only blocking synthesis when the queue is full. You can adjust this value with `--result-queue-size`.
|
||||
|
||||
If your long text is fixed-width with blank lines separating paragraphs like those from [Project Gutenberg](https://www.gutenberg.org/), use the `--process-on-blank-line` option so that sentences will not be broken at line boundaries. For example, you can listen to "Alice in Wonderland" like this:
|
||||
|
||||
```sh
|
||||
curl --output - 'https://www.gutenberg.org/files/11/11-0.txt' | \
|
||||
mimic3 --interactive --process-on-blank-line
|
||||
```
|
||||
|
||||
|
||||
#### Multiple WAV Output
|
||||
|
||||
With `--output-dir` set to a directory, Mimic 3 will output a separate WAV file for each sentence:
|
||||
|
||||
```sh
|
||||
mimic3 'Test 1. Test 2.' --output-dir /path/to/wavs
|
||||
```
|
||||
|
||||
By default, each WAV file will be named using the (slightly modified) text of the sentence. You can have WAV files named using a timestamp instead with `--output-naming time`. For full control of the output naming, the `--csv` command-line flag indicates that each sentence is of the form `id|text` where `id` will be the name of the WAV file.
|
||||
|
||||
```sh
|
||||
cat << EOF |
|
||||
s01|The birch canoe slid on the smooth planks.
|
||||
s02|Glue the sheet to the dark blue background.
|
||||
s03|It's easy to tell the depth of a well.
|
||||
s04|These days a chicken leg is a rare dish.
|
||||
s05|Rice is often served in round bowls.
|
||||
s06|The juice of lemons makes fine punch.
|
||||
s07|The box was thrown beside the parked truck.
|
||||
s08|The hogs were fed chopped corn and garbage.
|
||||
s09|Four hours of steady work faced us.
|
||||
s10|Large size in stockings is hard to sell.
|
||||
EOF
|
||||
mimic3 --csv --output-dir /path/to/wavs
|
||||
```
|
||||
|
||||
You can adjust the delimiter with `--csv-delimiter <delimiter>`.
|
||||
|
||||
Additionally, you can use the `--csv-voice` option to specify a different voice or speaker for each line:
|
||||
|
||||
```sh
|
||||
cat << EOF |
|
||||
s01|#awb|The birch canoe slid on the smooth planks.
|
||||
s02|#rms|Glue the sheet to the dark blue background.
|
||||
s03|#slt|It's easy to tell the depth of a well.
|
||||
s04|#ksp|These days a chicken leg is a rare dish.
|
||||
s05|#clb|Rice is often served in round bowls.
|
||||
s06|#aew|The juice of lemons makes fine punch.
|
||||
s07|#bdl|The box was thrown beside the parked truck.
|
||||
s08|#lnh|The hogs were fed chopped corn and garbage.
|
||||
s09|#jmk|Four hours of steady work faced us.
|
||||
s10|en_UK/apope_low|Large size in stockings is hard to sell.
|
||||
EOF
|
||||
mimic3 --voice 'en_US/cmu-arctic_low' --csv-voice --output-dir /path/to/wavs
|
||||
```
|
||||
|
||||
The second contain can contain a `#<speaker>` or an entirely different voice!
|
||||
|
||||
|
||||
#### Interactive Mode
|
||||
|
||||
With `--interactive`, Mimic 3 will switch into interactive mode. After entering a sentence, it will be played with `--play-program`.
|
||||
|
||||
```sh
|
||||
mimic3 --interactive
|
||||
Reading text from stdin...
|
||||
Hello world!<ENTER>
|
||||
```
|
||||
|
||||
Use `CTRL+D` or `CTRL+C` to exit.
|
||||
|
||||
|
||||
#### Noise and Length Settings
|
||||
|
||||
Synthesis has the following additional parameters:
|
||||
|
||||
* `--noise-scale` and `--noise-w`
|
||||
* Determine the speaker volatility during synthesis
|
||||
* 0-1, default is 0.667 and 0.8 respectively
|
||||
* `--length-scale` - makes the voice speaker slower (> 1) or faster (< 1)
|
||||
|
||||
Individual voices have default settings for these parameters in their `config.json` files (under `inference`).
|
||||
|
||||
|
||||
#### List Voices
|
||||
|
||||
```sh
|
||||
mimic3 --voices
|
||||
```
|
||||
|
||||
|
||||
#### CUDA Acceleration
|
||||
|
||||
If you have a GPU with support for CUDA, you can accelerate synthesis with the `--cuda` flag. This requires you to install the [onnxruntime-gpu](https://pypi.org/project/onnxruntime-gpu/) Python package.
|
||||
|
||||
Using [nvidia-docker](https://github.com/NVIDIA/nvidia-docker) is highly recommended. See the `Dockerfile.gpu` file in the parent repository for an example of how to build a compatible container.
|
||||
|
||||
|
||||
|
||||
### mimic3-download
|
||||
|
||||
Mimic 3 automatically downloads voices when they're first used, but you can manually download them too with `mimic3-download`.
|
||||
|
||||
For example:
|
||||
|
||||
``` sh
|
||||
mimic3-download 'en_US/*'
|
||||
```
|
||||
|
||||
will download all U.S. English voices to `${HOME}/.local/share/mimic3` (technically `${XDG_DATA_HOME}/mimic3`).
|
||||
|
||||
See `mimic3-download --help` for more options.
|
||||
|
||||
|
||||
## SSML
|
||||
|
||||
A subset of [SSML](https://www.w3.org/TR/speech-synthesis11/) (Speech Synthesis Markup Language) is supported:
|
||||
|
||||
* `<speak>` - wrap around SSML text
|
||||
* `lang` - set language for document
|
||||
* `<s>` - sentence (disables automatic sentence breaking)
|
||||
* `lang` - set language for sentence
|
||||
* `<w>` / `<token>` - word (disables automatic tokenization)
|
||||
* `<voice name="...">` - set voice of inner text
|
||||
* `voice` - name or language of voice
|
||||
* Name format is `tts:voice` (e.g., "glow-speak:en-us_mary_ann") or `tts:voice#speaker_id` (e.g., "coqui-tts:en_vctk#p228")
|
||||
* If one of the supported languages, a preferred voice is used (override with `--preferred-voice <lang> <voice>`)
|
||||
* `<prosody attribute="value">` - change speaking attributes
|
||||
* Supported `attribute` names:
|
||||
* `volume` - speaking volume
|
||||
* number in [0, 100] - 0 is silent, 100 is loudest (default)
|
||||
* +X, -X, +X%, -X% - absolute/percent offset from current volume
|
||||
* one of "default", "silent", "x-loud", "loud", "medium", "soft", "x-soft"
|
||||
* `rate` - speaking rate
|
||||
* number - 1 is default rate, < 1 is slower, > 1 is faster
|
||||
* X% - 100% is default rate, 50% is half speed, 200% is twice as fast
|
||||
* one of "default", "x-fast", "fast", "medium", "slow", "x-slow"
|
||||
* `<say-as interpret-as="">` - force interpretation of inner text
|
||||
* `interpret-as` one of "spell-out", "date", "number", "time", or "currency"
|
||||
* `format` - way to format text depending on `interpret-as`
|
||||
* number - one of "cardinal", "ordinal", "digits", "year"
|
||||
* date - string with "d" (cardinal day), "o" (ordinal day), "m" (month), or "y" (year)
|
||||
* `<break time="">` - Pause for given amount of time
|
||||
* time - seconds ("123s") or milliseconds ("123ms")
|
||||
* `<sub alias="">` - substitute `alias` for inner text
|
||||
* `<phoneme ph="">` - supply phonemes for inner text
|
||||
* See `phonemes.txt` in voice directory for available phonemes
|
||||
* Phonemes may need to be separated by whitespace
|
||||
|
||||
SSML `<say-as>` support varies between voice types:
|
||||
|
||||
* [gruut](https://github.com/rhasspy/gruut/#ssml)
|
||||
* [eSpeak-ng](http://espeak.sourceforge.net/ssml.html)
|
||||
* Character-based voices do not currently support `<say-as>`
|
||||
|
||||
|
||||
## Speech Dispatcher
|
||||
|
||||
Mimic 3 can be used with the [Orca screen reader](https://help.gnome.org/users/orca/stable/) for Linux via [speech-dispatcher](https://github.com/brailcom/speechd).
|
||||
|
||||
After [installing Mimic 3](https://github.com/MycroftAI/mimic3/#installation), make sure you also have speech-dispatcher installed:
|
||||
|
||||
``` sh
|
||||
sudo apt-get install speech-dispatcher
|
||||
```
|
||||
|
||||
Create the file `/etc/speech-dispatcher/modules/mimic3-generic.conf` with the contents:
|
||||
|
||||
``` text
|
||||
GenericExecuteSynth "printf %s \'$DATA\' | /path/to/mimic3 --remote --voice \'$VOICE\' --stdout | $PLAY_COMMAND"
|
||||
AddVoice "en-us" "MALE1" "en_UK/apope_low"
|
||||
```
|
||||
|
||||
You will need `sudo` access to do this. Make sure to change `/path/to/mimic3` to wherever you installed Mimic 3. Note that the `--remote` option is used to connect to a local Mimic 3 web server (use `--remote <URL>` if your server is somewhere besides `localhost`).
|
||||
|
||||
To change the voice later, you only need to replace `en_UK/apope_low`.
|
||||
|
||||
Next, edit the existing file `/etc/speech-dispatcher/speechd.conf` and ensure the following settings are present:
|
||||
|
||||
``` text
|
||||
DefaultVoiceType "MALE1"
|
||||
DefaultModule mimic3-generic
|
||||
```
|
||||
|
||||
Restart speech-dispatcher with:
|
||||
|
||||
``` sh
|
||||
sudo systemctl restart speech-dispatcher
|
||||
```
|
||||
|
||||
and test it out with:
|
||||
|
||||
``` sh
|
||||
spd-say 'Hello from speech dispatcher.'
|
||||
```
|
||||
|
||||
|
||||
### Systemd Service
|
||||
|
||||
To ensure that Mimic 3 runs at boot, create a systemd service at `$HOME/.config/systemd/user/mimic3.service` with the contents:
|
||||
|
||||
``` text
|
||||
[Unit]
|
||||
Description=Run Mimic 3 web server
|
||||
Documentation=https://github.com/MycroftAI/mimic3
|
||||
|
||||
[Service]
|
||||
ExecStart=/path/to/mimic3-server
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
```
|
||||
|
||||
Make sure to change `/path/to/mimic3-server` to wherever you installed Mimic 3.
|
||||
|
||||
Refresh the systemd services:
|
||||
|
||||
``` sh
|
||||
systemctl --user daemon-reload
|
||||
```
|
||||
|
||||
Now try starting the service:
|
||||
|
||||
``` sh
|
||||
systemctl --user start mimic3
|
||||
```
|
||||
|
||||
If that's successful, ensure it starts at boot:
|
||||
|
||||
``` sh
|
||||
systemctl --user enable mimic3
|
||||
```
|
||||
|
||||
|
||||
## Architecture
|
||||
|
||||
Mimic 3 uses the [VITS](https://arxiv.org/abs/2106.06103), a "Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech". VITS is a combination of the [GlowTTS duration predictor](https://arxiv.org/abs/2005.11129) and the [HiFi-GAN vocoder](https://arxiv.org/abs/2010.05646).
|
||||
|
||||
Our implementation is heavily based on [Jaehyeon Kim's PyTorch model](https://github.com/jaywalnut310/vits), with the addition of [Onnx runtime](https://onnxruntime.ai/) export for speed.
|
||||
|
||||

|
||||
|
||||
|
||||
### Phoneme Ids
|
||||
|
||||
At a high level, Mimic 3 performs two important tasks:
|
||||
|
||||
1. Converting raw text to numeric input for the VITS TTS model, and
|
||||
2. Using the model to transform numeric input into audio output
|
||||
|
||||
The second step is the same for every voice, but the first step (text to numbers) varies. There are currently three implementations of step 1, described below.
|
||||
|
||||
|
||||
### gruut Phoneme-based Voices
|
||||
|
||||
Voices that use [gruut](https://github.com/rhasspy/gruut/) for phonemization.
|
||||
|
||||
gruut normalizes text and phonemizes words according to a lexicon, with a pre-trained grapheme-to-phoneme model used to guess unknown word pronunciations.
|
||||
|
||||
|
||||
### eSpeak Phoneme-based Voices
|
||||
|
||||
Voices that use [eSpeak-ng](https://github.com/espeak-ng/espeak-ng) for phonemization (via [espeak-phonemizer](https://github.com/rhasspy/espeak-phonemizer)).
|
||||
|
||||
eSpeak-ng normalizes and phonemizes text using internal rules and lexicons. It supports a large number of languages, and can handle many textual forms.
|
||||
|
||||
|
||||
### Character-based Voices
|
||||
|
||||
Voices whose "phonemes" are characters from an alphabet, typically with some punctuation.
|
||||
|
||||
For voices whose orthography (writing system) is close enough to its spoken form, character-based voices allow for skipping the phonemization step. However, these voices do not support text normalization, so numbers, dates, etc. must be written out.
|
||||
|
||||
|
||||
### Epitran-based Voices
|
||||
|
||||
Voices that use [epitran](https://github.com/dmort27/epitran/) for phonemization.
|
||||
|
||||
epitran uses rules to generate phonetic pronunciations from text. It does not support text normalization, however, so numbers, dates, etc. must be written out.
|
||||
|
||||
|
||||
### Components of a Voice Model
|
||||
|
||||
Voice models are stored in a directory with a specific layout:
|
||||
|
||||
* `<language>_<region>` (e.g., `en_UK`)
|
||||
* `<voice-name>_<quality>` (e.g., `apope_low`)
|
||||
* `ALIASES` - alternative names for the voice, one per line (optional)
|
||||
* `config.json` - training/inference configuration (see [code](https://github.com/MycroftAI/mimic3/blob/master/mimic3-tts/mimic3_tts/config.py) for details)
|
||||
* `generator.onnx` - exported inference model (see `ids_to_audio` method in [`voice.py`](https://github.com/MycroftAI/mimic3/blob/master/mimic3-tts/mimic3_tts/voice.py))
|
||||
* `LICENSE` - text, name, or URL of voice model license
|
||||
* `phoneme_map.txt` - mapping from source phoneme to destination phoneme(s) (optional)
|
||||
* `phonemes.txt` - mapping from integer ids to phonemes (`_` = padding, `^` = beginning of utterance, `$` = end of utterance, `#` = word break)
|
||||
* `README.md` - description of the voice
|
||||
* `SOURCE` - URL(s) of the dataset(s) this voice was trained on
|
||||
* `VERSION` - version of the voice in the format "MAJOR.Minor.bugfix" (e.g. "1.0.2")
|
||||
|
||||
|
||||
## License
|
||||
|
||||
See [license file](LICENSE)
|
||||
1
mimic3_tts/VERSION
Normal file
1
mimic3_tts/VERSION
Normal file
|
|
@ -0,0 +1 @@
|
|||
0.1.8
|
||||
19
mimic3_tts/__init__.py
Normal file
19
mimic3_tts/__init__.py
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
from pathlib import Path
|
||||
|
||||
from opentts_abc import (
|
||||
AudioResult,
|
||||
BaseResult,
|
||||
BaseToken,
|
||||
MarkResult,
|
||||
Phonemes,
|
||||
SayAs,
|
||||
Voice,
|
||||
Word,
|
||||
)
|
||||
from opentts_abc.ssml import SSMLSpeaker
|
||||
|
||||
from ._resources import __version__
|
||||
from .const import DEFAULT_VOICE
|
||||
from .tts import Mimic3Settings, Mimic3TextToSpeechSystem
|
||||
|
||||
__author__ = "Michael Hansen"
|
||||
726
mimic3_tts/__main__.py
Normal file
726
mimic3_tts/__main__.py
Normal file
|
|
@ -0,0 +1,726 @@
|
|||
#!/usr/bin/env python3
|
||||
# Copyright 2022 Mycroft AI Inc.
|
||||
#
|
||||
# This program is free software: you can redistribute it and/or modify
|
||||
# it under the terms of the GNU Affero General Public License as published by
|
||||
# the Free Software Foundation, either version 3 of the License, or
|
||||
# (at your option) any later version.
|
||||
#
|
||||
# This program is distributed in the hope that it will be useful,
|
||||
# but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
# GNU Affero General Public License for more details.
|
||||
#
|
||||
# You should have received a copy of the GNU Affero General Public License
|
||||
# along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
#
|
||||
import argparse
|
||||
import csv
|
||||
import io
|
||||
import logging
|
||||
import os
|
||||
import shlex
|
||||
import shutil
|
||||
import string
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import threading
|
||||
import time
|
||||
import typing
|
||||
import wave
|
||||
from dataclasses import dataclass, field
|
||||
from enum import Enum
|
||||
from pathlib import Path
|
||||
from queue import Queue
|
||||
|
||||
from ._resources import _PACKAGE
|
||||
|
||||
if typing.TYPE_CHECKING:
|
||||
from . import BaseResult, Mimic3TextToSpeechSystem # noqa: F401
|
||||
|
||||
|
||||
_LOGGER = logging.getLogger(_PACKAGE)
|
||||
|
||||
_DEFAULT_PLAY_PROGRAMS = ["paplay", "play -q", "aplay -q"]
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
@dataclass
|
||||
class ResultToProcess:
|
||||
result: "BaseResult"
|
||||
line: str
|
||||
line_id: str = ""
|
||||
|
||||
|
||||
@dataclass
|
||||
class CommandLineInterfaceState:
|
||||
args: argparse.Namespace
|
||||
texts: typing.Optional[typing.Iterable[str]] = None
|
||||
mark_writer: typing.Optional[typing.TextIO] = None
|
||||
tts: typing.Optional["Mimic3TextToSpeechSystem"] = None
|
||||
text_from_stdin: bool = False
|
||||
|
||||
all_audio: bytes = field(default_factory=bytes)
|
||||
sample_rate_hz: int = 22050
|
||||
sample_width_bytes: int = 2
|
||||
num_channels: int = 1
|
||||
|
||||
result_queue: typing.Optional["Queue[typing.Optional[ResultToProcess]]"] = None
|
||||
result_thread: typing.Optional[threading.Thread] = None
|
||||
|
||||
|
||||
class OutputNaming(str, Enum):
|
||||
"""Format used for output file names"""
|
||||
|
||||
TEXT = "text"
|
||||
TIME = "time"
|
||||
ID = "id"
|
||||
|
||||
|
||||
class StdinFormat(str, Enum):
|
||||
"""Format of standard input"""
|
||||
|
||||
AUTO = "auto"
|
||||
"""Choose based on SSML state"""
|
||||
|
||||
LINES = "lines"
|
||||
"""Each line is a separate sentence/document"""
|
||||
|
||||
DOCUMENT = "document"
|
||||
"""Entire input is one document"""
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
def main():
|
||||
"""Main entry point"""
|
||||
args = get_args()
|
||||
|
||||
if args.debug:
|
||||
logging.basicConfig(level=logging.DEBUG)
|
||||
logging.getLogger().setLevel(logging.DEBUG)
|
||||
else:
|
||||
logging.basicConfig(level=logging.INFO)
|
||||
logging.getLogger().setLevel(logging.INFO)
|
||||
|
||||
if args.version:
|
||||
# Print version and exit
|
||||
from . import __version__
|
||||
|
||||
print(__version__)
|
||||
sys.exit(0)
|
||||
|
||||
state = CommandLineInterfaceState(args=args)
|
||||
initialize_args(state)
|
||||
initialize_tts(state)
|
||||
|
||||
try:
|
||||
if args.voices:
|
||||
# Print voices and exit
|
||||
print_voices(state)
|
||||
else:
|
||||
# Process user input
|
||||
if os.isatty(sys.stdin.fileno()):
|
||||
print("Reading text from stdin...", file=sys.stderr)
|
||||
|
||||
process_lines(state)
|
||||
finally:
|
||||
shutdown_tts(state)
|
||||
|
||||
|
||||
def initialize_args(state: CommandLineInterfaceState):
|
||||
"""Initialze CLI state from command-line arguments"""
|
||||
import numpy as np
|
||||
|
||||
args = state.args
|
||||
|
||||
# Create output directory
|
||||
if args.output_dir:
|
||||
args.output_dir = Path(args.output_dir)
|
||||
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# Open file for writing the names from <mark> tags in SSML.
|
||||
# Each name is printed on a single line.
|
||||
if args.mark_file and (args.mark_file != "-"):
|
||||
args.mark_file = Path(args.mark_file)
|
||||
args.mark_file.parent.mkdir(parents=True, exist_ok=True)
|
||||
state.mark_writer = open( # pylint: disable=consider-using-with
|
||||
args.mark_file, "w", encoding="utf-8"
|
||||
)
|
||||
elif args.stdout:
|
||||
state.mark_writer = sys.stderr
|
||||
else:
|
||||
state.mark_writer = sys.stdout
|
||||
|
||||
if args.seed is not None:
|
||||
_LOGGER.debug("Setting random seed to %s", args.seed)
|
||||
np.random.seed(args.seed)
|
||||
|
||||
if args.csv_voice:
|
||||
# --csv-voice implies --csv
|
||||
args.csv = True
|
||||
|
||||
if args.csv:
|
||||
args.output_naming = OutputNaming.ID
|
||||
elif args.ssml:
|
||||
# Avoid text mangling when using SSML
|
||||
args.output_naming = OutputNaming.TIME
|
||||
|
||||
# Read text from stdin or arguments
|
||||
if args.text:
|
||||
# Use arguments
|
||||
state.texts = args.text
|
||||
else:
|
||||
# Use stdin
|
||||
state.text_from_stdin = True
|
||||
stdin_format = StdinFormat.LINES
|
||||
|
||||
if (args.stdin_format == StdinFormat.AUTO) and args.ssml:
|
||||
# Assume SSML input is entire document
|
||||
stdin_format = StdinFormat.DOCUMENT
|
||||
|
||||
if stdin_format == StdinFormat.DOCUMENT:
|
||||
# One big line
|
||||
state.texts = [sys.stdin.read()]
|
||||
else:
|
||||
# Multiple lines
|
||||
state.texts = sys.stdin
|
||||
|
||||
assert state.texts is not None
|
||||
|
||||
if args.process_on_blank_line:
|
||||
|
||||
# Combine text until a blank line is encountered.
|
||||
# Good for line-wrapped books where
|
||||
# sentences are broken
|
||||
# up across multiple
|
||||
# lines.
|
||||
def process_on_blank_line(lines: typing.Iterable[str]):
|
||||
text = ""
|
||||
for line in lines:
|
||||
line = line.strip()
|
||||
if not line:
|
||||
if text:
|
||||
yield text
|
||||
|
||||
text = ""
|
||||
continue
|
||||
|
||||
text += " " + line
|
||||
|
||||
state.texts = process_on_blank_line(state.texts)
|
||||
|
||||
if args.remote and args.remote.endswith("/"):
|
||||
# Ensure no slash
|
||||
args.remote = args.remote[:-1]
|
||||
|
||||
if (not args.speaker) and args.voice and ("#" in args.voice):
|
||||
# Split apart voice
|
||||
args.voice, args.speaker = args.voice.split("#", maxsplit=1)
|
||||
|
||||
if args.deterministic:
|
||||
# Disable noise
|
||||
_LOGGER.debug("Disabling noise in deterministic mode")
|
||||
args.noise_scale = 0.0
|
||||
args.noise_w = 0.0
|
||||
|
||||
|
||||
def initialize_tts(state: CommandLineInterfaceState):
|
||||
"""Create Mimic 3 TTS from command-line arguments"""
|
||||
from mimic3_tts import Mimic3Settings, Mimic3TextToSpeechSystem # noqa: F811
|
||||
|
||||
args = state.args
|
||||
|
||||
if not args.remote:
|
||||
# Local TTS
|
||||
state.tts = Mimic3TextToSpeechSystem(
|
||||
Mimic3Settings(
|
||||
length_scale=args.length_scale,
|
||||
noise_scale=args.noise_scale,
|
||||
noise_w=args.noise_w,
|
||||
voices_directories=args.voices_dir,
|
||||
use_cuda=args.cuda,
|
||||
use_deterministic_compute=args.deterministic,
|
||||
)
|
||||
)
|
||||
|
||||
state.tts.voice = args.voice
|
||||
state.tts.speaker = args.speaker
|
||||
|
||||
if args.voices:
|
||||
# Don't bother with the rest of the initialization
|
||||
return
|
||||
|
||||
if state.tts:
|
||||
if state.args.voice:
|
||||
# Set default voice
|
||||
state.tts.voice = state.args.voice
|
||||
|
||||
if state.args.preload_voice:
|
||||
for voice_key in state.args.preload_voice:
|
||||
_LOGGER.debug("Preloading voice: %s", voice_key)
|
||||
state.tts.preload_voice(voice_key)
|
||||
|
||||
state.result_queue = Queue(maxsize=args.result_queue_size)
|
||||
|
||||
state.result_thread = threading.Thread(
|
||||
target=process_result, daemon=True, args=(state,)
|
||||
)
|
||||
state.result_thread.start()
|
||||
|
||||
|
||||
def process_result(state: CommandLineInterfaceState):
|
||||
try:
|
||||
from mimic3_tts import AudioResult, MarkResult
|
||||
|
||||
assert state.result_queue is not None
|
||||
args = state.args
|
||||
|
||||
while True:
|
||||
result_todo = state.result_queue.get()
|
||||
if result_todo is None:
|
||||
break
|
||||
|
||||
try:
|
||||
result = result_todo.result
|
||||
line = result_todo.line
|
||||
line_id = result_todo.line_id
|
||||
|
||||
if isinstance(result, AudioResult):
|
||||
if args.interactive or args.output_dir:
|
||||
# Convert to WAV audio
|
||||
wav_bytes: typing.Optional[bytes] = None
|
||||
if args.interactive:
|
||||
if args.stdout:
|
||||
# Write audio to stdout
|
||||
sys.stdout.buffer.write(result.audio_bytes)
|
||||
sys.stdout.buffer.flush()
|
||||
else:
|
||||
# Play sound
|
||||
if not wav_bytes:
|
||||
wav_bytes = result.to_wav_bytes()
|
||||
|
||||
if wav_bytes:
|
||||
play_wav_bytes(state.args, wav_bytes)
|
||||
|
||||
if args.output_dir:
|
||||
if not wav_bytes:
|
||||
wav_bytes = result.to_wav_bytes()
|
||||
|
||||
# Determine file name
|
||||
if args.output_naming == OutputNaming.TEXT:
|
||||
# Use text itself
|
||||
file_name = line.strip().replace(" ", "_")
|
||||
file_name = file_name.translate(
|
||||
str.maketrans(
|
||||
"", "", string.punctuation.replace("_", "")
|
||||
)
|
||||
)
|
||||
elif args.output_naming == OutputNaming.TIME:
|
||||
# Use timestamp
|
||||
file_name = str(time.time())
|
||||
elif args.output_naming == OutputNaming.ID:
|
||||
file_name = line_id
|
||||
|
||||
assert file_name, f"No file name for text: {line}"
|
||||
wav_path = args.output_dir / (file_name + ".wav")
|
||||
wav_path.write_bytes(wav_bytes)
|
||||
|
||||
_LOGGER.debug("Wrote %s", wav_path)
|
||||
else:
|
||||
# Combine all audio and output to stdout at the end
|
||||
state.all_audio += result.audio_bytes
|
||||
state.sample_rate_hz = result.sample_rate_hz
|
||||
state.sample_width_bytes = result.sample_width_bytes
|
||||
state.num_channels = result.num_channels
|
||||
elif isinstance(result, MarkResult):
|
||||
if state.mark_writer:
|
||||
print(result.name, file=state.mark_writer)
|
||||
except Exception:
|
||||
_LOGGER.exception("Error processing result")
|
||||
except Exception:
|
||||
_LOGGER.exception("process_result")
|
||||
|
||||
|
||||
def process_line(
|
||||
line: str,
|
||||
state: CommandLineInterfaceState,
|
||||
line_id: str = "",
|
||||
line_voice: typing.Optional[str] = None,
|
||||
):
|
||||
assert state.result_queue is not None
|
||||
args = state.args
|
||||
|
||||
if state.tts:
|
||||
# Local TTS
|
||||
from mimic3_tts import SSMLSpeaker
|
||||
|
||||
assert state.tts is not None
|
||||
|
||||
args = state.args
|
||||
|
||||
if line_voice:
|
||||
if line_voice.startswith("#"):
|
||||
# Same voice, but different speaker
|
||||
state.tts.speaker = line_voice[1:]
|
||||
else:
|
||||
# Different voice
|
||||
state.tts.voice = line_voice
|
||||
|
||||
if args.ssml:
|
||||
results = SSMLSpeaker(state.tts).speak(line)
|
||||
else:
|
||||
state.tts.begin_utterance()
|
||||
|
||||
# TODO: text language
|
||||
state.tts.speak_text(line)
|
||||
|
||||
results = state.tts.end_utterance()
|
||||
else:
|
||||
# Remote TTS
|
||||
from mimic3_tts import AudioResult
|
||||
|
||||
voice: typing.Optional[str] = None
|
||||
if line_voice:
|
||||
if line_voice.startswith("#"):
|
||||
# Same voice, but different speaker
|
||||
if args.voice:
|
||||
voice = f"{args.voice}{line_voice}"
|
||||
else:
|
||||
# Different voice
|
||||
voice = line_voice
|
||||
|
||||
# Get remote WAV data and repackage as AudioResult
|
||||
wav_bytes = get_remote_wav_bytes(state, line, voice=voice)
|
||||
with io.BytesIO(wav_bytes) as wav_io:
|
||||
wav_reader: wave.Wave_read = wave.open(wav_io, "rb")
|
||||
with wav_reader as wav_file:
|
||||
results = [
|
||||
AudioResult(
|
||||
sample_rate_hz=wav_file.getframerate(),
|
||||
sample_width_bytes=wav_file.getsampwidth(),
|
||||
num_channels=wav_file.getnchannels(),
|
||||
audio_bytes=wav_file.readframes(wav_file.getnframes()),
|
||||
)
|
||||
]
|
||||
|
||||
# Add results to processing queue
|
||||
for result in results:
|
||||
state.result_queue.put(
|
||||
ResultToProcess(
|
||||
result=result,
|
||||
line=line,
|
||||
line_id=line_id,
|
||||
)
|
||||
)
|
||||
|
||||
# Restore voice/speaker
|
||||
if state.tts:
|
||||
state.tts.voice = args.voice
|
||||
state.tts.speaker = args.speaker
|
||||
|
||||
|
||||
def process_lines(state: CommandLineInterfaceState):
|
||||
assert state.texts is not None
|
||||
|
||||
args = state.args
|
||||
|
||||
try:
|
||||
result_idx = 0
|
||||
|
||||
for line in state.texts:
|
||||
line_voice: typing.Optional[str] = None
|
||||
line_id = ""
|
||||
line = line.strip()
|
||||
if not line:
|
||||
continue
|
||||
|
||||
if args.output_naming == OutputNaming.ID:
|
||||
# Line has the format id|text instead of just text
|
||||
with io.StringIO(line) as line_io:
|
||||
reader = csv.reader(line_io, delimiter=args.csv_delimiter)
|
||||
row = next(reader)
|
||||
line_id, line = row[0], row[-1]
|
||||
if args.csv_voice:
|
||||
line_voice = row[1]
|
||||
|
||||
process_line(line, state, line_id=line_id, line_voice=line_voice)
|
||||
result_idx += 1
|
||||
|
||||
except KeyboardInterrupt:
|
||||
if state.result_queue is not None:
|
||||
# Draw audio playback queue
|
||||
while not state.result_queue.empty():
|
||||
state.result_queue.get()
|
||||
finally:
|
||||
# Wait for raw stream to finish
|
||||
if state.result_queue is not None:
|
||||
state.result_queue.put(None)
|
||||
|
||||
if state.result_thread is not None:
|
||||
state.result_thread.join()
|
||||
|
||||
# -------------------------------------------------------------------------
|
||||
|
||||
# Write combined audio to stdout
|
||||
if state.all_audio:
|
||||
_LOGGER.debug("Writing WAV audio to stdout")
|
||||
|
||||
if sys.stdout.isatty() and (not state.args.stdout):
|
||||
with io.BytesIO() as wav_io:
|
||||
wav_file_play: wave.Wave_write = wave.open(wav_io, "wb")
|
||||
with wav_file_play:
|
||||
wav_file_play.setframerate(state.sample_rate_hz)
|
||||
wav_file_play.setsampwidth(state.sample_width_bytes)
|
||||
wav_file_play.setnchannels(state.num_channels)
|
||||
wav_file_play.writeframes(state.all_audio)
|
||||
|
||||
play_wav_bytes(state.args, wav_io.getvalue())
|
||||
else:
|
||||
# Write output directly to stdout
|
||||
wav_file_write: wave.Wave_write = wave.open(sys.stdout.buffer, "wb")
|
||||
with wav_file_write:
|
||||
wav_file_write.setframerate(state.sample_rate_hz)
|
||||
wav_file_write.setsampwidth(state.sample_width_bytes)
|
||||
wav_file_write.setnchannels(state.num_channels)
|
||||
wav_file_write.writeframes(state.all_audio)
|
||||
|
||||
sys.stdout.buffer.flush()
|
||||
|
||||
|
||||
def shutdown_tts(state: CommandLineInterfaceState):
|
||||
if state.tts:
|
||||
state.tts.shutdown()
|
||||
state.tts = None
|
||||
|
||||
|
||||
def play_wav_bytes(args: argparse.Namespace, wav_bytes: bytes):
|
||||
with tempfile.NamedTemporaryFile(mode="wb+", suffix=".wav") as wav_file:
|
||||
wav_file.write(wav_bytes)
|
||||
wav_file.seek(0)
|
||||
|
||||
for play_program in reversed(args.play_program):
|
||||
play_cmd = shlex.split(play_program)
|
||||
if not shutil.which(play_cmd[0]):
|
||||
continue
|
||||
|
||||
play_cmd.append(wav_file.name)
|
||||
_LOGGER.debug("Playing WAV file: %s", play_cmd)
|
||||
subprocess.check_output(play_cmd)
|
||||
break
|
||||
|
||||
|
||||
def print_voices(state: CommandLineInterfaceState):
|
||||
if state.tts:
|
||||
# Local TTS
|
||||
voices = list(state.tts.get_voices())
|
||||
voices = sorted(voices, key=lambda v: v.key)
|
||||
else:
|
||||
# Remove TTS
|
||||
voices = get_remote_voices(state)
|
||||
|
||||
writer = csv.writer(sys.stdout, delimiter="\t")
|
||||
writer.writerow(("KEY", "LANGUAGE", "NAME", "DESCRIPTION", "LOCATION"))
|
||||
for voice in voices:
|
||||
writer.writerow(
|
||||
(voice.key, voice.language, voice.name, voice.description, voice.location)
|
||||
)
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
def get_remote_voices(state: CommandLineInterfaceState) -> typing.List:
|
||||
import requests
|
||||
|
||||
from mimic3_tts import Voice
|
||||
|
||||
args = state.args
|
||||
|
||||
url = f"{args.remote}/api/voices"
|
||||
_LOGGER.debug("Getting voices from remote server at %s", url)
|
||||
|
||||
voices_json = requests.get(url).json()
|
||||
|
||||
return [Voice(**voice_args) for voice_args in voices_json]
|
||||
|
||||
|
||||
def get_remote_wav_bytes(
|
||||
state: CommandLineInterfaceState,
|
||||
text: str,
|
||||
voice: typing.Optional[str] = None,
|
||||
) -> bytes:
|
||||
import requests
|
||||
|
||||
args = state.args
|
||||
|
||||
if args.ssml:
|
||||
headers = {"Content-Type": "application/ssml+xml"}
|
||||
else:
|
||||
headers = {"Content-Type": "text/plain"}
|
||||
|
||||
params: typing.Dict[str, str] = {}
|
||||
|
||||
if voice:
|
||||
params["voice"] = voice
|
||||
elif args.voice:
|
||||
if args.speaker:
|
||||
params["voice"] = f"{args.voice}#{args.speaker}"
|
||||
else:
|
||||
params["voice"] = args.voice
|
||||
|
||||
if args.length_scale:
|
||||
params["lengthScale"] = args.length_scale
|
||||
|
||||
if args.noise_scale:
|
||||
params["noiseScale"] = args.noise_scale
|
||||
|
||||
if args.noise_w:
|
||||
params["noiseW"] = args.noise_w
|
||||
|
||||
url = f"{args.remote}/api/tts"
|
||||
_LOGGER.debug("Synthesizing text remotely at %s", url)
|
||||
|
||||
wav_bytes = requests.post(url, headers=headers, params=params, data=text).content
|
||||
|
||||
return wav_bytes
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
def get_args(argv=None):
|
||||
"""Parse command-line arguments"""
|
||||
parser = argparse.ArgumentParser(
|
||||
prog=_PACKAGE, description="Mimic 3 command-line interface"
|
||||
)
|
||||
parser.add_argument(
|
||||
"text", nargs="*", help="Text to convert to speech (default: stdin)"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--remote",
|
||||
nargs="?",
|
||||
const="http://localhost:59125",
|
||||
help="Connect to Mimic 3 HTTP web server for synthesis (default: localhost)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--stdin-format",
|
||||
choices=[str(v.value) for v in StdinFormat],
|
||||
default=StdinFormat.AUTO,
|
||||
help="Format of stdin text (default: auto)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--voice",
|
||||
"-v",
|
||||
help="Name of voice (expected in <voices-dir>/<language>)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--speaker",
|
||||
"-s",
|
||||
help="Name or number of speaker (default: first speaker)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--voices-dir",
|
||||
action="append",
|
||||
help="Directory with voices (format is <language>/<voice_name>)",
|
||||
)
|
||||
parser.add_argument("--voices", action="store_true", help="List available voices")
|
||||
parser.add_argument("--output-dir", help="Directory to write WAV file(s)")
|
||||
parser.add_argument(
|
||||
"--output-naming",
|
||||
choices=[v.value for v in OutputNaming],
|
||||
default="text",
|
||||
help="Naming scheme for output WAV files (requires --output-dir)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--id-delimiter",
|
||||
default="|",
|
||||
help="Delimiter between id and text in lines (default: |). Requires --output-naming id",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--interactive",
|
||||
action="store_true",
|
||||
help="Play audio after each input line (see --play-program)",
|
||||
)
|
||||
parser.add_argument("--csv", action="store_true", help="Input format is id|text")
|
||||
parser.add_argument(
|
||||
"--csv-delimiter", default="|", help="Delimiter used with --csv (default: |)"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--csv-voice",
|
||||
action="store_true",
|
||||
help="Input format is id|voice|text or id|#speaker|text",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--mark-file",
|
||||
help="File to write mark names to as they're encountered (--ssml only)",
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
"--noise-scale",
|
||||
type=float,
|
||||
help="Noise scale [0-1], default is 0.667",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--length-scale",
|
||||
type=float,
|
||||
help="Length scale (1.0 is default speed, 0.5 is 2x faster)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--noise-w",
|
||||
type=float,
|
||||
help="Variation in cadence [0-1], default is 0.8",
|
||||
)
|
||||
|
||||
# Miscellaneous
|
||||
parser.add_argument(
|
||||
"--result-queue-size",
|
||||
default=5,
|
||||
help="Maximum number of sentences to maintain in output queue (default: 5)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--process-on-blank-line",
|
||||
action="store_true",
|
||||
help="Process text only after encountering a blank line",
|
||||
)
|
||||
parser.add_argument("--ssml", action="store_true", help="Input text is SSML")
|
||||
parser.add_argument(
|
||||
"--stdout",
|
||||
action="store_true",
|
||||
help="Force audio output to stdout even if a tty is detected",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--preload-voice", action="append", help="Preload voice when starting up"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--play-program",
|
||||
action="append",
|
||||
default=_DEFAULT_PLAY_PROGRAMS,
|
||||
help="Program(s) used to play WAV files",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--cuda",
|
||||
action="store_true",
|
||||
help="Use Onnx CUDA execution provider (requires onnxruntime-gpu)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--deterministic",
|
||||
action="store_true",
|
||||
help="Ensure that the same audio is always synthesized from the same text",
|
||||
)
|
||||
parser.add_argument("--seed", type=int, help="Set random seed (default: not set)")
|
||||
parser.add_argument("--version", action="store_true", help="Print version and exit")
|
||||
parser.add_argument(
|
||||
"--debug", action="store_true", help="Print DEBUG messages to the console"
|
||||
)
|
||||
|
||||
return parser.parse_args(args=argv)
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
51
mimic3_tts/_resources.py
Normal file
51
mimic3_tts/_resources.py
Normal file
|
|
@ -0,0 +1,51 @@
|
|||
# Copyright 2022 Mycroft AI Inc.
|
||||
#
|
||||
# This program is free software: you can redistribute it and/or modify
|
||||
# it under the terms of the GNU Affero General Public License as published by
|
||||
# the Free Software Foundation, either version 3 of the License, or
|
||||
# (at your option) any later version.
|
||||
#
|
||||
# This program is distributed in the hope that it will be useful,
|
||||
# but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
# GNU Affero General Public License for more details.
|
||||
#
|
||||
# You should have received a copy of the GNU Affero General Public License
|
||||
# along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
#
|
||||
"""Shared access to package resources"""
|
||||
import json
|
||||
import os
|
||||
import typing
|
||||
from pathlib import Path
|
||||
|
||||
try:
|
||||
import importlib.resources
|
||||
|
||||
files = importlib.resources.files
|
||||
except (ImportError, AttributeError):
|
||||
# Backport for Python < 3.9
|
||||
import importlib_resources # type: ignore
|
||||
|
||||
files = importlib_resources.files
|
||||
|
||||
_PACKAGE = "mimic3_tts"
|
||||
_DIR = Path(typing.cast(os.PathLike, files(_PACKAGE)))
|
||||
|
||||
__version__ = (_DIR / "VERSION").read_text(encoding="utf-8").strip()
|
||||
|
||||
# Load voices.json
|
||||
# {
|
||||
# "<lang>/<voice>": {
|
||||
# "files": {
|
||||
# "relative/path": {
|
||||
# "size_bytes": size in bytes,
|
||||
# "sha256_sum": sha256 hash
|
||||
# }
|
||||
# },
|
||||
# "speakers": [],
|
||||
# "properties": {}
|
||||
# }
|
||||
# }
|
||||
with open(_DIR / "voices.json", "r", encoding="utf-8") as voices_file:
|
||||
_VOICES = json.load(voices_file)
|
||||
356
mimic3_tts/config.py
Normal file
356
mimic3_tts/config.py
Normal file
|
|
@ -0,0 +1,356 @@
|
|||
# Copyright 2022 Mycroft AI Inc.
|
||||
#
|
||||
# This program is free software: you can redistribute it and/or modify
|
||||
# it under the terms of the GNU Affero General Public License as published by
|
||||
# the Free Software Foundation, either version 3 of the License, or
|
||||
# (at your option) any later version.
|
||||
#
|
||||
# This program is distributed in the hope that it will be useful,
|
||||
# but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
# GNU Affero General Public License for more details.
|
||||
#
|
||||
# You should have received a copy of the GNU Affero General Public License
|
||||
# along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
#
|
||||
"""Configuration classes"""
|
||||
import collections
|
||||
import json
|
||||
import typing
|
||||
from dataclasses import dataclass, field
|
||||
from enum import Enum
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
from dataclasses_json import DataClassJsonMixin
|
||||
from gruut_ipa import IPA
|
||||
from phonemes2ids import BlankBetween
|
||||
|
||||
|
||||
@dataclass
|
||||
class AudioConfig(DataClassJsonMixin):
|
||||
"""Audio input/output details"""
|
||||
|
||||
filter_length: int = 1024
|
||||
hop_length: int = 256
|
||||
win_length: int = 1024
|
||||
mel_channels: int = 80
|
||||
sample_rate: int = 22050
|
||||
sample_bytes: int = 2
|
||||
channels: int = 1
|
||||
mel_fmin: float = 0.0
|
||||
mel_fmax: typing.Optional[float] = None
|
||||
ref_level_db: float = 20.0
|
||||
spec_gain: float = 1.0
|
||||
|
||||
# Normalization
|
||||
signal_norm: bool = True
|
||||
min_level_db: float = -100.0
|
||||
max_norm: float = 1.0
|
||||
clip_norm: bool = True
|
||||
symmetric_norm: bool = True
|
||||
do_dynamic_range_compression: bool = True
|
||||
convert_db_to_amp: bool = True
|
||||
|
||||
do_trim_silence: bool = False
|
||||
trim_silence_db: float = 40.0
|
||||
trim_margin_sec: float = 0.01
|
||||
trim_keep_sec: float = 0.25
|
||||
|
||||
scale_mels: bool = False
|
||||
|
||||
def __post_init__(self):
|
||||
if self.mel_fmax is not None:
|
||||
assert self.mel_fmax <= self.sample_rate // 2
|
||||
|
||||
# -------------------------------------------------------------------------
|
||||
# Normalization
|
||||
# -------------------------------------------------------------------------
|
||||
|
||||
def normalize(self, mel_db: np.ndarray) -> np.ndarray:
|
||||
"""Put values in [0, max_norm] or [-max_norm, max_norm]"""
|
||||
mel_norm = ((mel_db - self.ref_level_db) - self.min_level_db) / (
|
||||
-self.min_level_db
|
||||
)
|
||||
if self.symmetric_norm:
|
||||
# Symmetric norm
|
||||
mel_norm = ((2 * self.max_norm) * mel_norm) - self.max_norm
|
||||
if self.clip_norm:
|
||||
mel_norm = np.clip(mel_norm, -self.max_norm, self.max_norm)
|
||||
else:
|
||||
# Asymmetric norm
|
||||
mel_norm = self.max_norm * mel_norm
|
||||
if self.clip_norm:
|
||||
mel_norm = np.clip(mel_norm, 0, self.max_norm)
|
||||
|
||||
return mel_norm
|
||||
|
||||
def denormalize(self, mel_db: np.ndarray) -> np.ndarray:
|
||||
"""Pull values out of [0, max_norm] or [-max_norm, max_norm]"""
|
||||
if self.symmetric_norm:
|
||||
# Symmetric norm
|
||||
if self.clip_norm:
|
||||
mel_denorm = np.clip(mel_db, -self.max_norm, self.max_norm)
|
||||
|
||||
mel_denorm = (
|
||||
(mel_denorm + self.max_norm) * -self.min_level_db / (2 * self.max_norm)
|
||||
) + self.min_level_db
|
||||
else:
|
||||
# Asymmetric norm
|
||||
if self.clip_norm:
|
||||
mel_denorm = np.clip(mel_db, 0, self.max_norm)
|
||||
|
||||
mel_denorm = (
|
||||
mel_denorm * -self.min_level_db / self.max_norm
|
||||
) + self.min_level_db
|
||||
|
||||
mel_denorm += self.ref_level_db
|
||||
|
||||
return mel_denorm
|
||||
|
||||
|
||||
@dataclass
|
||||
class ModelConfig(DataClassJsonMixin):
|
||||
"""TTS model hyperparameters"""
|
||||
|
||||
num_symbols: int = 0
|
||||
n_speakers: int = 1
|
||||
|
||||
inter_channels: int = 192
|
||||
hidden_channels: int = 192
|
||||
filter_channels: int = 768
|
||||
n_heads: int = 2
|
||||
n_layers: int = 6
|
||||
kernel_size: int = 3
|
||||
p_dropout: float = 0.1
|
||||
resblock: str = "1"
|
||||
resblock_kernel_sizes: typing.Tuple[int, ...] = (3, 7, 11)
|
||||
resblock_dilation_sizes: typing.Tuple[typing.Tuple[int, ...], ...] = (
|
||||
(1, 3, 5),
|
||||
(1, 3, 5),
|
||||
(1, 3, 5),
|
||||
)
|
||||
upsample_rates: typing.Tuple[int, ...] = (8, 8, 2, 2)
|
||||
upsample_initial_channel: int = 512
|
||||
upsample_kernel_sizes: typing.Tuple[int, ...] = (16, 16, 4, 4)
|
||||
n_layers_q: int = 3
|
||||
use_spectral_norm: bool = False
|
||||
gin_channels: int = 0 # single speaker
|
||||
use_sdp: bool = True # StochasticDurationPredictor
|
||||
|
||||
@property
|
||||
def is_multispeaker(self) -> bool:
|
||||
return self.n_speakers > 1
|
||||
|
||||
|
||||
@dataclass
|
||||
class PhonemesConfig(DataClassJsonMixin):
|
||||
"""Phonemes to ids configuration"""
|
||||
|
||||
phoneme_separator: str = " "
|
||||
"""Separator between individual phonemes in CSV input"""
|
||||
|
||||
word_separator: str = "#"
|
||||
"""Separator between word phonemes in CSV input (must not match phoneme_separator)"""
|
||||
|
||||
phoneme_to_id: typing.Optional[typing.Dict[str, int]] = None
|
||||
pad: typing.Optional[str] = "_"
|
||||
bos: typing.Optional[str] = None
|
||||
eos: typing.Optional[str] = None
|
||||
blank: typing.Optional[str] = "#"
|
||||
blank_word: typing.Optional[str] = None
|
||||
blank_between: typing.Union[str, BlankBetween] = BlankBetween.WORDS
|
||||
blank_at_start: bool = True
|
||||
blank_at_end: bool = True
|
||||
simple_punctuation: bool = True
|
||||
punctuation_map: typing.Optional[typing.Dict[str, str]] = None
|
||||
separate: typing.Optional[typing.List[str]] = None
|
||||
separate_graphemes: bool = False
|
||||
separate_tones: bool = False
|
||||
tone_before: bool = False
|
||||
phoneme_map: typing.Optional[typing.Dict[str, str]] = None
|
||||
auto_bos_eos: bool = False
|
||||
minor_break: typing.Optional[str] = IPA.BREAK_MINOR.value
|
||||
major_break: typing.Optional[str] = IPA.BREAK_MAJOR.value
|
||||
break_phonemes_into_graphemes: bool = False
|
||||
drop_stress: bool = False
|
||||
symbols: typing.Optional[typing.List[str]] = None
|
||||
|
||||
def split_word_phonemes(self, phonemes_str: str) -> typing.List[typing.List[str]]:
|
||||
"""Split phonemes string into a list of lists (outer is words, inner is individual phonemes in each word)"""
|
||||
return [
|
||||
word_phonemes_str.split(self.phoneme_separator)
|
||||
for word_phonemes_str in phonemes_str.split(self.word_separator)
|
||||
]
|
||||
|
||||
def join_word_phonemes(self, word_phonemes: typing.List[typing.List[str]]) -> str:
|
||||
"""Split phonemes string into a list of lists (outer is words, inner is individual phonemes in each word)"""
|
||||
return self.word_separator.join(
|
||||
self.phoneme_separator.join(wp) for wp in word_phonemes
|
||||
)
|
||||
|
||||
|
||||
class Phonemizer(str, Enum):
|
||||
"""Method used to convert text to phonemes"""
|
||||
|
||||
SYMBOLS = "symbols"
|
||||
GRUUT = "gruut"
|
||||
ESPEAK = "espeak"
|
||||
EPITRAN = "epitran"
|
||||
|
||||
|
||||
class Aligner(str, Enum):
|
||||
"""Text/audio aligner"""
|
||||
|
||||
KALDI_ALIGN = "kaldi_align"
|
||||
"""https://github.com/rhasspy/kaldi-align"""
|
||||
|
||||
|
||||
class TextCasing(str, Enum):
|
||||
"""Casing method applied to text"""
|
||||
|
||||
LOWER = "lower"
|
||||
UPPER = "upper"
|
||||
|
||||
|
||||
class MetadataFormat(str, Enum):
|
||||
"""Format of training metadata"""
|
||||
|
||||
TEXT = "text"
|
||||
PHONEMES = "phonemes"
|
||||
PHONEME_IDS = "ids"
|
||||
|
||||
|
||||
@dataclass
|
||||
class DatasetConfig:
|
||||
"""Training dataset configuration"""
|
||||
|
||||
name: str
|
||||
metadata_format: MetadataFormat = MetadataFormat.TEXT
|
||||
multispeaker: bool = False
|
||||
text_language: typing.Optional[str] = None
|
||||
audio_dir: typing.Optional[typing.Union[str, Path]] = None
|
||||
cache_dir: typing.Optional[typing.Union[str, Path]] = None
|
||||
|
||||
def get_cache_dir(self, output_dir: typing.Union[str, Path]) -> Path:
|
||||
if self.cache_dir is not None:
|
||||
cache_dir = Path(self.cache_dir)
|
||||
else:
|
||||
cache_dir = Path("cache") / self.name
|
||||
|
||||
if not cache_dir.is_absolute():
|
||||
cache_dir = Path(output_dir) / str(cache_dir)
|
||||
|
||||
return cache_dir
|
||||
|
||||
|
||||
@dataclass
|
||||
class AlignerConfig:
|
||||
"""Text/audio alignment configuration"""
|
||||
|
||||
aligner: typing.Optional[Aligner] = None
|
||||
casing: typing.Optional[TextCasing] = None
|
||||
|
||||
|
||||
@dataclass
|
||||
class InferenceConfig:
|
||||
"""Inference configuration"""
|
||||
|
||||
length_scale: float = 1.0
|
||||
noise_scale: float = 0.667
|
||||
noise_w: float = 0.8
|
||||
|
||||
minor_break_ms: typing.Optional[int] = None
|
||||
major_break_ms: typing.Optional[int] = None
|
||||
|
||||
|
||||
@dataclass
|
||||
class TrainingConfig(DataClassJsonMixin):
|
||||
"""Master configuration for training"""
|
||||
|
||||
seed: int = 1234
|
||||
epochs: int = 10000
|
||||
learning_rate: float = 2e-4
|
||||
betas: typing.Tuple[float, float] = field(default=(0.8, 0.99))
|
||||
eps: float = 1e-9
|
||||
batch_size: int = 32
|
||||
fp16_run: bool = False
|
||||
lr_decay: float = 0.999875
|
||||
segment_size: int = 8192
|
||||
init_lr_ratio: float = 1.0
|
||||
warmup_epochs: int = 0
|
||||
c_mel: int = 45
|
||||
c_kl: float = 1.0
|
||||
grad_clip: typing.Optional[float] = None
|
||||
|
||||
min_seq_length: typing.Optional[int] = None
|
||||
max_seq_length: typing.Optional[int] = None
|
||||
|
||||
min_spec_length: typing.Optional[int] = None
|
||||
max_spec_length: typing.Optional[int] = None
|
||||
|
||||
min_speaker_utterances: typing.Optional[int] = None
|
||||
|
||||
last_epoch: int = 1
|
||||
global_step: int = 1
|
||||
best_loss: typing.Optional[float] = None
|
||||
audio: AudioConfig = field(default_factory=AudioConfig)
|
||||
model: ModelConfig = field(default_factory=ModelConfig)
|
||||
phonemes: PhonemesConfig = field(default_factory=PhonemesConfig)
|
||||
text_aligner: AlignerConfig = field(default_factory=AlignerConfig)
|
||||
text_language: typing.Optional[str] = None
|
||||
phonemizer: typing.Optional[Phonemizer] = None
|
||||
datasets: typing.List[DatasetConfig] = field(default_factory=list)
|
||||
inference: InferenceConfig = field(default_factory=InferenceConfig)
|
||||
|
||||
version: int = 1
|
||||
git_commit: str = ""
|
||||
|
||||
@property
|
||||
def is_multispeaker(self):
|
||||
return self.model.is_multispeaker or any(d.multispeaker for d in self.datasets)
|
||||
|
||||
def save(self, config_file: typing.TextIO):
|
||||
"""Save config as JSON to a file"""
|
||||
json.dump(self.to_dict(), config_file, indent=4)
|
||||
|
||||
@staticmethod
|
||||
def load(config_file: typing.TextIO) -> "TrainingConfig":
|
||||
"""Load config from a JSON file"""
|
||||
return TrainingConfig.from_json(config_file.read())
|
||||
|
||||
@staticmethod
|
||||
def load_and_merge(
|
||||
config: "TrainingConfig",
|
||||
config_files: typing.Iterable[typing.Union[str, Path, typing.TextIO]],
|
||||
) -> "TrainingConfig":
|
||||
"""Loads one or more JSON configuration files and overlays them on top of an existing config"""
|
||||
base_dict = config.to_dict()
|
||||
for maybe_config_file in config_files:
|
||||
if isinstance(maybe_config_file, (str, Path)):
|
||||
# File path
|
||||
config_file = open(maybe_config_file, "r", encoding="utf-8")
|
||||
else:
|
||||
# File object
|
||||
config_file = maybe_config_file
|
||||
|
||||
with config_file:
|
||||
# Load new config and overlay on existing config
|
||||
new_dict = json.load(config_file)
|
||||
TrainingConfig.recursive_update(base_dict, new_dict)
|
||||
|
||||
return TrainingConfig.from_dict(base_dict)
|
||||
|
||||
@staticmethod
|
||||
def recursive_update(
|
||||
base_dict: typing.Dict[typing.Any, typing.Any],
|
||||
new_dict: typing.Mapping[typing.Any, typing.Any],
|
||||
) -> None:
|
||||
"""Recursively overwrites values in base dictionary with values from new dictionary"""
|
||||
for key, value in new_dict.items():
|
||||
if isinstance(value, collections.Mapping) and (
|
||||
base_dict.get(key) is not None
|
||||
):
|
||||
TrainingConfig.recursive_update(base_dict[key], value)
|
||||
else:
|
||||
base_dict[key] = value
|
||||
28
mimic3_tts/const.py
Normal file
28
mimic3_tts/const.py
Normal file
|
|
@ -0,0 +1,28 @@
|
|||
# Copyright 2022 Mycroft AI Inc.
|
||||
#
|
||||
# This program is free software: you can redistribute it and/or modify
|
||||
# it under the terms of the GNU Affero General Public License as published by
|
||||
# the Free Software Foundation, either version 3 of the License, or
|
||||
# (at your option) any later version.
|
||||
#
|
||||
# This program is distributed in the hope that it will be useful,
|
||||
# but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
# GNU Affero General Public License for more details.
|
||||
#
|
||||
# You should have received a copy of the GNU Affero General Public License
|
||||
# along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
#
|
||||
from pathlib import Path
|
||||
|
||||
from xdgenvpy import XDG
|
||||
|
||||
DEFAULT_VOICE = "en_UK/apope_low"
|
||||
DEFAULT_LANGUAGE = "en_UK"
|
||||
DEFAULT_VOICES_URL_FORMAT = (
|
||||
"https://github.com/MycroftAI/mimic3-voices/raw/master/voices/{lang}/{name}"
|
||||
)
|
||||
DEFAULT_VOICES_DOWNLOAD_DIR = Path(XDG().XDG_DATA_HOME) / "mimic3" / "voices"
|
||||
|
||||
DEFAULT_VOLUME = 100.0
|
||||
DEFAULT_RATE = 1.0
|
||||
240
mimic3_tts/download.py
Normal file
240
mimic3_tts/download.py
Normal file
|
|
@ -0,0 +1,240 @@
|
|||
# Copyright 2022 Mycroft AI Inc.
|
||||
#
|
||||
# This program is free software: you can redistribute it and/or modify
|
||||
# it under the terms of the GNU Affero General Public License as published by
|
||||
# the Free Software Foundation, either version 3 of the License, or
|
||||
# (at your option) any later version.
|
||||
#
|
||||
# This program is distributed in the hope that it will be useful,
|
||||
# but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
# GNU Affero General Public License for more details.
|
||||
#
|
||||
# You should have received a copy of the GNU Affero General Public License
|
||||
# along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
#
|
||||
"""A command-line tool for downloading Mimic 3 voices"""
|
||||
import argparse
|
||||
import itertools
|
||||
import json
|
||||
import logging
|
||||
import re
|
||||
import sys
|
||||
import typing
|
||||
import urllib.request
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from urllib.error import HTTPError
|
||||
|
||||
from ._resources import _PACKAGE, _VOICES
|
||||
from .const import DEFAULT_VOICES_DOWNLOAD_DIR, DEFAULT_VOICES_URL_FORMAT
|
||||
from .utils import file_sha256_sum, wildcard_to_regex
|
||||
|
||||
_LOGGER = logging.getLogger(__name__)
|
||||
|
||||
_WILDCARD = "*"
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
class VoiceDownloadError(Exception):
|
||||
"""Occurs when a voice fails to download"""
|
||||
|
||||
|
||||
@dataclass
|
||||
class VoiceFile:
|
||||
"""File associated with a voice to download"""
|
||||
|
||||
relative_path: str
|
||||
size_bytes: typing.Optional[int] = None
|
||||
sha256_sum: typing.Optional[str] = None
|
||||
|
||||
|
||||
def is_later_version(version1: str, version2: str) -> bool:
|
||||
"""True if version1 is later than version2"""
|
||||
v1_parts = [int(n) for n in version1.split(".")]
|
||||
v2_parts = [int(n) for n in version2.split(".")]
|
||||
|
||||
for p1, p2 in itertools.zip_longest(v1_parts, v2_parts, fillvalue=0):
|
||||
if p1 > p2:
|
||||
# 2.0 vs 1.0
|
||||
return True
|
||||
|
||||
if p1 < p2:
|
||||
# 1.0 vs 2.0
|
||||
return False
|
||||
|
||||
# 1.0 vs 1.0
|
||||
return False
|
||||
|
||||
|
||||
def download_voice(
|
||||
voice_key: str,
|
||||
url_base: str,
|
||||
voice_files: typing.Iterable[VoiceFile],
|
||||
voices_dir: typing.Union[str, Path],
|
||||
voice_version: str,
|
||||
chunk_bytes: int = 4096,
|
||||
redownload: bool = False,
|
||||
):
|
||||
"""Downloads a voice to a directory"""
|
||||
from tqdm.auto import tqdm
|
||||
|
||||
if url_base.endswith("/"):
|
||||
# Remove final slash
|
||||
url_base = url_base[:-1]
|
||||
|
||||
voice_dir = Path(voices_dir) / voice_key
|
||||
voice_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
_LOGGER.debug("Downloading voice %s to %s", voice_key, voice_dir)
|
||||
|
||||
version_path = voice_dir / "VERSION"
|
||||
if version_path.is_file():
|
||||
actual_version = version_path.read_text(encoding="utf-8").strip()
|
||||
if is_later_version(voice_version, actual_version):
|
||||
redownload = True
|
||||
_LOGGER.debug(
|
||||
"Replacing version %s of %s with version %s",
|
||||
actual_version,
|
||||
voice_key,
|
||||
voice_version,
|
||||
)
|
||||
|
||||
for voice_file in voice_files:
|
||||
file_url = f"{url_base}/{voice_file.relative_path}"
|
||||
file_path = voice_dir / voice_file.relative_path
|
||||
|
||||
if (not redownload) and voice_file.sha256_sum and file_path.is_file():
|
||||
# Check if file exists and has correct sha256
|
||||
expected_sha256 = voice_file.sha256_sum
|
||||
|
||||
with open(file_path, "rb") as check_file:
|
||||
actual_sha256 = file_sha256_sum(check_file)
|
||||
|
||||
if actual_sha256 == expected_sha256:
|
||||
_LOGGER.debug("Skipping download of %s (sha256 match)", file_path)
|
||||
continue
|
||||
|
||||
try:
|
||||
# Download file, show progress with tqdm
|
||||
with urllib.request.urlopen(file_url) as response:
|
||||
with open(file_path, mode="wb") as dest_file:
|
||||
with tqdm(
|
||||
unit="B",
|
||||
unit_scale=True,
|
||||
unit_divisor=1024,
|
||||
miniters=1,
|
||||
desc=voice_file.relative_path,
|
||||
total=int(response.headers.get("content-length", 0)),
|
||||
) as pbar:
|
||||
chunk = response.read(chunk_bytes)
|
||||
while chunk:
|
||||
dest_file.write(chunk)
|
||||
pbar.update(len(chunk))
|
||||
chunk = response.read(chunk_bytes)
|
||||
|
||||
_LOGGER.debug("Downloaded %s", file_path)
|
||||
except HTTPError as e:
|
||||
_LOGGER.exception("download_voice")
|
||||
raise VoiceDownloadError(
|
||||
f"Failed to download file for voice {voice_key} from {file_url}: {e}"
|
||||
) from e
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
def main(argv=None):
|
||||
"""Main entry point"""
|
||||
parser = argparse.ArgumentParser(
|
||||
prog=f"{_PACKAGE}.download", description="Download utility for Mimic 3 voices"
|
||||
)
|
||||
parser.add_argument(
|
||||
"key",
|
||||
nargs="*",
|
||||
help="Keys of voices to download (e.g., en_US/vctk_low). May contain wildcards (*)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--output-dir",
|
||||
default=DEFAULT_VOICES_DOWNLOAD_DIR,
|
||||
help="Path to output directory",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--url-format",
|
||||
default=DEFAULT_VOICES_URL_FORMAT,
|
||||
help="URL format string for voices (contains {key}, {lang}, {name})",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--redownload",
|
||||
action="store_true",
|
||||
help="Force re-downloading of files if they already exist",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--debug", action="store_true", help="Print DEBUG messages to console"
|
||||
)
|
||||
args = parser.parse_args(args=argv)
|
||||
|
||||
if args.debug:
|
||||
logging.basicConfig(level=logging.DEBUG)
|
||||
logging.getLogger().setLevel(logging.DEBUG)
|
||||
else:
|
||||
logging.basicConfig(level=logging.INFO)
|
||||
logging.getLogger().setLevel(logging.INFO)
|
||||
|
||||
_LOGGER.debug(args)
|
||||
|
||||
args.output_dir = Path(args.output_dir)
|
||||
args.key = args.key or []
|
||||
|
||||
if not args.key:
|
||||
# Print available voices and exit
|
||||
json.dump(_VOICES, sys.stdout, indent=4, ensure_ascii=False)
|
||||
sys.exit(0)
|
||||
|
||||
args.key = [
|
||||
wildcard_to_regex(key, wildcard=_WILDCARD) if _WILDCARD in key else key
|
||||
for key in args.key
|
||||
]
|
||||
|
||||
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
for key_or_pattern in args.key:
|
||||
if isinstance(key_or_pattern, re.Pattern):
|
||||
# Wildcards
|
||||
voice_keys = []
|
||||
for maybe_key in _VOICES.keys():
|
||||
if key_or_pattern.match(maybe_key):
|
||||
voice_keys.append(maybe_key)
|
||||
|
||||
_LOGGER.debug("%s matched %s", key_or_pattern, voice_keys)
|
||||
else:
|
||||
# No wildcards
|
||||
voice_keys = [key_or_pattern]
|
||||
|
||||
for voice_key in voice_keys:
|
||||
voice_lang, voice_name = voice_key.split("/", maxsplit=1)
|
||||
voice_info = _VOICES[voice_key]
|
||||
voice_url = str.format(
|
||||
args.url_format, key=voice_key, lang=voice_lang, name=voice_name
|
||||
)
|
||||
voice_files = voice_info["files"]
|
||||
|
||||
_LOGGER.info("Downloading %s", voice_key)
|
||||
download_voice(
|
||||
voice_key=voice_key,
|
||||
url_base=voice_url,
|
||||
voice_files=[
|
||||
VoiceFile(file_key, sha256_sum=file_info.get("sha256_sum"))
|
||||
for file_key, file_info in voice_files.items()
|
||||
],
|
||||
voice_version=voice_info["version"],
|
||||
voices_dir=args.output_dir,
|
||||
redownload=args.redownload,
|
||||
)
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
0
mimic3_tts/py.typed
Normal file
0
mimic3_tts/py.typed
Normal file
582
mimic3_tts/tts.py
Normal file
582
mimic3_tts/tts.py
Normal file
|
|
@ -0,0 +1,582 @@
|
|||
# Copyright 2022 Mycroft AI Inc.
|
||||
#
|
||||
# This program is free software: you can redistribute it and/or modify
|
||||
# it under the terms of the GNU Affero General Public License as published by
|
||||
# the Free Software Foundation, either version 3 of the License, or
|
||||
# (at your option) any later version.
|
||||
#
|
||||
# This program is distributed in the hope that it will be useful,
|
||||
# but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
# GNU Affero General Public License for more details.
|
||||
#
|
||||
# You should have received a copy of the GNU Affero General Public License
|
||||
# along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
#
|
||||
"""Implementation of OpenTTS for Mimic 3"""
|
||||
import audioop
|
||||
import itertools
|
||||
import logging
|
||||
import typing
|
||||
from copy import deepcopy
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
|
||||
from gruut_ipa import IPA
|
||||
from xdgenvpy import XDG
|
||||
|
||||
from opentts_abc import (
|
||||
AudioResult,
|
||||
BaseResult,
|
||||
BaseToken,
|
||||
MarkResult,
|
||||
Phonemes,
|
||||
SayAs,
|
||||
TextToSpeechSystem,
|
||||
Voice,
|
||||
Word,
|
||||
)
|
||||
|
||||
from ._resources import _VOICES
|
||||
from .config import TrainingConfig
|
||||
from .const import (
|
||||
DEFAULT_LANGUAGE,
|
||||
DEFAULT_RATE,
|
||||
DEFAULT_VOICE,
|
||||
DEFAULT_VOICES_DOWNLOAD_DIR,
|
||||
DEFAULT_VOICES_URL_FORMAT,
|
||||
DEFAULT_VOLUME,
|
||||
)
|
||||
from .download import VoiceFile, download_voice
|
||||
from .voice import SPEAKER_TYPE, BreakType, Mimic3Voice
|
||||
|
||||
_DIR = Path(__file__).parent
|
||||
|
||||
_LOGGER = logging.getLogger(__name__)
|
||||
|
||||
PHONEMES_LIST_TYPE = typing.List[typing.List[str]]
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
@dataclass
|
||||
class Mimic3Settings:
|
||||
"""Settings for Mimic 3 text to speech system"""
|
||||
|
||||
voice: typing.Optional[str] = None
|
||||
"""Default voice key"""
|
||||
|
||||
language: typing.Optional[str] = None
|
||||
"""Default language (e.g., "en_US")"""
|
||||
|
||||
voices_directories: typing.Optional[typing.Iterable[typing.Union[str, Path]]] = None
|
||||
"""Directories to search for voices (<lang>/<voice>)"""
|
||||
|
||||
voices_url_format: typing.Optional[str] = DEFAULT_VOICES_URL_FORMAT
|
||||
"""URL format string for a voice directory.
|
||||
|
||||
May contain:
|
||||
* {key} - unique voice key
|
||||
* {lang} - voice language
|
||||
* {name} - voice name
|
||||
"""
|
||||
|
||||
speaker: typing.Optional[SPEAKER_TYPE] = None
|
||||
"""Default speaker name or id"""
|
||||
|
||||
length_scale: typing.Optional[float] = None
|
||||
"""Default length scale (use voice config if None)"""
|
||||
|
||||
noise_scale: typing.Optional[float] = None
|
||||
"""Default noise scale (use voice config if None)"""
|
||||
|
||||
noise_w: typing.Optional[float] = None
|
||||
"""Default noise W (use voice config if None)"""
|
||||
|
||||
text_language: typing.Optional[str] = None
|
||||
"""Language of text (use voice language if None)"""
|
||||
|
||||
sample_rate: int = 22050
|
||||
"""Sample rate of silence from add_break() in Hertz"""
|
||||
|
||||
voices_download_dir: typing.Union[str, Path] = DEFAULT_VOICES_DOWNLOAD_DIR
|
||||
"""Directory to download voices to"""
|
||||
|
||||
no_download: bool = False
|
||||
"""Do not download voices automatically"""
|
||||
|
||||
use_cuda: bool = False
|
||||
"""Use CUDA GPU acceleration (requires onnxruntime-gpu)"""
|
||||
|
||||
share_onnx_models_between_threads: bool = True
|
||||
"""If True, Onnx models are shared between threads"""
|
||||
|
||||
volume: float = DEFAULT_VOLUME
|
||||
"""Voice volume in [0, 100]"""
|
||||
|
||||
rate: float = DEFAULT_RATE
|
||||
"""Voice speaking rate (< 1 is slower, > 1 is faster)"""
|
||||
|
||||
use_deterministic_compute: bool = False
|
||||
"""Force onnxruntime to use deterministic compute mode. For fully deterministic synthesis, also set noise_scale and noise_w to 0."""
|
||||
|
||||
|
||||
@dataclass
|
||||
class Mimic3Phonemes:
|
||||
"""Pending task to synthesize audio from phonemes with specific settings"""
|
||||
|
||||
current_settings: Mimic3Settings
|
||||
"""Settings used to synthesize audio"""
|
||||
|
||||
phonemes: typing.List[typing.List[str]] = field(default_factory=list)
|
||||
"""Phonemes for synthesis"""
|
||||
|
||||
is_utterance: bool = True
|
||||
"""True if this is the end of a full utterance"""
|
||||
|
||||
|
||||
class VoiceNotFoundError(Exception):
|
||||
"""Raised if a voice cannot be found"""
|
||||
|
||||
def __init__(self, voice: str):
|
||||
super().__init__(f"Voice not found: {voice}")
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
class Mimic3TextToSpeechSystem(TextToSpeechSystem):
|
||||
"""Convert text to speech using Mimic 3"""
|
||||
|
||||
def __init__(self, settings: Mimic3Settings):
|
||||
self.settings = settings
|
||||
|
||||
self._results: typing.List[typing.Union[BaseResult, Mimic3Phonemes]] = []
|
||||
self._loaded_voices: typing.Dict[str, Mimic3Voice] = {}
|
||||
|
||||
@staticmethod
|
||||
def get_default_voices_directories() -> typing.List[Path]:
|
||||
"""Get list of directories to search for voices by default.
|
||||
|
||||
On Linux, this is typically:
|
||||
- $HOME/.local/share/mimic3/voices
|
||||
- /usr/local/share/mimic3/voices
|
||||
- /usr/share/mimic3/voices
|
||||
"""
|
||||
return [Path(d) / "mimic3" / "voices" for d in XDG().XDG_DATA_DIRS.split(":")]
|
||||
|
||||
def get_voices(self) -> typing.Iterable[Voice]:
|
||||
"""Returns an iterable of all available voices"""
|
||||
voices_dirs: typing.Iterable[
|
||||
typing.Union[str, Path]
|
||||
] = Mimic3TextToSpeechSystem.get_default_voices_directories()
|
||||
|
||||
if self.settings.voices_directories is not None:
|
||||
voices_dirs = itertools.chain(self.settings.voices_directories, voices_dirs)
|
||||
|
||||
known_voices = set(_VOICES.keys())
|
||||
|
||||
# voices/<language>/<voice>/
|
||||
for voices_dir in voices_dirs:
|
||||
voices_dir = Path(voices_dir)
|
||||
|
||||
if not voices_dir.is_dir() or voices_dir.name.startswith("."):
|
||||
_LOGGER.debug("Skipping voice directory %s", voices_dir)
|
||||
continue
|
||||
|
||||
_LOGGER.debug("Searching %s for voices", voices_dir)
|
||||
|
||||
for lang_dir in voices_dir.iterdir():
|
||||
if not lang_dir.is_dir() or lang_dir.name.startswith("."):
|
||||
continue
|
||||
|
||||
for voice_dir in lang_dir.iterdir():
|
||||
if not voice_dir.is_dir() or voice_dir.name.startswith("."):
|
||||
continue
|
||||
|
||||
config_path = voice_dir / "config.json"
|
||||
if not config_path.is_file():
|
||||
continue
|
||||
|
||||
_LOGGER.debug("Voice found in %s", voice_dir)
|
||||
voice_lang = lang_dir.name
|
||||
|
||||
# Load config
|
||||
_LOGGER.debug("Loading config from %s", config_path)
|
||||
|
||||
with open(config_path, "r", encoding="utf-8") as config_file:
|
||||
config = TrainingConfig.load(config_file)
|
||||
|
||||
properties: typing.Dict[str, typing.Any] = {
|
||||
"length_scale": config.inference.length_scale,
|
||||
"noise_scale": config.inference.noise_scale,
|
||||
"noise_w": config.inference.noise_w,
|
||||
}
|
||||
|
||||
# Load speaker names
|
||||
voice_name = voice_dir.name
|
||||
speakers: typing.Optional[typing.Sequence[str]] = None
|
||||
|
||||
speakers_path = voice_dir / "speakers.txt"
|
||||
if speakers_path.is_file():
|
||||
speakers = []
|
||||
with open(
|
||||
speakers_path, "r", encoding="utf-8"
|
||||
) as speakers_file:
|
||||
for line in speakers_file:
|
||||
line = line.strip()
|
||||
if line:
|
||||
speakers.append(line)
|
||||
|
||||
# Load aliases
|
||||
aliases: typing.Optional[typing.Set[str]] = None
|
||||
aliases_path = voice_dir / "ALIASES"
|
||||
if aliases_path.is_file():
|
||||
aliases = set()
|
||||
|
||||
with open(aliases_path, "r", encoding="utf-8") as aliases_file:
|
||||
for line in aliases_file:
|
||||
line = line.strip()
|
||||
if line:
|
||||
aliases.add(line)
|
||||
|
||||
voice_key = f"{voice_lang}/{voice_name}"
|
||||
|
||||
yield Voice(
|
||||
key=voice_key,
|
||||
name=voice_name,
|
||||
language=voice_lang,
|
||||
description="",
|
||||
speakers=speakers,
|
||||
location=str(voice_dir.absolute()),
|
||||
properties=properties,
|
||||
aliases=aliases,
|
||||
)
|
||||
|
||||
known_voices.discard(voice_key)
|
||||
|
||||
# Yield voices that haven't yet been downloaded
|
||||
for voice_key in known_voices:
|
||||
voice_lang, voice_name = voice_key.split("/", maxsplit=1)
|
||||
voice_info = _VOICES.get(voice_key, {})
|
||||
speakers = voice_info.get("speakers", [])
|
||||
properties = voice_info.get("properties", {})
|
||||
|
||||
yield Voice(
|
||||
key=voice_key,
|
||||
name=voice_name,
|
||||
language=voice_lang,
|
||||
description="",
|
||||
speakers=speakers,
|
||||
location=str.format(
|
||||
self.settings.voices_url_format or DEFAULT_VOICES_URL_FORMAT,
|
||||
lang=voice_lang,
|
||||
name=voice_name,
|
||||
key=voice_key,
|
||||
),
|
||||
properties=properties,
|
||||
)
|
||||
|
||||
def preload_voice(self, voice_key: str):
|
||||
"""Ensure voice is loaded in memory before synthesis"""
|
||||
self._get_or_load_voice(voice_key)
|
||||
|
||||
# -------------------------------------------------------------------------
|
||||
|
||||
@property
|
||||
def voice(self) -> str:
|
||||
return self.settings.voice or DEFAULT_VOICE
|
||||
|
||||
@voice.setter
|
||||
def voice(self, new_voice: str):
|
||||
if new_voice != self.settings.voice:
|
||||
# Clear speaker on voice change
|
||||
self.speaker = None
|
||||
|
||||
self.settings.voice = new_voice or DEFAULT_VOICE
|
||||
|
||||
if "#" in self.settings.voice:
|
||||
# Split
|
||||
voice, speaker = self.settings.voice.split("#", maxsplit=1)
|
||||
self.settings.voice = voice
|
||||
self.speaker = speaker
|
||||
|
||||
@property
|
||||
def speaker(self) -> typing.Optional[SPEAKER_TYPE]:
|
||||
return self.settings.speaker
|
||||
|
||||
@speaker.setter
|
||||
def speaker(self, new_speaker: typing.Optional[SPEAKER_TYPE]):
|
||||
self.settings.speaker = new_speaker
|
||||
|
||||
@property
|
||||
def language(self) -> str:
|
||||
return self.settings.language or DEFAULT_LANGUAGE
|
||||
|
||||
@language.setter
|
||||
def language(self, new_language: str):
|
||||
self.settings.language = new_language
|
||||
|
||||
@property
|
||||
def volume(self) -> float:
|
||||
return self.settings.volume
|
||||
|
||||
@volume.setter
|
||||
def volume(self, new_volume: float):
|
||||
self.settings.volume = max(0, min(100, new_volume))
|
||||
|
||||
@property
|
||||
def rate(self) -> float:
|
||||
return self.settings.rate
|
||||
|
||||
@rate.setter
|
||||
def rate(self, new_rate: float):
|
||||
self.settings.rate = new_rate
|
||||
|
||||
def begin_utterance(self):
|
||||
pass
|
||||
|
||||
# pylint: disable=arguments-differ
|
||||
def speak_text(self, text: str, text_language: typing.Optional[str] = None):
|
||||
voice = self._get_or_load_voice(self.voice)
|
||||
|
||||
minor_break_ms = voice.config.inference.minor_break_ms
|
||||
major_break_ms = voice.config.inference.major_break_ms
|
||||
|
||||
for sent_phonemes, break_type in voice.text_to_phonemes(
|
||||
text, text_language=text_language
|
||||
):
|
||||
add_major_silence = (break_type == BreakType.MAJOR) and (
|
||||
major_break_ms is not None
|
||||
)
|
||||
add_minor_silence = (break_type == BreakType.MINOR) and (
|
||||
minor_break_ms is not None
|
||||
)
|
||||
|
||||
# Utterances have start/end meta phonemes (usually ^ and $)
|
||||
is_utterance = (
|
||||
(break_type == BreakType.UTTERANCE)
|
||||
or add_major_silence
|
||||
or add_minor_silence
|
||||
)
|
||||
|
||||
self._results.append(
|
||||
Mimic3Phonemes(
|
||||
current_settings=deepcopy(self.settings),
|
||||
phonemes=sent_phonemes,
|
||||
is_utterance=is_utterance,
|
||||
)
|
||||
)
|
||||
|
||||
# Add silence if using manual break intervals
|
||||
if add_major_silence:
|
||||
assert major_break_ms is not None
|
||||
self.add_break(major_break_ms)
|
||||
elif add_minor_silence:
|
||||
assert minor_break_ms is not None
|
||||
self.add_break(minor_break_ms)
|
||||
|
||||
# pylint: disable=arguments-differ
|
||||
def speak_tokens(
|
||||
self,
|
||||
tokens: typing.Iterable[BaseToken],
|
||||
text_language: typing.Optional[str] = None,
|
||||
):
|
||||
voice = self._get_or_load_voice(self.voice)
|
||||
token_phonemes: PHONEMES_LIST_TYPE = []
|
||||
|
||||
for token in tokens:
|
||||
if isinstance(token, Word):
|
||||
word_phonemes = voice.word_to_phonemes(
|
||||
token.text, word_role=token.role, text_language=text_language
|
||||
)
|
||||
token_phonemes.append(word_phonemes)
|
||||
elif isinstance(token, Phonemes):
|
||||
phoneme_str = token.text.strip()
|
||||
if " " in phoneme_str:
|
||||
token_phonemes.append(phoneme_str.split())
|
||||
else:
|
||||
token_phonemes.append(list(IPA.graphemes(phoneme_str)))
|
||||
elif isinstance(token, SayAs):
|
||||
say_as_phonemes = voice.say_as_to_phonemes(
|
||||
token.text,
|
||||
interpret_as=token.interpret_as,
|
||||
say_format=token.format,
|
||||
text_language=text_language,
|
||||
)
|
||||
token_phonemes.extend(say_as_phonemes)
|
||||
|
||||
if token_phonemes:
|
||||
self._results.append(
|
||||
Mimic3Phonemes(
|
||||
current_settings=deepcopy(self.settings),
|
||||
phonemes=token_phonemes,
|
||||
is_utterance=False,
|
||||
)
|
||||
)
|
||||
|
||||
def add_break(self, time_ms: int):
|
||||
# Generate silence (16-bit mono at sample rate)
|
||||
num_samples = int((time_ms / 1000.0) * self.settings.sample_rate)
|
||||
audio_bytes = bytes(num_samples * 2)
|
||||
|
||||
self._results.append(
|
||||
AudioResult(
|
||||
sample_rate_hz=self.settings.sample_rate,
|
||||
audio_bytes=audio_bytes,
|
||||
# 16-bit mono
|
||||
sample_width_bytes=2,
|
||||
num_channels=1,
|
||||
)
|
||||
)
|
||||
|
||||
def set_mark(self, name: str):
|
||||
self._results.append(MarkResult(name=name))
|
||||
|
||||
def end_utterance(self) -> typing.Iterable[BaseResult]:
|
||||
last_settings: typing.Optional[Mimic3Settings] = None
|
||||
|
||||
sent_phonemes: PHONEMES_LIST_TYPE = []
|
||||
|
||||
for result in self._results:
|
||||
if isinstance(result, Mimic3Phonemes):
|
||||
if result.is_utterance or (result.current_settings != last_settings):
|
||||
if sent_phonemes:
|
||||
yield self._speak_sentence_phonemes(
|
||||
sent_phonemes, settings=last_settings
|
||||
)
|
||||
sent_phonemes.clear()
|
||||
|
||||
sent_phonemes.extend(result.phonemes)
|
||||
last_settings = result.current_settings
|
||||
else:
|
||||
if sent_phonemes:
|
||||
yield self._speak_sentence_phonemes(
|
||||
sent_phonemes, settings=last_settings
|
||||
)
|
||||
sent_phonemes.clear()
|
||||
|
||||
yield result
|
||||
|
||||
if sent_phonemes:
|
||||
yield self._speak_sentence_phonemes(sent_phonemes, settings=last_settings)
|
||||
sent_phonemes.clear()
|
||||
|
||||
self._results.clear()
|
||||
|
||||
# -------------------------------------------------------------------------
|
||||
|
||||
def _speak_sentence_phonemes(
|
||||
self,
|
||||
sent_phonemes,
|
||||
settings: typing.Optional[Mimic3Settings] = None,
|
||||
) -> AudioResult:
|
||||
"""Synthesize audio from phonemes using given setings"""
|
||||
settings = settings or self.settings
|
||||
voice = self._get_or_load_voice(settings.voice or self.voice)
|
||||
sent_phoneme_ids = voice.phonemes_to_ids(sent_phonemes)
|
||||
|
||||
_LOGGER.debug("phonemes=%s, ids=%s", sent_phonemes, sent_phoneme_ids)
|
||||
|
||||
audio = voice.ids_to_audio(
|
||||
sent_phoneme_ids,
|
||||
speaker=settings.speaker,
|
||||
length_scale=settings.length_scale,
|
||||
noise_scale=settings.noise_scale,
|
||||
noise_w=settings.noise_w,
|
||||
rate=settings.rate,
|
||||
)
|
||||
|
||||
audio_bytes = audio.tobytes()
|
||||
|
||||
if settings.volume != DEFAULT_VOLUME:
|
||||
audio_bytes = audioop.mul(audio_bytes, 2, settings.volume / 100.0)
|
||||
|
||||
return AudioResult(
|
||||
sample_rate_hz=voice.config.audio.sample_rate,
|
||||
audio_bytes=audio_bytes,
|
||||
# 16-bit mono
|
||||
sample_width_bytes=2,
|
||||
num_channels=1,
|
||||
)
|
||||
|
||||
def _get_or_load_voice(self, voice_key: str) -> Mimic3Voice:
|
||||
"""Get a loaded voice or load from the file system"""
|
||||
existing_voice = self._loaded_voices.get(voice_key)
|
||||
if existing_voice is not None:
|
||||
return existing_voice
|
||||
|
||||
# Look up as substring of known voice
|
||||
model_dir: typing.Optional[Path] = None
|
||||
for maybe_voice in self.get_voices():
|
||||
if (voice_key == maybe_voice.key) or (
|
||||
maybe_voice.aliases and (voice_key in maybe_voice.aliases)
|
||||
):
|
||||
maybe_model_dir = Path(maybe_voice.location)
|
||||
|
||||
if (not maybe_model_dir.is_dir()) and (not self.settings.no_download):
|
||||
# Download voice
|
||||
maybe_model_dir = self._download_voice(voice_key)
|
||||
|
||||
if maybe_model_dir.is_dir():
|
||||
# Voice found
|
||||
model_dir = maybe_model_dir
|
||||
break
|
||||
|
||||
if model_dir is None:
|
||||
raise VoiceNotFoundError(voice_key)
|
||||
|
||||
voice_lang = model_dir.parent.name
|
||||
voice_name = model_dir.name
|
||||
canonical_key = f"{voice_lang}/{voice_name}"
|
||||
|
||||
existing_voice = self._loaded_voices.get(canonical_key)
|
||||
if existing_voice is not None:
|
||||
# Alias
|
||||
self._loaded_voices[voice_key] = existing_voice
|
||||
|
||||
return existing_voice
|
||||
|
||||
# https://onnxruntime.ai/docs/execution-providers/
|
||||
providers = None
|
||||
if self.settings.use_cuda:
|
||||
providers = ["CUDAExecutionProvider"]
|
||||
|
||||
voice = Mimic3Voice.load_from_directory(
|
||||
model_dir,
|
||||
providers=providers,
|
||||
share_models=self.settings.share_onnx_models_between_threads,
|
||||
use_deterministic_compute=self.settings.use_deterministic_compute,
|
||||
)
|
||||
|
||||
_LOGGER.info("Loaded voice from %s", model_dir)
|
||||
|
||||
# Add to cache
|
||||
self._loaded_voices[voice_key] = voice
|
||||
self._loaded_voices[canonical_key] = voice
|
||||
|
||||
return voice
|
||||
|
||||
def _download_voice(self, voice_key: str) -> Path:
|
||||
"""Downloads a voice by key"""
|
||||
voice_lang, voice_name = voice_key.split("/", maxsplit=1)
|
||||
voice_info = _VOICES[voice_key]
|
||||
voice_url = str.format(
|
||||
self.settings.voices_url_format or DEFAULT_VOICES_URL_FORMAT,
|
||||
key=voice_key,
|
||||
lang=voice_lang,
|
||||
name=voice_name,
|
||||
)
|
||||
voice_files = voice_info["files"]
|
||||
download_voice(
|
||||
voice_key=voice_key,
|
||||
url_base=voice_url,
|
||||
voice_files=[VoiceFile(file_key) for file_key in voice_files.keys()],
|
||||
voice_version=voice_info["version"],
|
||||
voices_dir=self.settings.voices_download_dir,
|
||||
)
|
||||
|
||||
voice_dir = Path(self.settings.voices_download_dir) / voice_key
|
||||
|
||||
return voice_dir
|
||||
63
mimic3_tts/utils.py
Normal file
63
mimic3_tts/utils.py
Normal file
|
|
@ -0,0 +1,63 @@
|
|||
# Copyright 2022 Mycroft AI Inc.
|
||||
#
|
||||
# This program is free software: you can redistribute it and/or modify
|
||||
# it under the terms of the GNU Affero General Public License as published by
|
||||
# the Free Software Foundation, either version 3 of the License, or
|
||||
# (at your option) any later version.
|
||||
#
|
||||
# This program is distributed in the hope that it will be useful,
|
||||
# but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
# GNU Affero General Public License for more details.
|
||||
#
|
||||
# You should have received a copy of the GNU Affero General Public License
|
||||
# along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
#
|
||||
"""Utility methods for Mimic 3"""
|
||||
import hashlib
|
||||
import re
|
||||
import typing
|
||||
|
||||
import numpy as np
|
||||
|
||||
|
||||
def audio_float_to_int16(
|
||||
audio: np.ndarray, max_wav_value: float = 32767.0
|
||||
) -> np.ndarray:
|
||||
"""Normalize audio and convert to int16 range"""
|
||||
audio_norm = audio * (max_wav_value / max(0.01, np.max(np.abs(audio))))
|
||||
audio_norm = np.clip(audio_norm, -max_wav_value, max_wav_value)
|
||||
audio_norm = audio_norm.astype("int16")
|
||||
return audio_norm
|
||||
|
||||
|
||||
def wildcard_to_regex(template: str, wildcard: str = "*") -> re.Pattern:
|
||||
"""Convert a string with wildcards into a regex pattern"""
|
||||
wildcard_escaped = re.escape(wildcard)
|
||||
|
||||
pattern_parts = ["^"]
|
||||
for i, template_part in enumerate(re.split(f"({wildcard_escaped})", template)):
|
||||
if (i % 2) == 0:
|
||||
# Fixed string
|
||||
pattern_parts.append(re.escape(template_part))
|
||||
else:
|
||||
# Wildcard separator
|
||||
pattern_parts.append(".*")
|
||||
|
||||
pattern_parts.append("$")
|
||||
pattern_str = "".join(pattern_parts)
|
||||
|
||||
return re.compile(pattern_str)
|
||||
|
||||
|
||||
def file_sha256_sum(fp: typing.BinaryIO, block_bytes: int = 4096) -> str:
|
||||
"""Return the sha256 sum of a (possibly large) file"""
|
||||
current_hash = hashlib.sha256()
|
||||
|
||||
# Read in blocks in case file is very large
|
||||
block = fp.read(block_bytes)
|
||||
while len(block) > 0:
|
||||
current_hash.update(block)
|
||||
block = fp.read(block_bytes)
|
||||
|
||||
return current_hash.hexdigest()
|
||||
762
mimic3_tts/voice.py
Normal file
762
mimic3_tts/voice.py
Normal file
|
|
@ -0,0 +1,762 @@
|
|||
# Copyright 2022 Mycroft AI Inc.
|
||||
#
|
||||
# This program is free software: you can redistribute it and/or modify
|
||||
# it under the terms of the GNU Affero General Public License as published by
|
||||
# the Free Software Foundation, either version 3 of the License, or
|
||||
# (at your option) any later version.
|
||||
#
|
||||
# This program is distributed in the hope that it will be useful,
|
||||
# but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
# GNU Affero General Public License for more details.
|
||||
#
|
||||
# You should have received a copy of the GNU Affero General Public License
|
||||
# along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
#
|
||||
import csv
|
||||
import logging
|
||||
import platform
|
||||
import threading
|
||||
import time
|
||||
import typing
|
||||
from abc import ABCMeta, abstractmethod
|
||||
from enum import Enum
|
||||
from pathlib import Path
|
||||
from xml.sax.saxutils import escape as xmlescape
|
||||
|
||||
import epitran
|
||||
import espeak_phonemizer
|
||||
import gruut
|
||||
import numpy as np
|
||||
import onnxruntime
|
||||
import phonemes2ids
|
||||
from gruut_ipa import IPA
|
||||
|
||||
from .config import Phonemizer, TrainingConfig
|
||||
from .const import DEFAULT_RATE
|
||||
from .utils import audio_float_to_int16
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
class BreakType(str, Enum):
|
||||
NONE = "none"
|
||||
MINOR = "minor"
|
||||
MAJOR = "major"
|
||||
UTTERANCE = "utterance"
|
||||
|
||||
|
||||
PHONEME_TYPE = str
|
||||
PHONEME_ID_TYPE = int
|
||||
WORD_PHONEMES_TYPE = typing.List[typing.List[PHONEME_TYPE]]
|
||||
PHONEME_MAP_TYPE = typing.Dict[PHONEME_TYPE, typing.List[PHONEME_TYPE]]
|
||||
TEXT_TO_PHONEMES_TYPE = typing.Iterable[typing.Tuple[WORD_PHONEMES_TYPE, BreakType]]
|
||||
|
||||
SPEAKER_NAME_TYPE = str
|
||||
SPEAKER_ID_TYPE = int
|
||||
SPEAKER_TYPE = typing.Union[SPEAKER_NAME_TYPE, SPEAKER_ID_TYPE]
|
||||
SPEAKER_MAP_TYPE = typing.Dict[SPEAKER_NAME_TYPE, SPEAKER_ID_TYPE]
|
||||
|
||||
DEFAULT_LANGUAGE = "en_US"
|
||||
|
||||
_LOGGER = logging.getLogger(__name__)
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
class Mimic3Voice(metaclass=ABCMeta):
|
||||
"""Base class for Mimic 3 voice implementations"""
|
||||
|
||||
_SHARED_MODELS: typing.Dict[str, onnxruntime.InferenceSession] = {}
|
||||
_SHARED_MODELS_LOCK = threading.Lock()
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
config: TrainingConfig,
|
||||
onnx_model: onnxruntime.InferenceSession,
|
||||
phoneme_to_id: typing.Dict[PHONEME_TYPE, int],
|
||||
phoneme_map: typing.Optional[PHONEME_MAP_TYPE] = None,
|
||||
speaker_map: typing.Optional[SPEAKER_MAP_TYPE] = None,
|
||||
):
|
||||
self.config = config
|
||||
self.onnx_model = onnx_model
|
||||
self.phoneme_to_id = phoneme_to_id
|
||||
self.phoneme_map = phoneme_map
|
||||
self.speaker_map = speaker_map
|
||||
|
||||
@abstractmethod
|
||||
def text_to_phonemes(
|
||||
self, text: str, text_language: typing.Optional[str] = None
|
||||
) -> TEXT_TO_PHONEMES_TYPE:
|
||||
"""Convert text into phonemes"""
|
||||
|
||||
def word_to_phonemes(
|
||||
self,
|
||||
word_text: str,
|
||||
word_role: typing.Optional[str] = None,
|
||||
text_language: typing.Optional[str] = None,
|
||||
) -> typing.List[PHONEME_TYPE]:
|
||||
"""Convert a single word (with optional role) into phonemes"""
|
||||
word_phonemes = []
|
||||
for sent_phonemes, _break_type in self.text_to_phonemes(
|
||||
word_text, text_language=text_language
|
||||
):
|
||||
for sent_word_phonemes in sent_phonemes:
|
||||
word_phonemes.extend(sent_word_phonemes)
|
||||
|
||||
return word_phonemes
|
||||
|
||||
def say_as_to_phonemes(
|
||||
self,
|
||||
text: str,
|
||||
interpret_as: str,
|
||||
say_format: typing.Optional[str] = None,
|
||||
text_language: typing.Optional[str] = None,
|
||||
) -> WORD_PHONEMES_TYPE:
|
||||
"""Speak a word or phrase with a specific interpretation/format"""
|
||||
word_phonemes = []
|
||||
for sent_phonemes, _break_type in self.text_to_phonemes(
|
||||
text, text_language=text_language
|
||||
):
|
||||
word_phonemes.extend(sent_phonemes)
|
||||
|
||||
return word_phonemes
|
||||
|
||||
def phonemes_to_ids(
|
||||
self, phonemes: WORD_PHONEMES_TYPE
|
||||
) -> typing.Sequence[PHONEME_ID_TYPE]:
|
||||
"""Convert phonemes to ids for a voice model (see phonemes.txt)"""
|
||||
phoneme_map = self.phoneme_map or self.config.phonemes.phoneme_map
|
||||
|
||||
return phonemes2ids.phonemes2ids(
|
||||
word_phonemes=phonemes,
|
||||
phoneme_to_id=self.phoneme_to_id,
|
||||
pad=self.config.phonemes.pad,
|
||||
bos=self.config.phonemes.bos,
|
||||
eos=self.config.phonemes.eos,
|
||||
auto_bos_eos=self.config.phonemes.auto_bos_eos,
|
||||
blank=self.config.phonemes.blank,
|
||||
blank_word=self.config.phonemes.blank_word,
|
||||
blank_between=self.config.phonemes.blank_between,
|
||||
blank_at_start=self.config.phonemes.blank_at_start,
|
||||
blank_at_end=self.config.phonemes.blank_at_end,
|
||||
simple_punctuation=self.config.phonemes.simple_punctuation,
|
||||
punctuation_map=self.config.phonemes.punctuation_map,
|
||||
separate=self.config.phonemes.separate,
|
||||
separate_graphemes=self.config.phonemes.separate_graphemes,
|
||||
separate_tones=self.config.phonemes.separate_tones,
|
||||
tone_before=self.config.phonemes.tone_before,
|
||||
phoneme_map=phoneme_map,
|
||||
fail_on_missing=False,
|
||||
)
|
||||
|
||||
def ids_to_audio(
|
||||
self,
|
||||
phoneme_ids: typing.Sequence[PHONEME_ID_TYPE],
|
||||
speaker: typing.Optional[
|
||||
typing.Union[SPEAKER_NAME_TYPE, SPEAKER_ID_TYPE]
|
||||
] = None,
|
||||
length_scale: typing.Optional[float] = None,
|
||||
noise_scale: typing.Optional[float] = None,
|
||||
noise_w: typing.Optional[float] = None,
|
||||
rate: float = DEFAULT_RATE,
|
||||
) -> np.ndarray:
|
||||
"""Synthesize audio from phoneme ids usng Onnx voice model (see generator.onnx)"""
|
||||
if length_scale is None:
|
||||
length_scale = self.config.inference.length_scale
|
||||
|
||||
# Scale length by rate
|
||||
if rate > 0:
|
||||
length_scale /= rate
|
||||
|
||||
if noise_scale is None:
|
||||
noise_scale = self.config.inference.noise_scale
|
||||
|
||||
if noise_w is None:
|
||||
noise_w = self.config.inference.noise_w
|
||||
|
||||
# Create model inputs
|
||||
text_array = np.expand_dims(np.array(phoneme_ids, dtype=np.int64), 0)
|
||||
text_lengths_array = np.array([text_array.shape[1]], dtype=np.int64)
|
||||
scales_array = np.array(
|
||||
[
|
||||
noise_scale,
|
||||
length_scale,
|
||||
noise_w,
|
||||
],
|
||||
dtype=np.float32,
|
||||
)
|
||||
|
||||
inputs = {
|
||||
"input": text_array,
|
||||
"input_lengths": text_lengths_array,
|
||||
"scales": scales_array,
|
||||
}
|
||||
|
||||
speaker_id = 0
|
||||
if self.config.is_multispeaker:
|
||||
if isinstance(speaker, SPEAKER_NAME_TYPE):
|
||||
if self.speaker_map:
|
||||
maybe_speaker_id = self.speaker_map.get(speaker)
|
||||
if maybe_speaker_id is None:
|
||||
try:
|
||||
# Interpret as speaker id
|
||||
speaker_id = int(speaker)
|
||||
except ValueError:
|
||||
_LOGGER.warning(
|
||||
"Unable to find a speaker with the name '%s'. Falling back to first speaker.",
|
||||
speaker,
|
||||
)
|
||||
pass
|
||||
else:
|
||||
speaker_id = maybe_speaker_id
|
||||
elif speaker is not None:
|
||||
speaker_id = speaker
|
||||
|
||||
speaker_id_array = np.array([speaker_id], dtype=np.int64)
|
||||
inputs["sid"] = speaker_id_array
|
||||
|
||||
_LOGGER.debug(
|
||||
"TTS settings: speaker-id=%s, length-scale=%s, noise-scale=%s, noise-w=%s",
|
||||
speaker_id,
|
||||
length_scale,
|
||||
noise_scale,
|
||||
noise_w,
|
||||
)
|
||||
|
||||
# Infer audio from phonemes
|
||||
start_time = time.perf_counter()
|
||||
audio = self.onnx_model.run(None, inputs)[0].squeeze()
|
||||
audio = audio_float_to_int16(audio)
|
||||
end_time = time.perf_counter()
|
||||
|
||||
# Compute real-time factor
|
||||
audio_duration_sec = audio.shape[-1] / self.config.audio.sample_rate
|
||||
infer_sec = end_time - start_time
|
||||
real_time_factor = (
|
||||
infer_sec / audio_duration_sec if audio_duration_sec > 0 else 0.0
|
||||
)
|
||||
|
||||
_LOGGER.debug("RTF: %s", real_time_factor)
|
||||
|
||||
return audio
|
||||
|
||||
@staticmethod
|
||||
def load_from_directory(
|
||||
voice_dir: typing.Union[str, Path],
|
||||
session_options: typing.Optional[onnxruntime.SessionOptions] = None,
|
||||
providers: typing.Optional[
|
||||
typing.Sequence[
|
||||
typing.Union[str, typing.Tuple[str, typing.Dict[str, typing.Any]]]
|
||||
]
|
||||
] = None,
|
||||
share_models: bool = True,
|
||||
use_deterministic_compute: bool = False,
|
||||
) -> "Mimic3Voice":
|
||||
"""Load a Mimic 3 voice from a directory"""
|
||||
voice_dir = Path(voice_dir)
|
||||
_LOGGER.debug("Loading voice from %s", voice_dir)
|
||||
|
||||
config_path = voice_dir / "config.json"
|
||||
_LOGGER.debug("Loading config from %s", config_path)
|
||||
|
||||
with open(config_path, "r", encoding="utf-8") as config_file:
|
||||
config = TrainingConfig.load(config_file)
|
||||
|
||||
# phoneme -> id
|
||||
phoneme_ids_path = voice_dir / "phonemes.txt"
|
||||
_LOGGER.debug("Loading model phonemes from %s", phoneme_ids_path)
|
||||
with open(phoneme_ids_path, "r", encoding="utf-8") as ids_file:
|
||||
phoneme_to_id = phonemes2ids.load_phoneme_ids(ids_file)
|
||||
|
||||
generator_path = voice_dir / "generator.onnx"
|
||||
|
||||
onnx_model: typing.Optional[onnxruntime.InferenceSession] = None
|
||||
|
||||
if share_models:
|
||||
with Mimic3Voice._SHARED_MODELS_LOCK:
|
||||
model_key = str(generator_path.absolute())
|
||||
onnx_model = Mimic3Voice._SHARED_MODELS.get(model_key)
|
||||
|
||||
if onnx_model is None:
|
||||
onnx_model = Mimic3Voice._load_model(
|
||||
generator_path,
|
||||
session_options=session_options,
|
||||
providers=providers,
|
||||
use_deterministic_compute=use_deterministic_compute,
|
||||
)
|
||||
|
||||
Mimic3Voice._SHARED_MODELS[model_key] = onnx_model
|
||||
else:
|
||||
_LOGGER.debug("Using shared Onnx model (%s)", model_key)
|
||||
else:
|
||||
onnx_model = Mimic3Voice._load_model(
|
||||
generator_path,
|
||||
session_options=session_options,
|
||||
providers=providers,
|
||||
use_deterministic_compute=use_deterministic_compute,
|
||||
)
|
||||
|
||||
# phoneme -> phoneme, phoneme, ...
|
||||
phoneme_map: typing.Optional[PHONEME_MAP_TYPE] = None
|
||||
phoneme_map_path = voice_dir / "phoneme_map.txt"
|
||||
if phoneme_map_path.is_file():
|
||||
_LOGGER.debug("Loading phoneme map from %s", phoneme_map_path)
|
||||
with open(phoneme_map_path, "r", encoding="utf-8") as map_file:
|
||||
phoneme_map = phonemes2ids.utils.load_phoneme_map(map_file)
|
||||
|
||||
# id -> speaker
|
||||
speaker_map: typing.Optional[SPEAKER_MAP_TYPE] = None
|
||||
speaker_map_path = voice_dir / "speaker_map.csv"
|
||||
if speaker_map_path.is_file():
|
||||
_LOGGER.debug("Loading speaker map from %s", speaker_map_path)
|
||||
with open(speaker_map_path, "r", encoding="utf-8") as map_file:
|
||||
# id | dataset | name | [alias] | [alias] ...
|
||||
reader = csv.reader(map_file, delimiter="|")
|
||||
speaker_map = {}
|
||||
for row in reader:
|
||||
speaker_id = int(row[0])
|
||||
for alias in row[2:]:
|
||||
speaker_map[alias] = speaker_id
|
||||
|
||||
if config.phonemizer == Phonemizer.GRUUT:
|
||||
# Phonemes from gruut: https://github.com/rhasspy/gruut/
|
||||
return GruutVoice(
|
||||
config=config,
|
||||
onnx_model=onnx_model,
|
||||
phoneme_to_id=phoneme_to_id,
|
||||
phoneme_map=phoneme_map,
|
||||
speaker_map=speaker_map,
|
||||
)
|
||||
|
||||
if config.phonemizer == Phonemizer.ESPEAK:
|
||||
# Phonemes from eSpeak-ng: https://github.com/espeak-ng/espeak-ng
|
||||
voice_class = EspeakVoice
|
||||
|
||||
if config.text_language == "fa":
|
||||
try:
|
||||
# Check if hazm is available
|
||||
# https://github.com/sobhe/hazm
|
||||
import hazm # noqa: F401
|
||||
|
||||
voice_class = HazmEspeakVoice
|
||||
except ImportError:
|
||||
_LOGGER.warning("hazm is highly recommended for language 'fa'")
|
||||
_LOGGER.warning("pip install 'hazm>=0.7.0'")
|
||||
|
||||
return voice_class(
|
||||
config=config,
|
||||
onnx_model=onnx_model,
|
||||
phoneme_to_id=phoneme_to_id,
|
||||
phoneme_map=phoneme_map,
|
||||
speaker_map=speaker_map,
|
||||
)
|
||||
|
||||
if config.phonemizer == Phonemizer.SYMBOLS:
|
||||
# Phonemes are characters from an alphabet
|
||||
return SymbolsVoice(
|
||||
config=config,
|
||||
onnx_model=onnx_model,
|
||||
phoneme_to_id=phoneme_to_id,
|
||||
phoneme_map=phoneme_map,
|
||||
speaker_map=speaker_map,
|
||||
)
|
||||
|
||||
if config.phonemizer == Phonemizer.EPITRAN:
|
||||
# Phonemes are from epitran: https://github.com/dmort27/epitran/
|
||||
return EpitranVoice(
|
||||
config=config,
|
||||
onnx_model=onnx_model,
|
||||
phoneme_to_id=phoneme_to_id,
|
||||
phoneme_map=phoneme_map,
|
||||
speaker_map=speaker_map,
|
||||
)
|
||||
|
||||
raise ValueError(f"Unsupported phonemizer: {config.phonemizer}")
|
||||
|
||||
@staticmethod
|
||||
def _load_model(
|
||||
generator_path: Path,
|
||||
session_options: typing.Optional[onnxruntime.SessionOptions] = None,
|
||||
providers: typing.Optional[
|
||||
typing.Sequence[
|
||||
typing.Union[str, typing.Tuple[str, typing.Dict[str, typing.Any]]]
|
||||
]
|
||||
] = None,
|
||||
use_deterministic_compute: bool = False,
|
||||
) -> onnxruntime.InferenceSession:
|
||||
_LOGGER.debug("Loading model from %s", generator_path)
|
||||
|
||||
# Load onnx model
|
||||
if session_options is None:
|
||||
session_options = onnxruntime.SessionOptions()
|
||||
|
||||
if platform.machine() == "armv7l":
|
||||
# Enabling optimizations on 32-bit ARM crashes
|
||||
session_options.graph_optimization_level = (
|
||||
onnxruntime.GraphOptimizationLevel.ORT_DISABLE_ALL
|
||||
)
|
||||
|
||||
session_options.use_deterministic_compute = use_deterministic_compute
|
||||
|
||||
onnx_model = onnxruntime.InferenceSession(
|
||||
str(generator_path), sess_options=session_options, providers=providers
|
||||
)
|
||||
|
||||
return onnx_model
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
class GruutVoice(Mimic3Voice):
|
||||
"""Voice whose phonemes come from gruut (https://github.com/rhasspy/gruut/)"""
|
||||
|
||||
def text_to_phonemes(
|
||||
self, text: str, text_language: typing.Optional[str] = None
|
||||
) -> TEXT_TO_PHONEMES_TYPE:
|
||||
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
|
||||
for sentence in gruut.sentences(text, lang=text_language):
|
||||
sent_phonemes = [w.phonemes for w in sentence if w.phonemes]
|
||||
if sent_phonemes:
|
||||
yield sent_phonemes, BreakType.UTTERANCE
|
||||
|
||||
def word_to_phonemes(
|
||||
self,
|
||||
word_text: str,
|
||||
word_role: typing.Optional[str] = None,
|
||||
text_language: typing.Optional[str] = None,
|
||||
) -> typing.List[PHONEME_TYPE]:
|
||||
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
|
||||
|
||||
word_role = xmlescape(word_role) if word_role else ""
|
||||
word_text = xmlescape(word_text)
|
||||
|
||||
sentence = next(
|
||||
iter(
|
||||
gruut.sentences(
|
||||
f'<w role="{word_role}">{word_text}</w>',
|
||||
ssml=True,
|
||||
lang=text_language,
|
||||
)
|
||||
)
|
||||
)
|
||||
|
||||
sentence_word = next(iter(sentence))
|
||||
|
||||
return sentence_word.phonemes
|
||||
|
||||
def say_as_to_phonemes(
|
||||
self,
|
||||
text: str,
|
||||
interpret_as: str,
|
||||
say_format: typing.Optional[str] = None,
|
||||
text_language: typing.Optional[str] = None,
|
||||
) -> WORD_PHONEMES_TYPE:
|
||||
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
|
||||
|
||||
word_text = xmlescape(text)
|
||||
interpret_as = xmlescape(interpret_as)
|
||||
format_attr = f'format="{xmlescape(say_format)}"' if say_format else ""
|
||||
|
||||
sentences = gruut.sentences(
|
||||
f'<say-as interpret-as="{interpret_as}" {format_attr}>{word_text}</say-as>',
|
||||
ssml=True,
|
||||
lang=text_language,
|
||||
)
|
||||
|
||||
sent_phonemes: WORD_PHONEMES_TYPE = []
|
||||
|
||||
for sentence in sentences:
|
||||
sent_phonemes.extend(w.phonemes for w in sentence if w.phonemes)
|
||||
|
||||
return sent_phonemes
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
class EspeakVoice(Mimic3Voice):
|
||||
"""Voice whose phonemes come from eSpeak-NG (https://github.com/espeak-ng/espeak-ng)"""
|
||||
|
||||
def __init__(self, *args, **kwargs):
|
||||
super().__init__(*args, **kwargs)
|
||||
self._phonemizer = espeak_phonemizer.Phonemizer()
|
||||
|
||||
def text_to_phonemes(
|
||||
self, text: str, text_language: typing.Optional[str] = None
|
||||
) -> TEXT_TO_PHONEMES_TYPE:
|
||||
phoneme_separator = ""
|
||||
word_separator = self.config.phonemes.word_separator
|
||||
|
||||
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
|
||||
|
||||
voice = self._language_to_voice(text_language)
|
||||
|
||||
phoneme_str = self._phonemizer.phonemize(
|
||||
text,
|
||||
voice=voice,
|
||||
keep_clause_breakers=True,
|
||||
phoneme_separator=phoneme_separator,
|
||||
word_separator=word_separator,
|
||||
punctuation_separator=phoneme_separator,
|
||||
)
|
||||
|
||||
all_word_phonemes = [
|
||||
list(IPA.graphemes(wp_str)) for wp_str in phoneme_str.split(word_separator)
|
||||
]
|
||||
|
||||
minor_break = self.config.phonemes.minor_break
|
||||
major_break = self.config.phonemes.major_break
|
||||
|
||||
if minor_break or major_break:
|
||||
# Split on breaks
|
||||
sent_phonemes = []
|
||||
for word_phonemes in all_word_phonemes:
|
||||
sent_phonemes.append(word_phonemes)
|
||||
|
||||
if minor_break and (word_phonemes[-1] == minor_break):
|
||||
yield sent_phonemes, BreakType.MINOR
|
||||
sent_phonemes = []
|
||||
elif major_break and (word_phonemes[-1] == major_break):
|
||||
yield sent_phonemes, BreakType.MAJOR
|
||||
sent_phonemes = []
|
||||
|
||||
if sent_phonemes:
|
||||
yield sent_phonemes, BreakType.MAJOR
|
||||
else:
|
||||
# No split
|
||||
yield all_word_phonemes, BreakType.UTTERANCE
|
||||
|
||||
def word_to_phonemes(
|
||||
self,
|
||||
word_text: str,
|
||||
word_role: typing.Optional[str] = None,
|
||||
text_language: typing.Optional[str] = None,
|
||||
) -> typing.List[PHONEME_TYPE]:
|
||||
phoneme_separator = ""
|
||||
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
|
||||
|
||||
word_role = xmlescape(word_role) if word_role else ""
|
||||
word_text = xmlescape(word_text)
|
||||
|
||||
voice = self._language_to_voice(text_language)
|
||||
|
||||
phoneme_str = self._phonemizer.phonemize(
|
||||
f'<w role="{word_role}">{word_text}</w>',
|
||||
voice=voice,
|
||||
keep_clause_breakers=True,
|
||||
phoneme_separator=phoneme_separator,
|
||||
punctuation_separator=phoneme_separator,
|
||||
ssml=True,
|
||||
)
|
||||
|
||||
word_phonemes = list(IPA.graphemes(phoneme_str))
|
||||
|
||||
return word_phonemes
|
||||
|
||||
def say_as_to_phonemes(
|
||||
self,
|
||||
text: str,
|
||||
interpret_as: str,
|
||||
say_format: typing.Optional[str] = None,
|
||||
text_language: typing.Optional[str] = None,
|
||||
) -> WORD_PHONEMES_TYPE:
|
||||
phoneme_separator = ""
|
||||
word_separator = self.config.phonemes.word_separator
|
||||
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
|
||||
|
||||
word_text = xmlescape(text)
|
||||
interpret_as = xmlescape(interpret_as)
|
||||
format_attr = f'format="{xmlescape(say_format)}"' if say_format else ""
|
||||
|
||||
voice = self._language_to_voice(text_language)
|
||||
|
||||
phoneme_str = self._phonemizer.phonemize(
|
||||
f'<say-as interpret-as="{interpret_as}" {format_attr}>{word_text}</say-as>',
|
||||
voice=voice,
|
||||
keep_clause_breakers=True,
|
||||
phoneme_separator=phoneme_separator,
|
||||
punctuation_separator=phoneme_separator,
|
||||
word_separator=word_separator,
|
||||
ssml=True,
|
||||
)
|
||||
|
||||
word_phonemes = [
|
||||
list(IPA.graphemes(wp_str)) for wp_str in phoneme_str.split(word_separator)
|
||||
]
|
||||
|
||||
return word_phonemes
|
||||
|
||||
def _language_to_voice(self, language: str) -> str:
|
||||
"""Make voice name from language name"""
|
||||
# en_US -> en-us
|
||||
return language.strip().lower().replace("_", "-")
|
||||
|
||||
|
||||
class HazmEspeakVoice(EspeakVoice):
|
||||
"""Persian espeak-ng voice that uses hazm (https://github.com/sobhe/hazm) for pre-processing"""
|
||||
|
||||
def __init__(self, *args, **kwargs):
|
||||
import gruut_lang_fa
|
||||
import hazm
|
||||
|
||||
super().__init__(*args, **kwargs)
|
||||
|
||||
self._normalizer = hazm.Normalizer()
|
||||
self._sent_tokenizer = hazm.SentenceTokenizer()
|
||||
self._word_tokenizer = hazm.WordTokenizer()
|
||||
|
||||
# Load part of speech tagger from gruut[fa]
|
||||
self._tagger = hazm.POSTagger(
|
||||
model=str(gruut_lang_fa.get_lang_dir() / "pos" / "postagger.model")
|
||||
)
|
||||
|
||||
def text_to_phonemes(
|
||||
self, text: str, text_language: typing.Optional[str] = None
|
||||
) -> TEXT_TO_PHONEMES_TYPE:
|
||||
phoneme_separator = ""
|
||||
word_separator = self.config.phonemes.word_separator
|
||||
|
||||
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
|
||||
voice = self._language_to_voice(text_language)
|
||||
|
||||
# Normalize with hazm
|
||||
sentences = self._preprocess_text(text)
|
||||
|
||||
for sentence in sentences:
|
||||
sent_text = " ".join(sentence)
|
||||
sent_phoneme_str = self._phonemizer.phonemize(
|
||||
sent_text,
|
||||
voice=voice,
|
||||
keep_clause_breakers=True,
|
||||
phoneme_separator=phoneme_separator,
|
||||
word_separator=word_separator,
|
||||
punctuation_separator=phoneme_separator,
|
||||
)
|
||||
|
||||
sent_word_phonemes = [
|
||||
list(IPA.graphemes(wp_str))
|
||||
for wp_str in sent_phoneme_str.split(word_separator)
|
||||
]
|
||||
|
||||
yield sent_word_phonemes, BreakType.UTTERANCE
|
||||
|
||||
def word_to_phonemes(
|
||||
self,
|
||||
word_text: str,
|
||||
word_role: typing.Optional[str] = None,
|
||||
text_language: typing.Optional[str] = None,
|
||||
) -> typing.List[PHONEME_TYPE]:
|
||||
word_text = self._fix_words([word_text])[0]
|
||||
|
||||
return super().word_to_phonemes(
|
||||
word_text, word_role=word_role, text_language=text_language
|
||||
)
|
||||
|
||||
def say_as_to_phonemes(
|
||||
self,
|
||||
text: str,
|
||||
interpret_as: str,
|
||||
say_format: typing.Optional[str] = None,
|
||||
text_language: typing.Optional[str] = None,
|
||||
) -> WORD_PHONEMES_TYPE:
|
||||
sentences = self._preprocess_text(text)
|
||||
text = " ".join(
|
||||
" ".join(word_text for word_text in words) for words in sentences
|
||||
)
|
||||
|
||||
return super().say_as_to_phonemes(
|
||||
text, interpret_as, say_format=say_format, text_language=text_language
|
||||
)
|
||||
|
||||
def _preprocess_text(self, text: str) -> typing.List[typing.List[str]]:
|
||||
"""Split/normalize text into sentences/words with hazm"""
|
||||
text = self._normalizer.normalize(text)
|
||||
processed_sentences = []
|
||||
|
||||
for sentence in self._sent_tokenizer.tokenize(text):
|
||||
words = self._word_tokenizer.tokenize(sentence)
|
||||
processed_words = self._fix_words(words)
|
||||
processed_sentences.append(processed_words)
|
||||
|
||||
return processed_sentences
|
||||
|
||||
def _fix_words(self, words: typing.List[str]) -> typing.List[str]:
|
||||
fixed_words = []
|
||||
|
||||
for word, pos in self._tagger.tag(words):
|
||||
if pos[-1] == "e":
|
||||
if word[-1] != "ِ":
|
||||
if (word[-1] == "ه") and (word[-2] != "ا"):
|
||||
word += "ی"
|
||||
word += "ِ"
|
||||
|
||||
fixed_words.append(word)
|
||||
|
||||
return fixed_words
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
class SymbolsVoice(Mimic3Voice):
|
||||
"""Voice whose phonemes are characters in an alphabet"""
|
||||
|
||||
def text_to_phonemes(
|
||||
self, text: str, text_language: typing.Optional[str] = None
|
||||
) -> TEXT_TO_PHONEMES_TYPE:
|
||||
word_separator = self.config.phonemes.word_separator
|
||||
word_phonemes = [
|
||||
list(IPA.graphemes(wp_str)) for wp_str in text.split(word_separator)
|
||||
]
|
||||
yield word_phonemes, BreakType.UTTERANCE
|
||||
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
|
||||
|
||||
class EpitranVoice(Mimic3Voice):
|
||||
"""Voice whose phonemes come from epitran (https://github.com/dmort27/epitran/)"""
|
||||
|
||||
def __init__(self, *args, **kwargs):
|
||||
super().__init__(*args, **kwargs)
|
||||
self._epis: typing.Dict[str, epitran.Epitran] = {}
|
||||
|
||||
def text_to_phonemes(
|
||||
self, text: str, text_language: typing.Optional[str] = None
|
||||
) -> TEXT_TO_PHONEMES_TYPE:
|
||||
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
|
||||
|
||||
epi = self._epis.get(text_language)
|
||||
if epi is None:
|
||||
epi = epitran.Epitran(text_language)
|
||||
self._epis[text_language] = epi
|
||||
|
||||
phoneme_str = epi.transliterate(text)
|
||||
all_word_phonemes = [
|
||||
list(IPA.graphemes(wp_str)) for wp_str in phoneme_str.split()
|
||||
]
|
||||
|
||||
minor_break = self.config.phonemes.minor_break
|
||||
major_break = self.config.phonemes.major_break
|
||||
|
||||
if minor_break or major_break:
|
||||
# Split on breaks
|
||||
sent_phonemes = []
|
||||
for word_phonemes in all_word_phonemes:
|
||||
sent_phonemes.append(word_phonemes)
|
||||
|
||||
if minor_break and (word_phonemes[-1] == minor_break):
|
||||
yield sent_phonemes, BreakType.MINOR
|
||||
sent_phonemes = []
|
||||
elif major_break and (word_phonemes[-1] == major_break):
|
||||
yield sent_phonemes, BreakType.MAJOR
|
||||
sent_phonemes = []
|
||||
|
||||
if sent_phonemes:
|
||||
yield sent_phonemes, BreakType.MAJOR
|
||||
else:
|
||||
# No split
|
||||
yield all_word_phonemes, BreakType.UTTERANCE
|
||||
1972
mimic3_tts/voices.json
Normal file
1972
mimic3_tts/voices.json
Normal file
File diff suppressed because it is too large
Load diff
Loading…
Add table
Add a link
Reference in a new issue