Refactor into a single package

This commit is contained in:
Michael Hansen 2022-05-02 15:24:11 -04:00
commit 5e866e4fca
94 changed files with 198 additions and 2867 deletions

348
mimic3_tts/README.md Normal file
View file

@ -0,0 +1,348 @@
# Mimic 3
A fast and local neural text to speech system for [Mycroft](https://mycroft.ai/) and the [Mark II](https://mycroft.ai/product/mark-ii/).
* [Available voices](https://github.com/MycroftAI/mimic3-voices)
* [Mimic 3 Architecture](#architecture)
## Command-Line Tools
### mimic3
#### Basic Synthesis
```sh
mimic3 --voice <voice> "<text>" > output.wav
```
where `<voice>` is a [voice key](https://github.com/MycroftAI/mimic3/#voice-keys) like `en_UK/apope_low`.
`<TEXT>` may contain multiple sentences, which will be combined in the final output WAV file. These can also be [split into separate WAV files](#multiple-wav-output).
#### SSML Synthesis
```sh
mimic3 --ssml --voice <voice> "<ssml>" > output.wav
```
where `<ssml>` is valid [SSML](https://www.w3.org/TR/speech-synthesis11/). Not all SSML features are supported, see [the documentation](#ssml) for details.
If your SSML contains `<mark>` tags, add `--mark-file <file>` to the command-line and use `--interactive` mode. As the marks are encountered, their names will be written on separate lines to the file:
```sh
mimic3 --ssml --interactive --mark-file - '<speak>Test 1. <mark name="here" /> Test 2.</speak>'
```
#### Long Texts
If your text is very long, and you would like to listen to it as its being synthesized, use `--interactive` mode:
```sh
mimic3 --interactive < long.txt
```
Each input line will be synthesized and played (see `--play-program`). By default, 5 sentences will be kept in an output queue, only blocking synthesis when the queue is full. You can adjust this value with `--result-queue-size`.
If your long text is fixed-width with blank lines separating paragraphs like those from [Project Gutenberg](https://www.gutenberg.org/), use the `--process-on-blank-line` option so that sentences will not be broken at line boundaries. For example, you can listen to "Alice in Wonderland" like this:
```sh
curl --output - 'https://www.gutenberg.org/files/11/11-0.txt' | \
mimic3 --interactive --process-on-blank-line
```
#### Multiple WAV Output
With `--output-dir` set to a directory, Mimic 3 will output a separate WAV file for each sentence:
```sh
mimic3 'Test 1. Test 2.' --output-dir /path/to/wavs
```
By default, each WAV file will be named using the (slightly modified) text of the sentence. You can have WAV files named using a timestamp instead with `--output-naming time`. For full control of the output naming, the `--csv` command-line flag indicates that each sentence is of the form `id|text` where `id` will be the name of the WAV file.
```sh
cat << EOF |
s01|The birch canoe slid on the smooth planks.
s02|Glue the sheet to the dark blue background.
s03|It's easy to tell the depth of a well.
s04|These days a chicken leg is a rare dish.
s05|Rice is often served in round bowls.
s06|The juice of lemons makes fine punch.
s07|The box was thrown beside the parked truck.
s08|The hogs were fed chopped corn and garbage.
s09|Four hours of steady work faced us.
s10|Large size in stockings is hard to sell.
EOF
mimic3 --csv --output-dir /path/to/wavs
```
You can adjust the delimiter with `--csv-delimiter <delimiter>`.
Additionally, you can use the `--csv-voice` option to specify a different voice or speaker for each line:
```sh
cat << EOF |
s01|#awb|The birch canoe slid on the smooth planks.
s02|#rms|Glue the sheet to the dark blue background.
s03|#slt|It's easy to tell the depth of a well.
s04|#ksp|These days a chicken leg is a rare dish.
s05|#clb|Rice is often served in round bowls.
s06|#aew|The juice of lemons makes fine punch.
s07|#bdl|The box was thrown beside the parked truck.
s08|#lnh|The hogs were fed chopped corn and garbage.
s09|#jmk|Four hours of steady work faced us.
s10|en_UK/apope_low|Large size in stockings is hard to sell.
EOF
mimic3 --voice 'en_US/cmu-arctic_low' --csv-voice --output-dir /path/to/wavs
```
The second contain can contain a `#<speaker>` or an entirely different voice!
#### Interactive Mode
With `--interactive`, Mimic 3 will switch into interactive mode. After entering a sentence, it will be played with `--play-program`.
```sh
mimic3 --interactive
Reading text from stdin...
Hello world!<ENTER>
```
Use `CTRL+D` or `CTRL+C` to exit.
#### Noise and Length Settings
Synthesis has the following additional parameters:
* `--noise-scale` and `--noise-w`
* Determine the speaker volatility during synthesis
* 0-1, default is 0.667 and 0.8 respectively
* `--length-scale` - makes the voice speaker slower (> 1) or faster (< 1)
Individual voices have default settings for these parameters in their `config.json` files (under `inference`).
#### List Voices
```sh
mimic3 --voices
```
#### CUDA Acceleration
If you have a GPU with support for CUDA, you can accelerate synthesis with the `--cuda` flag. This requires you to install the [onnxruntime-gpu](https://pypi.org/project/onnxruntime-gpu/) Python package.
Using [nvidia-docker](https://github.com/NVIDIA/nvidia-docker) is highly recommended. See the `Dockerfile.gpu` file in the parent repository for an example of how to build a compatible container.
### mimic3-download
Mimic 3 automatically downloads voices when they're first used, but you can manually download them too with `mimic3-download`.
For example:
``` sh
mimic3-download 'en_US/*'
```
will download all U.S. English voices to `${HOME}/.local/share/mimic3` (technically `${XDG_DATA_HOME}/mimic3`).
See `mimic3-download --help` for more options.
## SSML
A subset of [SSML](https://www.w3.org/TR/speech-synthesis11/) (Speech Synthesis Markup Language) is supported:
* `<speak>` - wrap around SSML text
* `lang` - set language for document
* `<s>` - sentence (disables automatic sentence breaking)
* `lang` - set language for sentence
* `<w>` / `<token>` - word (disables automatic tokenization)
* `<voice name="...">` - set voice of inner text
* `voice` - name or language of voice
* Name format is `tts:voice` (e.g., "glow-speak:en-us_mary_ann") or `tts:voice#speaker_id` (e.g., "coqui-tts:en_vctk#p228")
* If one of the supported languages, a preferred voice is used (override with `--preferred-voice <lang> <voice>`)
* `<prosody attribute="value">` - change speaking attributes
* Supported `attribute` names:
* `volume` - speaking volume
* number in [0, 100] - 0 is silent, 100 is loudest (default)
* +X, -X, +X%, -X% - absolute/percent offset from current volume
* one of "default", "silent", "x-loud", "loud", "medium", "soft", "x-soft"
* `rate` - speaking rate
* number - 1 is default rate, < 1 is slower, > 1 is faster
* X% - 100% is default rate, 50% is half speed, 200% is twice as fast
* one of "default", "x-fast", "fast", "medium", "slow", "x-slow"
* `<say-as interpret-as="">` - force interpretation of inner text
* `interpret-as` one of "spell-out", "date", "number", "time", or "currency"
* `format` - way to format text depending on `interpret-as`
* number - one of "cardinal", "ordinal", "digits", "year"
* date - string with "d" (cardinal day), "o" (ordinal day), "m" (month), or "y" (year)
* `<break time="">` - Pause for given amount of time
* time - seconds ("123s") or milliseconds ("123ms")
* `<sub alias="">` - substitute `alias` for inner text
* `<phoneme ph="">` - supply phonemes for inner text
* See `phonemes.txt` in voice directory for available phonemes
* Phonemes may need to be separated by whitespace
SSML `<say-as>` support varies between voice types:
* [gruut](https://github.com/rhasspy/gruut/#ssml)
* [eSpeak-ng](http://espeak.sourceforge.net/ssml.html)
* Character-based voices do not currently support `<say-as>`
## Speech Dispatcher
Mimic 3 can be used with the [Orca screen reader](https://help.gnome.org/users/orca/stable/) for Linux via [speech-dispatcher](https://github.com/brailcom/speechd).
After [installing Mimic 3](https://github.com/MycroftAI/mimic3/#installation), make sure you also have speech-dispatcher installed:
``` sh
sudo apt-get install speech-dispatcher
```
Create the file `/etc/speech-dispatcher/modules/mimic3-generic.conf` with the contents:
``` text
GenericExecuteSynth "printf %s \'$DATA\' | /path/to/mimic3 --remote --voice \'$VOICE\' --stdout | $PLAY_COMMAND"
AddVoice "en-us" "MALE1" "en_UK/apope_low"
```
You will need `sudo` access to do this. Make sure to change `/path/to/mimic3` to wherever you installed Mimic 3. Note that the `--remote` option is used to connect to a local Mimic 3 web server (use `--remote <URL>` if your server is somewhere besides `localhost`).
To change the voice later, you only need to replace `en_UK/apope_low`.
Next, edit the existing file `/etc/speech-dispatcher/speechd.conf` and ensure the following settings are present:
``` text
DefaultVoiceType "MALE1"
DefaultModule mimic3-generic
```
Restart speech-dispatcher with:
``` sh
sudo systemctl restart speech-dispatcher
```
and test it out with:
``` sh
spd-say 'Hello from speech dispatcher.'
```
### Systemd Service
To ensure that Mimic 3 runs at boot, create a systemd service at `$HOME/.config/systemd/user/mimic3.service` with the contents:
``` text
[Unit]
Description=Run Mimic 3 web server
Documentation=https://github.com/MycroftAI/mimic3
[Service]
ExecStart=/path/to/mimic3-server
[Install]
WantedBy=default.target
```
Make sure to change `/path/to/mimic3-server` to wherever you installed Mimic 3.
Refresh the systemd services:
``` sh
systemctl --user daemon-reload
```
Now try starting the service:
``` sh
systemctl --user start mimic3
```
If that's successful, ensure it starts at boot:
``` sh
systemctl --user enable mimic3
```
## Architecture
Mimic 3 uses the [VITS](https://arxiv.org/abs/2106.06103), a "Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech". VITS is a combination of the [GlowTTS duration predictor](https://arxiv.org/abs/2005.11129) and the [HiFi-GAN vocoder](https://arxiv.org/abs/2010.05646).
Our implementation is heavily based on [Jaehyeon Kim's PyTorch model](https://github.com/jaywalnut310/vits), with the addition of [Onnx runtime](https://onnxruntime.ai/) export for speed.
![mimic 3 architecture](img/mimic3-architecture.png)
### Phoneme Ids
At a high level, Mimic 3 performs two important tasks:
1. Converting raw text to numeric input for the VITS TTS model, and
2. Using the model to transform numeric input into audio output
The second step is the same for every voice, but the first step (text to numbers) varies. There are currently three implementations of step 1, described below.
### gruut Phoneme-based Voices
Voices that use [gruut](https://github.com/rhasspy/gruut/) for phonemization.
gruut normalizes text and phonemizes words according to a lexicon, with a pre-trained grapheme-to-phoneme model used to guess unknown word pronunciations.
### eSpeak Phoneme-based Voices
Voices that use [eSpeak-ng](https://github.com/espeak-ng/espeak-ng) for phonemization (via [espeak-phonemizer](https://github.com/rhasspy/espeak-phonemizer)).
eSpeak-ng normalizes and phonemizes text using internal rules and lexicons. It supports a large number of languages, and can handle many textual forms.
### Character-based Voices
Voices whose "phonemes" are characters from an alphabet, typically with some punctuation.
For voices whose orthography (writing system) is close enough to its spoken form, character-based voices allow for skipping the phonemization step. However, these voices do not support text normalization, so numbers, dates, etc. must be written out.
### Epitran-based Voices
Voices that use [epitran](https://github.com/dmort27/epitran/) for phonemization.
epitran uses rules to generate phonetic pronunciations from text. It does not support text normalization, however, so numbers, dates, etc. must be written out.
### Components of a Voice Model
Voice models are stored in a directory with a specific layout:
* `<language>_<region>` (e.g., `en_UK`)
* `<voice-name>_<quality>` (e.g., `apope_low`)
* `ALIASES` - alternative names for the voice, one per line (optional)
* `config.json` - training/inference configuration (see [code](https://github.com/MycroftAI/mimic3/blob/master/mimic3-tts/mimic3_tts/config.py) for details)
* `generator.onnx` - exported inference model (see `ids_to_audio` method in [`voice.py`](https://github.com/MycroftAI/mimic3/blob/master/mimic3-tts/mimic3_tts/voice.py))
* `LICENSE` - text, name, or URL of voice model license
* `phoneme_map.txt` - mapping from source phoneme to destination phoneme(s) (optional)
* `phonemes.txt` - mapping from integer ids to phonemes (`_` = padding, `^` = beginning of utterance, `$` = end of utterance, `#` = word break)
* `README.md` - description of the voice
* `SOURCE` - URL(s) of the dataset(s) this voice was trained on
* `VERSION` - version of the voice in the format "MAJOR.Minor.bugfix" (e.g. "1.0.2")
## License
See [license file](LICENSE)

1
mimic3_tts/VERSION Normal file
View file

@ -0,0 +1 @@
0.1.8

19
mimic3_tts/__init__.py Normal file
View file

@ -0,0 +1,19 @@
from pathlib import Path
from opentts_abc import (
AudioResult,
BaseResult,
BaseToken,
MarkResult,
Phonemes,
SayAs,
Voice,
Word,
)
from opentts_abc.ssml import SSMLSpeaker
from ._resources import __version__
from .const import DEFAULT_VOICE
from .tts import Mimic3Settings, Mimic3TextToSpeechSystem
__author__ = "Michael Hansen"

726
mimic3_tts/__main__.py Normal file
View file

@ -0,0 +1,726 @@
#!/usr/bin/env python3
# Copyright 2022 Mycroft AI Inc.
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
#
import argparse
import csv
import io
import logging
import os
import shlex
import shutil
import string
import subprocess
import sys
import tempfile
import threading
import time
import typing
import wave
from dataclasses import dataclass, field
from enum import Enum
from pathlib import Path
from queue import Queue
from ._resources import _PACKAGE
if typing.TYPE_CHECKING:
from . import BaseResult, Mimic3TextToSpeechSystem # noqa: F401
_LOGGER = logging.getLogger(_PACKAGE)
_DEFAULT_PLAY_PROGRAMS = ["paplay", "play -q", "aplay -q"]
# -----------------------------------------------------------------------------
@dataclass
class ResultToProcess:
result: "BaseResult"
line: str
line_id: str = ""
@dataclass
class CommandLineInterfaceState:
args: argparse.Namespace
texts: typing.Optional[typing.Iterable[str]] = None
mark_writer: typing.Optional[typing.TextIO] = None
tts: typing.Optional["Mimic3TextToSpeechSystem"] = None
text_from_stdin: bool = False
all_audio: bytes = field(default_factory=bytes)
sample_rate_hz: int = 22050
sample_width_bytes: int = 2
num_channels: int = 1
result_queue: typing.Optional["Queue[typing.Optional[ResultToProcess]]"] = None
result_thread: typing.Optional[threading.Thread] = None
class OutputNaming(str, Enum):
"""Format used for output file names"""
TEXT = "text"
TIME = "time"
ID = "id"
class StdinFormat(str, Enum):
"""Format of standard input"""
AUTO = "auto"
"""Choose based on SSML state"""
LINES = "lines"
"""Each line is a separate sentence/document"""
DOCUMENT = "document"
"""Entire input is one document"""
# -----------------------------------------------------------------------------
def main():
"""Main entry point"""
args = get_args()
if args.debug:
logging.basicConfig(level=logging.DEBUG)
logging.getLogger().setLevel(logging.DEBUG)
else:
logging.basicConfig(level=logging.INFO)
logging.getLogger().setLevel(logging.INFO)
if args.version:
# Print version and exit
from . import __version__
print(__version__)
sys.exit(0)
state = CommandLineInterfaceState(args=args)
initialize_args(state)
initialize_tts(state)
try:
if args.voices:
# Print voices and exit
print_voices(state)
else:
# Process user input
if os.isatty(sys.stdin.fileno()):
print("Reading text from stdin...", file=sys.stderr)
process_lines(state)
finally:
shutdown_tts(state)
def initialize_args(state: CommandLineInterfaceState):
"""Initialze CLI state from command-line arguments"""
import numpy as np
args = state.args
# Create output directory
if args.output_dir:
args.output_dir = Path(args.output_dir)
args.output_dir.mkdir(parents=True, exist_ok=True)
# Open file for writing the names from <mark> tags in SSML.
# Each name is printed on a single line.
if args.mark_file and (args.mark_file != "-"):
args.mark_file = Path(args.mark_file)
args.mark_file.parent.mkdir(parents=True, exist_ok=True)
state.mark_writer = open( # pylint: disable=consider-using-with
args.mark_file, "w", encoding="utf-8"
)
elif args.stdout:
state.mark_writer = sys.stderr
else:
state.mark_writer = sys.stdout
if args.seed is not None:
_LOGGER.debug("Setting random seed to %s", args.seed)
np.random.seed(args.seed)
if args.csv_voice:
# --csv-voice implies --csv
args.csv = True
if args.csv:
args.output_naming = OutputNaming.ID
elif args.ssml:
# Avoid text mangling when using SSML
args.output_naming = OutputNaming.TIME
# Read text from stdin or arguments
if args.text:
# Use arguments
state.texts = args.text
else:
# Use stdin
state.text_from_stdin = True
stdin_format = StdinFormat.LINES
if (args.stdin_format == StdinFormat.AUTO) and args.ssml:
# Assume SSML input is entire document
stdin_format = StdinFormat.DOCUMENT
if stdin_format == StdinFormat.DOCUMENT:
# One big line
state.texts = [sys.stdin.read()]
else:
# Multiple lines
state.texts = sys.stdin
assert state.texts is not None
if args.process_on_blank_line:
# Combine text until a blank line is encountered.
# Good for line-wrapped books where
# sentences are broken
# up across multiple
# lines.
def process_on_blank_line(lines: typing.Iterable[str]):
text = ""
for line in lines:
line = line.strip()
if not line:
if text:
yield text
text = ""
continue
text += " " + line
state.texts = process_on_blank_line(state.texts)
if args.remote and args.remote.endswith("/"):
# Ensure no slash
args.remote = args.remote[:-1]
if (not args.speaker) and args.voice and ("#" in args.voice):
# Split apart voice
args.voice, args.speaker = args.voice.split("#", maxsplit=1)
if args.deterministic:
# Disable noise
_LOGGER.debug("Disabling noise in deterministic mode")
args.noise_scale = 0.0
args.noise_w = 0.0
def initialize_tts(state: CommandLineInterfaceState):
"""Create Mimic 3 TTS from command-line arguments"""
from mimic3_tts import Mimic3Settings, Mimic3TextToSpeechSystem # noqa: F811
args = state.args
if not args.remote:
# Local TTS
state.tts = Mimic3TextToSpeechSystem(
Mimic3Settings(
length_scale=args.length_scale,
noise_scale=args.noise_scale,
noise_w=args.noise_w,
voices_directories=args.voices_dir,
use_cuda=args.cuda,
use_deterministic_compute=args.deterministic,
)
)
state.tts.voice = args.voice
state.tts.speaker = args.speaker
if args.voices:
# Don't bother with the rest of the initialization
return
if state.tts:
if state.args.voice:
# Set default voice
state.tts.voice = state.args.voice
if state.args.preload_voice:
for voice_key in state.args.preload_voice:
_LOGGER.debug("Preloading voice: %s", voice_key)
state.tts.preload_voice(voice_key)
state.result_queue = Queue(maxsize=args.result_queue_size)
state.result_thread = threading.Thread(
target=process_result, daemon=True, args=(state,)
)
state.result_thread.start()
def process_result(state: CommandLineInterfaceState):
try:
from mimic3_tts import AudioResult, MarkResult
assert state.result_queue is not None
args = state.args
while True:
result_todo = state.result_queue.get()
if result_todo is None:
break
try:
result = result_todo.result
line = result_todo.line
line_id = result_todo.line_id
if isinstance(result, AudioResult):
if args.interactive or args.output_dir:
# Convert to WAV audio
wav_bytes: typing.Optional[bytes] = None
if args.interactive:
if args.stdout:
# Write audio to stdout
sys.stdout.buffer.write(result.audio_bytes)
sys.stdout.buffer.flush()
else:
# Play sound
if not wav_bytes:
wav_bytes = result.to_wav_bytes()
if wav_bytes:
play_wav_bytes(state.args, wav_bytes)
if args.output_dir:
if not wav_bytes:
wav_bytes = result.to_wav_bytes()
# Determine file name
if args.output_naming == OutputNaming.TEXT:
# Use text itself
file_name = line.strip().replace(" ", "_")
file_name = file_name.translate(
str.maketrans(
"", "", string.punctuation.replace("_", "")
)
)
elif args.output_naming == OutputNaming.TIME:
# Use timestamp
file_name = str(time.time())
elif args.output_naming == OutputNaming.ID:
file_name = line_id
assert file_name, f"No file name for text: {line}"
wav_path = args.output_dir / (file_name + ".wav")
wav_path.write_bytes(wav_bytes)
_LOGGER.debug("Wrote %s", wav_path)
else:
# Combine all audio and output to stdout at the end
state.all_audio += result.audio_bytes
state.sample_rate_hz = result.sample_rate_hz
state.sample_width_bytes = result.sample_width_bytes
state.num_channels = result.num_channels
elif isinstance(result, MarkResult):
if state.mark_writer:
print(result.name, file=state.mark_writer)
except Exception:
_LOGGER.exception("Error processing result")
except Exception:
_LOGGER.exception("process_result")
def process_line(
line: str,
state: CommandLineInterfaceState,
line_id: str = "",
line_voice: typing.Optional[str] = None,
):
assert state.result_queue is not None
args = state.args
if state.tts:
# Local TTS
from mimic3_tts import SSMLSpeaker
assert state.tts is not None
args = state.args
if line_voice:
if line_voice.startswith("#"):
# Same voice, but different speaker
state.tts.speaker = line_voice[1:]
else:
# Different voice
state.tts.voice = line_voice
if args.ssml:
results = SSMLSpeaker(state.tts).speak(line)
else:
state.tts.begin_utterance()
# TODO: text language
state.tts.speak_text(line)
results = state.tts.end_utterance()
else:
# Remote TTS
from mimic3_tts import AudioResult
voice: typing.Optional[str] = None
if line_voice:
if line_voice.startswith("#"):
# Same voice, but different speaker
if args.voice:
voice = f"{args.voice}{line_voice}"
else:
# Different voice
voice = line_voice
# Get remote WAV data and repackage as AudioResult
wav_bytes = get_remote_wav_bytes(state, line, voice=voice)
with io.BytesIO(wav_bytes) as wav_io:
wav_reader: wave.Wave_read = wave.open(wav_io, "rb")
with wav_reader as wav_file:
results = [
AudioResult(
sample_rate_hz=wav_file.getframerate(),
sample_width_bytes=wav_file.getsampwidth(),
num_channels=wav_file.getnchannels(),
audio_bytes=wav_file.readframes(wav_file.getnframes()),
)
]
# Add results to processing queue
for result in results:
state.result_queue.put(
ResultToProcess(
result=result,
line=line,
line_id=line_id,
)
)
# Restore voice/speaker
if state.tts:
state.tts.voice = args.voice
state.tts.speaker = args.speaker
def process_lines(state: CommandLineInterfaceState):
assert state.texts is not None
args = state.args
try:
result_idx = 0
for line in state.texts:
line_voice: typing.Optional[str] = None
line_id = ""
line = line.strip()
if not line:
continue
if args.output_naming == OutputNaming.ID:
# Line has the format id|text instead of just text
with io.StringIO(line) as line_io:
reader = csv.reader(line_io, delimiter=args.csv_delimiter)
row = next(reader)
line_id, line = row[0], row[-1]
if args.csv_voice:
line_voice = row[1]
process_line(line, state, line_id=line_id, line_voice=line_voice)
result_idx += 1
except KeyboardInterrupt:
if state.result_queue is not None:
# Draw audio playback queue
while not state.result_queue.empty():
state.result_queue.get()
finally:
# Wait for raw stream to finish
if state.result_queue is not None:
state.result_queue.put(None)
if state.result_thread is not None:
state.result_thread.join()
# -------------------------------------------------------------------------
# Write combined audio to stdout
if state.all_audio:
_LOGGER.debug("Writing WAV audio to stdout")
if sys.stdout.isatty() and (not state.args.stdout):
with io.BytesIO() as wav_io:
wav_file_play: wave.Wave_write = wave.open(wav_io, "wb")
with wav_file_play:
wav_file_play.setframerate(state.sample_rate_hz)
wav_file_play.setsampwidth(state.sample_width_bytes)
wav_file_play.setnchannels(state.num_channels)
wav_file_play.writeframes(state.all_audio)
play_wav_bytes(state.args, wav_io.getvalue())
else:
# Write output directly to stdout
wav_file_write: wave.Wave_write = wave.open(sys.stdout.buffer, "wb")
with wav_file_write:
wav_file_write.setframerate(state.sample_rate_hz)
wav_file_write.setsampwidth(state.sample_width_bytes)
wav_file_write.setnchannels(state.num_channels)
wav_file_write.writeframes(state.all_audio)
sys.stdout.buffer.flush()
def shutdown_tts(state: CommandLineInterfaceState):
if state.tts:
state.tts.shutdown()
state.tts = None
def play_wav_bytes(args: argparse.Namespace, wav_bytes: bytes):
with tempfile.NamedTemporaryFile(mode="wb+", suffix=".wav") as wav_file:
wav_file.write(wav_bytes)
wav_file.seek(0)
for play_program in reversed(args.play_program):
play_cmd = shlex.split(play_program)
if not shutil.which(play_cmd[0]):
continue
play_cmd.append(wav_file.name)
_LOGGER.debug("Playing WAV file: %s", play_cmd)
subprocess.check_output(play_cmd)
break
def print_voices(state: CommandLineInterfaceState):
if state.tts:
# Local TTS
voices = list(state.tts.get_voices())
voices = sorted(voices, key=lambda v: v.key)
else:
# Remove TTS
voices = get_remote_voices(state)
writer = csv.writer(sys.stdout, delimiter="\t")
writer.writerow(("KEY", "LANGUAGE", "NAME", "DESCRIPTION", "LOCATION"))
for voice in voices:
writer.writerow(
(voice.key, voice.language, voice.name, voice.description, voice.location)
)
# -----------------------------------------------------------------------------
def get_remote_voices(state: CommandLineInterfaceState) -> typing.List:
import requests
from mimic3_tts import Voice
args = state.args
url = f"{args.remote}/api/voices"
_LOGGER.debug("Getting voices from remote server at %s", url)
voices_json = requests.get(url).json()
return [Voice(**voice_args) for voice_args in voices_json]
def get_remote_wav_bytes(
state: CommandLineInterfaceState,
text: str,
voice: typing.Optional[str] = None,
) -> bytes:
import requests
args = state.args
if args.ssml:
headers = {"Content-Type": "application/ssml+xml"}
else:
headers = {"Content-Type": "text/plain"}
params: typing.Dict[str, str] = {}
if voice:
params["voice"] = voice
elif args.voice:
if args.speaker:
params["voice"] = f"{args.voice}#{args.speaker}"
else:
params["voice"] = args.voice
if args.length_scale:
params["lengthScale"] = args.length_scale
if args.noise_scale:
params["noiseScale"] = args.noise_scale
if args.noise_w:
params["noiseW"] = args.noise_w
url = f"{args.remote}/api/tts"
_LOGGER.debug("Synthesizing text remotely at %s", url)
wav_bytes = requests.post(url, headers=headers, params=params, data=text).content
return wav_bytes
# -----------------------------------------------------------------------------
def get_args(argv=None):
"""Parse command-line arguments"""
parser = argparse.ArgumentParser(
prog=_PACKAGE, description="Mimic 3 command-line interface"
)
parser.add_argument(
"text", nargs="*", help="Text to convert to speech (default: stdin)"
)
parser.add_argument(
"--remote",
nargs="?",
const="http://localhost:59125",
help="Connect to Mimic 3 HTTP web server for synthesis (default: localhost)",
)
parser.add_argument(
"--stdin-format",
choices=[str(v.value) for v in StdinFormat],
default=StdinFormat.AUTO,
help="Format of stdin text (default: auto)",
)
parser.add_argument(
"--voice",
"-v",
help="Name of voice (expected in <voices-dir>/<language>)",
)
parser.add_argument(
"--speaker",
"-s",
help="Name or number of speaker (default: first speaker)",
)
parser.add_argument(
"--voices-dir",
action="append",
help="Directory with voices (format is <language>/<voice_name>)",
)
parser.add_argument("--voices", action="store_true", help="List available voices")
parser.add_argument("--output-dir", help="Directory to write WAV file(s)")
parser.add_argument(
"--output-naming",
choices=[v.value for v in OutputNaming],
default="text",
help="Naming scheme for output WAV files (requires --output-dir)",
)
parser.add_argument(
"--id-delimiter",
default="|",
help="Delimiter between id and text in lines (default: |). Requires --output-naming id",
)
parser.add_argument(
"--interactive",
action="store_true",
help="Play audio after each input line (see --play-program)",
)
parser.add_argument("--csv", action="store_true", help="Input format is id|text")
parser.add_argument(
"--csv-delimiter", default="|", help="Delimiter used with --csv (default: |)"
)
parser.add_argument(
"--csv-voice",
action="store_true",
help="Input format is id|voice|text or id|#speaker|text",
)
parser.add_argument(
"--mark-file",
help="File to write mark names to as they're encountered (--ssml only)",
)
parser.add_argument(
"--noise-scale",
type=float,
help="Noise scale [0-1], default is 0.667",
)
parser.add_argument(
"--length-scale",
type=float,
help="Length scale (1.0 is default speed, 0.5 is 2x faster)",
)
parser.add_argument(
"--noise-w",
type=float,
help="Variation in cadence [0-1], default is 0.8",
)
# Miscellaneous
parser.add_argument(
"--result-queue-size",
default=5,
help="Maximum number of sentences to maintain in output queue (default: 5)",
)
parser.add_argument(
"--process-on-blank-line",
action="store_true",
help="Process text only after encountering a blank line",
)
parser.add_argument("--ssml", action="store_true", help="Input text is SSML")
parser.add_argument(
"--stdout",
action="store_true",
help="Force audio output to stdout even if a tty is detected",
)
parser.add_argument(
"--preload-voice", action="append", help="Preload voice when starting up"
)
parser.add_argument(
"--play-program",
action="append",
default=_DEFAULT_PLAY_PROGRAMS,
help="Program(s) used to play WAV files",
)
parser.add_argument(
"--cuda",
action="store_true",
help="Use Onnx CUDA execution provider (requires onnxruntime-gpu)",
)
parser.add_argument(
"--deterministic",
action="store_true",
help="Ensure that the same audio is always synthesized from the same text",
)
parser.add_argument("--seed", type=int, help="Set random seed (default: not set)")
parser.add_argument("--version", action="store_true", help="Print version and exit")
parser.add_argument(
"--debug", action="store_true", help="Print DEBUG messages to the console"
)
return parser.parse_args(args=argv)
# -----------------------------------------------------------------------------
if __name__ == "__main__":
main()

51
mimic3_tts/_resources.py Normal file
View file

@ -0,0 +1,51 @@
# Copyright 2022 Mycroft AI Inc.
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
#
"""Shared access to package resources"""
import json
import os
import typing
from pathlib import Path
try:
import importlib.resources
files = importlib.resources.files
except (ImportError, AttributeError):
# Backport for Python < 3.9
import importlib_resources # type: ignore
files = importlib_resources.files
_PACKAGE = "mimic3_tts"
_DIR = Path(typing.cast(os.PathLike, files(_PACKAGE)))
__version__ = (_DIR / "VERSION").read_text(encoding="utf-8").strip()
# Load voices.json
# {
# "<lang>/<voice>": {
# "files": {
# "relative/path": {
# "size_bytes": size in bytes,
# "sha256_sum": sha256 hash
# }
# },
# "speakers": [],
# "properties": {}
# }
# }
with open(_DIR / "voices.json", "r", encoding="utf-8") as voices_file:
_VOICES = json.load(voices_file)

356
mimic3_tts/config.py Normal file
View file

@ -0,0 +1,356 @@
# Copyright 2022 Mycroft AI Inc.
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
#
"""Configuration classes"""
import collections
import json
import typing
from dataclasses import dataclass, field
from enum import Enum
from pathlib import Path
import numpy as np
from dataclasses_json import DataClassJsonMixin
from gruut_ipa import IPA
from phonemes2ids import BlankBetween
@dataclass
class AudioConfig(DataClassJsonMixin):
"""Audio input/output details"""
filter_length: int = 1024
hop_length: int = 256
win_length: int = 1024
mel_channels: int = 80
sample_rate: int = 22050
sample_bytes: int = 2
channels: int = 1
mel_fmin: float = 0.0
mel_fmax: typing.Optional[float] = None
ref_level_db: float = 20.0
spec_gain: float = 1.0
# Normalization
signal_norm: bool = True
min_level_db: float = -100.0
max_norm: float = 1.0
clip_norm: bool = True
symmetric_norm: bool = True
do_dynamic_range_compression: bool = True
convert_db_to_amp: bool = True
do_trim_silence: bool = False
trim_silence_db: float = 40.0
trim_margin_sec: float = 0.01
trim_keep_sec: float = 0.25
scale_mels: bool = False
def __post_init__(self):
if self.mel_fmax is not None:
assert self.mel_fmax <= self.sample_rate // 2
# -------------------------------------------------------------------------
# Normalization
# -------------------------------------------------------------------------
def normalize(self, mel_db: np.ndarray) -> np.ndarray:
"""Put values in [0, max_norm] or [-max_norm, max_norm]"""
mel_norm = ((mel_db - self.ref_level_db) - self.min_level_db) / (
-self.min_level_db
)
if self.symmetric_norm:
# Symmetric norm
mel_norm = ((2 * self.max_norm) * mel_norm) - self.max_norm
if self.clip_norm:
mel_norm = np.clip(mel_norm, -self.max_norm, self.max_norm)
else:
# Asymmetric norm
mel_norm = self.max_norm * mel_norm
if self.clip_norm:
mel_norm = np.clip(mel_norm, 0, self.max_norm)
return mel_norm
def denormalize(self, mel_db: np.ndarray) -> np.ndarray:
"""Pull values out of [0, max_norm] or [-max_norm, max_norm]"""
if self.symmetric_norm:
# Symmetric norm
if self.clip_norm:
mel_denorm = np.clip(mel_db, -self.max_norm, self.max_norm)
mel_denorm = (
(mel_denorm + self.max_norm) * -self.min_level_db / (2 * self.max_norm)
) + self.min_level_db
else:
# Asymmetric norm
if self.clip_norm:
mel_denorm = np.clip(mel_db, 0, self.max_norm)
mel_denorm = (
mel_denorm * -self.min_level_db / self.max_norm
) + self.min_level_db
mel_denorm += self.ref_level_db
return mel_denorm
@dataclass
class ModelConfig(DataClassJsonMixin):
"""TTS model hyperparameters"""
num_symbols: int = 0
n_speakers: int = 1
inter_channels: int = 192
hidden_channels: int = 192
filter_channels: int = 768
n_heads: int = 2
n_layers: int = 6
kernel_size: int = 3
p_dropout: float = 0.1
resblock: str = "1"
resblock_kernel_sizes: typing.Tuple[int, ...] = (3, 7, 11)
resblock_dilation_sizes: typing.Tuple[typing.Tuple[int, ...], ...] = (
(1, 3, 5),
(1, 3, 5),
(1, 3, 5),
)
upsample_rates: typing.Tuple[int, ...] = (8, 8, 2, 2)
upsample_initial_channel: int = 512
upsample_kernel_sizes: typing.Tuple[int, ...] = (16, 16, 4, 4)
n_layers_q: int = 3
use_spectral_norm: bool = False
gin_channels: int = 0 # single speaker
use_sdp: bool = True # StochasticDurationPredictor
@property
def is_multispeaker(self) -> bool:
return self.n_speakers > 1
@dataclass
class PhonemesConfig(DataClassJsonMixin):
"""Phonemes to ids configuration"""
phoneme_separator: str = " "
"""Separator between individual phonemes in CSV input"""
word_separator: str = "#"
"""Separator between word phonemes in CSV input (must not match phoneme_separator)"""
phoneme_to_id: typing.Optional[typing.Dict[str, int]] = None
pad: typing.Optional[str] = "_"
bos: typing.Optional[str] = None
eos: typing.Optional[str] = None
blank: typing.Optional[str] = "#"
blank_word: typing.Optional[str] = None
blank_between: typing.Union[str, BlankBetween] = BlankBetween.WORDS
blank_at_start: bool = True
blank_at_end: bool = True
simple_punctuation: bool = True
punctuation_map: typing.Optional[typing.Dict[str, str]] = None
separate: typing.Optional[typing.List[str]] = None
separate_graphemes: bool = False
separate_tones: bool = False
tone_before: bool = False
phoneme_map: typing.Optional[typing.Dict[str, str]] = None
auto_bos_eos: bool = False
minor_break: typing.Optional[str] = IPA.BREAK_MINOR.value
major_break: typing.Optional[str] = IPA.BREAK_MAJOR.value
break_phonemes_into_graphemes: bool = False
drop_stress: bool = False
symbols: typing.Optional[typing.List[str]] = None
def split_word_phonemes(self, phonemes_str: str) -> typing.List[typing.List[str]]:
"""Split phonemes string into a list of lists (outer is words, inner is individual phonemes in each word)"""
return [
word_phonemes_str.split(self.phoneme_separator)
for word_phonemes_str in phonemes_str.split(self.word_separator)
]
def join_word_phonemes(self, word_phonemes: typing.List[typing.List[str]]) -> str:
"""Split phonemes string into a list of lists (outer is words, inner is individual phonemes in each word)"""
return self.word_separator.join(
self.phoneme_separator.join(wp) for wp in word_phonemes
)
class Phonemizer(str, Enum):
"""Method used to convert text to phonemes"""
SYMBOLS = "symbols"
GRUUT = "gruut"
ESPEAK = "espeak"
EPITRAN = "epitran"
class Aligner(str, Enum):
"""Text/audio aligner"""
KALDI_ALIGN = "kaldi_align"
"""https://github.com/rhasspy/kaldi-align"""
class TextCasing(str, Enum):
"""Casing method applied to text"""
LOWER = "lower"
UPPER = "upper"
class MetadataFormat(str, Enum):
"""Format of training metadata"""
TEXT = "text"
PHONEMES = "phonemes"
PHONEME_IDS = "ids"
@dataclass
class DatasetConfig:
"""Training dataset configuration"""
name: str
metadata_format: MetadataFormat = MetadataFormat.TEXT
multispeaker: bool = False
text_language: typing.Optional[str] = None
audio_dir: typing.Optional[typing.Union[str, Path]] = None
cache_dir: typing.Optional[typing.Union[str, Path]] = None
def get_cache_dir(self, output_dir: typing.Union[str, Path]) -> Path:
if self.cache_dir is not None:
cache_dir = Path(self.cache_dir)
else:
cache_dir = Path("cache") / self.name
if not cache_dir.is_absolute():
cache_dir = Path(output_dir) / str(cache_dir)
return cache_dir
@dataclass
class AlignerConfig:
"""Text/audio alignment configuration"""
aligner: typing.Optional[Aligner] = None
casing: typing.Optional[TextCasing] = None
@dataclass
class InferenceConfig:
"""Inference configuration"""
length_scale: float = 1.0
noise_scale: float = 0.667
noise_w: float = 0.8
minor_break_ms: typing.Optional[int] = None
major_break_ms: typing.Optional[int] = None
@dataclass
class TrainingConfig(DataClassJsonMixin):
"""Master configuration for training"""
seed: int = 1234
epochs: int = 10000
learning_rate: float = 2e-4
betas: typing.Tuple[float, float] = field(default=(0.8, 0.99))
eps: float = 1e-9
batch_size: int = 32
fp16_run: bool = False
lr_decay: float = 0.999875
segment_size: int = 8192
init_lr_ratio: float = 1.0
warmup_epochs: int = 0
c_mel: int = 45
c_kl: float = 1.0
grad_clip: typing.Optional[float] = None
min_seq_length: typing.Optional[int] = None
max_seq_length: typing.Optional[int] = None
min_spec_length: typing.Optional[int] = None
max_spec_length: typing.Optional[int] = None
min_speaker_utterances: typing.Optional[int] = None
last_epoch: int = 1
global_step: int = 1
best_loss: typing.Optional[float] = None
audio: AudioConfig = field(default_factory=AudioConfig)
model: ModelConfig = field(default_factory=ModelConfig)
phonemes: PhonemesConfig = field(default_factory=PhonemesConfig)
text_aligner: AlignerConfig = field(default_factory=AlignerConfig)
text_language: typing.Optional[str] = None
phonemizer: typing.Optional[Phonemizer] = None
datasets: typing.List[DatasetConfig] = field(default_factory=list)
inference: InferenceConfig = field(default_factory=InferenceConfig)
version: int = 1
git_commit: str = ""
@property
def is_multispeaker(self):
return self.model.is_multispeaker or any(d.multispeaker for d in self.datasets)
def save(self, config_file: typing.TextIO):
"""Save config as JSON to a file"""
json.dump(self.to_dict(), config_file, indent=4)
@staticmethod
def load(config_file: typing.TextIO) -> "TrainingConfig":
"""Load config from a JSON file"""
return TrainingConfig.from_json(config_file.read())
@staticmethod
def load_and_merge(
config: "TrainingConfig",
config_files: typing.Iterable[typing.Union[str, Path, typing.TextIO]],
) -> "TrainingConfig":
"""Loads one or more JSON configuration files and overlays them on top of an existing config"""
base_dict = config.to_dict()
for maybe_config_file in config_files:
if isinstance(maybe_config_file, (str, Path)):
# File path
config_file = open(maybe_config_file, "r", encoding="utf-8")
else:
# File object
config_file = maybe_config_file
with config_file:
# Load new config and overlay on existing config
new_dict = json.load(config_file)
TrainingConfig.recursive_update(base_dict, new_dict)
return TrainingConfig.from_dict(base_dict)
@staticmethod
def recursive_update(
base_dict: typing.Dict[typing.Any, typing.Any],
new_dict: typing.Mapping[typing.Any, typing.Any],
) -> None:
"""Recursively overwrites values in base dictionary with values from new dictionary"""
for key, value in new_dict.items():
if isinstance(value, collections.Mapping) and (
base_dict.get(key) is not None
):
TrainingConfig.recursive_update(base_dict[key], value)
else:
base_dict[key] = value

28
mimic3_tts/const.py Normal file
View file

@ -0,0 +1,28 @@
# Copyright 2022 Mycroft AI Inc.
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
#
from pathlib import Path
from xdgenvpy import XDG
DEFAULT_VOICE = "en_UK/apope_low"
DEFAULT_LANGUAGE = "en_UK"
DEFAULT_VOICES_URL_FORMAT = (
"https://github.com/MycroftAI/mimic3-voices/raw/master/voices/{lang}/{name}"
)
DEFAULT_VOICES_DOWNLOAD_DIR = Path(XDG().XDG_DATA_HOME) / "mimic3" / "voices"
DEFAULT_VOLUME = 100.0
DEFAULT_RATE = 1.0

240
mimic3_tts/download.py Normal file
View file

@ -0,0 +1,240 @@
# Copyright 2022 Mycroft AI Inc.
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
#
"""A command-line tool for downloading Mimic 3 voices"""
import argparse
import itertools
import json
import logging
import re
import sys
import typing
import urllib.request
from dataclasses import dataclass
from pathlib import Path
from urllib.error import HTTPError
from ._resources import _PACKAGE, _VOICES
from .const import DEFAULT_VOICES_DOWNLOAD_DIR, DEFAULT_VOICES_URL_FORMAT
from .utils import file_sha256_sum, wildcard_to_regex
_LOGGER = logging.getLogger(__name__)
_WILDCARD = "*"
# -----------------------------------------------------------------------------
class VoiceDownloadError(Exception):
"""Occurs when a voice fails to download"""
@dataclass
class VoiceFile:
"""File associated with a voice to download"""
relative_path: str
size_bytes: typing.Optional[int] = None
sha256_sum: typing.Optional[str] = None
def is_later_version(version1: str, version2: str) -> bool:
"""True if version1 is later than version2"""
v1_parts = [int(n) for n in version1.split(".")]
v2_parts = [int(n) for n in version2.split(".")]
for p1, p2 in itertools.zip_longest(v1_parts, v2_parts, fillvalue=0):
if p1 > p2:
# 2.0 vs 1.0
return True
if p1 < p2:
# 1.0 vs 2.0
return False
# 1.0 vs 1.0
return False
def download_voice(
voice_key: str,
url_base: str,
voice_files: typing.Iterable[VoiceFile],
voices_dir: typing.Union[str, Path],
voice_version: str,
chunk_bytes: int = 4096,
redownload: bool = False,
):
"""Downloads a voice to a directory"""
from tqdm.auto import tqdm
if url_base.endswith("/"):
# Remove final slash
url_base = url_base[:-1]
voice_dir = Path(voices_dir) / voice_key
voice_dir.mkdir(parents=True, exist_ok=True)
_LOGGER.debug("Downloading voice %s to %s", voice_key, voice_dir)
version_path = voice_dir / "VERSION"
if version_path.is_file():
actual_version = version_path.read_text(encoding="utf-8").strip()
if is_later_version(voice_version, actual_version):
redownload = True
_LOGGER.debug(
"Replacing version %s of %s with version %s",
actual_version,
voice_key,
voice_version,
)
for voice_file in voice_files:
file_url = f"{url_base}/{voice_file.relative_path}"
file_path = voice_dir / voice_file.relative_path
if (not redownload) and voice_file.sha256_sum and file_path.is_file():
# Check if file exists and has correct sha256
expected_sha256 = voice_file.sha256_sum
with open(file_path, "rb") as check_file:
actual_sha256 = file_sha256_sum(check_file)
if actual_sha256 == expected_sha256:
_LOGGER.debug("Skipping download of %s (sha256 match)", file_path)
continue
try:
# Download file, show progress with tqdm
with urllib.request.urlopen(file_url) as response:
with open(file_path, mode="wb") as dest_file:
with tqdm(
unit="B",
unit_scale=True,
unit_divisor=1024,
miniters=1,
desc=voice_file.relative_path,
total=int(response.headers.get("content-length", 0)),
) as pbar:
chunk = response.read(chunk_bytes)
while chunk:
dest_file.write(chunk)
pbar.update(len(chunk))
chunk = response.read(chunk_bytes)
_LOGGER.debug("Downloaded %s", file_path)
except HTTPError as e:
_LOGGER.exception("download_voice")
raise VoiceDownloadError(
f"Failed to download file for voice {voice_key} from {file_url}: {e}"
) from e
# -----------------------------------------------------------------------------
def main(argv=None):
"""Main entry point"""
parser = argparse.ArgumentParser(
prog=f"{_PACKAGE}.download", description="Download utility for Mimic 3 voices"
)
parser.add_argument(
"key",
nargs="*",
help="Keys of voices to download (e.g., en_US/vctk_low). May contain wildcards (*)",
)
parser.add_argument(
"--output-dir",
default=DEFAULT_VOICES_DOWNLOAD_DIR,
help="Path to output directory",
)
parser.add_argument(
"--url-format",
default=DEFAULT_VOICES_URL_FORMAT,
help="URL format string for voices (contains {key}, {lang}, {name})",
)
parser.add_argument(
"--redownload",
action="store_true",
help="Force re-downloading of files if they already exist",
)
parser.add_argument(
"--debug", action="store_true", help="Print DEBUG messages to console"
)
args = parser.parse_args(args=argv)
if args.debug:
logging.basicConfig(level=logging.DEBUG)
logging.getLogger().setLevel(logging.DEBUG)
else:
logging.basicConfig(level=logging.INFO)
logging.getLogger().setLevel(logging.INFO)
_LOGGER.debug(args)
args.output_dir = Path(args.output_dir)
args.key = args.key or []
if not args.key:
# Print available voices and exit
json.dump(_VOICES, sys.stdout, indent=4, ensure_ascii=False)
sys.exit(0)
args.key = [
wildcard_to_regex(key, wildcard=_WILDCARD) if _WILDCARD in key else key
for key in args.key
]
args.output_dir.mkdir(parents=True, exist_ok=True)
for key_or_pattern in args.key:
if isinstance(key_or_pattern, re.Pattern):
# Wildcards
voice_keys = []
for maybe_key in _VOICES.keys():
if key_or_pattern.match(maybe_key):
voice_keys.append(maybe_key)
_LOGGER.debug("%s matched %s", key_or_pattern, voice_keys)
else:
# No wildcards
voice_keys = [key_or_pattern]
for voice_key in voice_keys:
voice_lang, voice_name = voice_key.split("/", maxsplit=1)
voice_info = _VOICES[voice_key]
voice_url = str.format(
args.url_format, key=voice_key, lang=voice_lang, name=voice_name
)
voice_files = voice_info["files"]
_LOGGER.info("Downloading %s", voice_key)
download_voice(
voice_key=voice_key,
url_base=voice_url,
voice_files=[
VoiceFile(file_key, sha256_sum=file_info.get("sha256_sum"))
for file_key, file_info in voice_files.items()
],
voice_version=voice_info["version"],
voices_dir=args.output_dir,
redownload=args.redownload,
)
# -----------------------------------------------------------------------------
if __name__ == "__main__":
main()

0
mimic3_tts/py.typed Normal file
View file

582
mimic3_tts/tts.py Normal file
View file

@ -0,0 +1,582 @@
# Copyright 2022 Mycroft AI Inc.
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
#
"""Implementation of OpenTTS for Mimic 3"""
import audioop
import itertools
import logging
import typing
from copy import deepcopy
from dataclasses import dataclass, field
from pathlib import Path
from gruut_ipa import IPA
from xdgenvpy import XDG
from opentts_abc import (
AudioResult,
BaseResult,
BaseToken,
MarkResult,
Phonemes,
SayAs,
TextToSpeechSystem,
Voice,
Word,
)
from ._resources import _VOICES
from .config import TrainingConfig
from .const import (
DEFAULT_LANGUAGE,
DEFAULT_RATE,
DEFAULT_VOICE,
DEFAULT_VOICES_DOWNLOAD_DIR,
DEFAULT_VOICES_URL_FORMAT,
DEFAULT_VOLUME,
)
from .download import VoiceFile, download_voice
from .voice import SPEAKER_TYPE, BreakType, Mimic3Voice
_DIR = Path(__file__).parent
_LOGGER = logging.getLogger(__name__)
PHONEMES_LIST_TYPE = typing.List[typing.List[str]]
# -----------------------------------------------------------------------------
@dataclass
class Mimic3Settings:
"""Settings for Mimic 3 text to speech system"""
voice: typing.Optional[str] = None
"""Default voice key"""
language: typing.Optional[str] = None
"""Default language (e.g., "en_US")"""
voices_directories: typing.Optional[typing.Iterable[typing.Union[str, Path]]] = None
"""Directories to search for voices (<lang>/<voice>)"""
voices_url_format: typing.Optional[str] = DEFAULT_VOICES_URL_FORMAT
"""URL format string for a voice directory.
May contain:
* {key} - unique voice key
* {lang} - voice language
* {name} - voice name
"""
speaker: typing.Optional[SPEAKER_TYPE] = None
"""Default speaker name or id"""
length_scale: typing.Optional[float] = None
"""Default length scale (use voice config if None)"""
noise_scale: typing.Optional[float] = None
"""Default noise scale (use voice config if None)"""
noise_w: typing.Optional[float] = None
"""Default noise W (use voice config if None)"""
text_language: typing.Optional[str] = None
"""Language of text (use voice language if None)"""
sample_rate: int = 22050
"""Sample rate of silence from add_break() in Hertz"""
voices_download_dir: typing.Union[str, Path] = DEFAULT_VOICES_DOWNLOAD_DIR
"""Directory to download voices to"""
no_download: bool = False
"""Do not download voices automatically"""
use_cuda: bool = False
"""Use CUDA GPU acceleration (requires onnxruntime-gpu)"""
share_onnx_models_between_threads: bool = True
"""If True, Onnx models are shared between threads"""
volume: float = DEFAULT_VOLUME
"""Voice volume in [0, 100]"""
rate: float = DEFAULT_RATE
"""Voice speaking rate (< 1 is slower, > 1 is faster)"""
use_deterministic_compute: bool = False
"""Force onnxruntime to use deterministic compute mode. For fully deterministic synthesis, also set noise_scale and noise_w to 0."""
@dataclass
class Mimic3Phonemes:
"""Pending task to synthesize audio from phonemes with specific settings"""
current_settings: Mimic3Settings
"""Settings used to synthesize audio"""
phonemes: typing.List[typing.List[str]] = field(default_factory=list)
"""Phonemes for synthesis"""
is_utterance: bool = True
"""True if this is the end of a full utterance"""
class VoiceNotFoundError(Exception):
"""Raised if a voice cannot be found"""
def __init__(self, voice: str):
super().__init__(f"Voice not found: {voice}")
# -----------------------------------------------------------------------------
class Mimic3TextToSpeechSystem(TextToSpeechSystem):
"""Convert text to speech using Mimic 3"""
def __init__(self, settings: Mimic3Settings):
self.settings = settings
self._results: typing.List[typing.Union[BaseResult, Mimic3Phonemes]] = []
self._loaded_voices: typing.Dict[str, Mimic3Voice] = {}
@staticmethod
def get_default_voices_directories() -> typing.List[Path]:
"""Get list of directories to search for voices by default.
On Linux, this is typically:
- $HOME/.local/share/mimic3/voices
- /usr/local/share/mimic3/voices
- /usr/share/mimic3/voices
"""
return [Path(d) / "mimic3" / "voices" for d in XDG().XDG_DATA_DIRS.split(":")]
def get_voices(self) -> typing.Iterable[Voice]:
"""Returns an iterable of all available voices"""
voices_dirs: typing.Iterable[
typing.Union[str, Path]
] = Mimic3TextToSpeechSystem.get_default_voices_directories()
if self.settings.voices_directories is not None:
voices_dirs = itertools.chain(self.settings.voices_directories, voices_dirs)
known_voices = set(_VOICES.keys())
# voices/<language>/<voice>/
for voices_dir in voices_dirs:
voices_dir = Path(voices_dir)
if not voices_dir.is_dir() or voices_dir.name.startswith("."):
_LOGGER.debug("Skipping voice directory %s", voices_dir)
continue
_LOGGER.debug("Searching %s for voices", voices_dir)
for lang_dir in voices_dir.iterdir():
if not lang_dir.is_dir() or lang_dir.name.startswith("."):
continue
for voice_dir in lang_dir.iterdir():
if not voice_dir.is_dir() or voice_dir.name.startswith("."):
continue
config_path = voice_dir / "config.json"
if not config_path.is_file():
continue
_LOGGER.debug("Voice found in %s", voice_dir)
voice_lang = lang_dir.name
# Load config
_LOGGER.debug("Loading config from %s", config_path)
with open(config_path, "r", encoding="utf-8") as config_file:
config = TrainingConfig.load(config_file)
properties: typing.Dict[str, typing.Any] = {
"length_scale": config.inference.length_scale,
"noise_scale": config.inference.noise_scale,
"noise_w": config.inference.noise_w,
}
# Load speaker names
voice_name = voice_dir.name
speakers: typing.Optional[typing.Sequence[str]] = None
speakers_path = voice_dir / "speakers.txt"
if speakers_path.is_file():
speakers = []
with open(
speakers_path, "r", encoding="utf-8"
) as speakers_file:
for line in speakers_file:
line = line.strip()
if line:
speakers.append(line)
# Load aliases
aliases: typing.Optional[typing.Set[str]] = None
aliases_path = voice_dir / "ALIASES"
if aliases_path.is_file():
aliases = set()
with open(aliases_path, "r", encoding="utf-8") as aliases_file:
for line in aliases_file:
line = line.strip()
if line:
aliases.add(line)
voice_key = f"{voice_lang}/{voice_name}"
yield Voice(
key=voice_key,
name=voice_name,
language=voice_lang,
description="",
speakers=speakers,
location=str(voice_dir.absolute()),
properties=properties,
aliases=aliases,
)
known_voices.discard(voice_key)
# Yield voices that haven't yet been downloaded
for voice_key in known_voices:
voice_lang, voice_name = voice_key.split("/", maxsplit=1)
voice_info = _VOICES.get(voice_key, {})
speakers = voice_info.get("speakers", [])
properties = voice_info.get("properties", {})
yield Voice(
key=voice_key,
name=voice_name,
language=voice_lang,
description="",
speakers=speakers,
location=str.format(
self.settings.voices_url_format or DEFAULT_VOICES_URL_FORMAT,
lang=voice_lang,
name=voice_name,
key=voice_key,
),
properties=properties,
)
def preload_voice(self, voice_key: str):
"""Ensure voice is loaded in memory before synthesis"""
self._get_or_load_voice(voice_key)
# -------------------------------------------------------------------------
@property
def voice(self) -> str:
return self.settings.voice or DEFAULT_VOICE
@voice.setter
def voice(self, new_voice: str):
if new_voice != self.settings.voice:
# Clear speaker on voice change
self.speaker = None
self.settings.voice = new_voice or DEFAULT_VOICE
if "#" in self.settings.voice:
# Split
voice, speaker = self.settings.voice.split("#", maxsplit=1)
self.settings.voice = voice
self.speaker = speaker
@property
def speaker(self) -> typing.Optional[SPEAKER_TYPE]:
return self.settings.speaker
@speaker.setter
def speaker(self, new_speaker: typing.Optional[SPEAKER_TYPE]):
self.settings.speaker = new_speaker
@property
def language(self) -> str:
return self.settings.language or DEFAULT_LANGUAGE
@language.setter
def language(self, new_language: str):
self.settings.language = new_language
@property
def volume(self) -> float:
return self.settings.volume
@volume.setter
def volume(self, new_volume: float):
self.settings.volume = max(0, min(100, new_volume))
@property
def rate(self) -> float:
return self.settings.rate
@rate.setter
def rate(self, new_rate: float):
self.settings.rate = new_rate
def begin_utterance(self):
pass
# pylint: disable=arguments-differ
def speak_text(self, text: str, text_language: typing.Optional[str] = None):
voice = self._get_or_load_voice(self.voice)
minor_break_ms = voice.config.inference.minor_break_ms
major_break_ms = voice.config.inference.major_break_ms
for sent_phonemes, break_type in voice.text_to_phonemes(
text, text_language=text_language
):
add_major_silence = (break_type == BreakType.MAJOR) and (
major_break_ms is not None
)
add_minor_silence = (break_type == BreakType.MINOR) and (
minor_break_ms is not None
)
# Utterances have start/end meta phonemes (usually ^ and $)
is_utterance = (
(break_type == BreakType.UTTERANCE)
or add_major_silence
or add_minor_silence
)
self._results.append(
Mimic3Phonemes(
current_settings=deepcopy(self.settings),
phonemes=sent_phonemes,
is_utterance=is_utterance,
)
)
# Add silence if using manual break intervals
if add_major_silence:
assert major_break_ms is not None
self.add_break(major_break_ms)
elif add_minor_silence:
assert minor_break_ms is not None
self.add_break(minor_break_ms)
# pylint: disable=arguments-differ
def speak_tokens(
self,
tokens: typing.Iterable[BaseToken],
text_language: typing.Optional[str] = None,
):
voice = self._get_or_load_voice(self.voice)
token_phonemes: PHONEMES_LIST_TYPE = []
for token in tokens:
if isinstance(token, Word):
word_phonemes = voice.word_to_phonemes(
token.text, word_role=token.role, text_language=text_language
)
token_phonemes.append(word_phonemes)
elif isinstance(token, Phonemes):
phoneme_str = token.text.strip()
if " " in phoneme_str:
token_phonemes.append(phoneme_str.split())
else:
token_phonemes.append(list(IPA.graphemes(phoneme_str)))
elif isinstance(token, SayAs):
say_as_phonemes = voice.say_as_to_phonemes(
token.text,
interpret_as=token.interpret_as,
say_format=token.format,
text_language=text_language,
)
token_phonemes.extend(say_as_phonemes)
if token_phonemes:
self._results.append(
Mimic3Phonemes(
current_settings=deepcopy(self.settings),
phonemes=token_phonemes,
is_utterance=False,
)
)
def add_break(self, time_ms: int):
# Generate silence (16-bit mono at sample rate)
num_samples = int((time_ms / 1000.0) * self.settings.sample_rate)
audio_bytes = bytes(num_samples * 2)
self._results.append(
AudioResult(
sample_rate_hz=self.settings.sample_rate,
audio_bytes=audio_bytes,
# 16-bit mono
sample_width_bytes=2,
num_channels=1,
)
)
def set_mark(self, name: str):
self._results.append(MarkResult(name=name))
def end_utterance(self) -> typing.Iterable[BaseResult]:
last_settings: typing.Optional[Mimic3Settings] = None
sent_phonemes: PHONEMES_LIST_TYPE = []
for result in self._results:
if isinstance(result, Mimic3Phonemes):
if result.is_utterance or (result.current_settings != last_settings):
if sent_phonemes:
yield self._speak_sentence_phonemes(
sent_phonemes, settings=last_settings
)
sent_phonemes.clear()
sent_phonemes.extend(result.phonemes)
last_settings = result.current_settings
else:
if sent_phonemes:
yield self._speak_sentence_phonemes(
sent_phonemes, settings=last_settings
)
sent_phonemes.clear()
yield result
if sent_phonemes:
yield self._speak_sentence_phonemes(sent_phonemes, settings=last_settings)
sent_phonemes.clear()
self._results.clear()
# -------------------------------------------------------------------------
def _speak_sentence_phonemes(
self,
sent_phonemes,
settings: typing.Optional[Mimic3Settings] = None,
) -> AudioResult:
"""Synthesize audio from phonemes using given setings"""
settings = settings or self.settings
voice = self._get_or_load_voice(settings.voice or self.voice)
sent_phoneme_ids = voice.phonemes_to_ids(sent_phonemes)
_LOGGER.debug("phonemes=%s, ids=%s", sent_phonemes, sent_phoneme_ids)
audio = voice.ids_to_audio(
sent_phoneme_ids,
speaker=settings.speaker,
length_scale=settings.length_scale,
noise_scale=settings.noise_scale,
noise_w=settings.noise_w,
rate=settings.rate,
)
audio_bytes = audio.tobytes()
if settings.volume != DEFAULT_VOLUME:
audio_bytes = audioop.mul(audio_bytes, 2, settings.volume / 100.0)
return AudioResult(
sample_rate_hz=voice.config.audio.sample_rate,
audio_bytes=audio_bytes,
# 16-bit mono
sample_width_bytes=2,
num_channels=1,
)
def _get_or_load_voice(self, voice_key: str) -> Mimic3Voice:
"""Get a loaded voice or load from the file system"""
existing_voice = self._loaded_voices.get(voice_key)
if existing_voice is not None:
return existing_voice
# Look up as substring of known voice
model_dir: typing.Optional[Path] = None
for maybe_voice in self.get_voices():
if (voice_key == maybe_voice.key) or (
maybe_voice.aliases and (voice_key in maybe_voice.aliases)
):
maybe_model_dir = Path(maybe_voice.location)
if (not maybe_model_dir.is_dir()) and (not self.settings.no_download):
# Download voice
maybe_model_dir = self._download_voice(voice_key)
if maybe_model_dir.is_dir():
# Voice found
model_dir = maybe_model_dir
break
if model_dir is None:
raise VoiceNotFoundError(voice_key)
voice_lang = model_dir.parent.name
voice_name = model_dir.name
canonical_key = f"{voice_lang}/{voice_name}"
existing_voice = self._loaded_voices.get(canonical_key)
if existing_voice is not None:
# Alias
self._loaded_voices[voice_key] = existing_voice
return existing_voice
# https://onnxruntime.ai/docs/execution-providers/
providers = None
if self.settings.use_cuda:
providers = ["CUDAExecutionProvider"]
voice = Mimic3Voice.load_from_directory(
model_dir,
providers=providers,
share_models=self.settings.share_onnx_models_between_threads,
use_deterministic_compute=self.settings.use_deterministic_compute,
)
_LOGGER.info("Loaded voice from %s", model_dir)
# Add to cache
self._loaded_voices[voice_key] = voice
self._loaded_voices[canonical_key] = voice
return voice
def _download_voice(self, voice_key: str) -> Path:
"""Downloads a voice by key"""
voice_lang, voice_name = voice_key.split("/", maxsplit=1)
voice_info = _VOICES[voice_key]
voice_url = str.format(
self.settings.voices_url_format or DEFAULT_VOICES_URL_FORMAT,
key=voice_key,
lang=voice_lang,
name=voice_name,
)
voice_files = voice_info["files"]
download_voice(
voice_key=voice_key,
url_base=voice_url,
voice_files=[VoiceFile(file_key) for file_key in voice_files.keys()],
voice_version=voice_info["version"],
voices_dir=self.settings.voices_download_dir,
)
voice_dir = Path(self.settings.voices_download_dir) / voice_key
return voice_dir

63
mimic3_tts/utils.py Normal file
View file

@ -0,0 +1,63 @@
# Copyright 2022 Mycroft AI Inc.
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
#
"""Utility methods for Mimic 3"""
import hashlib
import re
import typing
import numpy as np
def audio_float_to_int16(
audio: np.ndarray, max_wav_value: float = 32767.0
) -> np.ndarray:
"""Normalize audio and convert to int16 range"""
audio_norm = audio * (max_wav_value / max(0.01, np.max(np.abs(audio))))
audio_norm = np.clip(audio_norm, -max_wav_value, max_wav_value)
audio_norm = audio_norm.astype("int16")
return audio_norm
def wildcard_to_regex(template: str, wildcard: str = "*") -> re.Pattern:
"""Convert a string with wildcards into a regex pattern"""
wildcard_escaped = re.escape(wildcard)
pattern_parts = ["^"]
for i, template_part in enumerate(re.split(f"({wildcard_escaped})", template)):
if (i % 2) == 0:
# Fixed string
pattern_parts.append(re.escape(template_part))
else:
# Wildcard separator
pattern_parts.append(".*")
pattern_parts.append("$")
pattern_str = "".join(pattern_parts)
return re.compile(pattern_str)
def file_sha256_sum(fp: typing.BinaryIO, block_bytes: int = 4096) -> str:
"""Return the sha256 sum of a (possibly large) file"""
current_hash = hashlib.sha256()
# Read in blocks in case file is very large
block = fp.read(block_bytes)
while len(block) > 0:
current_hash.update(block)
block = fp.read(block_bytes)
return current_hash.hexdigest()

762
mimic3_tts/voice.py Normal file
View file

@ -0,0 +1,762 @@
# Copyright 2022 Mycroft AI Inc.
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU Affero General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU Affero General Public License for more details.
#
# You should have received a copy of the GNU Affero General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
#
import csv
import logging
import platform
import threading
import time
import typing
from abc import ABCMeta, abstractmethod
from enum import Enum
from pathlib import Path
from xml.sax.saxutils import escape as xmlescape
import epitran
import espeak_phonemizer
import gruut
import numpy as np
import onnxruntime
import phonemes2ids
from gruut_ipa import IPA
from .config import Phonemizer, TrainingConfig
from .const import DEFAULT_RATE
from .utils import audio_float_to_int16
# -----------------------------------------------------------------------------
class BreakType(str, Enum):
NONE = "none"
MINOR = "minor"
MAJOR = "major"
UTTERANCE = "utterance"
PHONEME_TYPE = str
PHONEME_ID_TYPE = int
WORD_PHONEMES_TYPE = typing.List[typing.List[PHONEME_TYPE]]
PHONEME_MAP_TYPE = typing.Dict[PHONEME_TYPE, typing.List[PHONEME_TYPE]]
TEXT_TO_PHONEMES_TYPE = typing.Iterable[typing.Tuple[WORD_PHONEMES_TYPE, BreakType]]
SPEAKER_NAME_TYPE = str
SPEAKER_ID_TYPE = int
SPEAKER_TYPE = typing.Union[SPEAKER_NAME_TYPE, SPEAKER_ID_TYPE]
SPEAKER_MAP_TYPE = typing.Dict[SPEAKER_NAME_TYPE, SPEAKER_ID_TYPE]
DEFAULT_LANGUAGE = "en_US"
_LOGGER = logging.getLogger(__name__)
# -----------------------------------------------------------------------------
class Mimic3Voice(metaclass=ABCMeta):
"""Base class for Mimic 3 voice implementations"""
_SHARED_MODELS: typing.Dict[str, onnxruntime.InferenceSession] = {}
_SHARED_MODELS_LOCK = threading.Lock()
def __init__(
self,
config: TrainingConfig,
onnx_model: onnxruntime.InferenceSession,
phoneme_to_id: typing.Dict[PHONEME_TYPE, int],
phoneme_map: typing.Optional[PHONEME_MAP_TYPE] = None,
speaker_map: typing.Optional[SPEAKER_MAP_TYPE] = None,
):
self.config = config
self.onnx_model = onnx_model
self.phoneme_to_id = phoneme_to_id
self.phoneme_map = phoneme_map
self.speaker_map = speaker_map
@abstractmethod
def text_to_phonemes(
self, text: str, text_language: typing.Optional[str] = None
) -> TEXT_TO_PHONEMES_TYPE:
"""Convert text into phonemes"""
def word_to_phonemes(
self,
word_text: str,
word_role: typing.Optional[str] = None,
text_language: typing.Optional[str] = None,
) -> typing.List[PHONEME_TYPE]:
"""Convert a single word (with optional role) into phonemes"""
word_phonemes = []
for sent_phonemes, _break_type in self.text_to_phonemes(
word_text, text_language=text_language
):
for sent_word_phonemes in sent_phonemes:
word_phonemes.extend(sent_word_phonemes)
return word_phonemes
def say_as_to_phonemes(
self,
text: str,
interpret_as: str,
say_format: typing.Optional[str] = None,
text_language: typing.Optional[str] = None,
) -> WORD_PHONEMES_TYPE:
"""Speak a word or phrase with a specific interpretation/format"""
word_phonemes = []
for sent_phonemes, _break_type in self.text_to_phonemes(
text, text_language=text_language
):
word_phonemes.extend(sent_phonemes)
return word_phonemes
def phonemes_to_ids(
self, phonemes: WORD_PHONEMES_TYPE
) -> typing.Sequence[PHONEME_ID_TYPE]:
"""Convert phonemes to ids for a voice model (see phonemes.txt)"""
phoneme_map = self.phoneme_map or self.config.phonemes.phoneme_map
return phonemes2ids.phonemes2ids(
word_phonemes=phonemes,
phoneme_to_id=self.phoneme_to_id,
pad=self.config.phonemes.pad,
bos=self.config.phonemes.bos,
eos=self.config.phonemes.eos,
auto_bos_eos=self.config.phonemes.auto_bos_eos,
blank=self.config.phonemes.blank,
blank_word=self.config.phonemes.blank_word,
blank_between=self.config.phonemes.blank_between,
blank_at_start=self.config.phonemes.blank_at_start,
blank_at_end=self.config.phonemes.blank_at_end,
simple_punctuation=self.config.phonemes.simple_punctuation,
punctuation_map=self.config.phonemes.punctuation_map,
separate=self.config.phonemes.separate,
separate_graphemes=self.config.phonemes.separate_graphemes,
separate_tones=self.config.phonemes.separate_tones,
tone_before=self.config.phonemes.tone_before,
phoneme_map=phoneme_map,
fail_on_missing=False,
)
def ids_to_audio(
self,
phoneme_ids: typing.Sequence[PHONEME_ID_TYPE],
speaker: typing.Optional[
typing.Union[SPEAKER_NAME_TYPE, SPEAKER_ID_TYPE]
] = None,
length_scale: typing.Optional[float] = None,
noise_scale: typing.Optional[float] = None,
noise_w: typing.Optional[float] = None,
rate: float = DEFAULT_RATE,
) -> np.ndarray:
"""Synthesize audio from phoneme ids usng Onnx voice model (see generator.onnx)"""
if length_scale is None:
length_scale = self.config.inference.length_scale
# Scale length by rate
if rate > 0:
length_scale /= rate
if noise_scale is None:
noise_scale = self.config.inference.noise_scale
if noise_w is None:
noise_w = self.config.inference.noise_w
# Create model inputs
text_array = np.expand_dims(np.array(phoneme_ids, dtype=np.int64), 0)
text_lengths_array = np.array([text_array.shape[1]], dtype=np.int64)
scales_array = np.array(
[
noise_scale,
length_scale,
noise_w,
],
dtype=np.float32,
)
inputs = {
"input": text_array,
"input_lengths": text_lengths_array,
"scales": scales_array,
}
speaker_id = 0
if self.config.is_multispeaker:
if isinstance(speaker, SPEAKER_NAME_TYPE):
if self.speaker_map:
maybe_speaker_id = self.speaker_map.get(speaker)
if maybe_speaker_id is None:
try:
# Interpret as speaker id
speaker_id = int(speaker)
except ValueError:
_LOGGER.warning(
"Unable to find a speaker with the name '%s'. Falling back to first speaker.",
speaker,
)
pass
else:
speaker_id = maybe_speaker_id
elif speaker is not None:
speaker_id = speaker
speaker_id_array = np.array([speaker_id], dtype=np.int64)
inputs["sid"] = speaker_id_array
_LOGGER.debug(
"TTS settings: speaker-id=%s, length-scale=%s, noise-scale=%s, noise-w=%s",
speaker_id,
length_scale,
noise_scale,
noise_w,
)
# Infer audio from phonemes
start_time = time.perf_counter()
audio = self.onnx_model.run(None, inputs)[0].squeeze()
audio = audio_float_to_int16(audio)
end_time = time.perf_counter()
# Compute real-time factor
audio_duration_sec = audio.shape[-1] / self.config.audio.sample_rate
infer_sec = end_time - start_time
real_time_factor = (
infer_sec / audio_duration_sec if audio_duration_sec > 0 else 0.0
)
_LOGGER.debug("RTF: %s", real_time_factor)
return audio
@staticmethod
def load_from_directory(
voice_dir: typing.Union[str, Path],
session_options: typing.Optional[onnxruntime.SessionOptions] = None,
providers: typing.Optional[
typing.Sequence[
typing.Union[str, typing.Tuple[str, typing.Dict[str, typing.Any]]]
]
] = None,
share_models: bool = True,
use_deterministic_compute: bool = False,
) -> "Mimic3Voice":
"""Load a Mimic 3 voice from a directory"""
voice_dir = Path(voice_dir)
_LOGGER.debug("Loading voice from %s", voice_dir)
config_path = voice_dir / "config.json"
_LOGGER.debug("Loading config from %s", config_path)
with open(config_path, "r", encoding="utf-8") as config_file:
config = TrainingConfig.load(config_file)
# phoneme -> id
phoneme_ids_path = voice_dir / "phonemes.txt"
_LOGGER.debug("Loading model phonemes from %s", phoneme_ids_path)
with open(phoneme_ids_path, "r", encoding="utf-8") as ids_file:
phoneme_to_id = phonemes2ids.load_phoneme_ids(ids_file)
generator_path = voice_dir / "generator.onnx"
onnx_model: typing.Optional[onnxruntime.InferenceSession] = None
if share_models:
with Mimic3Voice._SHARED_MODELS_LOCK:
model_key = str(generator_path.absolute())
onnx_model = Mimic3Voice._SHARED_MODELS.get(model_key)
if onnx_model is None:
onnx_model = Mimic3Voice._load_model(
generator_path,
session_options=session_options,
providers=providers,
use_deterministic_compute=use_deterministic_compute,
)
Mimic3Voice._SHARED_MODELS[model_key] = onnx_model
else:
_LOGGER.debug("Using shared Onnx model (%s)", model_key)
else:
onnx_model = Mimic3Voice._load_model(
generator_path,
session_options=session_options,
providers=providers,
use_deterministic_compute=use_deterministic_compute,
)
# phoneme -> phoneme, phoneme, ...
phoneme_map: typing.Optional[PHONEME_MAP_TYPE] = None
phoneme_map_path = voice_dir / "phoneme_map.txt"
if phoneme_map_path.is_file():
_LOGGER.debug("Loading phoneme map from %s", phoneme_map_path)
with open(phoneme_map_path, "r", encoding="utf-8") as map_file:
phoneme_map = phonemes2ids.utils.load_phoneme_map(map_file)
# id -> speaker
speaker_map: typing.Optional[SPEAKER_MAP_TYPE] = None
speaker_map_path = voice_dir / "speaker_map.csv"
if speaker_map_path.is_file():
_LOGGER.debug("Loading speaker map from %s", speaker_map_path)
with open(speaker_map_path, "r", encoding="utf-8") as map_file:
# id | dataset | name | [alias] | [alias] ...
reader = csv.reader(map_file, delimiter="|")
speaker_map = {}
for row in reader:
speaker_id = int(row[0])
for alias in row[2:]:
speaker_map[alias] = speaker_id
if config.phonemizer == Phonemizer.GRUUT:
# Phonemes from gruut: https://github.com/rhasspy/gruut/
return GruutVoice(
config=config,
onnx_model=onnx_model,
phoneme_to_id=phoneme_to_id,
phoneme_map=phoneme_map,
speaker_map=speaker_map,
)
if config.phonemizer == Phonemizer.ESPEAK:
# Phonemes from eSpeak-ng: https://github.com/espeak-ng/espeak-ng
voice_class = EspeakVoice
if config.text_language == "fa":
try:
# Check if hazm is available
# https://github.com/sobhe/hazm
import hazm # noqa: F401
voice_class = HazmEspeakVoice
except ImportError:
_LOGGER.warning("hazm is highly recommended for language 'fa'")
_LOGGER.warning("pip install 'hazm>=0.7.0'")
return voice_class(
config=config,
onnx_model=onnx_model,
phoneme_to_id=phoneme_to_id,
phoneme_map=phoneme_map,
speaker_map=speaker_map,
)
if config.phonemizer == Phonemizer.SYMBOLS:
# Phonemes are characters from an alphabet
return SymbolsVoice(
config=config,
onnx_model=onnx_model,
phoneme_to_id=phoneme_to_id,
phoneme_map=phoneme_map,
speaker_map=speaker_map,
)
if config.phonemizer == Phonemizer.EPITRAN:
# Phonemes are from epitran: https://github.com/dmort27/epitran/
return EpitranVoice(
config=config,
onnx_model=onnx_model,
phoneme_to_id=phoneme_to_id,
phoneme_map=phoneme_map,
speaker_map=speaker_map,
)
raise ValueError(f"Unsupported phonemizer: {config.phonemizer}")
@staticmethod
def _load_model(
generator_path: Path,
session_options: typing.Optional[onnxruntime.SessionOptions] = None,
providers: typing.Optional[
typing.Sequence[
typing.Union[str, typing.Tuple[str, typing.Dict[str, typing.Any]]]
]
] = None,
use_deterministic_compute: bool = False,
) -> onnxruntime.InferenceSession:
_LOGGER.debug("Loading model from %s", generator_path)
# Load onnx model
if session_options is None:
session_options = onnxruntime.SessionOptions()
if platform.machine() == "armv7l":
# Enabling optimizations on 32-bit ARM crashes
session_options.graph_optimization_level = (
onnxruntime.GraphOptimizationLevel.ORT_DISABLE_ALL
)
session_options.use_deterministic_compute = use_deterministic_compute
onnx_model = onnxruntime.InferenceSession(
str(generator_path), sess_options=session_options, providers=providers
)
return onnx_model
# -----------------------------------------------------------------------------
class GruutVoice(Mimic3Voice):
"""Voice whose phonemes come from gruut (https://github.com/rhasspy/gruut/)"""
def text_to_phonemes(
self, text: str, text_language: typing.Optional[str] = None
) -> TEXT_TO_PHONEMES_TYPE:
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
for sentence in gruut.sentences(text, lang=text_language):
sent_phonemes = [w.phonemes for w in sentence if w.phonemes]
if sent_phonemes:
yield sent_phonemes, BreakType.UTTERANCE
def word_to_phonemes(
self,
word_text: str,
word_role: typing.Optional[str] = None,
text_language: typing.Optional[str] = None,
) -> typing.List[PHONEME_TYPE]:
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
word_role = xmlescape(word_role) if word_role else ""
word_text = xmlescape(word_text)
sentence = next(
iter(
gruut.sentences(
f'<w role="{word_role}">{word_text}</w>',
ssml=True,
lang=text_language,
)
)
)
sentence_word = next(iter(sentence))
return sentence_word.phonemes
def say_as_to_phonemes(
self,
text: str,
interpret_as: str,
say_format: typing.Optional[str] = None,
text_language: typing.Optional[str] = None,
) -> WORD_PHONEMES_TYPE:
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
word_text = xmlescape(text)
interpret_as = xmlescape(interpret_as)
format_attr = f'format="{xmlescape(say_format)}"' if say_format else ""
sentences = gruut.sentences(
f'<say-as interpret-as="{interpret_as}" {format_attr}>{word_text}</say-as>',
ssml=True,
lang=text_language,
)
sent_phonemes: WORD_PHONEMES_TYPE = []
for sentence in sentences:
sent_phonemes.extend(w.phonemes for w in sentence if w.phonemes)
return sent_phonemes
# -----------------------------------------------------------------------------
class EspeakVoice(Mimic3Voice):
"""Voice whose phonemes come from eSpeak-NG (https://github.com/espeak-ng/espeak-ng)"""
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self._phonemizer = espeak_phonemizer.Phonemizer()
def text_to_phonemes(
self, text: str, text_language: typing.Optional[str] = None
) -> TEXT_TO_PHONEMES_TYPE:
phoneme_separator = ""
word_separator = self.config.phonemes.word_separator
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
voice = self._language_to_voice(text_language)
phoneme_str = self._phonemizer.phonemize(
text,
voice=voice,
keep_clause_breakers=True,
phoneme_separator=phoneme_separator,
word_separator=word_separator,
punctuation_separator=phoneme_separator,
)
all_word_phonemes = [
list(IPA.graphemes(wp_str)) for wp_str in phoneme_str.split(word_separator)
]
minor_break = self.config.phonemes.minor_break
major_break = self.config.phonemes.major_break
if minor_break or major_break:
# Split on breaks
sent_phonemes = []
for word_phonemes in all_word_phonemes:
sent_phonemes.append(word_phonemes)
if minor_break and (word_phonemes[-1] == minor_break):
yield sent_phonemes, BreakType.MINOR
sent_phonemes = []
elif major_break and (word_phonemes[-1] == major_break):
yield sent_phonemes, BreakType.MAJOR
sent_phonemes = []
if sent_phonemes:
yield sent_phonemes, BreakType.MAJOR
else:
# No split
yield all_word_phonemes, BreakType.UTTERANCE
def word_to_phonemes(
self,
word_text: str,
word_role: typing.Optional[str] = None,
text_language: typing.Optional[str] = None,
) -> typing.List[PHONEME_TYPE]:
phoneme_separator = ""
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
word_role = xmlescape(word_role) if word_role else ""
word_text = xmlescape(word_text)
voice = self._language_to_voice(text_language)
phoneme_str = self._phonemizer.phonemize(
f'<w role="{word_role}">{word_text}</w>',
voice=voice,
keep_clause_breakers=True,
phoneme_separator=phoneme_separator,
punctuation_separator=phoneme_separator,
ssml=True,
)
word_phonemes = list(IPA.graphemes(phoneme_str))
return word_phonemes
def say_as_to_phonemes(
self,
text: str,
interpret_as: str,
say_format: typing.Optional[str] = None,
text_language: typing.Optional[str] = None,
) -> WORD_PHONEMES_TYPE:
phoneme_separator = ""
word_separator = self.config.phonemes.word_separator
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
word_text = xmlescape(text)
interpret_as = xmlescape(interpret_as)
format_attr = f'format="{xmlescape(say_format)}"' if say_format else ""
voice = self._language_to_voice(text_language)
phoneme_str = self._phonemizer.phonemize(
f'<say-as interpret-as="{interpret_as}" {format_attr}>{word_text}</say-as>',
voice=voice,
keep_clause_breakers=True,
phoneme_separator=phoneme_separator,
punctuation_separator=phoneme_separator,
word_separator=word_separator,
ssml=True,
)
word_phonemes = [
list(IPA.graphemes(wp_str)) for wp_str in phoneme_str.split(word_separator)
]
return word_phonemes
def _language_to_voice(self, language: str) -> str:
"""Make voice name from language name"""
# en_US -> en-us
return language.strip().lower().replace("_", "-")
class HazmEspeakVoice(EspeakVoice):
"""Persian espeak-ng voice that uses hazm (https://github.com/sobhe/hazm) for pre-processing"""
def __init__(self, *args, **kwargs):
import gruut_lang_fa
import hazm
super().__init__(*args, **kwargs)
self._normalizer = hazm.Normalizer()
self._sent_tokenizer = hazm.SentenceTokenizer()
self._word_tokenizer = hazm.WordTokenizer()
# Load part of speech tagger from gruut[fa]
self._tagger = hazm.POSTagger(
model=str(gruut_lang_fa.get_lang_dir() / "pos" / "postagger.model")
)
def text_to_phonemes(
self, text: str, text_language: typing.Optional[str] = None
) -> TEXT_TO_PHONEMES_TYPE:
phoneme_separator = ""
word_separator = self.config.phonemes.word_separator
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
voice = self._language_to_voice(text_language)
# Normalize with hazm
sentences = self._preprocess_text(text)
for sentence in sentences:
sent_text = " ".join(sentence)
sent_phoneme_str = self._phonemizer.phonemize(
sent_text,
voice=voice,
keep_clause_breakers=True,
phoneme_separator=phoneme_separator,
word_separator=word_separator,
punctuation_separator=phoneme_separator,
)
sent_word_phonemes = [
list(IPA.graphemes(wp_str))
for wp_str in sent_phoneme_str.split(word_separator)
]
yield sent_word_phonemes, BreakType.UTTERANCE
def word_to_phonemes(
self,
word_text: str,
word_role: typing.Optional[str] = None,
text_language: typing.Optional[str] = None,
) -> typing.List[PHONEME_TYPE]:
word_text = self._fix_words([word_text])[0]
return super().word_to_phonemes(
word_text, word_role=word_role, text_language=text_language
)
def say_as_to_phonemes(
self,
text: str,
interpret_as: str,
say_format: typing.Optional[str] = None,
text_language: typing.Optional[str] = None,
) -> WORD_PHONEMES_TYPE:
sentences = self._preprocess_text(text)
text = " ".join(
" ".join(word_text for word_text in words) for words in sentences
)
return super().say_as_to_phonemes(
text, interpret_as, say_format=say_format, text_language=text_language
)
def _preprocess_text(self, text: str) -> typing.List[typing.List[str]]:
"""Split/normalize text into sentences/words with hazm"""
text = self._normalizer.normalize(text)
processed_sentences = []
for sentence in self._sent_tokenizer.tokenize(text):
words = self._word_tokenizer.tokenize(sentence)
processed_words = self._fix_words(words)
processed_sentences.append(processed_words)
return processed_sentences
def _fix_words(self, words: typing.List[str]) -> typing.List[str]:
fixed_words = []
for word, pos in self._tagger.tag(words):
if pos[-1] == "e":
if word[-1] != "ِ":
if (word[-1] == "ه") and (word[-2] != "ا"):
word += "‌ی"
word += "ِ"
fixed_words.append(word)
return fixed_words
# -----------------------------------------------------------------------------
class SymbolsVoice(Mimic3Voice):
"""Voice whose phonemes are characters in an alphabet"""
def text_to_phonemes(
self, text: str, text_language: typing.Optional[str] = None
) -> TEXT_TO_PHONEMES_TYPE:
word_separator = self.config.phonemes.word_separator
word_phonemes = [
list(IPA.graphemes(wp_str)) for wp_str in text.split(word_separator)
]
yield word_phonemes, BreakType.UTTERANCE
# -----------------------------------------------------------------------------
class EpitranVoice(Mimic3Voice):
"""Voice whose phonemes come from epitran (https://github.com/dmort27/epitran/)"""
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self._epis: typing.Dict[str, epitran.Epitran] = {}
def text_to_phonemes(
self, text: str, text_language: typing.Optional[str] = None
) -> TEXT_TO_PHONEMES_TYPE:
text_language = text_language or self.config.text_language or DEFAULT_LANGUAGE
epi = self._epis.get(text_language)
if epi is None:
epi = epitran.Epitran(text_language)
self._epis[text_language] = epi
phoneme_str = epi.transliterate(text)
all_word_phonemes = [
list(IPA.graphemes(wp_str)) for wp_str in phoneme_str.split()
]
minor_break = self.config.phonemes.minor_break
major_break = self.config.phonemes.major_break
if minor_break or major_break:
# Split on breaks
sent_phonemes = []
for word_phonemes in all_word_phonemes:
sent_phonemes.append(word_phonemes)
if minor_break and (word_phonemes[-1] == minor_break):
yield sent_phonemes, BreakType.MINOR
sent_phonemes = []
elif major_break and (word_phonemes[-1] == major_break):
yield sent_phonemes, BreakType.MAJOR
sent_phonemes = []
if sent_phonemes:
yield sent_phonemes, BreakType.MAJOR
else:
# No split
yield all_word_phonemes, BreakType.UTTERANCE

1972
mimic3_tts/voices.json Normal file

File diff suppressed because it is too large Load diff