A flexible, header-only text-to-speech (TTS) library designed specifically for microcontrollers and embedded systems. TinyTTSTools provides high-quality speech synthesis with minimal memory footprint and optional machine learning capabilities.
Unfortunately, microcontrollers do not have enough RAM and computational resources to implement high-quality TTS using machine learning directly. Therefore, we generate audio in the following three steps:
- Grapheme to Phoneme (G2P) - Converting text to phonetic representations
- Phoneme to Audio - Converting phonemes to PCM data
- Output of Audio - Output of PCM data to I2S, Analog Pins, PWM, PDM etc
For each of these steps, we provide different implementations with varying quality and resource requirements.
All implementations are based on the common G2PModelBase class and provide different approaches to convert text to phonemes:
G2PDictionaryModel- The simplest method using a dictionary lookup to translate words to phonemesG2PRuleBasedModelEN/DE/FR/ES- Rule-based letter-to-sound fallback, one per supported language, sharing the longest-match scaffolding inG2PRuleBasedModelBaseG2PNeuralModel- Dependency-free (no TensorFlow Lite) GRU neural network fallback for out-of-dictionary words (English weights shipped, ported from the sibling TinyTTS project; French -- 94.8% validation exact-match --, Spanish -- 98.7% -- and German -- 71.2% -- are all trained too, but not yet wired intoG2PNeuralModel.h's output table -- see Adding a Language)
Models can be combined using the G2PHybridModel class for improved accuracy:
G2PDictionaryAndRulesModel- Predefined combination of dictionary and rule-based models (English only; build the same combination for another language withG2PHybridModel+G2PDictionaryModel+ that language's rule model)G2PDictionaryNeuralAndRulesModel- Predefined combination of dictionary, neural fallback and rule-based models (seeexamples/G2PNeural)
Beyond English, small built-in dictionaries and rule-based fallbacks ship
for German, French, and Spanish (PhonemeDictionaryDE/FR/ES.h,
G2PRuleBasedModelDE/FR/ES.h), plus larger compressed dictionaries
(60k-93k words each, built from the OLaPh
pronunciation corpus, filtered to each language's top-100k most frequent
words per FrequencyWords so
they fit ESP32 PSRAM -- under 1MB-1.6MB each) for a desktop/SD-card/PSRAM
tier. The desktop CLI
(desktop/tinyttstools) demonstrates all four with --language en|de|fr|es.
See Adding a Language for the full picture --
what's already there for German/French/Spanish, and the steps to add a new
one, including the PhonemeModifier system (ˈstress, ːlength, nasalization,
palatalization, ...) that carries diacritics alongside the base Phone id.
In the second step, we generate audio data from phonemes. All implementations inherit from the VocoderBase class.
FormantVocoder- Generates PCM audio using formant synthesis rules- Pros: Minimal memory usage, no external audio data required
- Cons: Lower audio quality, more robotic sound
The following classes use pre-recorded (compressed) audio in PROGMEM:
PhonemeVocoder- Uses individual phoneme samples- Pros: Smaller dictionary, predictable output
- Cons: Less natural transitions between phonemes
DiphoneVocoder- Uses diphone samples (combinations of two phonemes)- Pros: More natural speech with smooth phoneme transitions
- Cons: Larger dictionary required
PSOLAVocoder- Re-synthesizes the same phoneme samples asPhonemeVocodervia TD-PSOLA (Time-Domain Pitch-Synchronous Overlap-Add), the technique the Praat phonetics software is best known for- Pros: Genuinely shifts pitch and stretches/compresses duration (
PhonemeSynthesisParams::pitchHz/speed), instead of only ever truncating pre-recorded audio - Cons: Higher CPU cost than plain playback; same dictionary size as
PhonemeVocoder
- Pros: Genuinely shifts pitch and stretches/compresses duration (
PhonemeSynthesisParams (duration, pitch contour, voicing -- see
VocoderBase.h) can be filled in per-phoneme by a
ProsodyModelBase
implementation instead of staying at its flat/default values. This is
independent of, and works with, every vocoder above -- it only ever
supplies prosody (never formants/audio data), so it changes how natural
the timing/intonation sounds, not which vocoder is doing the synthesis.
The shipped ProsodyNeuralModel predicts pitch contour for every
phoneme, plus duration and voicing for vowels only -- duration as one of
three discrete lengths (short/medium/long, where long is exactly the
vocoder's own default, unscaled, rather than the model's raw continuous
prediction); consonants keep the vocoder's own default duration/voicing
untouched (applying either to consonants distorted their spectral
identity badly enough in real listening tests to be worth avoiding
entirely; see ProsodyNeuralModel.h for the specifics), all learned from
real forced-aligned speech.
ProsodyNeuralModel- Dependency-free (no TensorFlow Lite) bidirectional-GRU model, trained offline on forced-aligned real speech (English weights shipped; seesetup/prosody-en/), predicting pitch contour for every phoneme, plus (vowels only) duration and voicing- Opt-in via
TinyTTSTools::setProsodyModel()-- omit it (the default) for the vocoder's own unchanged flat prosody; seeexamples/ProsodyNeural
Any subclass of Print can be used to output the audio. We recommend that you use the output classes provided by the Arduino Audio Tools:
- I2SStream for i2s
- AnalogAudioStream using the internal DAC
- PWMAudioStream using PWM
Alternatively you can also use the TTSAudioOutputCallback class to output the audio with the help of a callback or the audio output classes provided by your microcontroller.
- Tutorial - Full walkthrough: choosing a vocoder and G2P model, tuning synthesis, audio output
- Building on Desktop - CMake build instructions, including how to build and run the test suite
- Setup Tools - Regenerating the audio/dictionary data files (
setup/), including the neural G2P and prosody training pipelines - Adding a Language - Step-by-step: what German/French/Spanish already have, and how to add another language (phonemes, dictionary, rule-based G2P, audio, neural)
- Loadable Data - The same audio data as real
.wavfiles, forAudioDictionarySD/AudioEncodedDictionarySD(SD card/LittleFS) instead of PROGMEM - Phonemes - The phoneme set (ARPAbet + international/IPA extension) and
PhonemeModifierdiacritic system used throughout the library - Memory Usage - Flash/RAM cost of each vocoder, G2P model and dictionary format
- Class Documentation
- Header-only library - Easy integration, no separate compilation
- Multiple synthesis methods - Choose based on memory/quality requirements
- Compressed audio storage - Codecs reduce memory footprint
- Extensible architecture - Easy to add custom models and synthesizers
- Machine learning support - Dependency-free neural G2P fallback (
G2PNeuralModel) and prosody predictor (ProsodyNeuralModel) - Cross-platform - Works on Arduino, ESP32, and other microcontrollers
See the examples/ directory for complete usage examples:
AudioFormant/- Text-to-speech usingFormantVocoder(no audio data required)AudioPhoneme/- Text-to-speech usingPhonemeVocoder(pre-recorded phoneme samples)AudioDiphones/- Text-to-speech usingDiphoneVocoder(pre-recorded diphone samples)AudioPSOLA/- Text-to-speech usingPSOLAVocoder(TD-PSOLA re-synthesis with pitch/speed control)G2PCustomDictionary/- Adding custom pronunciations, phoneme conversion only (no audio)G2PNeural/- Dictionary + neural + rule-based G2P fallback chain, phoneme conversion only (no audio)ProsodyNeural/- Predicted vs. flat/default prosody throughFormantVocoderandPSOLAVocoder
- Download ZIP from GitHub or clone:
git clone https://github.com/pschatzmann/TinyTTSTools.git - Arduino IDE:
Sketch→Include Library→Add .ZIP Library... - Optionally install dependency: "Arduino Audio Tools"
lib_deps =
https://github.com/pschatzmann/TinyTTSTools.git
https://github.com/pschatzmann/arduino-audio-tools.git