FIELD GUIDE · PIPER TTS

From recordings
to a voice.

The essential steps for preparing voice data and training a Piper model. Kaggle limits can change, so check your quota before a long run.

01ARCHIVE STRUCTURE

Create an LJSpeech dataset for Piper TTS

LJSpeech is the standard format expected by Piper. Each sentence is a separate WAV file, paired with its transcript in metadata.csv.

data.zip
├── metadata.csv   # filename|transcript
└── wavs/
    ├── 000001.wav
    └── ...
  1. 01

    Collect the audio files

    Use recordings you are allowed to process and redistribute. Sources might include game wikis, media you own, or public datasets.

  2. 02

    Split the audio into sentences

    Create one WAV file per sentence. For automatic splitting or alignment, these tools can help:

    • audiosplitter · automatic splitting
    • Montreal Forced Aligner (MFA) · forced alignment
    • faster-whisper + VAD · transcription and timestamps
    • Audacity · manual edits for smaller collections
  3. 03

    Create metadata.csv

    Use one line per file. Separate the filename and transcript with a vertical bar; the filename must match the file in wavs/.

    000001.wav|This is an example sentence.
    000002.wav|Here is another training sentence.
  4. 04

    Zip the archive

    Place metadata.csv and the wavs/ directory at the archive root.

    zip -r data.zip metadata.csv wavs/
  5. 05

    Upload to Hugging Face

    Create a dataset repository, for example user/wheatley-english-ljspeech. Add data.zip and a README.md with this configuration:

    ---
    configs:
    - config_name: default
      data_files:
      - split: train
        path: data.zip
    ---
  6. 06

    Post in the Piper forum

    Start a thread in the datasets and training channel with the Hugging Face link and the appropriate tag.

    Open the Discord channel ↗
02KAGGLE · GPU

Train a voice on Kaggle

Kaggle assigns GPU quota to your account, and the allowance can vary with availability. Check the counter before starting a long session.

Weekly GPU quota30 hours · sometimes more

Kaggle lists a 30-hour weekly baseline, which may be higher depending on demand and available resources. The quota resets on Saturday at 00:00 UTC.

Maximum session length12 hours

Kaggle documentation lists a continuous 12-hour limit for CPU and GPU sessions. CPU-only notebooks do not use GPU quota, but their session length is still limited.

Available acceleratorT4 x2

Kaggle retired the Tesla P100 on September 15, 2026. T4 x2 remains available in notebooks; choose it for Piper training instead of following older P100 guides.

↻

Resume after an interruption

A full training run often exceeds a 12-hour session. This project’s Piper notebook resumes from a Hugging Face checkpoint: after a timeout or an exhausted quota, run it again to continue from the latest saved checkpoint.

Before a long training run

  • Check your remaining quota in notebook settings, your profile, or the accelerator usage page.
  • Stop GPU sessions when you are not using them.
  • Use T4 x2; P100 is no longer available on Kaggle.

Kaggle quotas, accelerators, and interfaces can change. Always check the information shown in your account before training.

03PIPER · LOCAL TTS

Use a Piper voice

Install Piper, get the two files for a voice, and generate a WAV locally. The commands below use the version maintained by the Open Home Foundation.

  1. 01

    Install Piper

    Piper is distributed as a Python package. Older guides using the rhasspy/piper repository or its binary archive refer to the historical release.

    python -m pip install piper-tts
    python -m piper.download_voices en_US-lessac-medium
  2. 02

    Download a voice

    Choose a voice from the catalogue and download its ONNX model and matching JSON configuration file. Keep both files together; the JSON path is the model path followed by .json.

    voice.onnx
    voice.onnx.json
  3. 03

    Generate a WAV file

    First download a voice from the Piper catalogue, or replace the name below with the path to your .onnx file.

    python -m piper -m ./voice.onnx -f output.wav -- "This is a speech synthesis test."
  4. 04

    Read text from a file

    python -m piper -m ./voice.onnx -f output.wav --input-file mon_texte.txt

    Stream raw audio to playback

    The sample rate depends on the voice. Check voice.onnx.json and change 22050 if needed.

    echo "Hello." | python -m piper -m ./voice.onnx --output-raw | aplay -r 22050 -f S16_LE -c 1 -t raw
  5. 05

    From Python

    import wave
    from piper import PiperVoice
    
    voice = PiperVoice.load("./voice.onnx")
    with wave.open("output.wav", "wb") as wav_file:
        voice.synthesize_wav("Hello world.", wav_file)
  6. 06

    With Home Assistant

    Home Assistant connects to Piper through Wyoming. Add Piper under Settings → Devices & services, then use the tts.speak action with the Piper TTS entity.

    Set up Piper in Home Assistant ↗

    With Node-RED

    The node-red-contrib-piper-tts package mentioned in some guides could not be confirmed in the public catalogue. One option is to call the Piper command from Node-RED’s built-in Exec node; adapt paths and permissions to your setup.

  7. 07

    Adjust the output

    Quality depends in part on the training data and selected voice. A higher length-scale slows speech; a lower value speeds it up. noise-scale and noise-w adjust variation in the generated audio.

    python -m piper -m ./voice.onnx --length-scale 1.5 --noise-scale 0.667 --noise-w 0.8 -f slow.wav -- "Slow speech."
    python -m piper -m ./voice.onnx --length-scale 0.7 --noise-scale 0.667 --noise-w 0.8 -f fast.wav -- "Fast speech."
04NODE.JS · TYPESCRIPT

Use Piper in a Node.js project with pipertts

pipertts brings Piper to Node.js, with file output or in-memory audio. The recommended mode is PiperNativeTTS: fully native in-process inference (ONNX plus a compiled espeak bridge), no Python, roughly 10x faster than the wrapper. Voices (.onnx plus .onnx.json) remain separate files.

  1. 01

    Install the dependencies

    The docs specify Node.js 18+ and TypeScript 5+. Native mode needs nothing else: plain npm install is enough. The classic wrapper mode (PiperTTS through python3 -m piper) additionally requires Python and the Piper module. Selecting a catalogue ID can download the model and its config into models/.

    npm install pipertts
    # wrapper mode only: python3 -m pip install piper-tts
  2. 02

    Quick start with a catalogue voice

    import { PiperNativeTTS, resolveModelPathFromOptions } from "pipertts";
    import * as fs from "node:fs";
    
    const modelPath = await resolveModelPathFromOptions({
      model: "en_US-lessac-medium",
      modelsDir: "./models",
    });
    const tts = await PiperNativeTTS.load({ modelPath });
    
    const { audio } = await tts.synthesize("Hello from Piper.");
    fs.writeFileSync("./hello.wav", audio);
  3. 03

    Use a local model

    For a voice you downloaded manually, pass the ONNX file path. Keep its matching .onnx.json file alongside it; set configPath if the config lives elsewhere.

    import { PiperNativeTTS } from "pipertts";
    
    const tts = await PiperNativeTTS.load({
      modelPath: "./models/my-voice.onnx",
    });
    const { audio, sampleRate } = await tts.synthesize("Hello.", { speakerId: 0 });
  4. 04

    API and output formats

    synthesize() returns audio (Buffer), sampleRate, and text; outputFile also writes to disk. Options: speakerId (multi-voice and multi-style), lengthScale (> 1 slows speech, < 1 speeds it up), noiseScale, noiseWScale, sentenceSilence, normalizeAudio, volume, and outputFormat (wav, raw, mp3, ogg, or opus). mp3 and opus have pure-JS fallbacks; ogg requires FFmpeg. phonemize() and synthesizeChunks() give low-level, sentence-by-sentence access.

  5. 05

    Legacy wrapper (PiperTTS)

    PiperTTS spawns python3 -m piper, one process per synthesis: about 10x slower, keep it for environments where the native engine cannot run. Requires Python and the Piper module.

    import { PiperTTS } from "pipertts";
    
    const tts = await PiperTTS.create({
      modelPath: "./models/en_US-lessac-medium.onnx",
    });
    await tts.synthesizeToFile("Hello from the legacy wrapper.", "./hello.wav");