Kaggle lists a 30-hour weekly baseline, which may be higher depending on demand and available resources. The quota resets on Saturday at 00:00 UTC.
FIELD GUIDE · PIPER TTS
From recordings
to a voice.
The essential steps for preparing voice data and training a Piper model. Kaggle limits can change, so check your quota before a long run.
Create an LJSpeech dataset for Piper TTS
LJSpeech is the standard format expected by Piper. Each sentence is a separate WAV file, paired with its transcript in metadata.csv.
data.zip
├── metadata.csv # filename|transcript
└── wavs/
├── 000001.wav
└── ...- 01
Collect the audio files
Use recordings you are allowed to process and redistribute. Sources might include game wikis, media you own, or public datasets.
Find LJSpeech datasets on Hugging Face ↗Download source media with yt-dlp - 02
Split the audio into sentences
Create one WAV file per sentence. For automatic splitting or alignment, these tools can help:
- audiosplitter · automatic splitting
- Montreal Forced Aligner (MFA) · forced alignment
- faster-whisper + VAD · transcription and timestamps
- Audacity · manual edits for smaller collections
- 03
Create metadata.csv
Use one line per file. Separate the filename and transcript with a vertical bar; the filename must match the file in wavs/.
000001.wav|This is an example sentence. 000002.wav|Here is another training sentence. - 04
Zip the archive
Place metadata.csv and the wavs/ directory at the archive root.
zip -r data.zip metadata.csv wavs/ - 05
Upload to Hugging Face
Create a dataset repository, for example user/wheatley-english-ljspeech. Add data.zip and a README.md with this configuration:
--- configs: - config_name: default data_files: - split: train path: data.zip --- - 06
Post in the Piper forum
Start a thread in the datasets and training channel with the Hugging Face link and the appropriate tag.
Open the Discord channel ↗
Train a voice on Kaggle
Kaggle assigns GPU quota to your account, and the allowance can vary with availability. Check the counter before starting a long session.
Kaggle documentation lists a continuous 12-hour limit for CPU and GPU sessions. CPU-only notebooks do not use GPU quota, but their session length is still limited.
Kaggle retired the Tesla P100 on September 15, 2026. T4 x2 remains available in notebooks; choose it for Piper training instead of following older P100 guides.
Resume after an interruption
A full training run often exceeds a 12-hour session. This project’s Piper notebook resumes from a Hugging Face checkpoint: after a timeout or an exhausted quota, run it again to continue from the latest saved checkpoint.
Before a long training run
- Check your remaining quota in notebook settings, your profile, or the accelerator usage page.
- Stop GPU sessions when you are not using them.
- Use T4 x2; P100 is no longer available on Kaggle.
Kaggle quotas, accelerators, and interfaces can change. Always check the information shown in your account before training.
Use a Piper voice
Install Piper, get the two files for a voice, and generate a WAV locally. The commands below use the version maintained by the Open Home Foundation.
- 01
Install Piper
Piper is distributed as a Python package. Older guides using the rhasspy/piper repository or its binary archive refer to the historical release.
python -m pip install piper-ttspython -m piper.download_voices en_US-lessac-medium - 02
Download a voice
Choose a voice from the catalogue and download its ONNX model and matching JSON configuration file. Keep both files together; the JSON path is the model path followed by .json.
voice.onnx
voice.onnx.json - 03
Generate a WAV file
First download a voice from the Piper catalogue, or replace the name below with the path to your .onnx file.
python -m piper -m ./voice.onnx -f output.wav -- "This is a speech synthesis test." - 04
Read text from a file
python -m piper -m ./voice.onnx -f output.wav --input-file mon_texte.txtStream raw audio to playback
The sample rate depends on the voice. Check voice.onnx.json and change 22050 if needed.
echo "Hello." | python -m piper -m ./voice.onnx --output-raw | aplay -r 22050 -f S16_LE -c 1 -t raw - 05
From Python
import wave from piper import PiperVoice voice = PiperVoice.load("./voice.onnx") with wave.open("output.wav", "wb") as wav_file: voice.synthesize_wav("Hello world.", wav_file) - 06
With Home Assistant
Home Assistant connects to Piper through Wyoming. Add Piper under Settings → Devices & services, then use the tts.speak action with the Piper TTS entity.
Set up Piper in Home Assistant ↗With Node-RED
The node-red-contrib-piper-tts package mentioned in some guides could not be confirmed in the public catalogue. One option is to call the Piper command from Node-RED’s built-in Exec node; adapt paths and permissions to your setup.
- 07
Adjust the output
Quality depends in part on the training data and selected voice. A higher length-scale slows speech; a lower value speeds it up. noise-scale and noise-w adjust variation in the generated audio.
python -m piper -m ./voice.onnx --length-scale 1.5 --noise-scale 0.667 --noise-w 0.8 -f slow.wav -- "Slow speech." python -m piper -m ./voice.onnx --length-scale 0.7 --noise-scale 0.667 --noise-w 0.8 -f fast.wav -- "Fast speech."
Use Piper in a Node.js project with pipertts
pipertts brings Piper to Node.js, with file output or in-memory audio. The recommended mode is PiperNativeTTS: fully native in-process inference (ONNX plus a compiled espeak bridge), no Python, roughly 10x faster than the wrapper. Voices (.onnx plus .onnx.json) remain separate files.
- 01
Install the dependencies
The docs specify Node.js 18+ and TypeScript 5+. Native mode needs nothing else: plain npm install is enough. The classic wrapper mode (PiperTTS through python3 -m piper) additionally requires Python and the Piper module. Selecting a catalogue ID can download the model and its config into models/.
npm install pipertts # wrapper mode only: python3 -m pip install piper-tts - 02
Quick start with a catalogue voice
import { PiperNativeTTS, resolveModelPathFromOptions } from "pipertts"; import * as fs from "node:fs"; const modelPath = await resolveModelPathFromOptions({ model: "en_US-lessac-medium", modelsDir: "./models", }); const tts = await PiperNativeTTS.load({ modelPath }); const { audio } = await tts.synthesize("Hello from Piper."); fs.writeFileSync("./hello.wav", audio); - 03
Use a local model
For a voice you downloaded manually, pass the ONNX file path. Keep its matching .onnx.json file alongside it; set configPath if the config lives elsewhere.
import { PiperNativeTTS } from "pipertts"; const tts = await PiperNativeTTS.load({ modelPath: "./models/my-voice.onnx", }); const { audio, sampleRate } = await tts.synthesize("Hello.", { speakerId: 0 }); - 04
API and output formats
synthesize() returns audio (Buffer), sampleRate, and text; outputFile also writes to disk. Options: speakerId (multi-voice and multi-style), lengthScale (> 1 slows speech, < 1 speeds it up), noiseScale, noiseWScale, sentenceSilence, normalizeAudio, volume, and outputFormat (wav, raw, mp3, ogg, or opus). mp3 and opus have pure-JS fallbacks; ogg requires FFmpeg. phonemize() and synthesizeChunks() give low-level, sentence-by-sentence access.
- 05
Legacy wrapper (PiperTTS)
PiperTTS spawns python3 -m piper, one process per synthesis: about 10x slower, keep it for environments where the native engine cannot run. Requires Python and the Piper module.
import { PiperTTS } from "pipertts"; const tts = await PiperTTS.create({ modelPath: "./models/en_US-lessac-medium.onnx", }); await tts.synthesizeToFile("Hello from the legacy wrapper.", "./hello.wav");
