Technology · Agentic → Voice Agents · Open Source · wiki:deep

Piper

Piper is a fast, local neural text-to-speech engine. Active development lives at OHF-Voice/piper1-gpl (Open Home Foundation; pip install piper-tts, GPL-3.0). It embeds espeak-ng for phonemization, ships voice models you run on-device, and exposes CLI, Python, HTTP, and C/C++ APIs. It is the speak substrate under voice agents — not a full STT→LLM→TTS product (that job stays with Pipecat / LiveKit Agents / Vocode).

The older rhasspy/piper README points development to the OHF repo; treat OHF as primary and the Rhasspy tree as legacy redirect.

Why it matters here

Whisper covers listen. Without a local TTS pack, “voice round-trip” only exists as a stage name inside pipeline frameworks. Piper is the research default when spoken replies must stay on-device (privacy, offline Studio, Home Assistant-class boxes) and still bind to the same gate rules as text agents: speech is channel chrome; approve / spend / mutate stay elsewhere.

How it works

Text enters Piper; espeak-ng phonemizes; a neural voice model synthesizes audio. Operators pick a voice ONNX (or trained voice), call the CLI / Python API / HTTP server, and stream or write WAV/raw audio to a transport that Pipecat or LiveKit Agents already owns.

  1. Install piper-tts (or build libpiper).
  2. Download or train a voice model.
  3. Synthesize text → audio (CLI, Python, or HTTP).
  4. Hand audio to the voice-agent transport; keep tool side effects on the gated path.

Related: whisper · pipecat · livekit-agents · vocode · topics/03-channels

Flow

When to reach for it

  • Use when: you need local neural TTS as the speak half of a voice agent (offline, privacy, or HA-style boxes) and already have STT + agent loop elsewhere.
  • Skip when: you only need a hosted voice API with no local weights, or you need a full telephony agent product — prefer Vocode / LiveKit Agents as the product and keep Piper as optional synthesizer.
  • Prefer instead: Whisper for STT; pipeline packs for STT→LLM→TTS composition; commercial TTS SaaS when local GPL voices are unacceptable.

Limits

  • License: piper1-gpl is GPL-3.0 — embedding in proprietary closed agents has copyleft implications; confirm counsel before shipping.
  • Maintainer signal: OHF is actively seeking maintainers; treat release cadence as an ops risk.
  • Quality / language: voice coverage and naturalness vary by ONNX; measure barge-in and latency on your transport, not demo WAVs alone.
  • Scope: TTS is not the ledger, not the approval authority, and not STT.

What we checked

Claims below are backed by science sources on disk.

Product / local TTS framing

OHF-Voice/piper1-gpl README · STRONG

“A fast and local neural text-to-speech engine that embeds espeak-ng for phonemization.”

Legacy redirect (Rhasspy → OHF)

rhasspy/piper redirect · STRONG

“Development has moved: https://github.com/OHF-Voice/piper1-gpl”
Research inventory

9 tags · 8 out · 9 in · 3 artifacts · 1 gaps · 0 corpus docs

Catalog tags

landscape.layer
Agentic
landscape.subcategory
Voice Agents
license_tag
Open Source
maps.dm
absent
maps.features
4
one_liner
Local neural TTS (speak half for voice agents)
review.depth
science
slug
piper
title
Piper

Artifacts

  • dm_map · absent
  • features_map · present · technologies/piper/features.md
  • readme · present · technologies/piper/README.md

Out · alternative_to

Out · maps_to

In · alternative_to

In · in_stack

In · maps_to