Documentation
Make your own voice
Backseat is built on an open-source toolkit. Give it a character idea and a voice, or your own recordings, and it produces a finished Waze voice pack: all 43 prompts, checked and sized to fit. For why it works the way it does, read How Backseat works.
What you need
- Python 3.10 or newer, and ffmpeg.
- Either a text-to-speech API key (OpenAI, ElevenLabs, Hume or Fish Audio), or recordings of your own voice.
- Waze on your phone, to test the result.
Install
git clone https://github.com/ammar-adam/waze-voice-sdk
cd waze-voice-sdk
python scripts/wvs.py doctor
doctor checks Python, ffmpeg and any API keys it can see, and tells you what's
missing. Set a key as an environment variable, for example:
export OPENAI_API_KEY=... # macOS / Linux
$env:OPENAI_API_KEY = "..." # Windows PowerShell
Your first pack
The quickest start is a preset: a character that's already written. This builds Eeyore, all 43 prompts in both metric and imperial, in about a minute:
python scripts/wvs.py preflight --preset eeyore # free: checks everything except the audio
python scripts/wvs.py pack new eeyore
python scripts/wvs.py quickstart --pack eeyore --preset eeyore --accept-voice-terms
The finished pack lands in packs/eeyore/audio/export/pack/. See every preset
with python scripts/wvs.py presets list.
Write a character
A character is one JSON file in presets/. It names a voice, describes how to
speak, and gives all 43 lines. The shortest route is to copy an existing preset and
rewrite it:
{
"label": "Captain Calm",
"description": "Unflappable airline captain. Everything is fine.",
"provider": "fish",
"voice": "<model id from fish.audio>",
"direction": "Speak slowly and warmly, like a captain over the intercom.",
"provider_options": { "temperature": 0.5, "top_p": 0.6 },
"units": "both",
"lines": {
"start_drive_1": "Good evening, folks. Sit back and relax.",
"turn_left": "Turn left.",
"in_quarter_mile": "In a quarter of a mile.",
"police_ahead": "Police ahead. Nice and smooth, now.",
"arrived": "We have arrived. Thank you for flying with us.",
"...": "all 43 lines"
}
}
The rules the checks enforce
- Direction words survive.
turn_leftmust say "left", and never "right". Distances must name exactly one distance. - Repeating prompts stay plain. Turns and short distances play at every junction, so a catchphrase there is maddening. Put personality in the greetings, alerts, reroute and arrival instead. Why.
- Greetings vary. Waze picks one of nine at random. No more than three may share an opening.
- Frequent prompts stay short, 70 characters at most.
Then check and build it:
python scripts/wvs.py presets check captain-calm
python scripts/wvs.py pack new captain-calm
python scripts/wvs.py quickstart --pack captain-calm --preset captain-calm --accept-voice-terms
Run python tests/run_tests.py before opening a pull request: it checks the rules above.
Check it
Listen to it as a route before you ship it. qa stitches clips together the way
Waze does, so you hear real sequences like "in a quarter of a mile, turn left":
python scripts/wvs.py qa --pack captain-calm
With a community voice model, check that every prompt sounds like the same person. This fingerprints each clip, flags the odd ones out, and re-takes them:
python scripts/voice_consistency.py captain-calm --fix --takes 3
Then rebuild so the kept takes reach the pack:
python scripts/build_all.py --only captain-calm --no-stage --reuse
Put it in Waze
Waze has no official way to add a custom voice, so packs go up through the community
waze-voicepack-links uploader.
Clone it next to this repo (or pass --uploader <path>), then stage your pack into it (the name you give is the name drivers see in Waze), run the
uploader, and it prints a link:
python scripts/stage_for_upload.py --preset captain-calm --name "Captain Calm"
# then, in waze-voicepack-links:
python mp3_upload/main.py
Open that link on your phone and Waze offers the voice. Then confirm what's live matches what you built:
python scripts/wvs.py verify-upload <UUID> --pack-dir packs/captain-calm/audio/export/pack
Rights
Your own voice, or a voice you have permission to use, is yours to publish. Characters and
real people belong to someone else. Every preset records its rights position in a required
rights block, so a pack can't quietly skip the question. The
presets guide
explains what each field means.
Reference
- README: the full pipeline, including building from your own recordings
- Presets: writing characters, and what the build enforces
- Text-to-speech providers: trade-offs between them
- Upload runbook: getting a pack into Waze, step by step