Notable changes. Dates are when the work landed on main.
Unreleased
Added
- Mickey Mouse replaces Hagrid. A full pack on the most-used Mickey model, run through the consistency pass and transcription, uploaded and verified. Hagrid's preset stays in the repo but is off the site and the build list.
- Faces, motion and a map. Tigger, Bugs, Paddington (blue duffle coat, fur texture), Pooh and Elmo redrawn, and Mickey drawn. Every character has its own four moves, played at random on hover and in the occasional fidget, never the same twice in a row. The page background is a faint street map that drifts slowly, and a dashed route draws itself to a pin across the hero. Reduced-motion users get none of it.
- Drive-test voice pass. New voices for Darth Vader (James Earl Jones: every
"Darth Vader" model measured 230 to 280 Hz, far too high) and Batman (Kevin
Conroy, the classic). "Recalculating" is gone from every reroute. Greetings
vary more (Bugs no longer opens every one with "Eh"), catchphrases repeat
less (Elmo, Tigger, Cookie Monster), and lines were reworked where they
didn't land. Every Fish voice runs at a lower temperature, and
scripts/voice_consistency.pyfingerprints each clip against a reference take so a pack sounds like one person. All twelve rebuilt and re-uploaded. - Site. A full-screen hero; faces that talk in five different ways and
fidget at random; redrawn Hagrid, Tigger and Bugs; install steps that match
what Waze actually does, per device; a "Who should ride next?" suggestion
board (
api/suggestions.js, needs a Redis store connected in Vercel); and two new pages: how it works, and make your own (the docs). - Catchphrases where a drive can't predict them, never where it repeats.
Waze plays one recording per prompt and only rotates the greetings, so a
catchphrase on the quarter mile or 400 m callout was heard at every junction
("if you please"). Those prompts are now plain in every pack. Every greeting
carries a catchphrase, a different one each, and so do the prompts that fire
when the road decides: alerts, reroutes, U-turns, far roundabout exits and
long-range callouts. All twelve packs rebuilt.
tests/test_repetition.pyenforces the plain prompts;docs/presets.mdexplains. - Five new voices: Hagrid, Gordon Ramsay, Darth Vader, Batman (the Bale register) and Eric Cartman. Each is a full 43-line preset on a Fish community model, built to the v3 pattern (maneuvers plain, character on distances and alerts, nine rotating greetings), uploaded, and on the site with its own face and colours. None of the five is ours (Ramsay is a real person, not a character), and their presets say so. Patrick Bateman shipped briefly and was replaced by Cartman; Vito Corleone and the Terminator were replaced by Hagrid and Gordon Ramsay.
- Voice models are checked against their cover art and description, not just their title. Two first picks were the wrong performance: "Batman" was Pattinson's (cover art from The Batman, 2022), and "Arnold Schwarzenegger" was his motivational-speaker voice rather than the T-800. Both packs were rebuilt on models whose metadata points at the right source, and re-uploaded.
- The launch film is rebuilt around the missed turn: one reroute moment,
eleven characters in escalating order. See
docs/launch-film.md. - The film renders itself.
film/is a Remotion project that turns the live packs' audio into four finished 9:16 videos with no hand edit;scripts/prepare_film.pyfeeds it. Seedocs/launch-film.md. - Vercel. The site deploys from
site/on every push, with no build step. The desktop QR codes are committed for that, and a test regenerates each one from voices.json and fails if it is stale. - backseatnav.com, the consumer site, in
site/and deployed to GitHub Pages. Each voice is a sticker face on its own colour with three real clips from its pack, the words it actually says, and one-tap install; desktop visitors get a QR code. Funnel events go to Plausible through onetrack()function, and a "Did it work?" prompt appears when the visitor comes back from Waze. site/voices.jsonis the only place a pack UUID lives. The page, the preview builder, the QR generator and the link checker all read it, and a test fails if a UUID appears anywhere else in the site.scripts/check_links.pyand a scheduled workflow that checks every voice resolves on Waze every six hours and opens an issue when one does not.scripts/build_film_audio.pyanddocs/film-line-sheet.md: the launch film's three beats cut from the live packs, so the film cannot say a line the product does not.scripts/build_previews.py: the site's preview clips, cut from existing pack audio with no API calls, with a loudness and duration check on each.
Removed
-
BUILD_PLAN.mdandPRD.md. Pre-build planning; what stayed true is indocs/and the README, and the rest is in history. -
wvs preflight. Everything checkable without an API call: preset validation, Waze filename mapping, clarity rules, and an estimated size against the cap. Prints what it cannot tell you. CI runs it. - Presets can carry
critical_provider_options, applied only to prompts a driver acts on, so a character can be fast in its greetings and unambiguous on "turn left". Tigger uses it: 1.15x normally, 1.0x for instructions.
Verified
- The distance filename mapping, from 11 real packs. Downloaded from the
community archive and transcribed offline.
200.mp3is the 0.1 mile callout,400.mp3a quarter mile,800.mp3half a mile,1500.mp3one mile: the readings the presets were written against are correct. Also confirmed the 43-filename list exactly, and that real packs sit at 53-94% of the size cap. Seedocs/waze-import-spike.md.
Changed
- Pack size is now measured from the encoded files rather than estimated, and the build reports how far measured drifted from the pre-flight estimate when that exceeds 10%.
- Clarity validation is word-boundary aware, so "leftover" no longer satisfies a left turn, and rejects lines that name two directions, two exits, or two distances.
-
Preset rights blocks now state Canada explicitly (public domain since 2007) alongside the US and the UK/EU 2027 date.
-
Character presets.
quickstart --preset eeyoreproduces a finished pack with no configuration: a licensed catalogue voice, a written delivery direction, and all 43 prompts rewritten in character. Shipseeyore,pooh, andtigger. Presets carry a mandatory rights block that is surfaced in the pack output, and the schema makes cloning a performance structurally impossible rather than merely discouraged.wvs presets list|show|check, and CI runspresets check.
Changed
-
The size allocator now aims for 85% of Waze's cap, not 100%. Packs were landing at 98.6%, which works until a voice with slower delivery produces slightly longer clips, and Waze's rejection is a greyed-out share button with no error. Builds fail above 92% and report utilisation on every run.
-
Hosted TTS providers, and
wvs quickstart. A complete pack from a voice id and an API key: no recording, no source media, no timestamps. ElevenLabs and OpenAI, over plain HTTPS from the standard library, so the fastest route to a finished pack is also the one that installs nothing.wvs voicesbrowses a provider's library;wvs doctorreports which keys are set. -
Multiple voice packs from one clone.
packs/<name>/holds a voice's own source list, audio tree, and optional config overrides. Every command takes--pack, or setWVS_PACKonce.wvs pack new|list|showmanages them. Packs share the Waze slot list and fall back toconfig/per file, so a pack usually needs only its ownsources.csv. -
CI on GitHub Actions: ruff and mypy, the test suite on Windows and Linux across Python 3.10 to 3.13, and an end-to-end pack build that uploads the resulting pack as an artifact.
CONTRIBUTING.md,SECURITY.md,CODE_OF_CONDUCT.md, issue templates (including one for real-device reports), and a pull request template.
Fixed
- Data loss.
export --export-dircleared its target recursively, so pointing it at a directory with anything else in it deleted that too. It now refuses to remove files it did not create;--forceis the explicit opt-out. max_kbpsbelow 32 withsample_rate_policy: fixedreturned 32 kbps anyway, quietly exceeding the ceiling that protects the size budget. The contradiction is now rejected when the allocation starts.- mypy is clean. Fixing it turned up
cmd_runreusing one variable for six different result types.
0.2.0
Added
- Mickey Mouse replaces Hagrid. A full pack on the most-used Mickey model, run through the consistency pass and transcription, uploaded and verified. Hagrid's preset stays in the repo but is off the site and the build list.
- Faces, motion and a map. Tigger, Bugs, Paddington (blue duffle coat, fur texture), Pooh and Elmo redrawn, and Mickey drawn. Every character has its own four moves, played at random on hover and in the occasional fidget, never the same twice in a row. The page background is a faint street map that drifts slowly, and a dashed route draws itself to a pin across the hero. Reduced-motion users get none of it.
- Drive-test voice pass. New voices for Darth Vader (James Earl Jones: every
"Darth Vader" model measured 230 to 280 Hz, far too high) and Batman (Kevin
Conroy, the classic). "Recalculating" is gone from every reroute. Greetings
vary more (Bugs no longer opens every one with "Eh"), catchphrases repeat
less (Elmo, Tigger, Cookie Monster), and lines were reworked where they
didn't land. Every Fish voice runs at a lower temperature, and
scripts/voice_consistency.pyfingerprints each clip against a reference take so a pack sounds like one person. All twelve rebuilt and re-uploaded. - Site. A full-screen hero; faces that talk in five different ways and
fidget at random; redrawn Hagrid, Tigger and Bugs; install steps that match
what Waze actually does, per device; a "Who should ride next?" suggestion
board (
api/suggestions.js, needs a Redis store connected in Vercel); and two new pages: how it works, and make your own (the docs). - The full pipeline: extract, clean, synth, normalize, qa, export, plus a
wvscommand that runs them end to end and adoctorthat reports what is missing and which step it blocks. - Waze pack building. All 43 filenames Waze recognises, both metric and imperial distance sets, and per-clip bitrate allocation to fit Waze's undocumented ~0.8 MB aggregate limit. Around 99.8% budget utilisation against a flat bitrate's 88.9%.
- Voice synthesis via Chatterbox, chosen over XTTS-v2 and F5-TTS because its weights are MIT rather than non-commercial. Optional throughout: the pipeline skips it cleanly and routes unfilled prompts to the checklist.
- Route-based QA that chains prompts the way Waze speaks them, with pass/fail verdicts and optional road-noise bed rendering.
- A test suite that generates its own audio, so a fresh clone can verify itself without any media.
Fixed
- Clip matching used a prefix glob, so
arriveclaimedarrived__take1.wavandtake10sorted beforetake2. - Demucs mode lost every phrase ID, because Demucs writes
<out>/<model>/<track>/vocals.wavrather than flat files. - Normalization used single-pass
loudnorm, which runs in dynamic mode and pumps on short material. - EBU R128 cannot measure a sub-two-second prompt. Clips are padded with silence before measurement; gating discards the padding.
- Extraction paired an input-side
-sswith an output-side one, which is applied after the filter graph, so edge fades silenced entire clips. - Cleaning could destroy a clip outright: spectral denoise can mistake a quiet, steady delivery for the noise it is removing. Loudness is now compared before and after and the clip reverts if too much was lost.
- Export numbering counted missing phrases, leaving gaps in the checklist.
- Synthesized clips were labelled as source media, because normalization rewrites
every status to
finalbefore export reads it.
Changed
- Step logic moved into a
waze_voice/package; the per-step scripts remain as thin wrappers with their original flags. - The phrase inventory was rebuilt against Waze's real prompt set. The previous list was invented and contained prompts Waze does not have.
0.1.0
- Initial scaffold: repository layout, phrase inventory, and a validator.