All docs

Reference

Pipeline design

How the six steps fit together, and why they are built the way they are.

Shape

sources.csv ─┐
             ├─> extract ──> audio/extracted/  phrase__take1.wav
your media ──┘                     │
                                   v
                             clean ──────> audio/processed/    phrase__take1.wav
                                   │
   phrases.json ──> synth ─────────┼──────> audio/synthesized/ phrase.wav
   (gaps only)                     │
                                   v
                            normalize ────> audio/master/      phrase.mp3
                                   │
                          ┌────────┴────────┐
                          v                 v
                         qa              export ──> audio/export/
                     (audition)                       pack/  <- upload this
                                                      UPLOAD_CHECKLIST.md
                                                      HOW-TO-UPLOAD.md
                                                      pack-manifest.json

Two files carry state between steps:

  • config/phrases.json is the contract. Every step reads it to know what the pack should contain. Normalization writes each phrase's status back so a fresh clone can see what is done.
  • audio/build-manifest.json is the record of what actually happened: which take a clip came from, whether it was cut or synthesized, what loudness it measured before and after, which stages it passed through. Later steps read it instead of guessing. Export uses it to mark synthetic prompts in the checklist; validation uses it to flag loudness outliers. It is regenerated freely, and is Git-ignored.

Why a library plus thin CLIs

Every step lives in waze_voice/steps/ and is exposed twice: as a subcommand of wvs, and as a standalone script in scripts/. Both call the same function.

The scaffold had six independent scripts, each with its own copy of "find the repo root", "load phrases.json", "check for ffmpeg", and "find the clip for this phrase". Four copies of a rule is four chances for them to disagree, and they had already started to: the clip-matching logic differed between normalize.py and clean.py, so a file one step produced was not necessarily a file the next step would find.

Step notes

extract

Validates the entire CSV, then checks every referenced media file exists, before running any ffmpeg. A typo in row 40 fails immediately rather than after 39 successful cuts.

Cutting uses a coarse input-side seek followed by an atrim inside the filter graph. Input seek is cheap on a two-hour file and lets a lossy decoder settle; the in-graph trim makes the cut accurate. The tempting alternative, input -ss paired with an output-side -ss/-t, is wrong: output seeking is applied after the filter graph, so a fade positioned relative to the clip gets applied to the wrong part of the stream, and the seek then selects a region the fade has already silenced. Silently. See waze_voice/media.py:cut.

Clips get 20 ms fades at both edges so a hard cut does not click.

clean

Three modes, because the right answer depends on the source and on what the user is willing to install:

  • copy for already-clean dialogue.
  • ffmpeg (default) band-limits and applies spectral denoise. No PyTorch, runs in seconds, handles room tone and hiss.
  • demucs runs real source separation and keeps the vocal stem. This is what rescues a line buried under a score.

Demucs is invoked once for all clips, not once per clip: the model loads on startup, so a per-file loop pays that cost every time. Its output layout is <out>/<model>/<track>/vocals.wav rather than flat files, so the step maps the results back to phrase IDs explicitly instead of globbing.

Cleaning is guarded. Spectral denoise decides what counts as noise from the clip's own content, and a quiet, steady delivery can look exactly like the noise it is removing. The step measures loudness before and after; if the clip lost more than clean.max_loss_lu, or went silent, the processed file is replaced with the original and the revert is reported. A noisy prompt is recoverable, a silent one is not.

The guard covers Demucs too, and there it fires for a different reason: Demucs keeps what it recognises as a vocal and discards everything else, so pointing it at material that is not speech returns near-silence for the whole batch. Reverts across most of a batch mean the source is wrong for the mode, not that the clips are bad.

synth

Fills phrases that have no audio anywhere. Default backend is Chatterbox zero-shot cloning: it takes the user's own cleaned clips as a reference and speaks new lines in that voice without any training run. A navigation pack yields well under a minute of usable source audio, far below what fine-tuning needs and comfortably above what zero-shot conditioning needs.

Chatterbox was chosen over XTTS-v2 and F5-TTS on licensing, not on quality or Python support. Both alternatives ship weights restricted to non-commercial use, and in Coqui's case the company that would have to grant a commercial licence no longer exists. Chatterbox is MIT including its weights. For a project whose premise is "bring audio you have the right to use", a default that silently caps every output at non-commercial would be the wrong trade. Both alternatives remain selectable via --backend. See tts.md.

A backend is just a callable: speak this text, in this voice, to this file. Adding one means writing a loader that returns such a callable, which keeps model-specific API differences out of the step logic. Backend-specific generation knobs go through synth.generate_options untouched, so a library renaming a parameter does not require a change here.

Synthesis is optional throughout. wvs run checks availability first and skips the step with one clear line rather than aborting the run, and any phrase left unfilled is carried into the export checklist as a prompt to record manually.

normalize

Measure, then apply one static gain. Not ffmpeg's single-pass loudnorm, which runs in dynamic mode: it compresses to hit the target as it goes, which pumps on short material and gives different results depending on where the speech sits in the clip.

The full order matters:

  1. Trim leading and trailing silence.
  2. Measure the trimmed length, then fade the new edges and add controlled padding.
  3. Measure loudness of that file, shaped exactly as it will ship.
  4. Apply the gain and a true-peak limiter, encode to MP3.
  5. Re-measure, and if the result drifted more than 0.3 LU, correct once and re-render.

Steps 2 and 3 are in that order for a reason. Measuring before shaping lets padding and fades move the result afterwards, which on a sub-second prompt is worth several LU. The correction pass in step 5 covers MP3 encoding and limiter engagement.

On the fixture, this brings clips whose inputs span 16 dB to within 0.1 LU of each other.

qa

Renders each route step as one continuous piece of audio. Real navigation chains "In 500 meters" onto "turn right" as a single instruction, then leaves a longer gap before the next one. Playing eighteen clips alphabetically tells you almost nothing; hearing them chained at driving pace is what surfaces the clip that is a beat too slow or lands at the wrong emphasis.

Interactive by default, recording pass/fail per instruction into audio/qa-report.json. --render writes the whole route to a file instead, optionally over a road-noise bed via --bed, so you can listen on the car audio you will actually navigate with.

export

Produces a directory Waze will accept: files named exactly as Waze expects, both unit systems, and every clip compressed so the pack total lands inside the aggregate size budget.

Two modules back it:

  • wazepack.py holds the 43 filenames, which of them are core, and which unit system each distance callout belongs to. Waze ignores unrecognised names without an error, so the list is data rather than something assembled at the call site, and the phrase validator rejects a name that is not on it.
  • budget.py decides how many bits each clip gets. See below.

After allocating, the step encodes, measures what landed on disk, and walks clips down a rung at a time until the real total fits. Predicted size is bitrate times duration; the real file also carries frame headers and padding, and that gap always runs the wrong way. Reporting a predicted figure the user then fails an upload with would be worse than useless, so the number printed is the number on disk.

Being over budget fails the run. --allow-missing forgives gaps in the pack; it does not forgive an oversized one, because that upload would be rejected silently.

The bitrate allocator

Waze's limit is aggregate, so the question is how to divide a fixed number of bytes between forty-odd clips. The community tooling uses one bitrate for everything, found by binary search. That is safe, and it spends the same bits per second on a nine-second greeting heard once per drive as on "turn left".

Model quality as rising with the log of bitrate, weight each clip by how much its quality matters, and maximise sum(w_i * log(b_i)) subject to sum(b_i * d_i) <= B. The Lagrangian gives b_i = lambda * w_i / d_i: bitrate proportional to weight over duration. Short clips get more bits per second than long ones, important clips more than incidental ones, on a defensible scale rather than a guess.

Lambda is found by bisection rather than the closed form. The closed form only holds while nothing is clamped, and in practice clips hit the floor, hit the ceiling, and get snapped to the discrete MP3 ladder. Total size is monotonic in lambda, so bisection converges on the largest feasible value with all of that folded into evaluating a candidate. An earlier freeze-as-you-go implementation had a subtle bug this avoids: a clip pinned to the floor early, while the budget still looked tight, never got revisited once other clips hit the ceiling and freed room up, so it stayed at 24 kbps while the pack ran 47% under budget.

Below 32 kbps MP3 has no rungs at 44.1 kHz, so those clips drop to 22.05 kHz. Almost nothing in a navigation prompt lives above 11 kHz, so that is a good trade.

On a full 43-prompt pack the weighted strategy reaches about 99.8% budget utilisation against the uniform strategy's 88.9%, with the extra bits going to the prompts actually heard. Both remain available via --strategy.

When even the floor overshoots, the plan says so and the export fails rather than producing an unusable pack. Silently dropping below the quality floor would trade a rejected upload for one that uploads and sounds terrible.

Packs

A pack is one voice. packs/<name>/ holds that voice's sources.csv, its audio/ tree, and optionally its own phrases.json, routes.json, or pipeline.json.

Everything resolves through waze_voice/paths.py, so selecting a pack is one call (set_active_pack) made before any path is looked up. Doing it late would leave half a run reading one tree and half another, which is why the CLI sets it as the first thing after parsing arguments.

Precedence for the audio tree:

  1. WVS_AUDIO_ROOT, so a test or a one-off can redirect without touching packs.
  2. The active pack, from --pack or WVS_PACK.
  3. The shared audio/ tree.

Config falls back per file rather than per pack: a pack without phrases.json uses the shared one, a pack with it uses its own. The Waze slot list does not change with the voice filling it, so most packs override nothing.

Pack names are validated against a strict pattern rather than escaped, because they become directory names and the export step deletes recursively inside whatever directory it is handed. --pack .. resolving to the repo root is not a mistake worth being clever about.

Two packs deliberately produce the same Waze filenames. That is required: Waze matches prompts by filename. Packs are separate uploads that you switch between on the device, not something that merges.

Take resolution

Which file represents a phrase is decided in one place, waze_voice/takes.py, searching processed then synthesized then extracted.

Matching is exact. A prefix match would let the phrase arrive claim arrived__take1.wav, and lexical sorting would put take10 ahead of take2. Both produce a finished pack that says the wrong thing with no error anywhere. The __take separator exists to make the split unambiguous, and both cases are pinned by tests.

Testing

python tests\run_tests.py

tests/fixtures.py generates synthetic source media with ffmpeg: modulated tone bursts at known timestamps, at deliberately different levels, optionally over a background bed. The end-to-end test runs the real pipeline over it and checks the output audio, not just that files appeared.

The bursts are modulated rather than pure tones on purpose. A steady sine looks exactly like noise to a spectral denoiser, so an unmodulated fixture would test the cleaner against a signal no real recording produces, and would pass while the real behaviour went unchecked.

Tests needing ffmpeg skip cleanly when it is absent. The suite uses unittest from the standard library, so it runs with nothing installed.

Edit on GitHub Source: docs/pipeline.md