Under the hood
How Backseat works
Waze has no API for custom voices and publishes nothing about the format. Everything here was worked out by taking real voice packs apart. This is how a character ends up in your car, and how we make sure it says the right thing at the right junction.
A voice is 43 audio files
Waze doesn't speak on the fly. A voice pack is a fixed set of 43 short MP3s, and Waze stitches them together as you drive: a distance clip, then a maneuver clip, and so on. Each file has an exact name, and a file Waze doesn't recognise is ignored without a word.
| Group | Files | What they say |
|---|---|---|
| Greetings | 9 | One plays at random when you start a route |
| Distances | 9 | Four imperial (a tenth, a quarter, half a mile, one mile), five metric |
| Maneuvers | 9 | Turn, keep left or right, exits, straight on, U-turn, "and then" |
| Roundabouts | 8 | "At the roundabout", then first to seventh exit |
| Alerts | 6 | Traffic, accident, hazard, police, speed and red-light cameras |
| Reroute, arrival | 2 | When you miss a turn, and when you get there |
The filenames lie
The imperial distance files are named in rough metres, not miles: 400.mp3 is
"in a quarter mile" and 1500.mp3 is "in one mile". The metric set is separate
(1500meters.mp3 is 1.5 km). We confirmed the mapping by downloading real packs and
transcribing every clip, because getting it wrong means a voice confidently announcing the
wrong distance. Metric and imperial ship together in every Backseat pack, so it works
whichever units your phone uses.
Where a character can talk
This was the biggest lesson from driving with the packs. A pack has one recording per prompt, and Waze plays that same file every single time. Approaching one turn you hear the maneuver clip at each distance callout: 800 m, 400 m, then at the junction. A catchphrase there is a catchphrase three times a minute, and it gets old fast.
The only thing Waze randomises is the greeting. So the writing follows how often each prompt comes round:
| Heard | Prompts | Written as |
|---|---|---|
| Every junction | Turns, short distances, roundabouts | Plain words, in the character's voice |
| At random, once a drive | Nine greetings | A catchphrase each, all different |
| Whenever the road decides | Alerts, reroutes, U-turns, long-range callouts, far exits | Catchphrases |
| Once, at the end | Arrival | A catchphrase |
The third row is what keeps a long drive fresh. Traffic, police and a missed turn happen at moments nobody can predict, so the character turns up when you don't expect it and never twice in a row at the same junction. A test fails the build if a catchphrase ever creeps onto a prompt that repeats.
Finding the right voice
Most of the voices are community models on Fish Audio. A model's title is whatever its uploader typed, and it's often wrong. So every candidate is checked three ways before it's used:
- Its cover art and description. A model called "Batman" turned out to have Robert Pattinson's Batman on the cover, not the one we wanted.
- Its pitch. Every model titled "Darth Vader" measured between 230 and 280 Hz: far too high for Vader. A model built on James Earl Jones's performance measures around 100 Hz. That's the one we use.
- Whether it can say the lines. Each finalist reads the same test lines, and a speech recogniser checks every word came out right.
The same person, every prompt
A voice model never gives the same performance twice. One take is warm, the next is flat, a third sounds like someone else. On a drive that reads as three characters sharing one voice. Two things fix it. Every Fish Audio voice runs at a lower sampling temperature, so takes vary less. Then every clip gets a fingerprint (its timbre, pitch, pitch range and brightness) and is compared against a take that's been judged right. Anything that drifts is regenerated several times, and the closest take is kept.
Every line is checked
Voice models sometimes mumble, skip a word or loop. Before any pack ships, every one of its 43 clips is transcribed by a speech recogniser and compared with its script. Fumbled takes are regenerated until they come back clean. It has caught a Godfather voice that mumbled 26 of its 43 lines, which sent that character off to find a new voice.
Fitting under the size cap
Waze reportedly rejects packs over about 0.8 MB, and says nothing when it does. Rather than squeeze every clip the same way, the exporter shares the byte budget by how often each prompt is heard: the turns you hear constantly stay crisp, and long lines you hear once a drive absorb the compression. Every clip is also levelled to the same loudness, so nothing jumps out at you mid-drive.
Shipping, and staying shipped
A finished pack is uploaded to Waze, which serves it from its own servers under a unique link. We then download it back and compare it byte for byte with what we built. Waze has no way to update a pack in place, so every fix means a new link. Every link on this site comes from one file, and a scheduled check confirms each one still works every six hours.