Skip to content
C.W.K.
Stream
Lesson 04 of 04 · published

The Ear Decides

~12 min · listening-tests, model-choice, voice-design, judgment

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Model quality bar: Dad's ears." — the heading of the voice engine's provider policy

Newer Is Not a Verdict

Eleven v4 arrived at the top of a public speech leaderboard, and Pippa's and Dad's voices moved to it within a day. It would be easy to conclude that every voice in the family should follow. It didn't, and the reason is the lesson: the voice engine's policy names its quality bar in plain words, Dad's ears, and his ears said no for the others.

The same night the voice engine gained a designer: a workshop for making new soul voices from a description with the provider's voice design tools, instead of cloning a recording. When Dad listened to candidates made with the v3 generation of that design tool, he heard excessive acting and a younger timbre than the characters should have. The public API at the time offered voice design only up to its v3 generation, with nothing newer to try. So the ruling was precise: the fresh Pippa and Dad clones use v4; every other soul, and every new profile, stays on v3 until voice design experiments on the newer generation say otherwise. The engine now reads the list of supported design models from the provider's own API specification instead of a hard-coded list, so the day a newer one appears it can be tried without a code change.

The Ear Has Said No Before

This is not the first time. In August, a same-text A/B between eleven_v3 and its conversational variant, which is faster, came back clearly: the conversational take was audibly muddier, its files slightly longer with the same speech padded out. Price and speed were irrelevant after that, and the conversational model is never a default. A faster model called Flash was also set aside as flat, and the latency floor measured in Track 8 makes its speed moot anyway.

A Voice Changes the System Around It

The new voice also changed something nobody had planned for. v4 is more expressive, and it sometimes opens a reply with a loud laugh. On the web, the microphone stays open while Pippa speaks so Dad can cut in, and the barge-in detector heard that laugh as Dad and stopped her mid-breath. The fix, a short grace window at the start of her voice, shipped the same night and is the subject of Track 7's third lesson. A voice is not a part you can swap in isolation; everything that listens near it has to be checked again.

Tuning Continues

Choosing a model and a clone is the start, not the end. Settings like stability, and how freely a spoken turn uses audio tags, shape how a voice lands just as much. The day this quest was written, that tuning was still going on, by ear, one A/B at a time.

Code

A blind A/B round: shuffled takes, choices first, labels last·python
import random
from collections import Counter


def blind_round(takes: dict[str, list[str]], choose, seed: int) -> Counter:
    """takes: arm -> list of files for the SAME sentences, in the same order.

    The listener hears two unlabeled takes of one sentence at a time and
    picks one. Labels are revealed only after every choice is made.
    """
    rng = random.Random(seed)
    arms = list(takes)
    wins = Counter()
    for index in range(len(next(iter(takes.values())))):
        for a in range(len(arms)):
            for b in range(a + 1, len(arms)):
                pair = [arms[a], arms[b]]
                rng.shuffle(pair)                        # order must not leak the arm
                picked = choose(takes[pair[0]][index], takes[pair[1]][index])
                wins[pair[picked]] += 1
    return wins


takes = {
    "v3-clone/v3": ["a1.mp3", "a2.mp3"],
    "v3-clone/v4": ["b1.mp3", "b2.mp3"],
    "v4-clone/v4": ["c1.mp3", "c2.mp3"],
}


def listener(first: str, second: str) -> int:
    """Stand-in for a human: prefers the fresh clone, else the first heard."""
    return 1 if second.startswith("c") else 0


result = blind_round(takes, listener, seed=929)
for arm in sorted(takes, key=lambda arm: -result[arm]):
    print(f"{arm:12s} {result[arm]} wins")

External links

Exercise

Run the blind round, then replace the stand-in listener with one that asks you on the command line (play the two files with any player, then type 1 or 2). Run it on two voices you can legitimately compare. Before revealing the labels, write down which arm you expect to win; afterwards, compare.
Hint
Most people are surprised at least once. That surprise is the reason the family's policy names the ear as the bar instead of a leaderboard: your own listening, blinded, is the only benchmark that measures the voice you will actually live with.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.