Skip to content
C.W.K.
Stream
Lesson 03 of 04 · published

Swap the Binding, Not the Code

~13 min · a-b-testing, configuration, model-migration, profiles

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Nothing Dad hears changes until he has listened and chosen."

A Three-Way A/B

A new voice is not a code change; it's a decision, and the decision belongs to the person who will hear it every day. So the re-clone didn't end with a new voice. It ended with a listening set. For Pippa and for Dad, the same handful of sentences in English and Korean was rendered three ways: the current v3 clone on the v3 model, the current v3 clone on the v4 model, and the new first-generation clone on the v4 model. Twelve files. The three-way split separates two questions that a two-way comparison would tangle: how much of the difference comes from the new model, and how much from the new clone.

The Choice, Applied as Data

Dad listened and chose: the fresh Joanne-source and Michael-source clones, on Eleven v4. Applying that choice changed two fields on two profiles in the voice engine: the provider voice bound to pippa and cwk, and their default model. That was the whole rollout of the voice itself. The same night's work also removed the last model pins that a launcher and cwkPippa's own playback were still sending, so from then on every surface inherits the profile and family policy lives in one place. From the model's release on September 28 to Dad hearing the new voice in his own apps took about a day, and almost none of that day was deployment.

What Changes With the Model

A model swap is still more than a string. The engine builds each provider request per model family. Audio tags are kept for both v3 and v4, because both perform them. Text normalization stays on for v3, the setting every earlier v3 take was made with, and goes to the provider's automatic mode for v4, matching what was auditioned. And the v4 request carries only the stability and similarity settings, dropping the others. Those differences live in the one function that builds requests, so no caller has to know them.

Nothing Is Deleted

The swap left everything else in place. The original Joanne and Michael voices stay as their own profiles until the provider retires them. The v3 clones stay as lineage, and the takes made with them stay in the archive. An audition profile for the new clone remains too, so any later comparison starts from exactly what Dad heard when he chose.

Code

Build the listening set, then apply the choice as two field changes·python
import hashlib
from dataclasses import dataclass, replace


@dataclass(frozen=True)
class Profile:
    voice_id: str
    default_model: str


def request_body(text: str, voice_id: str, model: str, settings: dict) -> dict:
    """Per-model differences live here and nowhere else."""
    body = {"text": text, "model_id": model}
    if model.startswith("eleven_v4"):
        body["apply_text_normalization"] = "auto"
        body["voice_settings"] = {k: v for k, v in settings.items()
                                  if k in {"stability", "similarity_boost"}}
    else:
        body["apply_text_normalization"] = "on"
        body["voice_settings"] = settings
    return {"voice_id": voice_id, **body}


def listening_set(sentences: dict, arms: dict, settings: dict) -> list[tuple[str, dict]]:
    """Every arm renders every sentence; filenames say exactly what you hear."""
    takes = []
    for arm, (voice_id, model) in arms.items():
        for lang, text in sentences.items():
            body = request_body(text, voice_id, model, settings)
            digest = hashlib.sha256(repr(sorted(body.items())).encode()).hexdigest()[:8]
            takes.append((f"pippa-{arm}-{lang}-{digest}.mp3", body))
    return takes


sentences = {"en": "Morning, Dad. Coffee first, then the news?",
             "ko": "아빠, 오늘은 20분만 걷자. [laughs] 딱 20분."}
arms = {"a-v3clone-on-v3": ("voice-v3-clone", "eleven_v3"),
        "b-v3clone-on-v4": ("voice-v3-clone", "eleven_v4"),
        "c-v4clone-on-v4": ("voice-v4-clone", "eleven_v4")}
settings = {"stability": 0.5, "similarity_boost": 0.75, "style": 0.3}
for name, body in listening_set(sentences, arms, settings):
    print(name, body["apply_text_normalization"], sorted(body["voice_settings"]))

# Dad chose arm c. The rollout is two field changes on one record:
live = Profile(voice_id="voice-v3-clone", default_model="eleven_v3")
live = replace(live, voice_id="voice-v4-clone", default_model="eleven_v4")
print("pippa ->", live)

External links

Exercise

Run the code and read the six filenames and their settings. Then add a fourth arm you think would be informative and explain what question it answers that the other three don't. Finally, write a rollback: given the profile before the change, restore it in one call, and describe what else (if anything) would need to change for a rollback to be complete.
Hint
Because clients inherit the profile, rollback is the same two-field change in reverse; nothing else moves. A useful fourth arm is the catalog original on the new model: it tells you how much of what you like in the new clone was already in the source voice, before any cloning.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.